From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria

Martín De Simone, Federico Tiberti, Maria Barron Rodriguez, Federico Manolio, Wuraola Mosuro, Eliot Jolomi Dikoru

Published:
ERCT Check Date:
DOI: 10.1596/1813-9450-11125
  • language arts
  • K12
  • Africa
  • blended learning
  • EdTech platform
0
  • C

    Randomization was at the student level, but the intervention is explicitly a personalized tutoring program (an AI chatbot acting as a virtual tutor), so the ERCT tutoring exception applies and student-level randomization is acceptable.

    "We evaluate a six-week after-school tutoring program in Nigeria that used a publicly available LLM (ChatGPT-4) to support students in learning English."

  • E

    The primary outcome was a custom multiple-choice assessment designed by experts specifically for this study, and the secondary outcome was a school-administered term exam, neither of which is a widely recognized standardized exam.

    "The assessment was administered in a traditional pencil-and-paper format and consisted of multiple-choice questions designed by experts based on the Nigerian curriculum."

  • T

    Outcomes were measured about six weeks after the intervention began (sessions 3 June to 11 July 2024, assessment 11-12 July 2024), which is well short of a full academic term.

    "The program was implemented over a six-week period between June and July 2024, targeting first-year senior secondary school students, who are typically 15 years old."

  • D

    The control group's size, demographics, baseline test scores, and business-as-usual condition are documented in detail, including a full balance table.

    "Once the period to express interest closed, the randomization was carried out ... to assign them either to the treatment group, which participated in the program, or to the control group, which did not receive any intervention but continued their regular learning in the classroom."

  • S

    Randomization was conducted among individual students within nine pre-selected schools, not among schools.

    "The randomization for the pilot program was conducted at the student level in the nine selected schools."

  • I

    The World Bank author team designed the intervention (prompts, toolkits, training) and also conducted the evaluation, with no statement of an independent third-party evaluator.

    "The lesson guides and their prompts were carefully crafted to position the LLM as a tutor, focusing on facilitating learning rather than simply providing direct answers."

  • Y

    The whole study, from intervention start to final measurement, lasted about six weeks, far short of 75 percent of an academic year, and criterion T already fails.

    "The program was implemented over a six-week period between June and July 2024..."

  • B

    The intervention's extra inputs (after-school sessions, computer lab time, teacher facilitation) are the integral treatment package explicitly being tested against a business-as-usual control, which the ERCT standard accepts as balanced by design.

    "The study analyzes the effects of an after-school program in which students interacted with a large language model twice per week to improve their English skills, following the national curriculum."

  • R

    No independent, peer-reviewed replication of this specific Nigerian Copilot tutoring RCT exists; the authors themselves call for replication, and only a computational reproducibility package (not an independent replication) is available.

    "Given the nascent application of LLMs in education, numerous questions remain unanswered, underscoring the importance of replicating this study, including with small variations."

  • A

    Criterion E fails, and the study measured only English plus AI and digital skills, not all main school subjects with standardized exams.

    "...a standardized assessment designed to measure three key outcomes: (a) English language proficiency ... (b) knowledge of AI, and (c) understanding of basic digital concepts."

  • G

    Measurement ended immediately after the six-week program with no follow-up toward graduation, and prerequisite criterion Y is not met.

    "Understanding the long-term impacts of the intervention is also crucial. Future studies should investigate whether the positive effects observed in the short term persist over time..."

  • P

    The paper contains no mention of a pre-registered protocol or registry entry, and no pre-registration for this trial was found in an external search.

Abstract

This study evaluates the impact of a program leveraging large language models for virtual tutoring in secondary education in Nigeria. Using a randomized controlled trial, the program deployed Microsoft Copilot (powered by GPT-4) to support first-year senior secondary students in English language learning over six weeks. The intervention demonstrated a significant improvement of 0.31 standard deviation on an assessment that included English topics aligned with the Nigerian curriculum, knowledge of artificial intelligence and digital skills. The effect on English, the main outcome of interest, was of 0.23 standard deviations. Cost-effectiveness analysis revealed substantial learning gains, equating to 1.5 to 2 years of 'business-as-usual' schooling, situating the intervention among some of the most cost-effective programs to improve learning outcomes. An analysis of heterogeneous effects shows that while the program benefits students across the baseline ability distribution, the largest effects are for female students, and those with higher initial academic performance. The findings highlight that artificial intelligence-powered tutoring, when designed and used properly, can have transformative impacts in the education sector in low-resource settings.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomization was at the student level, but the intervention is explicitly a personalized tutoring program (an AI chatbot acting as a virtual tutor), so the ERCT tutoring exception applies and student-level randomization is acceptable.
      • "We evaluate a six-week after-school tutoring program in Nigeria that used a publicly available LLM (ChatGPT-4) to support students in learning English."
      • Relevant Quotes: 1) "The randomization for the pilot program was conducted at the student level in the nine selected schools." (p. 8) 2) "Once the period to express interest closed, the randomization was carried out using simple random sampling without replacement among interested students to assign them either to the treatment group, which participated in the program, or to the control group, which did not receive any intervention but continued their regular learning in the classroom." (p. 9) 3) "We evaluate a six-week after-school tutoring program in Nigeria that used a publicly available LLM (ChatGPT-4) to support students in learning English." (p. 2) 4) "The intervention aimed to improve learning outcomes in English language classes using an AI chatbot as a virtual tutor." (p. 6) 5) "Students interacted with the LLM by asking questions, completing exercises, and receiving personalized feedback." (p. 8) 6) "First, the randomization was conducted at the student level rather than at the school level, a design feature that may have led to spillover effects, as students in the control group may have interacted with peers in the treatment group during regular school hours, potentially diffusing the impact of the intervention." (p. 12) Detailed Analysis: The unit of randomization was clearly the individual student within nine schools, not the class or the school, and the authors themselves acknowledge the resulting spillover risk. Ordinarily this would fail criterion C. However, the ERCT standard provides an explicit exception: if the intervention is designed for personal teaching like tutoring, student-level randomization is considered acceptable. This paper frames the program throughout as tutoring: an after-school "tutoring program" in which an LLM acts as a "virtual tutor" that adapts to each student's individual learning level and provides personalized feedback, positioned by the authors as a response to Bloom's two-sigma one-on-one tutoring problem. The core treatment is personalized (AI) tutoring rather than a classroom-level teaching method, so the tutoring exception applies. Criterion C is met because, although randomization was at the student level, the intervention is a personalized tutoring program and therefore falls under the ERCT tutoring exception.
    • E

      Exam-based Assessment

      • The primary outcome was a custom multiple-choice assessment designed by experts specifically for this study, and the secondary outcome was a school-administered term exam, neither of which is a widely recognized standardized exam.
      • "The assessment was administered in a traditional pencil-and-paper format and consisted of multiple-choice questions designed by experts based on the Nigerian curriculum."
      • Relevant Quotes: 1) "At the end of the six-week intervention, participating and non-participating students completed a standardized assessment designed to measure three key outcomes: (a) English language proficiency aligned with the Nigerian curriculum for the corresponding period (our main outcome of interest), (b) knowledge of AI, and (c) understanding of basic digital concepts." (p. 10) 2) "To minimize the risk of cheating, multiple versions of the assessment were created, each with a randomized order of questions." (p. 10) 3) "The assessment was administered in a traditional pencil-and-paper format and consisted of multiple-choice questions designed by experts based on the Nigerian curriculum." (p. 10) 4) "The weights were based on the ex-ante difficulty of each test item, which was determined by the test designers prior to the administration." (p. 10) 5) "In addition to the intervention-specific assessment, an additional dependent variable was derived from the student's final English exam scores. This exam, which was conducted independently by the school, covered the entire term's content, which extended beyond the six-week period of the after-school program." (p. 10) Detailed Analysis: Although the authors call their endline instrument a "standardized assessment," the quotes make clear it was a study-specific test created by the program's own test designers ("designed by experts based on the Nigerian curriculum," with multiple study versions and designer-set difficulty weights). It is not a named, widely recognised standardized exam such as a national curriculum examination (e.g., WAEC/NECO). The secondary outcome, the third-term English exam, is a regular curricular exam set and conducted by each school; it is a routine school exam rather than a recognised external standardized test, and it was not the study's primary outcome. The ERCT standard requires that outcomes be measured with standard, widely recognised standardized exams, which is not the case here. Criterion E is not met because the main outcome was a custom-made study assessment and the supplementary measure was an internal school term exam, not widely recognised standardized tests.
    • T

      Term Duration

      • Outcomes were measured about six weeks after the intervention began (sessions 3 June to 11 July 2024, assessment 11-12 July 2024), which is well short of a full academic term.
      • "The program was implemented over a six-week period between June and July 2024, targeting first-year senior secondary school students, who are typically 15 years old."
      • Relevant Quotes: 1) "The program was implemented over a six-week period between June and July 2024, targeting first-year senior secondary school students, who are typically 15 years old." (p. 6) 2) "At the end of the six-week intervention, participating and non-participating students completed a standardized assessment..." (p. 10) 3) "Pilot sessions 6/3/24 7/11/24 ... Standardized assessment 7/11/24 7/12/24 ... Third term examination 7/12/24 7/12/24 ... Endline questionnaire 7/14/24 7/14/24" (Table 14, p. 42) Detailed Analysis: The intervention started on 3 June 2024 and the primary outcome assessment was administered on 11-12 July 2024, immediately at the end of the six-week program; the school third-term exam was taken on 12 July 2024. The interval from intervention start to outcome measurement is therefore approximately six weeks. The ERCT standard requires outcomes to be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Six weeks is roughly half a term at best, and no later follow-up measurement is reported. Criterion T is not met because the interval from intervention start to outcome measurement was only about six weeks, less than one full academic term.
    • D

      Documented Control Group

      • The control group's size, demographics, baseline test scores, and business-as-usual condition are documented in detail, including a full balance table.
      • "Once the period to express interest closed, the randomization was carried out ... to assign them either to the treatment group, which participated in the program, or to the control group, which did not receive any intervention but continued their regular learning in the classroom."
      • Relevant Quotes: 1) "Once the period to express interest closed, the randomization was carried out using simple random sampling without replacement among interested students to assign them either to the treatment group, which participated in the program, or to the control group, which did not receive any intervention but continued their regular learning in the classroom." (p. 9) 2) "Initially, 657 students were assigned to the treatment group and 671 to the control group. However, only 422 students in the treatment group and 337 in the control group completed the final assessment, which constitutes the final sample used for the analysis." (p. 9) 3) "Table 1 provides summary statistics and balance tests for key observable characteristics of the two groups. Demographic variables include gender, age, and a socio-economic status (SES) index." (p. 9) 4) "The SES index, as well as other variables such as the proportion of female students and age, shows that the sample is balanced across the treatment and control groups, with differences that are small and not statistically significant." (p. 9) 5) "The difference between treatment and control group students in mean baseline scores for the First Term Exam is 0.131 (SE = 0.073), and for the Second Term Exam, it is 0.096 (SE = 0.073)." (pp. 9-10) Detailed Analysis: The paper documents the control group thoroughly: its size at assignment and endline, its condition (no intervention, regular classroom learning), and detailed baseline characteristics (gender, age, SES index, prior first- and second-term exam scores, and school composition) presented with balance tests in Table 1 (p. 29). Attrition is also analysed with Lee bounds and inverse probability weighting. The paper does transparently note that some control students inadvertently accessed sessions in early weeks, which is itself part of the documentation of the control condition. This level of description satisfies the requirement for a well-documented control group. Criterion D is met because the control group's composition, baseline performance, and treatment-as-usual condition are documented in detail with balance tests.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomization was conducted among individual students within nine pre-selected schools, not among schools.
      • "The randomization for the pilot program was conducted at the student level in the nine selected schools."
      • Relevant Quotes: 1) "The randomization for the pilot program was conducted at the student level in the nine selected schools." (p. 8) 2) "The selection of schools was based on the availability of computer labs." (p. 6) 3) "First, the randomization was conducted at the student level rather than at the school level, a design feature that may have led to spillover effects..." (p. 12) Detailed Analysis: Criterion S requires random assignment of entire schools (or equivalent implementing units) to conditions. Here the nine schools were purposively selected based on computer lab availability, and all nine schools contained both treatment and control students because randomization occurred among individual interested students within schools. The authors explicitly contrast their design with school-level randomization. No school-level assignment took place. Criterion S is not met because randomization was at the student level within schools, not at the school level.
    • I

      Independent Conduct

      • The World Bank author team designed the intervention (prompts, toolkits, training) and also conducted the evaluation, with no statement of an independent third-party evaluator.
      • "The lesson guides and their prompts were carefully crafted to position the LLM as a tutor, focusing on facilitating learning rather than simply providing direct answers."
      • Relevant Quotes: 1) "The team would like to thank Scherezad Latif and Halil Dundar, Education Practice Managers, World Bank. The team extends its appreciation to Dr. Joan Osa Oviawe and Jennifer Aisuan, for their collaboration throughout the implementation of the pilot, as well as Alex Twinomugisha, Robert Hawkins, and Cristobal Cobo for their support with the intervention." (p. 1, footnote) 2) "The lesson guides and their prompts were carefully crafted to position the LLM as a tutor, focusing on facilitating learning rather than simply providing direct answers. These prompts were informed by principles from the science of learning and were tailored to the cultural context of southern Nigeria..." (p. 7) 3) "Each teacher was provided with a three-part implementation toolkit which included: a) curated online learning resources on the use of Copilot and LLMs; b) a handbook focused on AI literacy and potential risks and benefits; and c) session guidelines, including suggested initial prompts and potential follow-up questions to assist students if needed." (p. 7) 4) "To ensure the fidelity of program implementation, monitors were first trained, provided with monitoring guidelines, and then assigned to track student attendance and gather information about each session using Kobo Toolbox." (p. 8) 5) "The team acknowledges the financial support received from the Mastercard Foundation." (p. 1, footnote) Detailed Analysis: Criterion I requires that the study be conducted independently of the intervention's designers, or at least that data collection and analysis be performed by a third party with documented independence. Here the World Bank author team designed the intervention package (session guides, prompts, teacher toolkit, training, and monitoring system), oversaw its implementation with local partners, and carried out the data analysis and reporting themselves. Although the underlying LLM (Microsoft Copilot) is a third-party product, the intervention as tested was designed by the same team that evaluated it. Monitors and teachers were part of the program apparatus rather than an independent evaluation agency, and the paper contains no statement of external evaluation, blinded outcome assessment, or third-party oversight of data collection and analysis of the kind accepted in ERCT exception examples. Criterion I is not met because the same team designed, implemented, and evaluated the intervention with no documented independent conduct or oversight.
    • Y

      Year Duration

      • The whole study, from intervention start to final measurement, lasted about six weeks, far short of 75 percent of an academic year, and criterion T already fails.
      • "The program was implemented over a six-week period between June and July 2024..."
      • Relevant Quotes: 1) "The program was implemented over a six-week period between June and July 2024..." (p. 6) 2) "Pilot sessions 6/3/24 7/11/24 ... Standardized assessment 7/11/24 7/12/24" (Table 14, p. 42) 3) "Further analysis predicts substantial gains with extended program duration, estimating an increase of between 1.2 and 2.2 standard deviations for a full academic year of participation, depending on attendance rates." (p. 3) Detailed Analysis: Criterion Y requires outcomes to be measured at least 75 percent of an academic year (roughly 9-10 months) after the intervention begins. The entire study window here is six weeks (3 June to 12 July 2024). The authors themselves only extrapolate what a full-year program might achieve, explicitly because their study did not run that long. Additionally, since the weaker criterion T (term duration) is not met, criterion Y cannot be met. Criterion Y is not met because tracking lasted only about six weeks, far below 75 percent of an academic year.
    • B

      Balanced Control Group

      • The intervention's extra inputs (after-school sessions, computer lab time, teacher facilitation) are the integral treatment package explicitly being tested against a business-as-usual control, which the ERCT standard accepts as balanced by design.
      • "The study analyzes the effects of an after-school program in which students interacted with a large language model twice per week to improve their English skills, following the national curriculum."
      • Relevant Quotes: 1) "The study analyzes the effects of an after-school program in which students interacted with a large language model twice per week to improve their English skills, following the national curriculum." (p. 6) 2) "Each student was allowed to participate in a maximum of two 1.5-hour after-school sessions per week." (p. 6) 3) "...or to the control group, which did not receive any intervention but continued their regular learning in the classroom." (p. 9) 4) "In other words, we interpret that the intervention as a whole -which includes the interaction with the LLM and teacher guidance with specific prompts- is driving the results." (p. 21) 5) "We have reasons to believe that the effects are not driven solely by the additional time with teachers, given that the impact of human tutoring tends to be very low when is not one-on-one or in small groups." (pp. 21-22) 6) "Similarly, additional arms can help disentangle the multiple causal mechanisms that might be driving the effect, including additional instructional time and the interaction with the chatbot with teacher support." (p. 20) Detailed Analysis: The treatment group clearly received extra educational inputs relative to the control group: up to twelve 90-minute after-school sessions, computer and internet access, and supervising teachers, while the control group continued business-as-usual classroom learning with no compensating activity. Following the ERCT decision tree: extra resources are present and non-negligible, so the deciding question is whether these resources are integral to the treatment being tested. The paper defines the intervention as the after-school AI-tutoring program itself, i.e., the additional sessions with the LLM and teacher guidance ARE the treatment package whose effect is being estimated against business-as-usual schooling, analogous to the ERCT exception examples (e.g., after-school para-instructor classes and DPL tools tested against business-as-usual). The authors are transparent that additional instructional time is one possible mechanism, argue it is unlikely to be the sole driver, and propose future arms to disentangle it, but as designed the extra time and technology are integral, not a separable confounding add-on. It should be clearly noted that the intervention involved additional after-school instructional time, hardware, internet, and teacher supervision not provided to controls. This was re-checked against the current ERCT decision-tree definition of criterion B and the conclusion is unchanged. Criterion B is met because the additional after-school time and resources are integral to the tutoring intervention explicitly tested against a business-as-usual control.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent, peer-reviewed replication of this specific Nigerian Copilot tutoring RCT exists; the authors themselves call for replication, and only a computational reproducibility package (not an independent replication) is available.
      • "Given the nascent application of LLMs in education, numerous questions remain unanswered, underscoring the importance of replicating this study, including with small variations."
      • Relevant Quotes: 1) "Given the nascent application of LLMs in education, numerous questions remain unanswered, underscoring the importance of replicating this study, including with small variations." (p. 6) 2) "A verified reproducibility package for this paper is available at http://reproducibility.worldbank.org" (p. 1) 3) "In Ghana, students who were given access to a phone for one hour a week and were allowed to use an AI-powered math tutor via a messaging app to independently study math improved their scores much more than those without access, with an effect size of 0.36 (Henkel et al., 2024)." (p. 5) Detailed Analysis: Criterion R requires that this specific study be independently replicated by a different team in a different context and published in a peer-reviewed journal. The paper itself frames the program as one of the first of its kind in a developing country and explicitly calls for future replication. The World Bank "verified reproducibility package" (hosted at reproducibility.worldbank.org, catalog entry 419) is a computational reproduction of the paper's own analysis from its own data, not an independent replication by another team. Related studies (e.g., Henkel et al. 2024 in Ghana; Bastani et al. 2024 in Turkiye) are different interventions that predate or parallel this study and do not replicate this Nigerian Copilot-based English tutoring trial. A fresh internet search (July 2026, more than a year after publication) for citing papers, replication studies, and follow-ups of this specific trial found only news coverage, commentary (VoxDev, AITrends.ng, AEI podcast), the reproducibility package, and related-but-distinct studies, with no independent peer-reviewed replication of this specific study. Criterion R is not met because no independent replication of this specific study has been published.
    • A

      All-subject Exams

      • Criterion E fails, and the study measured only English plus AI and digital skills, not all main school subjects with standardized exams.
      • "...a standardized assessment designed to measure three key outcomes: (a) English language proficiency ... (b) knowledge of AI, and (c) understanding of basic digital concepts."
      • Relevant Quotes: 1) "At the end of the six-week intervention, participating and non-participating students completed a standardized assessment designed to measure three key outcomes: (a) English language proficiency aligned with the Nigerian curriculum for the corresponding period (our main outcome of interest), (b) knowledge of AI, and (c) understanding of basic digital concepts." (p. 10) 2) "In addition to the intervention-specific assessment, an additional dependent variable was derived from the student's final English exam scores." (p. 10) 3) "For example, future studies could examine whether familiarity with AI tools in English language lessons enhances student performance in other subjects, such as mathematics or science." (p. 20) Detailed Analysis: Criterion A requires standardized exam-based assessment of all main subjects taught at the educational level, and criterion E is an explicit prerequisite. Criterion E is not met here (the assessment was custom-made), so criterion A automatically fails. Moreover, outcomes covered only English language plus AI and digital skills; core senior secondary subjects such as mathematics and science were not assessed, and the authors themselves flag cross-subject effects as future research. No justified specialisation exception applies since English is one of several core subjects at this level. Criterion A is not met because criterion E fails and only English-related outcomes (plus AI/digital skills) were measured, not all main subjects.
    • G

      Graduation Tracking

      • Measurement ended immediately after the six-week program with no follow-up toward graduation, and prerequisite criterion Y is not met.
      • "Understanding the long-term impacts of the intervention is also crucial. Future studies should investigate whether the positive effects observed in the short term persist over time..."
      • Relevant Quotes: 1) "Understanding the long-term impacts of the intervention is also crucial. Future studies should investigate whether the positive effects observed in the short term persist over time, contributing to lasting improvements in students' academic trajectories." (p. 20) 2) "Standardized assessment 7/11/24 7/12/24 ... Third term examination 7/12/24 7/12/24 ... Endline questionnaire 7/14/24 7/14/24 ... Evaluation analysis 7/15/24 8/30/24" (Table 14, p. 42) 3) "Considering that our participants are in Senior Secondary 1, we use t = 3 as the remaining expected years in school." (p. 19, footnote 14) Detailed Analysis: Participants were first-year senior secondary students with about three years remaining until graduation, but all data collection concluded in July 2024, immediately after the six-week program. The authors explicitly identify long-term follow-up as future work, confirming that no graduation tracking was conducted. A fresh internet search (July 2026) for follow-up publications tracking this same cohort of Benin City students (by De Simone, Tiberti, Barron Rodriguez, Manolio, Mosuro, or Dikoru) found only coverage and commentary on the original pilot; no subsequent paper reporting on this cohort's later outcomes was found. Additionally, since criterion Y is not met, criterion G cannot be met under the ranking instructions. Criterion G is not met because tracking stopped at the end of the six-week intervention and students were not followed to graduation, and no subsequent tracking paper was found.
    • P

      Pre-Registered

      • The paper contains no mention of a pre-registered protocol or registry entry, and no pre-registration for this trial was found in an external search.
      • Relevant Quotes: 1) "Randomization was conducted without stratification using a computerized system. Although the randomization process did not incorporate a fixed random seed, the allocation results were documented and saved, ensuring the assignments are reproducible and transparent, as recommended by Bruhn and McKenzie (2009)." (p. 9, footnote 6) 2) "A verified reproducibility package for this paper is available at http://reproducibility.worldbank.org" (p. 1) Detailed Analysis: Criterion P requires the full study protocol (hypotheses, methods, planned analyses) to be registered on a public registry before data collection began. The paper never mentions a trial registry (e.g., AEA RCT Registry, ClinicalTrials.gov, RIDIE), a registration ID, or a pre-analysis plan. The transparency measures it does describe (saved randomization allocations and a post-publication World Bank reproducibility package) are ex-post and do not constitute pre-registration. A fresh internet search of the AEA RCT Registry and general web for a pre-registration of this trial (by these authors, this intervention, Benin City schools) found none. Criterion P is not met because there is no evidence of a pre-registered protocol before data collection.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.