Abstract
Using a randomised experiment in 200 Bangladeshi villages, we evaluate the impact of an over-the-phone learning support intervention (telementoring) among primary school children and their mothers during Covid-19 school closures. Post-intervention, treated children scored 35% higher on a standardised test, and the homeschooling involvement of treated mothers increased by 22 minutes per day (26%). We also found that the intervention forestalled treated children's learning losses. When we returned to the participants one year later, after schools briefly reopened, we found that the treatment effects had persisted. Academically weaker children benefited the most from the intervention that only cost USD20 per child.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was at the individual household level, but the intervention is one-on-one telementoring/tutoring, so the personal-teaching exception applies and student-level RCT is acceptable.
- "The intervention involves one-on-one direct tutoring of children, which is recognised as an effective method to boost learning (Nickow et al., 2024)." (p. 2420)
Relevant Quotes:
1) "Our unit of randomisation was at the individual level." (p. 2423)
2) "We randomly assigned half of 838 households (419) to the treatment arm—those who received weekly telementoring—and the remaining half (419) to the control arm—no telementoring was provided." (p. 2424)
3) "The intervention involves one-on-one direct tutoring of children, which is recognised as an effective method to boost learning (Nickow et al., 2024)." (p. 2420)
4) "Each mentor was randomly assigned to two primary school children in the same grade and their mothers (mentees)." (p. 2425)
5) "Our sample consists of 838 mother–child dyads distributed across 200 villages in five subdistricts of the Khulna Division." (p. 2423)
Detailed Analysis:
Randomisation was explicitly conducted at the individual (household) level rather than at the class or school level. Normally this would fail criterion C. However, the ERCT standard provides an exception: if the intervention is designed for personal teaching such as tutoring, then student-level randomisation is acceptable. This intervention is precisely one-on-one telementoring: each mentor tutored a child and mentored the child's mother over the phone in weekly 30-minute one-to-one sessions, with a mentor-to-mentee ratio of 1:2. In addition, schools were completely closed during the intervention, and the 838 households were spread across 200 villages, so classroom-based contamination between treatment and control children was structurally unlikely.
Criterion C is met because the intervention is one-on-one personal tutoring/mentoring, which qualifies for the tutoring exception that permits individual-level randomisation.
-
E
Exam-based Assessment
- Outcomes were measured with tests created by the researchers to mirror textbook content, not with a widely recognised standardised exam.
- "All test questions were created by closely following existing textbooks developed by the National Curriculum and Textbook Board, Bangladesh." (p. 2425)
Relevant Quotes:
1) "Learning outcomes were measured using a standardised one-on-one assessment test: word translation, fill-in-the-blanks, additions, etc." (p. 2425)
2) "All test questions were created by closely following existing textbooks developed by the National Curriculum and Textbook Board, Bangladesh." (p. 2425)
3) "Due to school closures and exam cancellations in Bangladesh for over two years, we could not use school administrative data. Instead, we designed our tests to mirror primary school exams, which are closely following the textbooks (but not copied directly)." (footnote 7, p. 2425)
4) "We conducted our own pandemic-adjusted assessment as schools were closed for around two years and exams were cancelled, so administrative data from schools were not available." (p. 2432)
Detailed Analysis:
Although the abstract describes the assessment as a "standardised test", the standardisation refers only to uniform administration and scoring across children, not to the use of a recognised external exam. The authors explicitly state that they designed the test questions themselves, following the national textbooks but "not copied directly", because school exams were cancelled and administrative data were unavailable. Under the ERCT standard, criterion E requires a standard, widely recognised exam (e.g., a national or state-wide standardised test) that was not created for the purposes of the study. A researcher-constructed instrument, even one aligned with the national curriculum, risks being overly aligned with the tutored content; indeed the paper notes "the tutoring component of the intervention directly maps into our main learning outcomes" (p. 2425), which is exactly the alignment problem this criterion guards against.
Criterion E is not met because the outcome assessment was a custom, researcher-designed test rather than a widely recognised standardised exam.
-
T
Term Duration
- The intervention began in early September 2020 and outcomes were measured in January 2021 (about 4.5 months later) and again in December 2021, exceeding one full academic term.
- "The intervention ran for 13 weeks in late 2020 when all schools were closed. One month after the intervention ended (in January 2021), we conducted standardised learning assessments among children and surveys among mothers to evaluate the immediate impact." (p. 2420)
Relevant Quotes:
1) "We recruited student volunteers from various local public universities as mentors to provide learning support to primary school children (grades 1–3) and homeschooling support to their mothers every week for 13 consecutive weeks (from early September to late December 2020)." (p. 2423)
2) "The intervention ran for 13 weeks in late 2020 when all schools were closed. One month after the intervention ended (in January 2021), we conducted standardised learning assessments among children and surveys among mothers to evaluate the immediate impact." (p. 2420)
3) "We then returned to the participants one year later (in December 2021)—when schools briefly reopened—and conducted a second round of standardised learning assessments and surveys to evaluate the medium-term impact." (p. 2420)
Detailed Analysis:
The intervention started in early September 2020. Primary outcomes were first measured in January 2021, about 4.5 months after the intervention began, which exceeds the minimum of one full academic term (approximately 3-4 months). Moreover, a second round of outcome measurement occurred in December 2021, roughly 15 months after the intervention start, which also satisfies the stronger year-duration requirement and therefore automatically satisfies this weaker term-duration criterion.
Criterion T is met because the interval from intervention start (early September 2020) to outcome measurement (January 2021 and December 2021) covers at least one full academic term.
-
D
Documented Control Group
- The control group is thoroughly documented with baseline demographics, baseline test scores, balance tests and a clear statement that it received no support or alternative learning opportunities.
- "In the treatment group (419 households), mother-child dyads received weekly telementoring, while those in the control group (419 households) did not receive any support." (p. 2420)
Relevant Quotes:
1) "In the treatment group (419 households), mother–child dyads received weekly telementoring, while those in the control group (419 households) did not receive any support. Note that the control group did not have access to alternative learning opportunities, as online/over-the-phone teachings were unavailable and access to television and radio was very limited in rural areas." (p. 2420)
2) "Table 1 reports our baseline sample characteristics by treatment and control status, where children are about 7.5 years old and 50% are female, and parents have about 6 years of education and earn BDT11,500 (USD135) per month... Importantly, these characteristics are balanced across the two arms (joint F-test p = 0.60)." (p. 2426)
3) "Table 1. Baseline Sample Characteristics and Balance Checks... Control N = 419... Child age in years 7.40... English literacy score of children (out of 30) 16.24 ... Numeracy score of children (out of 20) 14.75" (p. 2427)
4) "Finally, attritions across the two arms are statistically similar (T-test p > 0.10)." (p. 2426)
Detailed Analysis:
The paper documents the control group in detail: its size (419 households), demographic composition (child age, gender, parental education, income, religion, siblings, land size), baseline English literacy and numeracy scores, homeschooling time and private tutoring access, all reported separately by arm in Table 1 with balance tests. The conditions of the control group are clearly stated: no telementoring, no support, and no access to alternative remote learning during the closures. Attrition is also analysed by arm. This satisfies the requirement for detailed control group documentation enabling proper comparison.
Criterion D is met because the control group's size, demographics, baseline performance and (absence of) treatment are clearly and quantitatively documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation was performed at the individual household level, not at the school or institutional level.
- "Our unit of randomisation was at the individual level." (p. 2423)
Relevant Quotes:
1) "Our unit of randomisation was at the individual level." (p. 2423)
2) "We randomly assigned half of 838 households (419) to the treatment arm—those who received weekly telementoring—and the remaining half (419) to the control arm—no telementoring was provided." (p. 2424)
3) "Not all villages include both treatment and control households. As a result, we use union council fixed effects..." (footnote 8, p. 2427)
Detailed Analysis:
Criterion S requires randomisation among schools or equivalent implementing institutions. Here, individual households (mother-child dyads) were randomised within a pool drawn from 200 villages; neither schools nor villages were the unit of assignment. The tutoring exception that rescues criterion C applies only to the class-level criterion, not to the stronger school-level requirement. No quote indicates school-level or site-level random assignment.
Criterion S is not met because randomisation occurred at the individual household level rather than at the school or institutional level.
-
I
Independent Conduct
- The co-authors designed the intervention, trained the mentors, and led the implementation and analysis themselves; blinded enumerators help data quality but there was no independent evaluation body.
- "Two co-authors of this study, Hashibul Hassan and Asad Islam, conducted the training." (p. 2424)
Relevant Quotes:
1) "We partnered with a research NGO, Global Development and Research Initiative (GDRI), to implement and evaluate our telementoring intervention using a randomised controlled trial (RCT) in rural Bangladesh." (p. 2423)
2) "Two co-authors of this study, Hashibul Hassan and Asad Islam, conducted the training." (p. 2424)
3) "A small team from GDRI did support Hashibul Hassan and Asad Islam in the implementation, but their role was limited to distributing printed copies of textbook solutions to treatment households and conducting a rapid survey on the mentors only." (p. 2424)
4) "To ensure data quality, enumerators that measured outcomes were kept blind to the treatment status." (p. 2425)
5) "This is because the programme leveraged the unpaid involvements of two primary investigators and volunteers during implementation." (p. 2422)
6) "As we consider the scalability of these programmes, we encourage that future research be conducted through independent implementing bodies." (p. 2436)
Detailed Analysis:
The intervention was designed by the authors, and two co-authors personally trained the mentors and led the implementation; the same author team conducted the analysis. GDRI's implementation role was explicitly "limited", and GDRI is the authors' local field partner rather than an independent third-party evaluator commissioned to run the trial. The positive elements—outcome enumerators blind to treatment status and non-overlapping with the 2019 survey staff—improve measurement quality but do not amount to independent conduct of the study: unlike the ERCT exception examples, there is no statement that the intervention designers were removed from data collection, analysis, or the formulation of conclusions. The authors themselves recommend that future research use independent implementing bodies, implicitly acknowledging this limitation.
Criterion I is not met because the intervention designers themselves implemented and analysed the study without an independent evaluation team or third-party oversight.
-
Y
Year Duration
- Outcomes were re-measured in December 2021, about 15 months after the intervention began in early September 2020, exceeding 75% of an academic year of tracking.
- "We then returned to the participants one year later (in December 2021)—when schools briefly reopened—and conducted a second round of standardised learning assessments and surveys to evaluate the medium-term impact." (p. 2420)
Relevant Quotes:
1) "We recruited student volunteers from various local public universities as mentors to provide learning support to primary school children (grades 1–3) and homeschooling support to their mothers every week for 13 consecutive weeks (from early September to late December 2020)." (p. 2423)
2) "We then returned to the participants one year later (in December 2021)—when schools briefly reopened—and conducted a second round of standardised learning assessments and surveys to evaluate the medium-term impact." (p. 2420)
3) "The positive impacts persisted one year after the intervention ended: 0.30 SD (19%) higher in English and 0.44 SD (20%) higher in numeracy." (p. 2420)
Detailed Analysis:
Criterion Y requires that outcomes be measured at least 75% of a full academic year after the intervention begins, even if the intervention itself is shorter. The intervention itself lasted only 13 weeks, but the study tracked the same participants and re-administered the full learning assessments and surveys in December 2021, approximately 15 months after the intervention started in early September 2020. This tracking interval comfortably exceeds one full academic year, satisfying the year-duration requirement (and by implication the weaker term-duration criterion T).
Criterion Y is met because the interval from intervention start (early September 2020) to the second endline measurement (December 2021) is about 15 months, exceeding a full academic year.
-
B
Balanced Control Group
- The additional tutoring/mentoring time is itself the treatment variable being tested against a business-as-usual control during school closures, so the unbalanced control is by design.
- "Using a randomised experiment in 200 Bangladeshi villages, we evaluate the impact of an over-the-phone learning support intervention (telementoring) among primary school children and their mothers during Covid-19 school closures." (p. 2418)
Relevant Quotes:
1) "Using a randomised experiment in 200 Bangladeshi villages, we evaluate the impact of an over-the-phone learning support intervention (telementoring) among primary school children and their mothers during Covid-19 school closures." (p. 2418)
2) "Children received weekly tutoring (30 minutes per session) on mathematics and English... and mothers received homeschooling mentoring over the phone (telementoring hereinafter), which was not otherwise available to them." (p. 2419)
3) "In the treatment group (419 households), mother–child dyads received weekly telementoring, while those in the control group (419 households) did not receive any support. Note that the control group did not have access to alternative learning opportunities, as online/over-the-phone teachings were unavailable and access to television and radio was very limited in rural areas." (p. 2420)
4) "Through GDRI, treated mothers were also provided with printed solutions to textbook problems and a study plan..." (p. 2423)
5) "our intervention only costs USD20 per mother–child dyad, making it low-cost and policy relevant." (p. 2422)
Detailed Analysis (Criterion B decision-tree applied):
The treatment group received additional educational resources: 13 weekly 30-minute mentoring calls (6.5 hours of direct dosage), weekly text messages, and printed textbook solutions with a study plan, while the control group received nothing beyond business as usual, i.e., extra resources are present and non-negligible. The key question is whether these resources are integral to, and themselves constitute, the treatment being tested. They are: the study's explicit stated purpose is to evaluate the impact of providing over-the-phone learning support (telementoring) during school closures when no other support was available; the additional tutoring and mentoring time and accompanying materials are the intervention itself, not a separable confounding add-on. This matches the ERCT exception in which the control group may remain at the standard business-as-usual level when additional resources are the primary treatment variable being tested (RESOURCES_ARE_ TREATMENT branch of the decision tree). The printed solutions and text messages are documented components of the telementoring package rather than optional extras.
Criterion B is met because the extra tutoring and mentoring inputs are explicitly the treatment variable being tested against a business-as-usual control, so the resource difference is integral to the study design.
-
Level 3 Criteria
-
R
Reproduced
- No independent team has replicated this specific telementoring study in another context in a peer-reviewed journal; the paper's own reproducibility check is a code/data verification, not a field replication, and related phone-tutoring trials are distinct studies rather than replications.
- "Our key contribution, relative to these existing studies, is that we show volunteer-delivered learning support via basic mobile phones can be particularly effective in addressing learning losses in poor environments." (p. 2422)
Relevant Quotes:
1) "Our key contribution, relative to these existing studies, is that we show volunteer-delivered learning support via basic mobile phones can be particularly effective in addressing learning losses in poor environments." (p. 2422)
2) "For instance, Angrist et al. (2022) show that weekly phone calls and test messages from an non-governmental organisation (NGO) to parents of primary school-aged children in Botswana, over five weeks, improved the learning outcomes of children by 0.12 SD." (p. 2421)
3) "Crawfurd et al. (2023) find that 15-minute weekly tutoring calls with children from their school teachers in Sierra Leone, over 16 weeks, increased educational engagement by parents (0.31 SD) and children (0.34 SD), but did not affect test scores." (p. 2421)
4) "However, there are several programme features that are different from Angrist et al. (2022) and Crawfurd et al. (2023). First, our primary focus was on mentoring and guiding mothers to enhance homeschooling quality and engagement... whereas other studies emphasised directly tutoring children." (footnote 3, p. 2422)
5) "The data and codes for this paper are available on the Journal repository. They were checked for their ability to reproduce the results presented in the paper. The replication package for this paper is available at the following address: https://doi.org/10.5281/zenodo.10696444." (p. 2418)
Detailed Analysis:
The paper positions itself as novel relative to the closest comparable studies. The phone-based trials it cites (Angrist et al. 2022 in Botswana; Crawfurd et al. 2023 in Sierra Leone; Wang et al. 2023 in Bangladesh) are contemporaneous or prior studies of different interventions, and the authors explicitly enumerate the design differences (mother-focused mentoring, volunteer mentors, texting, dosage); Crawfurd et al. also found no test-score impact, so they cannot be read as confirming replications. The paper itself documents only a computational reproducibility check of its own data and code (quote 5, p. 2418), performed as part of the Economic Journal's publication process; this is a code/data verification exercise, not an independent field replication of the intervention in a new context, and does not satisfy criterion R. An independent citation search (OpenAlex, July 2026) of the six papers currently citing this study found: Carlana and La Ferrara (2025, AER) on online tutoring in Italy; Gallego, Molina and Neilson (2025) on television-based information provision in Peru; Hassan, Islam, Kayes and Wang (2025, Economics of Education Review) on parenting style and out-of-school learning in rural Bangladesh; Anger et al. (2026, European Economic Review) on online tutoring in Germany; and Dong and Lin (2026) on head teachers in China. None of these is an independent replication of this specific telementoring intervention by a different research team in a different context; they are distinct interventions and study designs. No independent field replication of this specific telementoring study was found.
Criterion R is not met because no independent, peer-reviewed field replication of this specific study in a different context exists; related studies are distinct experiments rather than replications, and the paper's own reproducibility check is limited to code/data verification.
-
A
All-subject Exams
- Although all four core subjects were assessed, the assessments were custom researcher-made tests, so the prerequisite criterion E fails and criterion A fails with it.
- "There were four segments in the test: English (6 questions, 30 points), numeracy (5 questions, 30 points), Bangla (4 questions, 20 points) and general knowledge (4 questions, 20 points)." (p. 2425)
Relevant Quotes:
1) "There were four segments in the test: English (6 questions, 30 points), numeracy (5 questions, 30 points), Bangla (4 questions, 20 points) and general knowledge (4 questions, 20 points)." (p. 2425)
2) "We also find positive spillovers on two other core subjects taught in Bangladeshi schools, Bangla and general knowledge, which were not targeted by the intervention." (p. 2420)
3) "All test questions were created by closely following existing textbooks developed by the National Curriculum and Textbook Board, Bangladesh." (p. 2425)
Detailed Analysis:
The study commendably assessed both targeted subjects (English, numeracy) and non-targeted core subjects (Bangla, general knowledge), which is the breadth of coverage that criterion A seeks. However, criterion A explicitly requires that the all-subject assessment use standard standardised exam-based assessments, and criterion E is a prerequisite: if E is not met, A cannot be met. As established under criterion E, all four test segments were custom instruments created by the researchers because school exams had been cancelled; they are not widely recognised standardised exams.
Criterion A is not met because the multi-subject assessment relied on custom researcher-designed tests, failing the standardised-exam prerequisite (criterion E).
-
G
Graduation Tracking
- Participants in grades 1-3 were tracked for only about one year after the intervention, not until graduation from primary school.
- "We then returned to the participants one year later (in December 2021)—when schools briefly reopened—and conducted a second round of standardised learning assessments and surveys to evaluate the medium-term impact." (p. 2420)
Relevant Quotes:
1) "We recruited student volunteers from various local public universities as mentors to provide learning support to primary school children (grades 1–3)..." (p. 2423)
2) "We then returned to the participants one year later (in December 2021)—when schools briefly reopened—and conducted a second round of standardised learning assessments and surveys to evaluate the medium-term impact." (p. 2420)
3) "A further novelty of our study is that we demonstrate both immediate and one-year impacts of an intervention that was implemented and evaluated amid the pandemic." (p. 2422)
Detailed Analysis:
The children were in grades 1-3 at baseline, so primary school graduation (end of grade 5 in Bangladesh) would occur two to four years after the intervention. Measurement stopped at the December 2021 endline, roughly one year after the intervention ended, when children were in grades 2-4 at most. No quote describes tracking the cohort through the end of primary education. The authors' companion paper (Hassan et al., 2023, AEA Papers and Proceedings) examines child mental health at the same two endlines and does not extend tracking to graduation. An independent citation search (OpenAlex, July 2026) of papers citing this study found one further paper by three of the four original authors (Hassan, Islam and Wang, with Imrul Kayes replacing Abu Siddique), "Building education resilience through parenting style and out-of-school learning: Field experimental evidence from rural Bangladesh" (Economics of Education Review, 2025); its full text/abstract could not be retrieved to confirm whether it follows the same cohort of children, and no quote confirming graduation tracking of this cohort was found. No follow-up publication tracking this specific cohort through primary-school graduation was located.
Criterion G is not met because follow-up ended about one year after the intervention, well before the participants graduated from primary school, and no later publication confirming graduation tracking was found.
-
P
Pre-Registered
- The trial was registered at the AEA RCT registry (AEARCTR-0006395) on 21 September 2020, confirmed directly against the registry record, before endline outcome data collection began in January 2021, so the criterion is met.
- "This trial is registered at the AEA RCT registry: AEARCTR-0006395, and the ethical clearance was received from Monash University, Australia: project no. 25039."
Relevant Quotes:
1) "This trial is registered at the AEA RCT registry: AEARCTR-0006395, and the ethical clearance was received from Monash University, Australia: project no. 25039." (p. 2418, footnote)
2) "Therefore, we pre-registered that our intervention was expected to have positive impacts on several parenting perceptions that are related to the weekly themes." (p. 2433)
3) "In Hassan et al. (2023), we explored an additional, un(pre)registered outcome related to children's mental health..." (footnote 12, p. 2435)
4) "The intervention ran for 13 weeks in late 2020 when all schools were closed. One month after the intervention ended (in January 2021), we conducted standardised learning assessments among children and surveys among mothers." (p. 2420)
Internet Search (Registry Verification):
The AEA RCT registry entry AEARCTR-0006395, "Educational inequality and parental involvement during the COVID-19 pandemic: Randomised controlled experiment of a telementoring program in rural Bangladesh," was directly checked at socialscienceregistry.org. It confirms a first registration date of 21 September 2020, IRB approval on 13 July 2020, an intervention start date of 4 September 2020, and pre-specified primary outcomes (children's literacy/numeracy and parental involvement) and secondary outcomes (children's social preferences and parenting perceptions) consistent with the outcomes reported in the paper.
Detailed Analysis:
Criterion P requires pre-registration of the protocol, including hypotheses, methods and planned analyses, before data collection begins. The paper explicitly cites its registration at the AEA RCT registry (AEARCTR-0006395) and repeatedly distinguishes pre-registered from un(pre)registered outcomes, indicating a substantive registered protocol. The directly-verified AEA registry entry confirms a first registration date of 21 September 2020. The intervention began in early September 2020 (registry: 4 September 2020), so registration occurred about two weeks after the intervention started; however, it clearly preceded all outcome data collection, which began at the first endline in January 2021 (the baseline measures came from a pre-existing 2019 survey). The standard's stated check is that registration "must be before data collection began"; the outcome data used to evaluate the intervention were collected months after registration, so hypotheses and analyses were fixed before any endline outcomes existed, satisfying the criterion's anti-selective-reporting purpose.
Criterion P is met because the trial was registered at the AEA RCT registry (AEARCTR-0006395, confirmed 21 September 2020) before outcome data collection began in January 2021, as verified directly against the registry record.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.