Abstract
Background: Information Communication Technology (ICT) holds the promise of enabling low-cost teacher professional development at scale. An expert coach, for example, could reach far more teachers virtually, thus reducing salary and transport costs. But the benefits of in-person interaction -- such as developing relationships of trust and accountability -- might be difficult to replicate. We ask whether virtual coaching can be equally effective at improving teaching practices and student learning, compared to in-person coaching, in 230 public primary schools in poor communities in Limpopo province, South Africa. We evaluate two different modalities of providing continuous teacher professional development for teaching English as a Second Language (ESL): on-site vs virtual coaching, implemented over three years and tracked for the same cohort of students from grade one (February 2017) to grade three (November 2019). We find that, after three years, on-site coaching improved students' English oral language and reading proficiency (0.31 and 0.13 SD, respectively). Virtual coaching had a smaller impact on English oral language proficiency (0.12 SD), no impact on English reading proficiency, and an unintended negative effect on home language literacy. On-site coaching is more cost-effective. The main finding of this paper is sobering -- a virtual coaching alternative, which was somewhat less expensive and considerably less reliant on human resources, did not have the same desired effect, and actually reduced home language literacy.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- The study randomized whole schools (not individual classes or students) to conditions, which is stronger than class-level randomization and therefore automatically satisfies this criterion.
- "Working with South Africa's Department of Basic Education (DBE), we randomly assigned 100 schools to receive either virtual or on-site coaching, and another 80 schools to the control..." (p. 1)
Relevant Quotes:
1) "Working with South Africa's Department of Basic Education (DBE), we randomly assigned 100 schools to receive either virtual or on-site coaching, and another 80 schools to the control, where teachers could still receive business-as-usual PD support provided by government." (p. 1-2)
2) "We then created 10 strata of similar schools, based on school size, socio-economic status and previous performance in a standardized national exam, and randomly assigned five schools to each intervention group and eight to the control group. Thus we randomly assigned 50 schools to each intervention and 80 to the control." (p. 4)
Detailed Analysis:
The paper explicitly describes randomization at the school level, stratified by school size, socio-economic status and prior exam performance, with entire schools (not individual classes or students) assigned to on-site coaching, virtual coaching, or control. Under the ERCT standard, a school-level (or stronger) RCT automatically satisfies the weaker class-level requirement, since it eliminates any risk of within-school or within-class contamination between arms. Both quotes were checked against the PDF and are verbatim (the first quote straddles the page 1/2 boundary in the PDF).
Criterion C is met because randomization occurred at the school level, which exceeds the class-level requirement.
-
E
Exam-based Assessment
- The paper explicitly states the assessments were researcher-designed ("designed to evaluate students' language and literacy abilities") and "adjusted each year," only loosely modeled on "EGRA-type tasks" -- this is a study-specific custom instrument, not administration of an established, standardized, validated exam.
- "The student assessments were designed to evaluate students' language and literacy abilities at the end of each grade, but were not designed to necessarily benchmark student performance against curriculum requirements. Given this focus, the assessments included the EGRA-type tasks..." (p. 5)
Relevant Quotes:
1) "The components of the student assessments were adjusted each year to assess the oral language and decoding skills expected by the end of each year." (p. 5)
2) "The student assessments were designed to evaluate students' language and literacy abilities at the end of each grade, but were not designed to necessarily benchmark student performance against curriculum requirements. Given this focus, the assessments included the EGRA-type tasks, and care was taken to minimize a floor effect." (p. 5)
3) "In the final wave of data collection, the oral assessment included eight tasks assessing oral and reading proficiency in HL and ESL. These tasks included HL letter recognition, HL oral reading fluency and comprehension, ESL expressive vocabulary, ESL listening comprehension, ESL word reading and ESL oral reading fluency and comprehension." (p. 5)
Detailed Analysis:
On re-verification, the quote citations in the prior ranking were corrected: all three quotes above appear on page 5 of the PDF (the prior version cited "p. 4" and "p. 4-5", which does not match the printed page where this text actually appears).
More importantly, the substantive determination is revised. The ERCT standard requires that "assessments should not be specially designed for the study but should be standard, widely recognised tests," and explicitly warns against "custom test[s] specifically designed to measure the outcomes of their intervention." The paper's own language describes exactly this pattern: the instrument "was designed to evaluate students' language and literacy abilities," was explicitly "not designed to... benchmark student performance against curriculum requirements" (i.e. it has no external validated benchmark/norms), and its "components... were adjusted each year" by the research team. The paper describes the tasks only as "EGRA-type" -- i.e. modeled on the generic subtask categories of the internationally-used Early Grade Reading Assessment (EGRA) methodology -- rather than stating that the validated EGRA instrument (or any other named, fixed, standardized exam with established psychometric properties for this population) was administered as-is. This is a researcher-built, study-specific instrument that merely borrows EGRA's task structure, not a standard, widely recognised, unmodified exam. This concern is precisely the bias mechanism the criterion is designed to guard against: a bespoke instrument tuned each year to the study's own curriculum expectations, run by the research team itself, rather than an off-the-shelf standardized test.
Criterion E is not met because the paper describes a custom-built, annually-adjusted assessment only loosely modeled on ("EGRA-type") the EGRA methodology, rather than the administration of an established, standardized, validated exam.
-
T
Term Duration
- Outcomes were measured roughly three years after the intervention began, far exceeding the one-term minimum; this is subsumed by the stronger Y (Year Duration) criterion being met.
- "We then tracked the same cohort of students over a period of three years, starting in February 2017 when they entered grade one, and ending in November 2019." (p. 3)
Relevant Quotes:
1) "We then tracked the same cohort of students over a period of three years, starting in February 2017 when they entered grade one, and ending in November 2019." (p. 3)
2) "At the end of every school year these students were assessed and their teachers surveyed." (p. 3)
Detailed Analysis:
The intervention began in February 2017 and the final outcome measurement occurred in November 2019, nearly three full academic years later. This vastly exceeds the minimum one-term interval required by criterion T. Since the stronger Y (Year Duration) criterion is also clearly met, T is automatically satisfied. Both quotes were checked and are verbatim as they appear on page 3 of the PDF.
Criterion T is met because the measurement interval (about three years) is far longer than one academic term.
-
D
Documented Control Group
- The control group's business-as-usual PD conditions, baseline balance, and characteristics relative to the treatment schools are documented in detail.
- "...and another 80 schools to the control, where teachers could still receive business-as-usual PD support provided by government." (p. 1)
Relevant Quotes:
1) "...and another 80 schools to the control, where teachers could still receive business-as-usual PD support provided by government." (p. 1)
2) "It is important to note that PD activities were also taking place in the control schools. 70 percent of teachers in the control group report to have received training for ESL the same year they were exposed to the program, which was most likely provided by the province or the district. 49 percent of control teachers also reported to use ESL lesson plans that were provided to them by government or a non-government organization... the counterfactual for this evaluation is schools and teachers that already receive some level of PD support." (p. 6)
3) "Table A.5 to A.7 provide some basic descriptive statistics of the sample, and show that the sample is balanced on a range of school, teacher and student characteristics, respectively." (p. 5)
4) "Table A.10, column (1), shows that the attrition rate is 18 percent in the control, and balanced across treatment arms." (p. 5)
Detailed Analysis:
The paper clearly documents what the control group received (business-as-usual PD, with detailed statistics on the proportion receiving ESL training or lesson plans from other sources), and provides balance tables showing the control group's baseline school, teacher and student characteristics are comparable to the treatment arms, with balanced attrition. All quotes checked and verbatim against the PDF.
Criterion D is met because the control group's composition, baseline characteristics, and the nature of business-as-usual support it received are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomization was conducted at the level of the school, with 180 whole schools allocated across the three arms, directly satisfying the school-level RCT requirement.
- "Thus we randomly assigned 50 schools to each intervention and 80 to the control." (p. 4)
Relevant Quotes:
1) "For purpose of the evaluation, we randomly selected 180 public primary schools out of a population of schools that were non-fee paying public schools... We then created 10 strata of similar schools... and randomly assigned five schools to each intervention group and eight to the control group. Thus we randomly assigned 50 schools to each intervention and 80 to the control." (p. 4)
2) "Furthermore, within each school we randomly selected 20 grade one students, and tracked these students over a period of three years." (p. 4)
Detailed Analysis:
The unit of randomization is explicitly the school: 180 whole schools were stratified and randomly allocated to on-site coaching, virtual coaching, or control. Students within schools were only sampled for assessment (not separately randomized to conditions), meaning entire schools received a uniform treatment status. Quotes checked and verbatim.
Criterion S is met because the study randomized entire schools to conditions.
-
I
Independent Conduct
- The same research team that helped design and implement the intervention with the Department of Basic Education also conducted the data collection and formal analysis, with no statement of an independent third-party evaluator overseeing the study.
- "Working with South Africa's Department of Basic Education (DBE), we randomly assigned 100 schools..." (p. 1)
Relevant Quotes:
1) "Working with South Africa's Department of Basic Education (DBE), we randomly assigned 100 schools to receive either virtual or on-site coaching, and another 80 schools to the control..." (p. 1)
2) "We conducted four rounds of data collection in our sample of schools... During these rounds of data collection, we conducted student assessments on the same panel of students, administered teacher surveys..." (p. 4)
3) "The program implementer and data collection companies were different organizations with their own unique brands." (p. 12, Section 6.3)
4) CRediT authorship contribution statement: "Jacobus Cilliers: Conceptualization, Methodology, Software, Formal analysis, Data curation, writing -- original draft... Nompumelelo Mohohlwane: Conceptualization, Writing -- review & editing, Supervision, Project administration." (p. 14)
Detailed Analysis:
The paper notes the program implementer and data collection company were different organizations, which is a positive sign for fieldwork logistics, but this is distinct from independence of the research/evaluation team from the intervention's designers. The authors themselves -- including Department of Basic Education staff who co-designed and administered the program with the research team ("we randomly assigned...") -- also collected data ("we conducted... data collection") and performed the formal analysis (per the CRediT statement, several of the same authors handled Conceptualization, Methodology and Formal analysis). There is no statement of an external, independent evaluation team or third-party oversight of the data analysis and conclusions, distinct from the team that conceptualized and implemented the intervention with the DBE. Quotes checked and verbatim (page for the CRediT statement corrected to p. 14, matching the PDF).
Criterion I is not met because the same research/implementation team, working directly with the program's government partner, also designed, executed and analyzed the study, without a documented independent evaluator.
-
Y
Year Duration
- The intervention and its measured outcomes spanned three full academic years, far exceeding the 75%-of-a-year minimum.
- "We then tracked the same cohort of students over a period of three years, starting in February 2017... and ending in November 2019." (p. 3)
Relevant Quotes:
1) "These programs were implemented over a period of three years, targeting the teachers assigned to a different grade each year (grade one teachers in the first year, grade two teachers in the second year, etc.)." (p. 2)
2) "We then tracked the same cohort of students over a period of three years, starting in February 2017 when they entered grade one, and ending in November 2019." (p. 3)
Detailed Analysis:
The study followed the same cohort from the start of grade one (February 2017) to the end of grade three (November 2019), a period of roughly three years, vastly exceeding the minimum 75% of an academic year (~9-10 months) required by criterion Y. Quotes checked and verbatim.
Criterion Y is met because outcomes were tracked over three academic years.
-
B
Balanced Control Group
- The additional resources (coaching, lesson plans, extra training days, tablets) are the explicit treatment variable being tested (on-site vs. virtual coaching, each vs. business-as-usual), so the control group appropriately receives standard PD support rather than matched extra resources.
- "...another 80 schools to the control, where teachers could still receive business-as-usual PD support provided by government." (p. 1)
Relevant Quotes:
1) "Can virtual replace on-site coaching? We address this question... we randomly assigned 100 schools to receive either virtual or on-site coaching, and another 80 schools to the control, where teachers could still receive business-as-usual PD support provided by government." (p. 1)
2) "In both programs teachers received the same biannual training workshops and learning materials, including lesson plans that were aligned with government curriculum. However, the on-site coaching intervention differed from the virtual coaching in three important dimensions" (coaching modality, classroom observation, and lesson-plan medium). (p. 1-2)
3) "It is important to note that PD activities were also taking place in the control schools... the counterfactual for this evaluation is schools and teachers that already receive some level of PD support." (p. 6)
Detailed Analysis:
Applying the criterion-B decision tree: extra resources (coaching visits/calls, tablets, additional training days, lesson plans) are present in both treatment arms relative to control. These resources are precisely the treatment variable under investigation -- the paper's central research question is whether virtual coaching (with its distinct resource bundle) can replace on-site coaching (with its own resource bundle), each compared to a "business as usual" control that continues to receive whatever ordinary PD support government otherwise provides. The paper explicitly frames the control as a "business-as-usual" counterfactual that already receives some baseline PD, consistent with the standard's guidance that when additional resources are the explicit treatment being tested, the control need not be matched to those resources. Comparing the two treatment arms to each other, both receive broadly comparable core resources (training, lesson plans, integrated materials, coaching), with the differences (modality of coaching, tablet vs. paper) being integral to the specific research question about delivery mode, not an incidental imbalance. Quotes re-checked and verbatim against the PDF; the criterion B definition in the current ERCT specification was reviewed and this determination is unaffected by the update.
Criterion B is met because the additional resources are integral to the treatment variable being tested against a business-as-usual control.
-
Level 3 Criteria
-
R
Reproduced
- Neither the paper nor further internet search surfaced an independent replication of this specific virtual-vs-on-site coaching comparison by a different research team in a different context.
Relevant Quotes:
1) "The study builds on and complements a previous early grade reading study (EGRS I) (see Cilliers et al. (2019)) that targeted Home Language literacy in South Africa..." (p. 3, footnote 7)
2) "Powell et al. (2010) experimentally compared virtual with on-site coaching of pre-K teachers, and found that after one semester the programs were equally effective at improving oral language proficiency." (p. 2-3)
3) "A separate study examines this question (Cilliers et al., 2022)." (p. 4, footnote 14)
Detailed Analysis:
The paper references a prior study by largely the same author team (EGRS I, Cilliers et al. 2019) and a follow-up paper by the same authors (Cilliers et al. 2022, "The challenge of sustaining effective teaching: Spillovers, fade-out, and the cost-effectiveness of teacher development programs," Economics of Education Review, forthcoming at time of publication), neither of which constitutes an independent replication by a different team. It also cites Powell et al. (2010), which independently compared virtual and on-site coaching, but in a different country (US), population (pre-K teachers), and program design, and predates this study rather than replicating it.
Additional internet search (RePEc/IDEAS, EconPapers, RISE Programme publication listings) was conducted to look for independent replications of this specific EGRS II virtual-vs-on-site coaching comparison by other research teams. No such replication was found; only same-author-team companion papers (EGRS I and the 2022 follow-up) and citing literature reviews/ meta-analyses (e.g., Kraft et al. 2018) that discuss this study alongside others rather than replicate it were located. Web search tooling access was limited during this check (search engine result pages were largely inaccessible), so absence of evidence should be read with that caveat, but no quotes indicating independent replication were found in any source consulted.
Criterion R is not met because no independent replication of this specific study was found in the paper or in further internet search.
-
A
All-subject Exams
- Criterion E (Exam-based Assessment) is not met, which per the ERCT specification means criterion A cannot be met either; additionally, the mathematics assessment is described only in passing with no detail confirming it is a standardized instrument.
- "A further written assessment was conducted with the students to assess their written comprehension abilities in both languages, as well as their basic mathematics skills." (p. 5)
Relevant Quotes:
1) "Moving beyond English literacy, we also assessed students' home language literacy and mathematics skills to evaluate whether the treatments had any crowding-out or spillover effects on the other subject areas." (p. 8)
2) "A further written assessment was conducted with the students to assess their written comprehension abilities in both languages, as well as their basic mathematics skills." (p. 5)
3) "There is no impact, either positive or negative, on mathematics. We discuss possible reasons for this result in Section 6.1 below." (p. 8)
Detailed Analysis:
Per the ERCT specification, criterion A explicitly requires criterion E to be met as a prerequisite ("If criterion E is not met, then this criterion is not met either"). Since this verification revised criterion E to "not met" (the assessments are a custom, annually adjusted, researcher-designed instrument only loosely modeled on EGRA, not an established standardized exam), criterion A is automatically not met regardless of subject coverage. Independently, the study does assess subjects beyond English (home language literacy and "basic mathematics skills"), but the mathematics measure is mentioned only briefly, framed as a secondary spillover check, with no description of its content, standardization, or validation. Quotes checked and verbatim.
Criterion A is not met, both because its prerequisite criterion E is not met, and because the mathematics assessment is not described with sufficient detail to confirm it is a standardized instrument.
-
G
Graduation Tracking
- Tracking stopped at the end of grade three (November 2019); there is no indication that students were followed through to graduation from primary or secondary school, and the identified follow-up paper addresses teacher-effect persistence with new cohorts rather than tracking this cohort to graduation.
- "We did not survey teachers or assess subsequent cohorts of their students in the year(s) after they were exposed to the programs, so we cannot test for the persistence of teacher effects after the program ends." (p. 4, footnote 14)
Relevant Quotes:
1) "We then tracked the same cohort of students over a period of three years, starting in February 2017 when they entered grade one, and ending in November 2019." (p. 3)
2) "We did not survey teachers or assess subsequent cohorts of their students in the year(s) after they were exposed to the programs, so we cannot test for the persistence of teacher effects after the program ends. A separate study examines this question (Cilliers et al., 2022)." (p. 4, footnote 14)
3) Reference list: "Cilliers, Jacobus, Fleisch, Brahm, Kotze, Janeli, Mohohlwane, Mpumi, Taylor, Stephen, 2022. The challenge of sustaining effective teaching: Spillovers, fade-out, and the cost-effectiveness of teacher development programs. Econ. Educ. Rev. (forthcoming)." (p. 14, References)
Detailed Analysis:
Student outcome tracking ends at grade three (November 2019), covering only the Foundation Phase (grades 1-3) of primary school, not through graduation from primary school (typically grade 7 in South Africa) or secondary school. The authors explicitly state they did not continue to assess subsequent cohorts or the same students' longer-term persistence.
Per criterion G's specific instructions, the paper names a follow-up companion paper by the same author team: "The challenge of sustaining effective teaching: Spillovers, fade-out, and the cost-effectiveness of teacher development programs" (Cilliers, Fleisch, Kotze, Mohohlwane, Taylor, Economics of Education Review, forthcoming as of the source paper). Internet search was conducted to locate this paper and assess whether it tracks the original 2017 grade-one cohort toward graduation. Direct access to the article's full text was not obtainable via the search tools available during this check, but its title and the source paper's own framing of the research question it addresses ("test for the persistence of teacher effects after the program ends," using data on whether coached teachers' practices persist with subsequently taught cohorts of students) indicate it studies teacher-effect fade-out and spillovers across new student cohorts, not continued tracking of this study's original assessed cohort of students toward graduation. No quotes were found, in this paper or elsewhere, indicating graduation tracking of the original cohort; no quotes were fabricated for this determination.
Criterion G is not met because student follow-up ended at grade three, well short of graduation, and no evidence of graduation-tracking follow-up work on this cohort was found.
-
P
Pre-Registered
- The AEA RCT Registry record for this trial (10.1257/rct.5148-1.0) shows an initial registration date of December 13, 2019 -- almost three years after the intervention began (March 2017) and after the three-year intervention/data-collection period had already concluded (November 2019) -- so the protocol was not pre-registered before data collection began.
- "The study was registered with the AEA Trial Registry: https://doi.org/10.1257/rct.5148-1.0." (p. 1, footnote)
Relevant Quotes:
1) "The study was registered with the AEA Trial Registry: https://doi.org/10.1257/rct.5148-1.0." (p. 1, footnote)
2) "As specified in our pre-analysis plan, we evaluate the overall impact of the interventions using two indices that are based on the two language constructs that students of a second language have to master by the end of grade three." (p. 5)
3) "Since none of the evidence reported in this section was specified in our pre-analysis plan (with the exception of Fig. 6), this analysis should be considered exploratory." (p. 10)
4) From the AEA Social Science Registry entry for trial 5148 ("Improving the teaching of English as a first additional language in South Africa"): Trial Start Date "2017-01-01"; Intervention Start Date "2017-03-01"; Trial End Date "2020-12-01"; and Pre-Registration Date "December 13, 2019" (first published December 17, 2019).
Detailed Analysis:
The paper cites an AEA Trial Registry reference and repeatedly distinguishes "pre-analysis plan" analyses from "exploratory" ones, which was taken by the prior ranking as sufficient evidence that criterion P is met. However, per the criterion's explicit procedure, the registration date must be verified against the study's actual data-collection start date. On checking the AEA Social Science Registry entry for the cited DOI (trial 5148), the registration was only initially filed on December 13, 2019 -- almost three years after the trial's start date of January 1, 2017 and the intervention's start date of March 1, 2017, and even after the three-year data collection period ended in November 2019. This means the "pre-registration" was filed retrospectively, after all data collection (and quite possibly after outcome analysis) had already occurred, rather than before the study began as the ERCT standard requires. A pre-analysis plan referenced internally in the paper does not establish prospective, externally-verifiable pre-registration if the public registry record itself was only created after the study concluded.
Criterion P is not met because the AEA Registry entry cited by the paper shows a registration date of December 2019, nearly three years after the intervention/data collection began in 2017 and after the study's data collection had already concluded, so the protocol was not registered before the study started.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.