Abstract
Integrating Information and Communication Technology (ICT) is considered a promising educational input to help disadvantaged students. The overall goal of this paper is twofold: (1) to evaluate the impact of a government-implemented computer-assisted learning (CAL) program by comparing it to a CAL program implemented by a research team; and (2) to understand whether a computer-assisted instruction (CAI) program is effective in raising learning outcomes and to compare the relative effectiveness of a CAI program against a CAL program. In a clustered RCT in 127 rural primary schools in Qinghai Province, China, students whose CAL treatment was implemented by the research group improved significantly in English test scores (0.18 SD), while the government-implemented CAL program had no impact; government schools were more likely to replace regular English classes with CAL sessions. In a second study, the CAI program (integrated into teaching) proved more effective than the CAL program at raising students' English test scores (0.08 SD), benefiting both low- and high-performing students.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was conducted at the school level (127 schools randomly assigned), which is stronger than class-level and therefore satisfies the class-level RCT criterion.
- "We randomly divided the 127 sample schools into two treatment groups and one control group." (p. 6)
Relevant Quotes:
1) "We conducted a clustered RCT of CAL and CAI in rural schools during the 2013/14 academic year. A total of 120 primary schools in poor minority areas in China's Qinghai Province are included in our study." (p. 5)
2) "We randomly divided the 127 sample schools into two treatment groups and one control group." (p. 6)
3) "Based on power calculations, we randomly chose 22 schools to receive the CAI intervention (i.e. the program sessions were instructed by an English teacher). We randomly chose another 22 schools to receive the CAL treatment (i.e. the program sessions were supervised by a non-English teacher). The final 83 schools were assigned to the control group." (p. 6)
4) "After identifying the 127 sample schools, we randomly chose 40 schools to receive the CAL intervention implemented by our research team (treatment group 1), 40 schools to receive the CAL intervention implemented by the government (treatment group 2), and 40 schools to have no intervention (the control group)." (p. 7)
Detailed Analysis:
The report describes two clustered RCTs (Study 1 on CAL implementer effects with fourth graders; Study 2 comparing CAI vs CAL with fifth graders), both randomised at the school level. Entire schools, not individual students within a class, were assigned to treatment or control conditions. The unit of randomisation, sample sizes, and allocation numbers are clearly documented in the text and in Figures 1 and 2. Under the ERCT standard, school-level randomisation is stronger than class-level randomisation and automatically satisfies criterion C.
Criterion C is met because randomisation was clearly performed at the school level, which exceeds the class-level requirement.
-
E
Exam-based Assessment
- The English test was constructed, validated and piloted by the research team for the study rather than being a widely recognised, standardised national exam.
- "These test items went through several rounds of selection, validation and piloting by the research team and teacher experts from primary schools." (p. 11, footnote 5)
Relevant Quotes:
1) "In the first part students were given a standardized English test. The English test included 90 questions. Students were required to finish tests in each subject in 30 minutes." (p. 11)
2) "All the questions were multiple choice and the tests were graded using STATA. Therefore, the test scores were unlikely to be biased in the grading process. The test questions were based on the standard national curriculum for English language. These test items went through several rounds of selection, validation and piloting by the research team and teacher experts from primary schools." (p. 11, footnote 5)
3) "All of the questions on the English test in the endline survey were different from the questions in the baseline survey. We only chose the questions that did not overlap with the exercises in our software package." (p. 13)
Detailed Analysis:
Although the authors repeatedly call their instrument a "standardized English test", footnote 5 makes clear that the test items were selected, validated and piloted by the research team itself (with teacher experts), and item choice was tailored to avoid overlap with the intervention software. No named, widely recognised standardised exam (e.g. a national or provincial examination) was used. Under the ERCT standard, curriculum-aligned tests assembled by the researchers for the purposes of the study count as custom assessments, not standardised exams, even when carefully piloted. Appendix A confirms the items are bespoke multiple-choice questions.
Criterion E is not met because outcomes were measured with a researcher-constructed test rather than a widely recognised standardised exam.
-
T
Term Duration
- Outcomes were measured about nine months after the intervention began (September 2013 baseline to June 2014 endline), which exceeds one full academic term.
- "The length of the program was one academic year, or a little over nine months." (p. 11)
Relevant Quotes:
1) "The first round of surveys was a baseline survey conducted with all students in 127 schools in September 2013 at the beginning of the autumn semester. It was before any implementation of our program had begun. The second round survey was an evaluation survey conducted at the end of the ICT program in June 2014, a time that coincided with the end of the 2014 spring semester." (p. 12)
2) "The length of the program was one academic year, or a little over nine months." (p. 11)
3) "Baseline (Sept. 2013) ... Evaluation (June 2014)" (Figure 1, p. 22)
Detailed Analysis:
The intervention began after the September 2013 baseline survey and endline outcomes were collected in June 2014, an interval of roughly nine to ten months. This far exceeds the minimum one academic term (3-4 months) required by criterion T. (Section 3.1.5 contains a typographical inconsistency referring to fall 2012/spring 2013, but the experiment profiles in Figures 1 and 2 and Section 3.1.6 consistently date the study to September 2013 - June 2014.) Since the stronger year-duration criterion is satisfied, the weaker term-duration criterion is automatically met.
Criterion T is met because the interval from intervention start to outcome measurement covered a full academic year.
-
D
Documented Control Group
- The control group's size, conditions and baseline characteristics are documented in detail, with balance tests across demographic and baseline performance variables.
- "A total of 3,996 fifth-grade students in 83 control schools constituted the control group. During the program, students in the control group did not receive any intervention." (p. 11)
Relevant Quotes:
1) "A total of 3,996 fifth-grade students in 83 control schools constituted the control group. During the program, students in the control group did not receive any intervention. To avoid any form of spillover effects and Hawthorne effect (Landsberger, 1958), our program team did not visit or contact any control schools except during the baseline and endline surveys." (p. 11)
2) "The students in control schools took their regular classes at school as before." (p. 11)
3) "We used a set of student characteristics to check the validity of the random assignment (Tables 1 and 8) ... We found that the pair-wise differences among the three groups ... were all statistically insignificant for all the student characteristics." (p. 6)
4) "The 1,823 fourth-grade students in the rest of the 40 schools served as the control group" (p. 7)
5) "At the endline of the experiment, no control group school was found to have started using any educational software." (p. 8, footnote 1)
Detailed Analysis:
For both studies the control groups are precisely quantified (1,823 fourth graders in 40 schools for Study 1; 3,996 fifth graders in 83 schools for Study 2), their treatment status is explicitly described (business as usual, no intervention, no educational software use), and detailed baseline demographic and test-score comparisons between treatment and control groups are reported in Tables 1, 3, 8 and 10, together with attrition analyses (Tables 2 and 9). This satisfies the requirement for a well-documented control group.
Criterion D is met because the control group's composition, size, baseline performance and conditions are thoroughly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Entire schools were the unit of randomisation in both studies, satisfying the school-level RCT criterion.
- "After identifying the 127 sample schools, we randomly chose 40 schools to receive the CAL intervention implemented by our research team (treatment group 1), 40 schools to receive the CAL intervention implemented by the government (treatment group 2), and 40 schools to have no intervention (the control group)." (p. 7)
Relevant Quotes:
1) "We randomly divided the 127 sample schools into two treatment groups and one control group." (p. 6)
2) "Based on power calculations, we randomly chose 22 schools to receive the CAI intervention ... We randomly chose another 22 schools to receive the CAL treatment ... The final 83 schools were assigned to the control group." (p. 6)
3) "After identifying the 127 sample schools, we randomly chose 40 schools to receive the CAL intervention implemented by our research team (treatment group 1), 40 schools to receive the CAL intervention implemented by the government (treatment group 2), and 40 schools to have no intervention (the control group)." (p. 7)
Detailed Analysis:
Both experiments assigned whole primary schools to treatment or control arms: 40/40/40 schools in Study 1 and 22/22/83 schools in Study 2. The randomisation unit is the school (the educational institution implementing the intervention), the number of schools is stated, and school-clustered standard errors are used throughout the analysis, confirming the clustered school-level design.
Criterion S is met because randomisation was explicitly conducted among schools.
-
I
Independent Conduct
- The same research team designed the CAL/CAI programs, software, protocols and training and also conducted the surveys and analysis, with no independent third-party evaluator.
- "To standardize the program implementation, we designed a detailed CAL curriculum and implementation protocol." (p. 7)
Relevant Quotes:
1) "Our program team developed another piece of the software package. It provided a large number of additional exercise questions. We worked with teachers and experts from the organization Teaching English to Speakers of Other Languages (TESOL) to choose the questions." (p. 7)
2) "To standardize the program implementation, we designed a detailed CAL curriculum and implementation protocol." (p. 7)
3) "The implementer of the CAL program in treatment group 1 was our research team, while the implementer in treatment group 2 was the county education bureau." (p. 7)
4) "In order to facilitate the incorporation of the ICT program into English teaching practices during CAI classes, we carefully designed and compiled a CAI implementation protocol." (p. 9)
5) "The research team conducted two rounds of surveys in the 127 sample schools." (p. 11)
6) "These test items went through several rounds of selection, validation and piloting by the research team" (p. 11, footnote 5)
Detailed Analysis:
Criterion I requires that the evaluation be conducted independently of the intervention's designers. Here the research team (Stanford REAP / CEEE) designed the CAL and CAI curricula, co-developed part of the software, wrote the implementation protocols, ran the teacher training, directly implemented treatment group 1, constructed the outcome test, and conducted the baseline and endline surveys and the analysis. Although treatment group 2 was implemented by the county government and enumerators were recruited from local universities to monitor and survey, the overall evaluation (design, measurement instruments, data collection organisation and analysis) remained in the hands of the team that designed the intervention, with no statement of independent third-party oversight or an external evaluation agency responsible for data collection and analysis.
Criterion I is not met because the intervention designers themselves conducted the implementation, data collection and analysis without documented independent oversight.
-
Y
Year Duration
- The program and outcome tracking spanned one full academic year (September 2013 to June 2014), a little over nine months.
- "The length of the program was one academic year, or a little over nine months." (p. 11)
Relevant Quotes:
1) "The length of the program was one academic year, or a little over nine months." (p. 11)
2) "The first round of surveys was a baseline survey conducted with all students in 127 schools in September 2013 at the beginning of the autumn semester ... The second round survey was an evaluation survey conducted at the end of the ICT program in June 2014, a time that coincided with the end of the 2014 spring semester." (p. 12)
3) "We conducted a clustered RCT of CAL and CAI in rural schools during the 2013/14 academic year." (p. 5)
Detailed Analysis:
Criterion Y requires outcomes to be measured at least 75% of one full academic year after the intervention begins. The intervention ran across the entire 2013/14 academic year, from the start of the autumn semester (September 2013) to the end of the spring semester (June 2014), and the endline test was administered at the conclusion of that year. The authors themselves characterise the program length as "one academic year, or a little over nine months", which fully covers the Chinese academic year.
Criterion Y is met because outcome measurement occurred a full academic year (about nine months) after the intervention started.
-
B
Balanced Control Group
- The CAL/CAI sessions ran during existing computer class time against business-as-usual control schools with the same computer facilities, and the software, protocol, training and teacher stipends were integral parts of the treatment being tested.
- "The evaluation is designed to determine whether a computer room with a CAL program can improve student academics when compared to computer rooms without CAL." (p. 5)
Relevant Quotes:
1) "Under the supervision of a local teacher-supervisor, trained by either our research group or the government, the students in the treatment groups were supposed to have two 40-minute CAL sessions per week during the computer classes. The sessions were mandatory." (p. 8)
2) "We also suggested that since the CAL classes were implemented during computer class time, the computer teacher should conduct the CAL sessions." (p. 8)
3) "The computer:student ratio is uniform across the treatment and control groups (0.13 computers per student)." (p. 11, footnote 4)
4) "Although the education bureau requires that in these rural public school computer classes need to be scheduled weekly, we found no school in the baseline that had employed computers and/or educational software for instructional purposes in core academic subjects." (p. 8, footnote 1)
5) "We compensated each teacher-supervisor with an allowance of RMB500 (USD82) every semester." (p. 8)
6) "The evaluation is designed to determine whether a computer room with a CAL program can improve student academics when compared to computer rooms without CAL." (p. 5)
7) "During the program, students in the control group did not receive any intervention ... The students in control schools took their regular classes at school as before." (p. 11)
Detailed Analysis:
Applying the ERCT criterion B decision tree: extra resources are present (CAL/CAI software package, teacher training, protocols, monitoring support and a RMB500-per-semester teacher stipend go only to the treatment schools). However, for students the CAL/CAI sessions occupied the existing weekly computer-class slot rather than adding new instructional time, and the computer:student ratio (0.13) was uniform across treatment and control schools, so the incremental student time and hardware burden was negligible. For the remaining budget items (software, training, teacher stipend), the study's own stated research question is whether "a computer room WITH a CAL program" outperforms "computer rooms without CAL" (and, for Study 1, which implementer makes the package work) - i.e. the software/protocol/training package is explicitly the treatment variable being tested against a business-as-usual control, not a separable add-on. Under the decision tree this satisfies RESOURCES_ARE_TREATMENT, so the criterion is met regardless of whether the control group received an equivalent budget. The extra inputs (software, training, RMB500 stipend) are integral to the intervention package, not separable optional add-ons.
Criterion B is met because the added software, training and stipend are integral to the treatment variable being tested against business-as-usual controls, and sessions occupied existing computer class time with uniform computer facilities across groups.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study by a different research team was found; a citation search turned up only two unrelated CAL/CAI evaluations in India by other teams and a methods paper, none of which replicate this Qinghai school-level CAL/CAI implementer study.
- "The Center for Experimental Economics in Education (CEEE) conducted three large-scale randomized control trials (RCTs) to evaluate a game-based computer remedial tutoring program ... (Lai et al, 2013)." (p. 4)
Relevant Quotes:
1) "The Center for Experimental Economics in Education (CEEE) conducted three large-scale randomized control trials (RCTs) to evaluate a game-based computer remedial tutoring program designed to remediate learning in two core subjects of the national curriculum - mathematics and Chinese ... (Lai et al, 2013)." (p. 4)
2) "The only known study to have compared a government-implemented program with an NGO-implemented program is an experiment that involved contract teachers in Kenya (Bold et al., 2013)." (p. 1)
3) "no rigorous studies have tested whether integrating ICT into English education is effective in improving English learning in rural China." (p. 4)
Detailed Analysis:
Criterion R requires independent replication of the study by a different research team in a different context, published in a peer-reviewed journal. The prior and subsequent CAL trials in rural China (Lai et al. 2013; Mo et al. 2014a, 2014b; Huang et al. 2014; the Journal of Development Economics version of this evaluation) all come from the same Stanford REAP / CEEE research group and are therefore not independent replications.
A citation-index search (OpenAlex, work ID W2917152041) for papers citing this 3ie report as of July 2026 returned only four records: (i) Jimenez, Waddington, Goel, Prost, Pullin, White, Lahiri and Narain (2018), "Mixing and matching: using qualitative methods to improve quantitative impact evaluations (IEs) and systematic reviews (SRs) of development outcomes" - a methods paper, not a replication; (ii) Muralidharan, Singh and Ganimian (2019), "Disrupting Education? Experimental Evidence on Technology-Aided Instruction in India" - an RCT of the "Mindspark" computer-adaptive learning software in Delhi middle/secondary schools; and (iii) de Barros and Ganimian (2023), "Which Students Benefit from Computer-Based Individualized Instruction? Experimental Evidence from Public Schools in India" - a related Mindspark study. Both India studies test a different software product, age group, subject mix and country context, and neither attempts to reproduce the specific Qinghai/Haidong school-level CAL-vs-CAI-vs-implementer design or its English-language outcome measures. No other CAL RCTs by unrelated teams replicating this specific study's design were identified.
Criterion R is not met because no independent replication of this specific study by a different research team was found.
-
A
All-subject Exams
- Only English outcomes were measured with a researcher-constructed test, so neither all main subjects nor the prerequisite standardised-exam criterion E is satisfied.
- "We used the English test scores as the measure of English academic performance." (p. 11)
Relevant Quotes:
1) "We used the English test scores as the measure of English academic performance." (p. 11)
2) "In the first part students were given a 30-minute standardized English test and we used the scores of the students as our measure of student academic performance." (p. 12)
3) "The primary outcome variable of our analysis is the student academic outcome, measured by the student standardized test scores in English." (p. 13)
Detailed Analysis:
Criterion A requires standardised exam-based assessment of all main subjects taught at the educational level, with criterion E as a prerequisite. This study measured academic outcomes only in English; mathematics, Chinese and other core primary-school subjects were not assessed, so potential negative spillovers on non-target subjects (a salient risk here, given that 39 per cent of government-implemented treatment schools replaced regular English classes with CAL sessions) cannot be evaluated. Moreover, criterion E is not met because the English test was researcher-constructed, which by itself makes criterion A fail. No exception for a specialised upper-secondary or vocational intervention applies to this primary-school program.
Criterion A is not met because only English was assessed, with a custom test, and criterion E is not met.
-
G
Graduation Tracking
- Measurement stopped at the June 2014 endline when students were in grades four and five, with no tracking of the cohort through primary school graduation, and no follow-up publications tracking this cohort were located.
- "The second round survey was an evaluation survey conducted at the end of the ICT program in June 2014, a time that coincided with the end of the 2014 spring semester." (p. 12)
Relevant Quotes:
1) "The second round survey was an evaluation survey conducted at the end of the ICT program in June 2014, a time that coincided with the end of the 2014 spring semester." (p. 12)
2) "Phase II was a final evaluation survey conducted at the conclusion of the program at the end of the spring semester" (p. 11)
3) "Evaluation survey (June 2014) and analysis" (Figure 2, p. 23)
Detailed Analysis:
Criterion G requires participants to be tracked until graduation from the relevant educational stage. Participants were fourth graders (Study 1) and fifth graders (Study 2) in primary schools, and data collection ended with the endline survey in June 2014, when the students were still one to two years away from completing primary school (grade six). The report describes no follow-up beyond the endline.
A search of available sources (OpenAlex works by Di Mo, Yu Bai, Matthew Boswell and Scott Rozelle, and REAP/CEEE working papers) found only "Computer Technology in Education: Evidence from a Pooled Study of Computer Assisted Learning Programs among Rural Students in China" (Huang, Mo, Shi, Zhang, Boswell and Rozelle), which pools baseline/endline results from several within-one-year CAL evaluations rather than tracking any cohort to graduation. No paper following this specific Qinghai/Haidong 2013/14 fourth- or fifth-grade cohort through to primary-school graduation was found in any available source.
Criterion G is not met because tracking stopped at the end-of-year evaluation and no subsequent graduation-tracking publication for this cohort could be identified.
-
P
Pre-Registered
- The report contains no mention of a trial registry, registration ID or pre-registered protocol with a date before data collection, and no registry entry for this study was found.
Relevant Quotes:
No quotes referencing pre-registration exist in the paper. The methods section states only: "The research team conducted two rounds of surveys in the 127 sample schools." (p. 11) and describes the design without reference to any registry.
Detailed Analysis:
Criterion P requires the full study protocol to be pre-registered on a public registry before data collection begins, with the registration identifiable and dated. A search of the full report found no mention of a registry platform (e.g. RIDIE, AEA RCT Registry, ClinicalTrials.gov, ISRCTN), no registration number, and no statement that hypotheses and analysis plans were registered before the September 2013 baseline. An additional search of citation indices and 3ie's own RIDIE registry site for this title and author set did not surface a pre-registration record. While 3ie-funded studies are often encouraged to register on RIDIE, no quoted or independently located evidence of pre-registration or its timing was found for this study.
Criterion P is not met because no pre-registration statement, registry ID or registration date is provided in the paper or located through independent search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.