The Effect of In-Service Teacher Training on Student Learning of English as a Second Language

Rosangela Bando, Xia Li

Published:
ERCT Check Date:
DOI: 10.18235/0011651
  • L2 languages
  • K12
  • Latam
0
  • C

    Teachers (and thus their entire classes) were the unit of random allocation, satisfying class-level randomisation.

    "Within each stratum, teachers were randomly allocated either to treatment or control groups." (p. 10)

  • E

    The student outcome used a DIME instrument specifically designed for the study, not a widely recognised standardised exam.

    "Five randomly selected students per teacher were tested with a DIME English test specifically designed for students." (p. 10)

  • T

    Student outcomes were measured about seven and a half months after exposure began, exceeding one academic term.

    "After seven and a half months of exposure to a trained teacher, students improved their English." (Abstract)

  • D

    The control group's size, demographics and baseline proficiency are documented in detail in Tables 2 and 4.

    "In the evaluation sample, the treatment and control groups are equal on average before intervention." (pp. 12-13)

  • S

    Randomisation was at the individual teacher level, not at the school level.

    "Within each stratum, teachers were randomly allocated either to treatment or control groups." (p. 10)

  • I

    The IDB evaluators were independent of the Worldfund/Dartmouth team that designed and delivered the intervention.

    "The teacher selection process is a joint effort of Dartmouth College's Rassias Center, Worldfund, and Mexican state governments." (p. 5)

  • Y

    The average seven-and-a-half-month tracking exceeds 75% of the 40-week academic year defined in the paper.

    "The students of trained teachers improved... in an average of seven and a half months of exposure." (p. 1)

  • B

    Teacher training is the explicit treatment variable tested against a business-as-usual control, with students in both arms receiving equivalent class time.

    "A randomized experiment was conducted in Mexico to test whether teacher training could increase teacher efficiency in public secondary schools." (Abstract)

  • R

    No independent peer-reviewed replication of this specific teacher-training RCT is reported or found.

    "This study is one of the first to identify the causal effect of in-service teacher training alone on student learning in a developing country." (p. 4)

  • A

    Only English was assessed, no other core subjects, and criterion E (standardised exam) is not met.

    "The main evaluation question is whether or not the program improves English in students." (p. 17)

  • G

    Students were measured only at the end of the school year, with no tracking to graduation, and no follow-up study tracking the cohort was found.

    "It is important to explore whether changes in teacher and student behavior are sustainable over time and, if so, for how long." (p. 36)

  • P

    The paper contains no reference to any pre-registration registry, ID, or pre-registered analysis plan, and no such record was found through internet search.

Abstract

In-service teacher training aims to improve the supply of public education. A randomized experiment was conducted in Mexico to test whether teacher training could increase teacher efficiency in public secondary schools. After seven and a half months of exposure to a trained teacher, students improved their English. This paper explores two mechanisms through which training can affect student learning. First, trained teachers improved their English by 0.35 standard deviations in the short run. Teachers in the control group caught up with treatment teachers by the end of the school year in part because teachers in the treatment group reduced out-of-pocket expenditures to learn English in 53 percent. Second, teachers changed classroom practices by providing more opportunities for students to actively engage in learning. This evidence suggests that teacher training may be effective at improving student learning and that teacher incentives may play a role in mediating its effects.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Teachers (and thus their entire classes) were the unit of random allocation, satisfying class-level randomisation.
      • "Within each stratum, teachers were randomly allocated either to treatment or control groups." (p. 10)
      • Relevant Quotes: 1) "Random allocation of 77 teachers to a treatment group and 67 to a control group allow for identification of causal effects on student learning as measured by standardized test scores." (p. 1) 2) "The 205 eligible teachers identified before school visits were randomly allocated to receive training or not. Within each of the three cohorts recruited, teachers were stratified by whether they reported teaching in lower or upper secondary schools and by state. Within each stratum, teachers were randomly allocated either to treatment or control groups." (p. 10) 3) "Errors are clustered at the teacher level." (p. 15) Detailed Analysis: The unit of randomisation was the teacher. Because the intervention is teacher training and each teacher delivers instruction to their own class(es), randomising teachers means every student of a given teacher receives a uniformly trained or untrained teacher. Randomisation was not carried out among individual students within a single class; instead entire teacher/class clusters were assigned to treatment or control, with errors clustered at the teacher level. This class-level (teacher-level) assignment isolates treatment and control at the class unit and prevents within-class contamination, satisfying the class-level RCT requirement (the study does not, however, randomise at the stronger school level, see criterion S). Criterion C is met because randomisation was conducted at the teacher/class level, with each class uniformly assigned to treatment or control.
    • E

      Exam-based Assessment

      • The student outcome used a DIME instrument specifically designed for the study, not a widely recognised standardised exam.
      • "Five randomly selected students per teacher were tested with a DIME English test specifically designed for students." (p. 10)
      • Relevant Quotes: 1) "Five randomly selected students per teacher were tested with a DIME English test specifically designed for students." (p. 10) 2) "Baseline data were collected on the English proficiency and general characteristics of the teachers before training sessions for each cohort started with the Diagnostic Instrument for the Measurement of English (DIME) test... The DIME test was developed using best practices in test item development and reflects rigorous English language proficiency measurement protocols. DIME was originally developed to assess English teachers in the Mexican Ministry of Education and has been validated for the purposes of this study." (p. 9) 3) "The research base used in the development of the test is that used in the development of the ELPA test developed by the University of Michigan, the TOEIC Bridge test developed by the Educational Testing Service, and the battery of tests developed by Cambridge English. The validation was against the Stanford battery of tests published by NCS Pearson." (footnote 7, p. 9) Detailed Analysis: Criterion E requires a standardised, widely recognised exam that was not specially created for the study. The primary student outcome was measured with "a DIME English test specifically designed for students," i.e. an instrument purpose-built for this evaluation. Although the DIME development drew on the research base of recognised tests (ELPA, TOEIC Bridge, Cambridge, Stanford battery) and the teacher version originated in the Ministry of Education, the student instrument was specifically designed and "validated for the purposes of this study." It is a bespoke diagnostic instrument rather than an externally administered, widely recognised standardised exam such as a national curriculum test or TOEFL. Under the standard, a study-specific instrument does not satisfy the exam-based assessment requirement. Criterion E is not met because the primary student outcome relied on a test specifically designed for the study rather than a widely recognised standardised exam.
    • T

      Term Duration

      • Student outcomes were measured about seven and a half months after exposure began, exceeding one academic term.
      • "After seven and a half months of exposure to a trained teacher, students improved their English." (Abstract)
      • Relevant Quotes: 1) "After seven and a half months of exposure to a trained teacher, students improved their English." (Abstract) 2) "The students of trained teachers improved by around 0.16 standard deviations when compared to students of non-trained teachers in an average of seven and a half months of exposure." (p. 1) 3) "Follow-up data were collected at the end of the school year, in May to June of 2013." (p. 9) Detailed Analysis: A term is roughly 3-4 months. Teachers were trained across three cohorts during 2012 (May, September and November) and student outcomes were collected at the end of the school year (May-June 2013), giving an average of seven and a half months of student exposure to a trained teacher. This interval from intervention start to outcome measurement clearly exceeds one full academic term. Criterion T is met because outcomes were measured after an average of seven and a half months of exposure, well beyond one academic term.
    • D

      Documented Control Group

      • The control group's size, demographics and baseline proficiency are documented in detail in Tables 2 and 4.
      • "In the evaluation sample, the treatment and control groups are equal on average before intervention." (pp. 12-13)
      • Relevant Quotes: 1) "Random allocation of 77 teachers to a treatment group and 67 to a control group..." (p. 1) 2) "Table 4 shows conditional and unconditional balance on observable teacher characteristics on the sample of 144 teachers for which there is information to evaluate. In the evaluation sample, the treatment and control groups are equal on average before intervention." (pp. 12-13) 3) Table 2 "Baseline Means over Randomization Sample" and Table 4 "Baseline Means on Teacher Characteristics in Evaluation Sample" report, for treatment vs control, DIME scores, CEFR level, oral interview score, lower/upper secondary status, school size, gender, age, education, full-time status, marital status, socioeconomic status and cohort (pp. 13, 16). 4) "Random allocation creates two groups of teachers with students that are not different before program intervention." (p. 10) Detailed Analysis: The control group (67 teachers / 81 in the randomisation sample) is documented in detail. Tables 2 and 4 give side-by-side baseline demographic, school and English proficiency characteristics for control and treatment groups, with standardised differences and p-values confirming baseline equivalence. The control condition is described as business-as-usual (untrained teachers), and its composition and baseline performance are clearly reported. Criterion D is met because the control group's size, demographic and baseline characteristics are thoroughly documented in Tables 2 and 4.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was at the individual teacher level, not at the school level.
      • "Within each stratum, teachers were randomly allocated either to treatment or control groups." (p. 10)
      • Relevant Quotes: 1) "The 205 eligible teachers identified before school visits were randomly allocated to receive training or not... Within each stratum, teachers were randomly allocated either to treatment or control groups." (p. 10) 2) "Errors are clustered at the teacher level." (p. 15) 3) "The evaluation sample consists of 144 teachers actively working at the secondary level (grades 7 to 12) in public schools in the states of Puebla and Tlaxcala." (p. 7) Detailed Analysis: Criterion S requires randomisation at the level of the school (the institution/unit implementing the intervention). Here randomisation was performed on individual teachers, stratified only by school level (lower/upper secondary) and state, not by school. Teachers within the same school could therefore be split across treatment and control conditions; there is no description of whole schools being randomly assigned. The unit of assignment is the teacher/class, not the school. Criterion S is not met because randomisation was carried out at the teacher level, not at the school level.
    • I

      Independent Conduct

      • The IDB evaluators were independent of the Worldfund/Dartmouth team that designed and delivered the intervention.
      • "The teacher selection process is a joint effort of Dartmouth College's Rassias Center, Worldfund, and Mexican state governments." (p. 5)
      • Relevant Quotes: 1) "This study evaluated the impact of a ten-day component of the Inter-American Partnership for Education (IAPE) program, a Clinton Global Initiative Commitment that aims to empower classroom teachers in Mexican public schools... The teacher selection process is a joint effort of Dartmouth College's Rassias Center, Worldfund, and Mexican state governments." (p. 5) 2) "We are indebted to the Worldfund team and the Ministry of Education in Puebla and Tlaxcala, who provided logistical support and provided the necessary information to make the evaluation possible... Steve Marban, Raúl Abreu, and Armando Loera provided support to make data collection possible." (footnote 1, p. 1) 3) The authors are affiliated with the Inter-American Development Bank (title page and corresponding author: "Rosangela Bando. 1300 New York Avenue, NW, Washington, DC... Email: rosangelab@iadb.org"). Detailed Analysis: The intervention (the IAPE training program, based on the Dartmouth Rassias method) was designed and delivered by Worldfund and Dartmouth College's Rassias Center. The evaluation was designed and conducted by economists at the Inter-American Development Bank (Bando and Li), who are distinct from the program's designers and providers. The program provider (Worldfund) supplied logistics and information but the causal evaluation was carried out independently by IDB staff, and data collection was supported by separate personnel. This separation of the evaluators from the intervention designers satisfies the independent conduct requirement. Criterion I is met because the trial was evaluated by independent IDB researchers, separate from the Worldfund/Dartmouth team that designed and delivered the intervention.
    • Y

      Year Duration

      • The average seven-and-a-half-month tracking exceeds 75% of the 40-week academic year defined in the paper.
      • "The students of trained teachers improved... in an average of seven and a half months of exposure." (p. 1)
      • Relevant Quotes: 1) "After seven and a half months of exposure to a trained teacher, students improved their English." (Abstract) 2) "The students of trained teachers improved by around 0.16 standard deviations when compared to students of non-trained teachers in an average of seven and a half months of exposure." (p. 1) 3) "A school year is 40 weeks of class time." (footnote 10, p. 18) 4) "Follow-up data were collected at the end of the school year, in May to June of 2013." (p. 9) Detailed Analysis: Criterion Y requires outcome measurement at least 75% of a full academic year after the intervention begins. The paper defines the academic year as 40 weeks of class time (approximately 9.2 months); 75% of that is roughly 6.9 months. Student outcomes were measured at the end of the school year after an average of seven and a half months of exposure to a trained teacher, which exceeds the 75% threshold of the academic year. The measurement therefore spans essentially the full academic year for the average student. Criterion Y is met because outcomes were measured after an average of seven and a half months of exposure, exceeding 75% of the 40-week (about 9.2 month) academic year.
    • B

      Balanced Control Group

      • Teacher training is the explicit treatment variable tested against a business-as-usual control, with students in both arms receiving equivalent class time.
      • "A randomized experiment was conducted in Mexico to test whether teacher training could increase teacher efficiency in public secondary schools." (Abstract)
      • Relevant Quotes: 1) "A randomized experiment was conducted in Mexico to test whether teacher training could increase teacher efficiency in public secondary schools." (Abstract) 2) "The program provides 100 hours of intensive training, 80 of which are devoted to intensive English instruction and 20 hours to pedagogical training." (p. 6) 3) "The main contribution of this paper is that it provides quantitative evidence of whether or not in-service teacher training alone can change teacher and student behavior to improve student learning through a randomized controlled trial..." (p. 1) 4) "This study is also one of the few that isolate the effects of teacher training from changes in other inputs or factors that usually accompany training, such as changes to curricula, provision of didactic materials, or technology." (p. 4) Detailed Analysis: The additional resource provided to the treatment group is the teacher training itself (100 hours of training given to treatment teachers; control teachers received business-as-usual with no training). Applying the decision tree: extra resources are present (training), but those resources ARE the explicit treatment variable being tested - the study is designed "to test whether teacher training could increase teacher efficiency" and explicitly isolates "the effects of teacher training alone" from other inputs such as curricula, materials, or technology. For the students, the outcome of interest, no extra instructional time or budget was added: both treatment and control students received their normal English class time; only the trained versus untrained status of the teacher differed. Because the training is the integral, central treatment variable and students received equivalent class time, the business-as-usual control is appropriate by design. Criterion B is met because the additional resource (teacher training) is the explicit treatment variable being tested against a business-as-usual control, and students in both arms received equivalent instructional time.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent peer-reviewed replication of this specific teacher-training RCT is reported or found.
      • "This study is one of the first to identify the causal effect of in-service teacher training alone on student learning in a developing country." (p. 4)
      • Relevant Quotes: 1) "Harris and Sass (2011) review quantitative studies aiming to assess the effects of teacher training on student learning. The authors find only three randomized controlled trials on the effect of in-service teacher training on student learning, all carried out in the United States." (p. 2) 2) "This study is one of the first to identify the causal effect of in-service teacher training alone on student learning in a developing country." (p. 4) Detailed Analysis: Criterion R requires independent replication of this specific study (the IAPE/Rassias-method teacher-training RCT among ESL teachers in Puebla and Tlaxcala) by a different research team, in a different context, published in a peer-reviewed outlet. The paper presents itself as one of the first RCTs of in-service teacher training in a developing country and does not report any replication of its own findings. A dedicated internet search (searches for the authors, title, DOI, and the IAPE/Rassias program, including citation trackers) found no independent peer-reviewed reproduction of this specific Mexico teacher-training RCT by a separate research team. The IAPE/Rassias method has since expanded to thousands of Mexican teachers and other states (e.g., Baja California) through continued program delivery, but no subsequent peer-reviewed randomised evaluation of the program by an independent team was located. Other teacher-training RCTs cited in the literature (e.g., Muralidharan and Sundararaman, 2010, in India; Angrist and Lavy, 2001, in Jerusalem) test different interventions in different contexts and are not replications of this particular study or intervention. Criterion R is not met because no independent replication of this specific study has been reported or found.
    • A

      All-subject Exams

      • Only English was assessed, no other core subjects, and criterion E (standardised exam) is not met.
      • "The main evaluation question is whether or not the program improves English in students." (p. 17)
      • Relevant Quotes: 1) "Five randomly selected students per teacher were tested with a DIME English test specifically designed for students." (p. 10) 2) "The main evaluation question is whether or not the program improves English in students. Table 5 shows estimated program effects for the DIME test." (p. 17) Detailed Analysis: Criterion A requires that impact be measured across all main subjects using standardised exam-based assessments, and it explicitly depends on criterion E being met. The study measured only English (the intervention subject) via the DIME student test; no other core subjects such as mathematics or science were assessed. Furthermore, criterion E is not met because the outcome instrument was a study-specific test rather than a widely recognised standardised exam, which independently causes criterion A to fail. Criterion A is not met because only English was assessed (no other core subjects), and because criterion E is not met.
    • G

      Graduation Tracking

      • Students were measured only at the end of the school year, with no tracking to graduation, and no follow-up study tracking the cohort was found.
      • "It is important to explore whether changes in teacher and student behavior are sustainable over time and, if so, for how long." (p. 36)
      • Relevant Quotes: 1) "Follow-up data were collected at the end of the school year, in May to June of 2013." (p. 9) 2) "It is important to explore whether changes in teacher and student behavior are sustainable over time and, if so, for how long." (p. 36) 3) "If the program provides benefits to students less than 24 months and teachers forget within one school year and do not benefit any other generations, then the investment is a loss." (pp. 35-36) Detailed Analysis: Criterion G requires tracking participants through to graduation from the educational stage. Student outcomes were measured only once, at the end of the school year, about seven and a half months after exposure began. The authors explicitly note the absence of longer-term follow-up and flag sustainability over time as an open question, and frame the program's long-run return as contingent on unverified assumptions about whether gains persist for multiple years or fade within one school year. A dedicated internet search for subsequent publications by Bando, Li, or related IDB/Worldfund/IAPE evaluation teams that track this same cohort of teachers or their students to graduation found no such follow-up study. No peer-reviewed or working-paper follow-up reporting graduation-stage outcomes for this Puebla/Tlaxcala cohort was identified. Criterion G is not met because measurement stopped at the end of the school year, with no tracking of students to graduation, and no follow-up publication tracking this cohort to graduation was found.
    • P

      Pre-Registered

      • The paper contains no reference to any pre-registration registry, ID, or pre-registered analysis plan, and no such record was found through internet search.
      • Relevant Quotes: 1) "The 205 eligible teachers identified before school visits were randomly allocated to receive training or not." (p. 10) 2) No statement referencing a pre-registration registry, registration identifier, or a pre-registered analysis plan appears anywhere in the paper. Detailed Analysis: Criterion P requires that the full study protocol (hypotheses, methods and planned analyses) be pre-registered on a public registry before data collection began, with a verifiable date. This 2014 IDB working paper contains no reference to any trial registry (e.g. AEA RCT Registry, ClinicalTrials.gov), no registration ID, and no mention of a pre-registered analysis plan. A dedicated internet search of the AEA RCT Registry and general web sources for a pre-registration record tied to this study (by authors, title, or the IAPE/Rassias intervention in Puebla/Tlaxcala) found no matching entry. This is consistent with the timing of the study: baseline data collection began with the first training cohort in May 2012, before the AEA RCT Registry existed in its current form (launched in 2013), making pre-registration on that platform not feasible at the time and no equivalent registry record was found elsewhere. Criterion P is not met because the paper provides no reference to a pre-registered protocol or registry entry, and no such record was found through internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.