Improving Literacy Instruction in Kenya Through Teacher Professional Development and Text Messages Support: A Cluster Randomized Trial

Matthew C. H. Jukes, Elizabeth L. Turner, Margaret M. Dubeck, Katherine E. Halliday, Hellen N. Inyega, Sharon Wolf, Stephanie Simmons Zuilkowski, and Simon J. Brooker

Published:
ERCT Check Date:
DOI: 10.1080/19345747.2016.1221487
  • reading
  • language arts
  • K12
  • Africa
  • mobile learning
1
  • C

    Randomisation occurred at the TAC tutor zone level (clusters of 3-6 schools), which is stronger than class-level randomisation and therefore satisfies this criterion.

    "the first stage of randomization involved random allocation by TAC tutor zone so that each of the 26 TAC tutor zones were randomly allocated to receive either the literacy intervention or to serve as a literacy control." (p. 453)

  • E

    The literacy battery was built on the internationally recognised EGRA and PALS assessment frameworks, widely used across sub-Saharan Africa.

    "Literacy tests were adapted from previous work in Kenya (Jukes, Vagh, & Kim, 2006), and the Early Grade Reading Assessment (EGRA; RTI International, 2009), both of which have been widely used in the least developed countries in sub-Saharan Africa and elsewhere..." (p. 456)

  • T

    Outcomes were measured at 9 and 24 months after the intervention began, far exceeding one academic term.

    "Children from Grade 1 were randomly selected and followed up for 24 months to assess the impact of the interventions." (p. 452)

  • D

    Table 3 provides detailed baseline demographic, school, and household characteristics for the 50 control schools compared with intervention schools.

    "Table 3. Baseline school, child, and household characteristics of 2,539 Kenyan Grade-1 children enrolled in the trial in the 51 intervention and 50 control schools." (p. 463)

  • S

    Schools (organised in randomly-allocated clusters of TAC tutor zones) were the unit of random allocation to literacy intervention or control.

    "Overall, 51 schools were randomized to literacy intervention and 50 schools to literacy control." (p. 453)

  • I

    The paper explicitly states, and lists as a limitation, that the same team that designed and implemented the intervention also conducted the evaluation.

    "One potential limitation of this study is that the evaluation was conducted by the same team of people that designed the intervention." (p. 476)

  • Y

    The study tracked outcomes for 24 months, well beyond the required 75% of an academic year.

    "Children from Grade 1 were randomly selected and followed up for 24 months to assess the impact of the interventions." (p. 452)

  • B

    The additional teacher training, materials, and SMS support are themselves the explicit treatment variable being tested against a business-as-usual control, which satisfies the balance requirement by design.

    "We evaluated a program to improve literacy instruction on the Kenyan coast using training workshops, semiscripted lesson plans, and weekly text-message support for teachers..." (p. 450)

  • R

    No independent peer-reviewed replication of this specific cluster-randomized HALI study was found; later related programs are similar in concept but distinct interventions/designs, and a review of the citing literature confirms no independent reproduction exists.

    "we are not aware of any previous published evaluation of a teacher professional development program using mobile phones to coach and support teachers in a low-income country." (p. 451)

  • A

    In addition to literacy, the study assessed numeracy outcomes using an EGMA-based battery, covering the main subjects taught at this early-grade level.

    "We also included assessments of early numeracy skills, sustained attention, and nonverbal reasoning to assess whether the literacy intervention had spillover effects in other domains." (p. 456)

  • G

    Tracking stopped at the end of Grade 2 (24 months), with no evidence of follow-up through graduation from primary school, and a search of the authors' subsequent publications found no such extended tracking.

    "Children from Grade 1 were randomly selected and followed up for 24 months to assess the impact of the interventions." (p. 452)

  • P

    The trial was registered on ClinicalTrials.gov (NCT00878007) in April 2009, about nine months before data collection began in January 2010, as confirmed by independent verification of the registry record.

    "The trial protocol was registered with the U. S. National Institute of Health Clinical Trials Registry (ClinicalTrials.gov; Identifier: NCT00878007)." (p. 460)

Abstract

We evaluated a program to improve literacy instruction on the Kenyan coast using training workshops, semiscripted lesson plans, and weekly text-message support for teachers to understand its impact on students' literacy outcomes and on the classroom practices leading to those outcomes. The evaluation ran from the beginning of Grade 1 to the end of Grade 2 in 51 government primary schools chosen at random, with 50 schools acting as controls. The intervention had an impact on classroom practices with effect sizes from 0.57 to 1.15. There was more instruction with written text and more focus on letters and sounds. There was a positive impact on three of four primary measures of children's literacy after two years, with effect sizes up to 0.64, and school dropout reduced from 5.3% to 2.1%. This approach to literacy instruction is sustainable, and affordable and a similar approach has subsequently been adopted nationally in Kenya.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation occurred at the TAC tutor zone level (clusters of 3-6 schools), which is stronger than class-level randomisation and therefore satisfies this criterion.
      • "the first stage of randomization involved random allocation by TAC tutor zone so that each of the 26 TAC tutor zones were randomly allocated to receive either the literacy intervention or to serve as a literacy control." (p. 453)
      • Relevant Quotes: 1) "The evaluation involved a cluster randomized trial (Brooker et al., 2010), in which 101 public primary schools were randomly allocated to one of four arms receiving either: (a) the malaria intervention alone; (b) the literacy intervention alone; (c) both interventions combined; or (d) neither intervention." (p. 452) 2) "Therefore, the first stage of randomization involved random allocation by TAC tutor zone so that each of the 26 TAC tutor zones were randomly allocated to receive either the literacy intervention or to serve as a literacy control." (p. 453) 3) "This ensured that no TAC tutor zone contained both intervention and control schools." (p. 453) 4) "Overall, 51 schools were randomized to literacy intervention and 50 schools to literacy control." (p. 453) Detailed Analysis: Randomisation was conducted at the level of TAC tutor zones (clusters of three to six schools each), which were then randomly allocated in their entirety to either literacy intervention or literacy control. This is a cluster/school-level (indeed, supra-school) randomisation unit, explicitly designed to avoid "leakage of the intervention between schools." Because this exceeds the class-level requirement (and in fact satisfies the stronger "S" criterion, see below), the weaker class-level criterion is automatically considered met per the ERCT standard. Criterion C is met because randomisation occurred at the level of school clusters (TAC tutor zones), exceeding the class-level requirement.
    • E

      Exam-based Assessment

      • The literacy battery was built on the internationally recognised EGRA and PALS assessment frameworks, widely used across sub-Saharan Africa.
      • "Literacy tests were adapted from previous work in Kenya (Jukes, Vagh, & Kim, 2006), and the Early Grade Reading Assessment (EGRA; RTI International, 2009), both of which have been widely used in the least developed countries in sub-Saharan Africa and elsewhere..." (p. 456)
      • Relevant Quotes: 1) "Literacy tests were adapted from previous work in Kenya (Jukes, Vagh, & Kim, 2006), and the Early Grade Reading Assessment (EGRA; RTI International, 2009), both of which have been widely used in the least developed countries in sub-Saharan Africa and elsewhere and from the Phonological Awareness Literacy Screening (Invernizzi, Juel, Swank, & Meier, 2007)." (p. 456) 2) "Before the study began, all instruments were adapted to the Kenyan context over a period of five months to ensure face validity and appropriate stimuli." (p. 458) 3) "Spelling. A group test in English only, which assesses phonemic awareness and letter knowledge. Students spell five words with consonant-vowel- consonant syllable patterns (e.g., sad)." (p. 456) Detailed Analysis: The core outcome measures (letter reading fluency, word reading fluency, passage reading fluency and comprehension) are explicitly adapted from EGRA, a widely recognised, internationally used standardised assessment toolkit, rather than being invented purely for this study's purposes. Some subtests (e.g., spelling, beginning sound awareness) were custom items, but these sit within the same validated, widely-used battery framework (EGRA/PALS) adapted for local language/context, consistent with how similarly adapted EGRA/EGMA batteries have been treated as meeting this criterion in other ERCT evaluations. Criterion E is met because the primary outcome measures derive from the widely recognised, validated EGRA/PALS assessment frameworks.
    • T

      Term Duration

      • Outcomes were measured at 9 and 24 months after the intervention began, far exceeding one academic term.
      • "Children from Grade 1 were randomly selected and followed up for 24 months to assess the impact of the interventions." (p. 452)
      • Relevant Quotes: 1) "The HALI literacy intervention was evaluated between January 2010 and March 2012..." (p. 452) 2) "Children's educational outcomes were assessed at baseline, and at 9 months (FU1) and 24 months (FU2) follow-up." (p. 455) 3) "Children from Grade 1 were randomly selected and followed up for 24 months to assess the impact of the interventions." (p. 452) Detailed Analysis: Outcomes were assessed at 9 months and again at 24 months following the start of the intervention, both of which substantially exceed the minimum one-term (3-4 month) requirement. This duration is in fact long enough to also satisfy the stronger Y (Year Duration) criterion. Criterion T is met because outcome measurement occurred at 9 and 24 months post-intervention start, well beyond one academic term.
    • D

      Documented Control Group

      • Table 3 provides detailed baseline demographic, school, and household characteristics for the 50 control schools compared with intervention schools.
      • "Table 3. Baseline school, child, and household characteristics of 2,539 Kenyan Grade-1 children enrolled in the trial in the 51 intervention and 50 control schools." (p. 463)
      • Relevant Quotes: 1) "Table 3 shows that children had broadly similar characteristics in each of the literacy intervention study arms." (p. 461) 2) "However, a higher proportion of schools in the control arm had school feeding programs when surveyed in January 2010." (p. 461) 3) "Overall, most children (95.5%) had attended preschool, a third (32.1%) had previously failed Grade 1, and most (86.9%) read aloud in class." (p. 461) 4) Table 3 (pp. 463): detailed breakdown of school characteristics (exam scores, school size, feeding/ deworming/malaria programs, facilities), child characteristics (age, sex, nutritional status, school experience) and household characteristics (parental education, SES, household size, language spoken at home) separately for the 51 intervention and 50 control schools. Detailed Analysis: The paper provides an extensive, explicitly tabulated breakdown of control-group characteristics at multiple levels (school, child, household), including baseline exam performance, demographics, and socioeconomic status, and explicitly flags where imbalances existed (e.g., school feeding programs, SES, parental education). This level of detail satisfies the documentation requirement. Criterion D is met because the control group's characteristics are extensively documented in Table 3 and the accompanying text.
  • Level 2 Criteria

    • S

      School-level RCT

      • Schools (organised in randomly-allocated clusters of TAC tutor zones) were the unit of random allocation to literacy intervention or control.
      • "Overall, 51 schools were randomized to literacy intervention and 50 schools to literacy control." (p. 453)
      • Relevant Quotes: 1) "the first stage of randomization involved random allocation by TAC tutor zone so that each of the 26 TAC tutor zones were randomly allocated to receive either the literacy intervention or to serve as a literacy control." (p. 453) 2) "Overall, 51 schools were randomized to literacy intervention and 50 schools to literacy control." (p. 453) 3) "Inclusion of control schools and intervention schools from the same TAC tutor zone may have led to leakage of the intervention between schools... Such contamination was documented in a literacy instruction evaluation in a nearby district." (p. 453) Detailed Analysis: The unit implementing the intervention (or serving as control) was the individual school, with schools grouped and randomised by TAC tutor zone specifically to prevent within-zone contamination between schools. Every school in a given zone received the same condition, and the resulting sample of 51 intervention and 50 control schools is described throughout the paper as school-level allocation. This meets (and arguably exceeds) the school-level RCT requirement. Criterion S is met because entire schools (grouped by randomly allocated tutor zone clusters) were assigned to intervention or control condition.
    • I

      Independent Conduct

      • The paper explicitly states, and lists as a limitation, that the same team that designed and implemented the intervention also conducted the evaluation.
      • "One potential limitation of this study is that the evaluation was conducted by the same team of people that designed the intervention." (p. 476)
      • Relevant Quotes: 1) "The evaluation was conducted by the same team that implemented the HALI program." (p. 452) 2) "To avoid any bias in findings that this might entail, a number of measures were put in place. An independent data-monitoring committee was set up to scrutinize data collection and analysis... Impact analyses were conducted by a statistician working independently from the HALI implementation team." (p. 452) 3) "To avoid bias in data collection, the assessment team were blind to the intervention status of the school and were highly trained so that they responded to all students in the same scripted way..." (p. 452) 4) "One potential limitation of this study is that the evaluation was conducted by the same team of people that designed the intervention." (p. 476) Detailed Analysis: The paper is explicit and unambiguous that the same research/implementation team that designed and delivered the HALI intervention also conducted its evaluation, and the authors themselves list this as a limitation in their Discussion section. Several bias-mitigation measures were implemented -- an independent data-monitoring committee, a pre-specified analysis plan, an independently-working statistician for impact analyses, and blinded assessors for data collection -- which reduce (but by the authors' own framing do not eliminate) the risk of bias from lack of independence. The ERCT standard requires that the study be "conducted independently from the authors who designed the intervention," and per the standard's procedure, if the same authors both designed and carried out the study this criterion fails unless there is clear third-party oversight of the whole study; here, independent oversight was partial (data collection and analysis only), not full study conduct, and the authors' own limitation statement confirms this shortfall. Criterion I is not met because the study explicitly states that the same team that designed the intervention also conducted the evaluation, despite partial independence measures for analysis and data collection.
    • Y

      Year Duration

      • The study tracked outcomes for 24 months, well beyond the required 75% of an academic year.
      • "Children from Grade 1 were randomly selected and followed up for 24 months to assess the impact of the interventions." (p. 452)
      • Relevant Quotes: 1) "The HALI literacy intervention was evaluated between January 2010 and March 2012..." (p. 452) 2) "Of the 2,539 study children, 2,516 (99.1%) took part in baseline assessments, 2,238 (88.1%) in the first follow-up at 9 months and 2,030 (80.0%) in the second follow-up at 24 months." (p. 461) Detailed Analysis: The evaluation followed children from the start of Grade 1 through the end of Grade 2, a period of 24 months, which comfortably exceeds a full academic year (let alone 75% of one). Since Y is met, the weaker T criterion is automatically also considered met. Criterion Y is met because the study tracked outcomes over a 24-month period, well beyond the one-year requirement.
    • B

      Balanced Control Group

      • The additional teacher training, materials, and SMS support are themselves the explicit treatment variable being tested against a business-as-usual control, which satisfies the balance requirement by design.
      • "We evaluated a program to improve literacy instruction on the Kenyan coast using training workshops, semiscripted lesson plans, and weekly text-message support for teachers..." (p. 450)
      • Relevant Quotes: 1) "We evaluated a program to improve literacy instruction on the Kenyan coast using training workshops, semiscripted lesson plans, and weekly text-message support for teachers to understand its impact on students' literacy outcomes..." (p. 450, Abstract) 2) "140 sequential, semiscripted lesson plans for literacy sessions... Training, including a three-day initial workshop... Ongoing support for teachers for two years through weekly text messages providing brief instructional tips and motivation... Teachers also received credit of $0.50... each week for their mobile phones." (p. 451) 3) "51 schools were randomized to literacy intervention and 50 schools to literacy control." (p. 453) [with no parallel professional-development package described for control schools] 4) "A key concern was the increased time taken to prepare and conduct the intervention lessons compared with the standard curriculum..." (p. 475) Detailed Analysis: Applying the criterion B decision procedure: extra resources (teacher training workshops, lesson materials, SMS support, small phone credit) are clearly present and were given only to intervention-arm teachers, and control-arm teachers received no equivalent alternative package. However, this teacher professional-development package is not a supplementary add-on -- it is explicitly the primary treatment variable under evaluation, as stated directly in the paper's title and abstract ("teacher professional development and text message support"). The control arm represents the standard "business as usual" teaching condition the intervention is being tested against, which is the legitimate comparator when additional resources are the object of study rather than a confound to be balanced out. Criterion B is met because the additional resources (training, materials, SMS support) are integral to and constitute the explicit treatment variable being tested, with the control arm appropriately representing business-as-usual practice.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent peer-reviewed replication of this specific cluster-randomized HALI study was found; later related programs are similar in concept but distinct interventions/designs, and a review of the citing literature confirms no independent reproduction exists.
      • "we are not aware of any previous published evaluation of a teacher professional development program using mobile phones to coach and support teachers in a low-income country." (p. 451)
      • Relevant Quotes: 1) "However, we are not aware of any previous published evaluation of a teacher professional development program using mobile phones to coach and support teachers in a low-income country." (p. 451) 2) "Subsequent to our intervention, evaluations have emerged showing a positive impact of a teacher training program involving text-message communication (Piper, Zuilkowski, & Mugenda, 2014) and also of a program of teacher support using tablets (Piper, Jepkemei, Kwayumba, & Kibukho, 2015; Piper, Zuilkowski, Kwayumba, & Strigel, 2016)." (p. 451) 3) "Subsequent to the evaluation of the HALI project the Kenya government evaluated the PRIMR initiative -- a pilot of a comprehensive approach to early-grade instruction in literacy and mathematics. The pilot adopted the approach of supporting teachers with text messages and also used instructional principles similar to the HALI project." (pp. 475-476) Detailed Analysis: The paper itself states this was a novel type of program at the time, and later cites related but distinct evaluations (PRIMR, tablet-based teacher support) that adopt broadly similar principles (SMS teacher support, evidence-based literacy pedagogy). An internet review of the 89 papers citing this study (via the Semantic Scholar citation graph) found no independent peer-reviewed reproduction of this specific cluster-randomized HALI trial's design (same intervention package, same cluster-randomized structure, same outcome battery) by a fully independent team. The closest analogue identified, a Malawi case study on SMS-based remote teacher support ("Short Message Service (SMS)-Based Remote Support and Teacher Retention of Training Gains in Malawi," Slade, Kipp, Cummings, & Nyirongo, 2018), evaluates a different, unrelated teacher-training program as a descriptive case study rather than a controlled replication of this study's design and findings. PRIMR in particular is a different, larger-scale, government-run program combining literacy and mathematics instruction, evaluated by an overlapping set of researchers (RTI International, shared with some HALI authors) rather than a fully independent team replicating this precise study. No quote in the text, nor in the broader citing literature, indicates a direct, independent peer-reviewed reproduction of this specific trial's design and findings. Criterion R is not met because no independent replication of this specific study (as opposed to related programs testing similar general principles) was identified in the text or in a review of the subsequent citing literature.
    • A

      All-subject Exams

      • In addition to literacy, the study assessed numeracy outcomes using an EGMA-based battery, covering the main subjects taught at this early-grade level.
      • "We also included assessments of early numeracy skills, sustained attention, and nonverbal reasoning to assess whether the literacy intervention had spillover effects in other domains." (p. 456)
      • Relevant Quotes: 1) "We also included assessments of early numeracy skills, sustained attention, and nonverbal reasoning to assess whether the literacy intervention had spillover effects in other domains." (p. 456) 2) "Numeracy assessments were based on previous work in Kenya (Jukes, Vagh, & Kim, 2006) and the Early Grade Mathematics Assessment (Reubens, 2009)." (p. 457) 3) Table 5: reports standardized adjusted mean differences for "Numeracy (score: 0-20)" at 9 and 24 months alongside the literacy outcomes. (p. 467) 4) "In fact, the direction of effect on numeracy score at 24 months was negative (Adj. MD: -0.46, 95% CI: -0.93, 0.01, p = 0.058)." (p. 468) Detailed Analysis: For Grade 1-2 children in this Kenyan context, the main curriculum subjects assessed in early-grade evaluations are literacy (English and Swahili) and numeracy/ mathematics. The study measured literacy through the EGRA-based battery and separately measured numeracy through an EGMA-based battery, explicitly to check for unintended (including negative) spillover effects on the non-target subject -- precisely the concern this criterion is designed to guard against. This is comparable to prior ERCT evaluations of early-grade programs where EGRA plus EGMA together were judged to cover the main subjects for that grade level. Criterion A is met because both main subjects for this grade level (literacy and numeracy) were assessed using recognised assessment frameworks, revealing that the intervention had no positive (and possibly a negative) effect on the non-target subject.
    • G

      Graduation Tracking

      • Tracking stopped at the end of Grade 2 (24 months), with no evidence of follow-up through graduation from primary school, and a search of the authors' subsequent publications found no such extended tracking.
      • "Children from Grade 1 were randomly selected and followed up for 24 months to assess the impact of the interventions." (p. 452)
      • Relevant Quotes: 1) "The evaluation ran from the beginning of Grade 1 to the end of Grade 2 in 51 government primary schools..." (p. 450, Abstract) 2) "Children from Grade 1 were randomly selected and followed up for 24 months to assess the impact of the interventions." (p. 452) 3) "the intervention supported teachers in developing the literacy skills of one cohort of children through the first two years of primary school..." (p. 451) Detailed Analysis: Follow-up ended at 24 months, coinciding with the end of Grade 2, out of a primary schooling cycle that typically runs through Grade 8 in Kenya. An internet search of subsequent publications by the same author team (Jukes, Dubeck, Turner, Wolf, and colleagues) identified two related papers that reuse HALI-trial data: "Changing Literacy Instruction in Kenyan Classrooms" (Wolf, Turner, Jukes, & Dubeck, 2018, International Journal of Educational Development) and "Literacy Acquisition in Multilingual Educational Contexts: Evidence from Coastal Kenya" (Jasinska, Wolf, Jukes, & Dubeck, 2019, Developmental Science). Both re-analyze classroom-process and psycholinguistic data collected within the original 24-month HALI trial period rather than extending outcome tracking of the same cohort into later grades or to primary-school graduation. No subsequent publication tracking this cohort through graduation was identified in any available source. Criterion G is not met because tracking ended at the end of Grade 2, well short of graduation, with no evidence of subsequent follow-up publications extending tracking to graduation despite a search of the authors' later work.
    • P

      Pre-Registered

      • The trial was registered on ClinicalTrials.gov (NCT00878007) in April 2009, about nine months before data collection began in January 2010, as confirmed by independent verification of the registry record.
      • "The trial protocol was registered with the U. S. National Institute of Health Clinical Trials Registry (ClinicalTrials.gov; Identifier: NCT00878007)." (p. 460)
      • Relevant Quotes: 1) "The trial protocol was registered with the U. S. National Institute of Health Clinical Trials Registry (ClinicalTrials.gov; Identifier: NCT00878007)." (p. 460) 2) "Primary and secondary outcomes were prespecified in a statistical analysis plan and approved by an independent data monitoring committee (DMC) separately for 9-month (FU1) and 24-month (FU2) analyses." (p. 460) 3) "The HALI literacy intervention was evaluated between January 2010 and March 2012..." (p. 452) External Verification: 4) The ClinicalTrials.gov registry record for NCT00878007 lists a "First Submitted Date" of 2009-04-07 and a "First Posted Date" of 2009-04-08 (estimate), with a study start date of January 2010. Detailed Analysis: The paper confirms registration on a recognised public trial registry (ClinicalTrials.gov, NCT00878007) and describes a prespecified analysis plan reviewed by an independent DMC, both positive indicators of transparency. The ERCT standard requires quoted evidence of the registration date and confirmation that it preceded the start of data collection; the paper itself does not state this date. Direct verification of the ClinicalTrials.gov registry, however, shows the protocol was first submitted on 7 April 2009 and first posted 8 April 2009, approximately nine months before baseline data collection began in January 2010. This satisfies the requirement that pre-registration occur before data collection began. Criterion P is met because independent verification of the ClinicalTrials.gov registry confirms the trial protocol (NCT00878007) was registered in April 2009, well before data collection began in January 2010.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.