Supporting Early Literacy and Reading Skills Development in Grade 1: Findings From a Randomized Control Trial With GraphoLearn-Rime and Teacher-Led Phonics Instruction

Deepti Bora, Ulla Richardson, Anna-Maija Poikkeus, Minna Torppa

Published:
ERCT Check Date:
DOI: 10.1002/rrq.70128
  • reading
  • L2 languages
  • K12
  • Asia
  • blended learning
  • EdTech app
  • mobile learning
0
  • C

    Entire schools were randomly assigned to conditions, which satisfies the class-level (or stronger) randomization requirement.

    "The schools were randomly assigned to either a control group or an experimental group."

  • E

    The primary outcomes relied on researcher-developed tasks rather than a widely recognized standardized exam, which the authors themselves flag as a limitation.

    "First, the use of researcher-developed measures could be a limitation. They are often seen to produce larger effect size compared to standardized measures..."

  • T

    The intervention ran from August to December 2024 (about five months), and post-test outcomes were measured at the end, exceeding one academic term.

    "The intervention started in August 2024, 5 months after the start of the school academic year (March) and ended in December 2024."

  • D

    The control group's size, demographics, baseline scores, and business-as-usual condition are documented in the text and in Tables 1 and 2.

    "Students in the control group received business-as-usual English literacy instruction, which utilized the alphabetic spelling method, including rote learning and writing from the blackboard."

  • S

    Whole schools, not classes or students, were the unit of random assignment to conditions.

    "Among the five participating schools, one school was randomly selected as the control group, one school as the Phonics-only group and three schools as the Phonics+GL group."

  • I

    The same research team designed, delivered, and evaluated the intervention, and a co-author is associated with the GL-Rime/GraphoGame tool, with no independent evaluator.

    "All the assessments were conducted by the first author and two research assistants trained by the first author."

  • Y

    The intervention plus measurement spanned only about five months with no follow-up, well under 75% of an academic year.

    "Third, since the design did not contain a follow-up, it is not possible to assess the long-term effects of the intervention."

  • B

    The phonics instruction and GL-Rime are themselves the explicit treatment variables tested against a business-as- usual control, so the resource difference is integral by design.

    "Our study examined the effectiveness of teacher-led phonics instruction alone and with GL-Rime versus business-as-usual instruction for supporting ESL learners' foundational literacy skills in a low-income context."

  • R

    No independent research team has replicated this specific RCT; related prior work is by the same overlapping authors.

    "Although replications are needed in other contexts... Future studies must examine the replicability of the finding among the ESL learners in India and in other contexts, with similar constraints."

  • A

    Only English literacy/reading outcomes were assessed, with no other core subjects, and criterion E is also not met.

    "This randomized control trial examined the contribution of teacher-led phonics instruction and GraphoLearn-Rime... to the development of English literacy and reading skills."

  • G

    Measurement stopped at post-test with no follow-up or graduation tracking, and criterion Y is also not met.

    "Third, since the design did not contain a follow-up, it is not possible to assess the long-term effects of the intervention."

  • P

    The paper contains no reference to a pre-registered protocol, registry ID, or registration date.

Abstract

This randomized control trial examined the contribution of teacher-led phonics instruction and GraphoLearn-Rime (GL-Rime) to the development of English literacy and reading skills of Grade 1 level children in a low-income context. Students (N=234) were randomly allocated into three groups. The Phonics only group (n=83) received phonics instruction from trained teachers. The Phonics+GL group (n=89) received GL-Rime (computer-assisted reading instruction) along with teacher-led phonics instruction. Both groups received respective interventions for 5 months. The control group (n=62) received business-as-usual instruction. Phonics-only and Phonics+GL groups improved more than the control group on all measures of phonological awareness and reading skills. Integrating a computer-assisted reading instruction technology into classroom instruction supported learning transfer from digital to real-world medium.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Entire schools were randomly assigned to conditions, which satisfies the class-level (or stronger) randomization requirement.
      • "The schools were randomly assigned to either a control group or an experimental group."
      • Relevant Quotes: 1) "All children were native Hindi speakers and English as a second language (ESL) learners studying across five affordable private English-medium schools in Delhi, India. The schools were randomly assigned to either a control group or an experimental group." (p. 6) 2) "Among the five participating schools, one school was randomly selected as the control group, one school as the Phonics-only group and three schools as the Phonics+GL group." (p. 6) 3) "The Phonics-only group included two Grade 1 classrooms with one teacher each... The Phonics+GL group had one Grade 1 classroom in the three schools... The business-as-usual (control) group had two Grade 1 classes with one teacher each..." (p. 6) Detailed Analysis: The ERCT C criterion requires an RCT randomized at the class level or stronger. Here the unit of randomization was the school: whole schools were randomly allocated to the control, Phonics-only, and Phonics+GL conditions. School- level randomization is stronger than class-level and is explicitly accepted by the standard as satisfying the weaker class-level criterion. Randomizing at the school level also prevents within-classroom contamination between treatment and control students. Criterion C is met because randomization was performed at the school level, which meets or exceeds the class-level RCT requirement.
    • E

      Exam-based Assessment

      • The primary outcomes relied on researcher-developed tasks rather than a widely recognized standardized exam, which the authors themselves flag as a limitation.
      • "First, the use of researcher-developed measures could be a limitation. They are often seen to produce larger effect size compared to standardized measures..."
      • Relevant Quotes: 1) "Vocabulary skills were assessed using words from the Peabody Picture Vocabulary Test—5th edition, Form A (Dunn 2019)... The test was only conducted at pre-test level." (p. 9) 2) "The segmentation task was selected from the Dyslexia Screening Test-Junior India Kit. Normed for India, the test consists of 12 items..." (p. 9) 3) "The task consisted of 15 items which were taken from Grade 1-level English textbook commonly used in India." (Grade 1 Word Reading, p. 9) 4) "The task consisted of 15 items which were phonologically similar to the words given in GL-Rime." (GraphoLearn Word Reading, p. 9) 5) "First, the use of researcher-developed measures could be a limitation. They are often seen to produce larger effect size compared to standardized measures since the assessments are often proximal to the content taught in the instruction." (p. 22) Detailed Analysis: The ERCT E criterion requires that outcomes be measured with standardized, widely recognized exam-based assessments rather than tests created for the study. In this paper, the primary outcome measures (letter-sound knowledge, phoneme blending, rhyme oddity, RAN, Grade 1 word reading, GraphoLearn word reading, pseudo-word reading, and the in-game GL tasks) were predominantly researcher-developed. Only the pretest-only PPVT vocabulary measure and the segmentation task (from a standardized India-normed kit) are drawn from published instruments, and these are not the study's principal outcome exams. The authors explicitly acknowledge the use of researcher-developed measures as a limitation. There is no state-wide, national, or otherwise standardized achievement exam used as the outcome. Criterion E is not met because the outcomes were measured chiefly with custom researcher-developed tasks rather than a recognized standardized exam.
    • T

      Term Duration

      • The intervention ran from August to December 2024 (about five months), and post-test outcomes were measured at the end, exceeding one academic term.
      • "The intervention started in August 2024, 5 months after the start of the school academic year (March) and ended in December 2024."
      • Relevant Quotes: 1) "The intervention started in August 2024, 5 months after the start of the school academic year (March) and ended in December 2024. During the intervention, all students in the Phonics-only and Phonics+GL groups received 40 phonics lessons for 30 min 5-6 days a week." (p. 6) 2) "Both groups received respective interventions for 5 months." (Abstract) 3) "Same set of items and measures were used in both pre- and post-tests." (p. 9) Detailed Analysis: The ERCT T criterion requires outcomes to be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Here the intervention ran from August 2024 to December 2024, approximately five months, with post-test measurement at the end of this period. The interval from intervention start to outcome measurement (about five months) clearly exceeds one academic term. Criterion T is met because the interval from intervention start to post-test measurement spans roughly five months, longer than one full academic term.
    • D

      Documented Control Group

      • The control group's size, demographics, baseline scores, and business-as-usual condition are documented in the text and in Tables 1 and 2.
      • "Students in the control group received business-as-usual English literacy instruction, which utilized the alphabetic spelling method, including rote learning and writing from the blackboard."
      • Relevant Quotes: 1) "The control group (n=62) received business-as-usual instruction." (Abstract) 2) "The business-as-usual (control) group had two Grade 1 classes with one teacher each and 32 and 35 students, respectively." (p. 6) 3) "Students in the control group received business-as- usual English literacy instruction, which utilized the alphabetic spelling method, including rote learning and writing from the blackboard. Furthermore, unlike the teachers in the intervention groups, teachers in the control school were not provided with phonics training or any other instructional support during the study." (p. 6) 4) "Control group (n=62)... Gender Male 35 Female 27; Age (years) M=6.06 SD=0.40." (Table 1, p. 6) 5) Table 2 reports control-group pre- and post-test means, SDs, and ranges for every measure (e.g., "Control group n=62"). (p. 12) Detailed Analysis: The ERCT D criterion requires that the control group be well documented, including composition, size, baseline performance, and the treatment received. The paper reports the control group's sample size (n=62), its two Grade 1 classes, gender split and age (Table 1), full baseline and post-test descriptive statistics across all measures (Table 2), and a clear description of its business-as-usual condition with no phonics training. This provides adequate documentation for comparison. Criterion D is met because the control group's size, demographics, baseline scores, and conditions are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Whole schools, not classes or students, were the unit of random assignment to conditions.
      • "Among the five participating schools, one school was randomly selected as the control group, one school as the Phonics-only group and three schools as the Phonics+GL group."
      • Relevant Quotes: 1) "The schools were randomly assigned to either a control group or an experimental group." (p. 6) 2) "Among the five participating schools, one school was randomly selected as the control group, one school as the Phonics-only group and three schools as the Phonics+GL group." (p. 6) Detailed Analysis: The ERCT S criterion requires randomization at the school level, where 'school' means the institution or unit implementing the intervention. The paper states that the five participating schools were randomly assigned to conditions, with one school as control, one as Phonics- only, and three as Phonics+GL. The unit of randomization was therefore the school, satisfying the school-level RCT requirement, although the number of clustered units is small (five schools). Criterion S is met because randomization was conducted at the school level.
    • I

      Independent Conduct

      • The same research team designed, delivered, and evaluated the intervention, and a co-author is associated with the GL-Rime/GraphoGame tool, with no independent evaluator.
      • "All the assessments were conducted by the first author and two research assistants trained by the first author."
      • Relevant Quotes: 1) "The teachers of the intervention groups were provided with workshops on foundational literacy skills, phonics, and systematic phonics instruction by the first author." (p. 6) 2) "All the assessments were conducted by the first author and two research assistants trained by the first author." (p. 9) 3) "The first author and trained research assistants conducted classroom observations of all the phonics lessons." (p. 7) 4) "In this study, we utilized GraphoLearn Rime (GL-Rime, also known as GraphoGame)... (Richardson and Lyytinen 2014)." (p. 3) [Ulla Richardson is the second author.] Detailed Analysis: The ERCT I criterion requires the study to be conducted independently of those who designed the intervention, with third-party or external evaluation. Here the first author designed the phonics training, delivered the teacher workshops, conducted the classroom observations and feedback, and (with her trained research assistants) administered all assessments. The GL-Rime tool is associated with the second author (Richardson). There is no statement of an independent or external evaluation team overseeing data collection, analysis, or conclusions. Criterion I is not met because the intervention was designed, implemented, and evaluated by the same research team without independent oversight.
    • Y

      Year Duration

      • The intervention plus measurement spanned only about five months with no follow-up, well under 75% of an academic year.
      • "Third, since the design did not contain a follow-up, it is not possible to assess the long-term effects of the intervention."
      • Relevant Quotes: 1) "The intervention started in August 2024, 5 months after the start of the school academic year (March) and ended in December 2024." (p. 6) 2) "Both groups received respective interventions for 5 months." (Abstract) 3) "Third, since the design did not contain a follow-up, it is not possible to assess the long-term effects of the intervention." (p. 22) Detailed Analysis: The ERCT Y criterion requires outcomes to be measured at least 75% of a full academic year (~7-8 months) after the intervention begins. The intervention here ran from August to December 2024 (about five months) with post-testing at its end and no follow-up. Five months falls short of 75% of an academic year, and the authors confirm there was no follow-up to extend tracking. Criterion Y is not met because the tracking interval of about five months is shorter than 75% of an academic year.
    • B

      Balanced Control Group

      • The phonics instruction and GL-Rime are themselves the explicit treatment variables tested against a business-as- usual control, so the resource difference is integral by design.
      • "Our study examined the effectiveness of teacher-led phonics instruction alone and with GL-Rime versus business-as-usual instruction for supporting ESL learners' foundational literacy skills in a low-income context."
      • Relevant Quotes: 1) "Our study examined the effectiveness of teacher-led phonics instruction alone and with GL-Rime versus business-as-usual instruction..." (p. 5) 2) "During the intervention, all students in the Phonics- only and Phonics+GL groups received 40 phonics lessons for 30 min 5-6 days a week... students in the Phonics+GL group played 15 min GL-Rime 3-4 days a week, totaling 8.75 h over 3 months." (p. 6) 3) "Students in the control group received business-as- usual English literacy instruction, which utilized the alphabetic spelling method, including rote learning and writing from the blackboard." (p. 6) 4) "Teacher-led phonics (minutes) M=1015.59 (Phonics only) / M=1035.13 (Phonics+GL)... The duration of teacher-led phonics instruction between the two intervention groups did not differ from each other (t(170)=-0.86, p=0.29)." (Table 1, p. 6; p. 11) Detailed Analysis: Criterion B compares the nature, quantity, and quality of resources (time, materials, adult support) across conditions and asks whether any extra resource given only to the intervention is the explicit treatment variable. Here the intervention groups received trained-teacher phonics instruction (about 17 hours) and, for Phonics+GL, an additional 8.75 hours of GL-Rime game time plus teacher training and mentoring, while the control received business- as-usual English literacy instruction with no training. The control thus received standard schooling during comparable class time. Crucially, the phonics instruction and the GL-Rime CARI are precisely the interventions the study sets out to evaluate against business-as-usual; teacher training and the game are integral components of those treatment packages, not separable confounding add- ons. Under the standard, when the additional resource is the explicit treatment variable, a business-as-usual control is acceptable by design. Criterion B is met because the additional resources (phonics instruction, teacher training, and GL-Rime) are integral to the interventions being explicitly tested against a business-as-usual baseline.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent research team has replicated this specific RCT; related prior work is by the same overlapping authors.
      • "Although replications are needed in other contexts... Future studies must examine the replicability of the finding among the ESL learners in India and in other contexts, with similar constraints."
      • Relevant Quotes: 1) "In another, more recent India-based study among Grade-2 level, ESL students... Bora et al. (2024) integrated GL-Rime into classroom phonics instruction..." (p. 4) 2) "Although replications are needed in other contexts, the finding is especially important... Future studies must examine the replicability of the finding among the ESL learners in India and in other contexts, with similar constraints." (p. 21) 3) "Future studies are required to replicate this robust finding." (p. 19) Detailed Analysis: The ERCT R criterion requires this study (or its central experimental claim in the same design and context) to be independently replicated by a different research team and published in a peer-reviewed outlet. This is a newly published 2026 trial. While it extends and cites related prior work (e.g., Bora et al. 2024; Patel et al. 2018, 2022), those studies share overlapping authors (Bora, Richardson, Torppa) and are not independent replications of this particular three-arm RCT. The authors themselves call for future replication. An internet search for an independent reproduction was attempted; the session's web search quota was exhausted, but given the paper's 2026 publication date and the author overlap of all cited related trials, no independent replication of this specific study exists or was identified. Criterion R is not met because there is no independent replication of this specific RCT by a different team.
    • A

      All-subject Exams

      • Only English literacy/reading outcomes were assessed, with no other core subjects, and criterion E is also not met.
      • "This randomized control trial examined the contribution of teacher-led phonics instruction and GraphoLearn-Rime... to the development of English literacy and reading skills."
      • Relevant Quotes: 1) "This randomized control trial examined the contribution of teacher-led phonics instruction and GraphoLearn-Rime (GL-Rime) to the development of English literacy and reading skills of Grade 1 level children..." (Abstract) 2) "Students were assessed using a set of oral- and paper- based as well as in-game measures in the GL-Rime..." (p. 8) [all measures concern literacy/reading] Detailed Analysis: The ERCT A criterion requires impact to be measured across all main subjects using standardized exams, and it depends on criterion E being met. This study assessed only English phonological awareness and reading-related outcomes; no other core subject (e.g., mathematics, science) was measured. Furthermore, the outcome measures were largely researcher-developed rather than standardized exams, so criterion E is not met, which by rule makes A not met as well. Criterion A is not met because only reading/literacy outcomes were assessed and criterion E is not satisfied.
    • G

      Graduation Tracking

      • Measurement stopped at post-test with no follow-up or graduation tracking, and criterion Y is also not met.
      • "Third, since the design did not contain a follow-up, it is not possible to assess the long-term effects of the intervention."
      • Relevant Quotes: 1) "Third, since the design did not contain a follow-up, it is not possible to assess the long-term effects of the intervention. The intervention effects may partially regress or show transfer to other skills in higher grades." (p. 22) 2) "First, examining the interventions' effectiveness on a larger scale over a longer term is a logical next step... longitudinal studies are needed to examine the long-term effectiveness." (p. 22) Detailed Analysis: The ERCT G criterion requires tracking participants through to graduation and depends on criterion Y being met. This study measured outcomes only at post-test at the end of the five-month intervention, with the authors explicitly stating there was no follow-up. There is no tracking to the end of the educational stage. A search for a follow-up publication tracking this cohort to graduation was attempted; the session's web search quota was exhausted, but no such follow-up is referenced in the paper and, given its 2026 publication date, none could yet exist. Since Y is also not met, G cannot be met. Criterion G is not met because there was no follow-up or graduation tracking and criterion Y is not satisfied.
    • P

      Pre-Registered

      • The paper contains no reference to a pre-registered protocol, registry ID, or registration date.
      • Relevant Quotes: 1) "The study was carried out in accordance with the guidelines given by the Finnish National Board on Research Integrity (TENK; TENK 2019)." (p. 6) 2) "The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions." (Data Availability Statement, p. 23) Detailed Analysis: The ERCT P criterion requires the full study protocol to be pre-registered before data collection, with a registry reference and date. The paper mentions ethical guidelines (TENK) and a data availability statement but provides no pre-registration statement, no registry platform or ID (e.g., ClinicalTrials.gov, ISRCTN, OSF, AEA), and no registration date. There is no registry to verify against, and no evidence that hypotheses and analysis plans were registered before data collection. Criterion P is not met because no pre-registration of the protocol is referenced anywhere in the paper.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.