Early language screening and intervention can be delivered successfully at scale: evidence from a cluster randomized controlled trial

Gillian West, Margaret J. Snowling, Arne Lervag, Elizabeth Buchanan-Worster, Mihaela Duta, Alexandra Hall, Henrietta McLachlan, and Charles Hulme

Published:
ERCT Check Date:
DOI: 10.1111/jcpp.13415
  • reading
  • language arts
  • kindergarten
  • UK
  • digital assessment
1
  • C

    Randomisation occurred at the school level (193 schools), which is a stronger unit than class-level and therefore automatically satisfies the class-level RCT requirement.

    "Following recruitment, schools were randomly allocated, within geographical area, to either a 20-week oral language intervention group or a business-as-usual control group." (p. 1426)

  • E

    The study used well-established standardized language assessments (CELF Preschool, Renfrew APT), not custom-built instruments.

    "As outlined in the trial preregistration, the primary outcome measures were the four standardized tests of language ability administered at t1 and t2 to children identified as eligible for the NELI programme." (p. 1427)

  • T

    Outcomes were measured roughly 7-8 months after the intervention began, far exceeding the one-term minimum.

    "Assessments took place before the start of the intervention at screening (t0)...and at pretest (t1)...and immediately following the intervention (post-test, t2)." (p. 1426)

  • D

    The control group's size, demographics, baseline scores and the provision it received are documented in detail in the text, CONSORT diagram and Table 1.

    "Schools in the control group delivered their usual school provision and received payment to purchase the programme at the end of the trial if they wished." (p. 1426)

  • S

    The trial randomly allocated entire schools (193 schools) to intervention or control, directly satisfying the school-level criterion.

    "A cluster randomized controlled trial (RCT) was conducted in 193 state primary schools (containing 238 Reception classrooms)..." (p. 1426)

  • I

    Despite independent randomisation and blinded testers, several co-authors are the intervention's original designers or hold direct financial stakes in commercial products central to the trial.

    "C.H., M.S., and G.W. are Directors and M.D. is a shareholder of OxEd and Assessment Ltd, a University of Oxford spin-out company founded to distribute LanguageScreen as a commercial product." (Acknowledgements, p. 1433)

  • Y

    Tracking from intervention start to the final outcome measurement spans roughly 8 months, exceeding 75% of the UK academic year.

    Figure 1 timeline shows intervention training beginning "Nov 2018" with final "in-depth posttest (t2, June - July 2019)." (p. 1427)

  • B

    The additional teaching-assistant time and training is the treatment variable being tested; the control group's business-as-usual provision is the appropriate comparison, not an imbalance.

    "Schools in the control group delivered their usual school provision and received payment to purchase the programme at the end of the trial if they wished." (p. 1426)

  • R

    No independent replication of this specific large-scale cluster RCT by a different research team was found in the paper or via further internet search.

  • A

    Only oral language and early word reading were assessed; no other core subject (e.g. mathematics) was measured, with no stated rationale for this narrow scope.

    "Secondary outcome measures were the school administered LanguageScreen scores and word reading ability (YARC Early Word Reading subtest)." (p. 1427)

  • G

    Same-team follow-ups extend tracking to only about two years post-intervention (around ages 6-7); no evidence of tracking through graduation was found.

    "Further research is needed to assess the possible long-term effects of such language interventions and their implications for educational and social policy." (Conclusion, p. 1433)

  • P

    The trial was preregistered on the ISRCTN registry on 2018-06-05, before screening (t0) began in September 2018, and the paper states its analyses followed this preregistered plan.

    "The trial was preregistered on the ISRCTN registry (https://doi.org/10.1186/ ISRCTN12991126)." (p. 1426)

Abstract

Background: It is well established that oral language skills provide a critical foundation for formal education. This study evaluated the effectiveness of the Nuffield Early Language Intervention (NELI) programme in ameliorating language difficulties in the first year of school when delivered at scale. Methods: We conducted a cluster randomized controlled trial (RCT) in 193 primary schools (containing 238 Reception classrooms). Schools were randomly allocated to either a 20-week oral language intervention or a business-as-usual control group. All classes (N = 5,879 children) in participating schools were screened by school staff using an automated App to assess children's oral language skills. Screening identified 1,173 children as eligible for language intervention: schools containing 571 of these children were allocated to the control group and 569 to the intervention group. Results: Children receiving the NELI programme made significantly larger gains than the business-as-usual control group on a latent variable reflecting standardized measures of language ability (d = .26) and on the school-administered automated assessment of receptive and expressive language skills (d = .32). The effects of intervention did not vary as a function of home language background or gender. Conclusions: This study provides strong evidence for the effectiveness of a school-based language intervention programme (NELI) delivered at scale. These findings demonstrate that language difficulties can be identified by school-based testing and ameliorated by a TA delivered intervention; this has important implications for educational and social policy.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation occurred at the school level (193 schools), which is a stronger unit than class-level and therefore automatically satisfies the class-level RCT requirement.
      • "Following recruitment, schools were randomly allocated, within geographical area, to either a 20-week oral language intervention group or a business-as-usual control group." (p. 1426)
      • Relevant Quotes: 1) "A cluster randomized controlled trial (RCT) was conducted in 193 state primary schools (containing 238 Reception classrooms) from 13 geographical areas in the UK..." (p. 1426) 2) "Following recruitment, schools were randomly allocated, within geographical area, to either a 20-week oral language intervention group or a business-as-usual control group." (p. 1426) 3) "After completion of t0 and t1 testing, schools were randomized to intervention or control group by an independent evaluator. Randomization was stratified by geographical area and the number of classes participating in each school (dichotomized: 1, or more than 1, class)." (p. 1427) Detailed Analysis: The paper explicitly and repeatedly states that the unit of randomisation was the school, not individual classes or students. Entire schools (193 of them) were allocated to intervention or control, with stratification by geographical area and number of participating classes. Under the ERCT Standard, a school-level RCT is a stronger design than a class-level RCT and, where present, automatically satisfies the weaker class-level criterion, since randomising whole schools eliminates any risk of within-class or within-school contamination between treatment and control pupils. Criterion C is met because randomisation was conducted at the school level, which exceeds the class-level requirement.
    • E

      Exam-based Assessment

      • The study used well-established standardized language assessments (CELF Preschool, Renfrew APT), not custom-built instruments.
      • "As outlined in the trial preregistration, the primary outcome measures were the four standardized tests of language ability administered at t1 and t2 to children identified as eligible for the NELI programme." (p. 1427)
      • Relevant Quotes: 1) "As outlined in the trial preregistration, the primary outcome measures were the four standardized tests of language ability administered at t1 and t2 to children identified as eligible for the NELI programme." (p. 1427) 2) "Language skills were assessed with the Expressive Vocabulary subtest from the Child Evaluation of Language Fundamentals (CELF) Preschool IIUK (Semel, Wiig, & Secord, 2006), The Renfrew Action Picture Test (APT; Renfrew, 2003; information and grammar scores) and the Recalling sentences subtest from the Child Evaluation of Language Fundamentals (CELF) Preschool IIUK (Semel et al., 2006)." (p. 1427) Detailed Analysis: The primary outcome measures are drawn from the CELF Preschool (Semel, Wiig, & Secord, 2006) and the Renfrew Action Picture Test (Renfrew, 2003), both long-established, commercially published, psychometrically validated standardized language assessments used widely in research and clinical practice, not instruments created for this study. The authors themselves label these "standardized tests of language ability." Criterion E is met because the primary outcome measures are recognised, published standardized assessments rather than custom-built tests.
    • T

      Term Duration

      • Outcomes were measured roughly 7-8 months after the intervention began, far exceeding the one-term minimum.
      • "Assessments took place before the start of the intervention at screening (t0)...and at pretest (t1)...and immediately following the intervention (post-test, t2)." (p. 1426)
      • Relevant Quotes: 1) Figure 1 timeline: "Screening (t0, Sept 2018)", "In-depth pretest (t1, Oct 2018)", "2-day training (Nov 2018)" marking the start of intervention delivery, through to "Concurrent screening & in-depth posttest (t2, June - July 2019)." (p. 1427) 2) "Assessments took place before the start of the intervention at screening (t0) for all children in participating classrooms and at pretest (t1) for children selected via screening...and immediately following the intervention (post-test, t2). The timeline is presented in Figure 1." (p. 1426) Detailed Analysis: Intervention delivery (TA training and NELI sessions) began around November 2018, and the final outcome measurement (t2) took place in June-July 2019, an interval of roughly 7-8 months. This comfortably exceeds the minimum one academic term (approximately 3-4 months) required by this criterion. Criterion T is met because the interval from intervention start to outcome measurement substantially exceeds one academic term.
    • D

      Documented Control Group

      • The control group's size, demographics, baseline scores and the provision it received are documented in detail in the text, CONSORT diagram and Table 1.
      • "Schools in the control group delivered their usual school provision and received payment to purchase the programme at the end of the trial if they wished." (p. 1426)
      • Relevant Quotes: 1) "Schools in the control group delivered their usual school provision and received payment to purchase the programme at the end of the trial if they wished." (p. 1426) 2) "Allocated to Control group: School n = 96; mean cluster size = 7.20; cluster variance = 9.07; children n = 592" (Figure 2, CONSORT diagram, p. 1428) 3) Table 1 reports, for the Control Group (n = 592), baseline and post-test means (SD) for age, LanguageScreen subtests, CELF, APT and YARC word reading measures alongside the intervention group. (p. 1429-1430) 4) "Critically, there were no significant differences at pretest in gender...age...or language factor scores derived from standardized language tests...between children who completed the study and those who dropped out at post-test." (p. 1429) Detailed Analysis: The control condition (business-as-usual teaching, with deferred access to the programme) is clearly described, and the control group's size, cluster structure, demographic composition and detailed baseline test scores are reported in Table 1 and the CONSORT diagram, alongside confirmation that completers and dropouts did not differ significantly at baseline. Criterion D is met because the control group is thoroughly documented with size, baseline characteristics and the treatment it received.
  • Level 2 Criteria

    • S

      School-level RCT

      • The trial randomly allocated entire schools (193 schools) to intervention or control, directly satisfying the school-level criterion.
      • "A cluster randomized controlled trial (RCT) was conducted in 193 state primary schools (containing 238 Reception classrooms)..." (p. 1426)
      • Relevant Quotes: 1) "A cluster randomized controlled trial (RCT) was conducted in 193 state primary schools (containing 238 Reception classrooms) from 13 geographical areas in the UK..." (p. 1426) 2) "After completion of t0 and t1 testing, schools were randomized to intervention or control group by an independent evaluator. Randomization was stratified by geographical area and the number of classes participating in each school (dichotomized: 1, or more than 1, class)." (p. 1427) Detailed Analysis: The school, defined as the educational institution implementing the programme, is explicitly the unit of random allocation here: 193 schools were randomised as clusters. This is precisely the design the S criterion requires, and it captures whole-school, real-world implementation rather than isolated classes. Criterion S is met because randomisation was conducted at the school (cluster) level.
    • I

      Independent Conduct

      • Despite independent randomisation and blinded testers, several co-authors are the intervention's original designers or hold direct financial stakes in commercial products central to the trial.
      • "C.H., M.S., and G.W. are Directors and M.D. is a shareholder of OxEd and Assessment Ltd, a University of Oxford spin-out company founded to distribute LanguageScreen as a commercial product." (Acknowledgements, p. 1433)
      • Relevant Quotes: 1) "After completion of t0 and t1 testing, schools were randomized to intervention or control group by an independent evaluator." (p. 1427) 2) "Post-testing (t2) using the App was completed by teachers, but individual post-testing by trained testers was done blind to treatment arm." (p. 1427) 3) "C.H., M.S., and G.W. are Directors and M.D. is a shareholder of OxEd and Assessment Ltd, a University of Oxford spin-out company founded to distribute LanguageScreen as a commercial product." (Acknowledgements, p. 1433) 4) "H. M. and A. H. work for Elklan Training Ltd, an independent training company which provides training for school staff using the Nuffield Early Language Intervention programme." (Acknowledgements, p. 1433) 5) "Schools in the intervention group delivered the Nuffield Early Language Intervention programme with training and delivery support provided by an independent training consultancy (Elklan...)." (p. 1426) Detailed Analysis: Two specific procedural steps were independent: the allocation of schools to arms was carried out by "an independent evaluator," and individual post-testing used testers blind to treatment arm. However, criterion I concerns the independence of the study's conduct as a whole from the designers of the intervention, and here the overlap is substantial. Three co-authors (C.H., M.S., G.W.) are Directors, and a fourth (M.D.) a shareholder, of OxEd and Assessment Ltd, the commercial company distributing LanguageScreen, the very App used both to screen and select eligible children and to generate the secondary outcome measure. Two further co-authors (H.M., A.H.) are staff of Elklan Training Ltd, the company contracted to train and support the teaching assistants who delivered the NELI intervention under evaluation. C.H. and M.S. are additionally co-authors of the published NELI programme itself (Fricke, Bowyer-Crane, Snowling, & Hulme, 2018) and of the earlier efficacy trials that established it (Fricke et al., 2013, 2017), meaning the same core team designed, refined, trained delivery for, and evaluated the intervention, while holding direct commercial stakes in instruments central to its measurement and delivery pipeline. This is a disclosed author conflict of interest (commercial and financial stakes in the LanguageScreen App used to screen, select and measure participants), and it is exactly the pattern of overlapping design, delivery-support and evaluation roles, combined with disclosed financial interests, that the Independent Conduct criterion is designed to flag, notwithstanding the isolated independent randomisation step. Per the ERCT Standard, a disclosed conflict of interest does not by itself automatically determine the verdict on this criterion; the verdict here rests on the documented overlap between intervention designers/beneficiaries and the evaluation team, of which the financial conflict is one clear symptom. Criterion I is not met because several authors are the intervention's original designers or hold financial stakes in the commercial products used to screen participants and measure outcomes, and other co-authors are employees of the firm delivering the intervention's training.
    • Y

      Year Duration

      • Tracking from intervention start to the final outcome measurement spans roughly 8 months, exceeding 75% of the UK academic year.
      • Figure 1 timeline shows intervention training beginning "Nov 2018" with final "in-depth posttest (t2, June - July 2019)." (p. 1427)
      • Relevant Quotes: 1) Figure 1 timeline: "2-day training (Nov 2018)" marking intervention onset, followed by NELI Part 1 (sessions 1-28) and Part 2 (sessions 29-57) delivered across the spring and summer terms, through to "Concurrent screening & in-depth posttest (t2, June - July 2019)." (p. 1427) 2) "Screening (t0, Sept 2018)" through to post-test in "June - July 2019" spans the trial's overall observation window. (p. 1427) Detailed Analysis: The UK academic year runs from roughly September to July (approximately 9-10 months). Here, tracking from the intervention's onset (training in November 2018) to the final outcome assessment (June-July 2019) covers approximately 8 months, which exceeds 75% of a 9-10 month academic year (a threshold of roughly 7-7.5 months). Considered from initial screening (September 2018), the trial's overall window spans essentially the full 2018-19 academic year. Criterion Y is met because the tracking interval from intervention start to final measurement covers at least 75% of a full academic year.
    • B

      Balanced Control Group

      • The additional teaching-assistant time and training is the treatment variable being tested; the control group's business-as-usual provision is the appropriate comparison, not an imbalance.
      • "Schools in the control group delivered their usual school provision and received payment to purchase the programme at the end of the trial if they wished." (p. 1426)
      • Relevant Quotes: 1) "Schools in the intervention group delivered the Nuffield Early Language Intervention programme with training and delivery support provided by an independent training consultancy (Elklan; https://www.elklan.co.uk/)." (p. 1426) 2) "Schools in the control group delivered their usual school provision and received payment to purchase the programme at the end of the trial if they wished." (p. 1426) 3) "NELI is a 20-week programme...The programme includes a total of 57 small group sessions, each lasting 30 min and 37 individual sessions, each lasting 15 min (total intervention time: small group sessions 28.5 h; individual sessions 9.25 h)." (p. 1427) 4) "This study evaluated the effectiveness of the Nuffield Early Language Intervention (NELI) programme when delivered at scale in educationally realistic circumstances." (Discussion, p. 1431) Detailed Analysis: Applying the criterion B decision tree: extra resources are present (the intervention group received substantial additional TA time — 28.5 h of small-group and 9.25 h of individual sessions per child over 20 weeks — plus specialist training and ongoing delivery support from Elklan). These additional resources are explicitly framed as the treatment package whose effectiveness this trial was designed to test ("evaluated the effectiveness of the...NELI programme...delivered at scale"), not as a supplementary add-on to some separate core intervention. Per the ERCT decision logic, when extra resources are themselves the treatment variable under investigation, the control group may legitimately receive standard "business-as-usual" provision without a matched resource, provided this is the explicit study design (as it clearly is here: intervention vs. "business-as-usual control group"). The deferred offer to purchase the programme after the trial further indicates the control condition was genuine business-as-usual rather than a deliberately under-resourced comparator. Criterion B is met because the additional TA-delivered sessions and training are the explicit treatment variable being tested against a business-as-usual control, consistent with the study's stated design intent (RESOURCES_ARE_ TREATMENT branch of the criterion B decision tree), and this integral-resource status is clearly documented above.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific large-scale cluster RCT by a different research team was found in the paper or via further internet search.
      • Detailed Analysis: The paper itself situates this trial as an extension of earlier NELI efficacy studies (Bowyer-Crane et al., 2008; Fricke et al., 2013, 2017), but those studies share overlapping authorship with the present trial (C.H. and M.S. are co-authors across all of them) and, in any case, precede rather than independently replicate this specific 2021 large-scale trial. A citation search (via OpenAlex, covering the 47 works citing this paper as of the current check) for independent replication found no study by a wholly different research team replicating this specific cluster-RCT design and findings in a new context. All NELI-related citing works identified continue to share authors with the original team (C. Hulme, M. Snowling and/or A. Lervag): West, Lervag, Birchenough, Korell, Rios Diaz, Duta, Cripps, Gardner, Fairhurst, & Hulme (2024), "Oral language enrichment in preschool improves children's language skills: a cluster randomised controlled trial," JCPP, doi:10.1111/jcpp.13947, is a related but distinct preschool-enrichment RCT, not a replication of this study. Esposito, Lervag, & Hulme (2024), "Oral language intervention in the late primary school years is effective: evidence from a randomised control trial," JCPP, doi:10.1111/jcpp.14084, again shares authorship and tests a different age-group/intervention variant rather than replicating this trial. West, Lervag, Snowling, Buchanan-Worster, Duta, & Hulme (2022) and Hulme, West, Rios Diaz, Hearne, Korell, Duta, & Snowling (2025) are same-cohort follow-ups of this exact trial, not independent replications (see Criterion G). Two studies by unrelated teams were also identified but do not replicate this trial: Acosta Rodriguez, Ramirez Santana, & Hernandez Exposito (2021), an oral-language comprehension intervention for Spanish preschool children with developmental language disorder, is a different intervention in a different population; Donolato et al. (2023) is a systematic review/meta-analysis, not a primary replication study. No evidence of an independent research team replicating this specific trial's cluster-RCT design and findings in a different context was identified. Criterion R is not met because no independent, peer-reviewed replication of this specific study by a different research team was found.
    • A

      All-subject Exams

      • Only oral language and early word reading were assessed; no other core subject (e.g. mathematics) was measured, with no stated rationale for this narrow scope.
      • "Secondary outcome measures were the school administered LanguageScreen scores and word reading ability (YARC Early Word Reading subtest)." (p. 1427)
      • Relevant Quotes: 1) "Our primary outcome measure was a language latent variable defined by the standardized measures of language ability (i.e. CELF expressive vocabulary, CELF recalling sentences and APT information and grammar scores)." (p. 1429) 2) "Secondary outcome measures were the school administered LanguageScreen scores and word reading ability (YARC Early Word Reading subtest)." (p. 1427) Detailed Analysis: All outcome measures in this trial (CELF, APT, LanguageScreen, YARC word reading) assess oral language and early reading only. No other core curriculum area for Reception-age children (most notably early mathematics/numeracy, a specific area of the Early Years Foundation Stage curriculum) was assessed. The paper offers no rationale for restricting assessment to language and reading beyond the intervention's own focus, and this is not a specialised upper-secondary or vocational context to which the ERCT exception for narrow assessment would apply. Criterion A is not met because the study assessed only language and reading outcomes, with no measurement of, or justification for omitting, other main subjects such as mathematics.
    • G

      Graduation Tracking

      • Same-team follow-ups extend tracking to only about two years post-intervention (around ages 6-7); no evidence of tracking through graduation was found.
      • "Further research is needed to assess the possible long-term effects of such language interventions and their implications for educational and social policy." (Conclusion, p. 1433)
      • Relevant Quotes: 1) "Further research is needed to assess the possible long-term effects of such language interventions and their implications for educational and social policy." (Conclusion, p. 1433) 2) "The current paper reports a large cluster randomized trial in which language intervention was delivered in some 100 schools and confirms that implementation at scale is practicable. However, inevitably there were some limitations." (Limitations, p. 1432) Detailed Analysis: Outcome data in this paper are reported only up to the immediate post-test (t2, June-July 2019), at the end of children's Reception year, and the authors explicitly flag long-term effects as an open question requiring future research. A citation search (via OpenAlex) identified two same-cohort follow-up publications by overlapping authors: West, Lervag, Snowling, Buchanan-Worster, Duta, & Hulme (2022), "Early language intervention improves behavioral adjustment in school: Evidence from a cluster randomized trial," Journal of School Psychology, 92, 334-345, doi:10.1016/j.jsp.2022.04.006, reporting behavioural-adjustment outcomes for this trial's cohort; and Hulme, West, Rios Diaz, Hearne, Korell, Duta, & Snowling (2025), "The Nuffield Early Language Intervention (NELI) programme is associated with lasting improvements in children's language and reading skills," JCPP, doi:10.1111/jcpp.14157, which (per its indexed metadata) reassessed language and reading approximately two years after the intervention ended, when children were around 6-7 years old. I was not able to retrieve full verbatim text from either follow-up paper (only abstracts/ metadata were accessible), so no direct quotes from them are included here, per instructions not to fabricate quotes. Neither follow-up's indexed abstract/metadata indicates tracking continued through the end of primary schooling or graduation from that educational stage, which for Reception-age children (school entry around age 4-5) would be several more years away. Criterion G is not met because tracking, even including identified follow-up publications extending to roughly two years post-intervention, stops well short of graduation from the relevant educational stage.
    • P

      Pre-Registered

      • The trial was preregistered on the ISRCTN registry on 2018-06-05, before screening (t0) began in September 2018, and the paper states its analyses followed this preregistered plan.
      • "The trial was preregistered on the ISRCTN registry (https://doi.org/10.1186/ ISRCTN12991126)." (p. 1426)
      • Relevant Quotes: 1) "The trial was preregistered on the ISRCTN registry (https://doi.org/10.1186/ ISRCTN12991126)." (p. 1426) 2) "The analyses followed the preregistered plan (https://doi.org/10.1186/ISRCTN12991126)." (p. 1428) 3) "As outlined in the trial preregistration, the primary outcome measures were the four standardized tests of language ability administered at t1 and t2 to children identified as eligible for the NELI programme." (p. 1427) Detailed Analysis: The paper provides a specific, resolvable registry identifier (ISRCTN12991126) and explicitly states that both the primary outcome measures and the statistical analysis plan followed the preregistered protocol. I checked the ISRCTN registry directly: the trial record for ISRCTN12991126 shows a registration date of 2018-06-05, with a recruitment start date of 2018-06-11. Screening (t0), the trial's earliest data collection activity, began in September 2018 per the paper's timeline (Figure 1). Registration in June 2018 therefore clearly precedes the start of data collection in September 2018. The trial was also funded by the Education Endowment Foundation, a funder that requires prospective registration of the evaluations it commissions. Criterion P is met because the study documents a public registry entry, independently confirmed to predate data collection by several months, and explicitly states that its outcome measures and analysis plan followed the preregistered protocol.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.