Abstract
The Nuffield Early Language Intervention (NELI) is designed to improve the language skills of reception pupils (aged four to five) through scripted individual and small-group teaching sessions delivered by school staff. As part of the Department for Education's efforts to support education recovery following Covid-19, NELI was made available at no cost to state-funded schools; around 4,000 additional schools registered for wave two in 2021/22. This study, conducted by NFER, evaluated the impact of NELI delivered at national scale on pupils' oral language skills using a quasi-experimental Fuzzy Regression Discontinuity (FRD) design, since the evaluation did not lend itself to a randomised controlled trial. Five hundred and forty-eight schools (19,212 pupils) agreed to take part; the final analytical sample comprised 10,759 pupils in 356 schools. Pupils who received NELI made the equivalent of four additional months' progress in language skills, on average, compared to pupils who did not receive NELI (effect size 0.297, moderate to high security rating). FSM-eligible pupils made an additional seven months' progress; EAL pupils made an additional four months' progress, though this was not statistically significant.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- The study explicitly states it is not a randomised controlled trial but a quasi-experimental Fuzzy Regression Discontinuity design, with pupils selected for NELI by schools using LanguageScreen scores plus subjective teacher/TA judgement rather than random assignment.
- "the evaluation did not lend itself to a randomised controlled trial. We therefore adopted a quasi-experimental approach using FRD" (p. 13)
Relevant Quotes:
1) "As implementation of the second wave of the programme had already begun when the evaluation was commissioned, and one of the objectives of the evaluation was to understand the impact of NELI delivered at scale under real-world conditions, the evaluation did not lend itself to a randomised controlled trial." (p. 13)
2) "We therefore adopted a quasi-experimental approach using FRD—this was contingent upon the way in which pupils were selected by schools to receive NELI." (p. 13)
3) "Whereas the efficacy and effectiveness trials used randomised controlled trials to evaluate the impact of NELI, this quasi-experimental evaluation used a Fuzzy Regression Discontinuity design that leveraged the treatment assignment rule created by the LanguageScreen cutoff score." (p. 40)
4) "However, early findings from the IPE suggested that only 33% of school staff (of 181 surveyed in December 2021) selected pupils based on LanguageScreen scores alone: 66% selected pupils based on other factors in addition to LanguageScreen." (p. 14)
Detailed Analysis:
Criterion C requires that the study be a Randomised Controlled Trial with randomisation clearly described at least at the class level (or an equivalent stronger level, or a valid individual-tutoring exception). This paper explicitly and repeatedly disclaims RCT status: it states the design "did not lend itself to a randomised controlled trial" and instead used a quasi-experimental Fuzzy Regression Discontinuity design. Treatment assignment was determined by where a pupil's LanguageScreen baseline score fell relative to an implicit, statistically modelled class cutoff, combined in the majority of schools with additional non-random selection criteria (behavioural factors, SEN, EAL status, teacher judgement; Figure 1). No entire classes, schools, or individual pupils were randomly allocated to treatment or comparison conditions at any point. While NELI targets pupils individually (potentially engaging the tutoring exception), the exception requires the underlying design still be a genuine (student-level) RCT with random assignment; here assignment was non-random (need/score-based), so the exception does not rescue the criterion.
Final sentence: Criterion C is not met because the study is explicitly a quasi-experimental Regression Discontinuity design, not a randomised controlled trial, with treatment determined by baseline need/score rather than random allocation.
-
E
Exam-based Assessment
- Pupils' oral language outcomes were measured using LanguageScreen, an independently developed, widely used, standardised tablet-based assessment with strong reported reliability, not created for this study.
- "LanguageScreen is a tablet-based standardised assessment... provided to schools new to NELI as part of the DfE-funded offer" (p. 13)
Relevant Quotes:
1) "LanguageScreen is a tablet-based standardised assessment (see Outcome Measures section below) provided to schools new to NELI as part of the DfE-funded offer and hence was already being used as part of the wave two implementation to undertake baseline assessments." (p. 13)
2) "West et al. (2021) reported that LanguageScreen reliability was high in the effectiveness RCT (pre-test screening Cronbach's alpha = 0.84) with good concurrent validity, while the standardisation paper, Hulme et al. (submitted), reported that the LanguageScreen total score has excellent reliability (Cronbach's alpha = 0.92; Person Separation Reliability = 0.94)." (p. 17)
3) "the standardised score being based on a sample of 348,944 children that was used for standardisation" (p. 16)
Detailed Analysis:
LanguageScreen is not a bespoke instrument created to showcase this particular study's results; it is a separately developed, independently standardised, widely deployed screening tool used across the whole NELI programme nationally, with a very large standardisation sample and documented high reliability. This satisfies the requirement for a standard, validated, widely recognised assessment rather than a custom test aligned too closely to the intervention.
Final sentence: Criterion E is met because LanguageScreen is a widely used, independently validated standardised assessment, not a custom-built test for this study.
-
T
Term Duration
- Outcomes were measured around 20 weeks after intervention delivery began, exceeding the minimum one-term (roughly 3-4 month) requirement.
- "measured in the summer of 2022 (endline), around 20 weeks after the expected start of the delivery of the intervention to pupils" (p. 16)
Relevant Quotes:
1) "Following completion of training, intervention delivery was scheduled to begin in schools in January 2022 and to be completed by end of the summer term of 2022." (p. 9)
2) "The evaluation included one primary outcome, a standardised score measuring pupils' oral language skills measured in the summer of 2022 (endline), around 20 weeks after the expected start of the delivery of the intervention to pupils." (p. 16)
3) "Schools were invited to take part in the evaluation in June 2022 and data collection was completed by September that year." (p. 5)
Detailed Analysis:
The intervention began in January 2022 and the primary (endline) outcome was measured approximately 20 weeks later, in the summer term of 2022, with data collection finishing in September 2022. Twenty weeks (roughly 4.5-5 months) comfortably exceeds the minimum requirement of one academic term (approximately 3-4 months) between intervention start and outcome measurement.
Final sentence: Criterion T is met because the interval from intervention start to outcome measurement (around 20 weeks) exceeds one academic term.
-
D
Documented Control Group
- The comparison group (pupils not selected for NELI, receiving usual teaching) is extensively documented by size and by demographic, regional, and school-level characteristics in multiple tables.
- "the remaining 8,430 pupils did not [receive NELI], likely receiving teaching as usual" (p. 26)
Relevant Quotes:
1) "Our final sample for analysis therefore comprised both baseline and endline data for 10,759 pupils in 510 classes (356 schools): 2,329 pupils received NELI while the remaining 8,430 pupils did not, likely receiving teaching as usual." (p. 26)
2) Table 8 reports covariate balance for the comparison group by region, rural/urban status, Ofsted rating, gender, FSM-eligibility, EAL status, and SEN status, both for all pupils entering the model and for those within the MSE-optimal bandwidth. (p. 30-31)
3) Table 24 (Appendix J) further characterises the comparison group of 12,514 pupils recruited to the evaluation, including school governance type, region, Ofsted rating, and pupil demographics. (p. 69-71)
Detailed Analysis:
The paper provides detailed, multi-table documentation of the comparison group's size, demographic composition, school characteristics, and baseline status, comparable to the level of detail given for the intervention group. This allows a reader to assess comparability between groups at baseline (also used explicitly for covariate-balance checks).
Final sentence: Criterion D is met because the comparison group's size, characteristics, and status (usual teaching, no NELI) are documented in detail across several tables.
-
Level 2 Criteria
-
S
School-level RCT
- As with criterion C, there was no randomisation at all (let alone at school level); pupils were selected for NELI within schools/classes based on LanguageScreen scores and teacher discretion.
- "the evaluation did not lend itself to a randomised controlled trial" (p. 13)
Relevant Quotes:
1) "the evaluation did not lend itself to a randomised controlled trial. We therefore adopted a quasi-experimental approach using FRD" (p. 13)
2) "Schools selected pupils to receive the intervention on the basis of the LanguageScreen baseline assessment scores and other criteria described below." (p. 14)
3) "66% selected pupils based on other factors in addition to LanguageScreen" (p. 14), referring to behavioural factors, SEN, EAL, and teacher/TA assessments (Figure 1).
Detailed Analysis:
Criterion S requires random assignment at the school level (a stronger requirement than class-level randomisation). Here, whole schools were not randomised to intervention/comparison conditions at all; instead, individual pupils within the same schools and classes were assigned to NELI or comparison status based on a score-based cutoff (itself subject to considerable additional discretionary/needs-based selection). There is no random allocation mechanism at any unit of analysis (student, class, or school), so the stronger school-level requirement cannot be satisfied.
Final sentence: Criterion S is not met because there was no random allocation to treatment and comparison conditions at the school level, or indeed at any level.
-
I
Independent Conduct
- The impact evaluation was conducted independently by NFER, separate from OxEd (the delivery partner) and the University of Oxford-based developers of NELI.
- "This independent impact evaluation of the Nuffield Early Language Intervention (NELI) wave 2 scale-up was undertaken by the National Foundation for Educational Research (NFER)." (p. 3)
Relevant Quotes:
1) "This independent impact evaluation of the Nuffield Early Language Intervention (NELI) wave 2 scale-up was undertaken by the National Foundation for Educational Research (NFER). The evaluation team was led by Jack Worth, Lead Economist." (p. 3)
2) "NELI was developed by researchers led by Professors Charles Hulme and Maggie Snowling (now at the University of Oxford) including Silke Fricke and Claudine Bowyer-Crane (now at the University of Sheffield)... Recruitment of schools for the scale-up impact evaluation was conducted by the NFER." (p. 8-9)
3) "The DfE is the data controller and makes decisions about how personal data is used in the evaluation. The EEF and OxEd and Assessment Ltd are the data processors and the NFER is the data sub-processor." (p. 12)
Detailed Analysis:
NFER, an independent research organisation with no role in designing or delivering NELI, led the impact evaluation, including school recruitment, data analysis, and reporting, while OxEd (delivery partner) and the intervention's academic developers are organisationally and financially separate. This matches the intent of criterion I: implementation and analysis were conducted by a party independent of the intervention's design and delivery.
Final sentence: Criterion I is met because the evaluation was led by the independent NFER, distinct from the NELI developers and delivery partner OxEd.
-
Y
Year Duration
- Outcomes were measured only around 20 weeks (roughly five months) after intervention start, well short of 75% of a full academic year.
- "around 20 weeks after the expected start of the delivery of the intervention to pupils" (p. 16)
Relevant Quotes:
1) "intervention delivery was scheduled to begin in schools in January 2022 and to be completed by end of the summer term of 2022." (p. 9)
2) "measured in the summer of 2022 (endline), around 20 weeks after the expected start of the delivery of the intervention to pupils." (p. 16)
3) "The 20-week intervention consists of two 15-minute individual sessions and three 30-minute small group sessions each week" (p. 4), confirming the programme itself (and hence the tracked interval) spans 20 weeks, not a full year.
Detailed Analysis:
Criterion Y requires outcome measurement at least 75% of a full academic year (roughly 9-10 months) after intervention start. Here, the interval from intervention start (January 2022) to outcome measurement (summer 2022, around 20 weeks / 4.5-5 months later) covers only around half of a typical academic year, well below the 75% threshold. There is no indication of a longer follow-up period for the primary outcome.
Final sentence: Criterion Y is not met because the approximately 20-week interval between intervention start and outcome measurement is well short of 75% of an academic year.
-
B
Balanced Control Group
- NELI's additional small-group/individual sessions are the explicit treatment variable being tested; the comparison group received standard, business-as-usual teaching, which is an acceptable baseline under this exception.
- "the remaining 8,430 pupils did not [receive NELI], likely receiving teaching as usual" (p. 26)
Relevant Quotes:
1) "The 20-week intervention consists of two 15-minute individual sessions and three 30-minute small group sessions each week, delivered to the three to six pupils with the weakest language skills." (p. 4)
2) "While NELI sessions were delivered during normal classroom hours, pupils selected to receive the programme were taken out of classes." (p. 9)
3) "Our final sample for analysis therefore comprised both baseline and endline data for 10,759 pupils... 2,329 pupils received NELI while the remaining 8,430 pupils did not, likely receiving teaching as usual." (p. 26)
4) "This study, conducted by the NFER, evaluated the impact of NELI delivered at national scale... on pupil's oral language skills" (p. 5), confirming the study's explicit purpose is to test the effect of providing the NELI sessions themselves.
Detailed Analysis:
Following the criterion B decision procedure: extra resources are present (NELI provides additional, targeted individual and group teaching time, materials, and TA training beyond ordinary classroom instruction). However, these additional resources are precisely the treatment variable under investigation — the whole evaluation exists to determine whether providing this extra, targeted language teaching improves outcomes relative to ordinary classroom teaching. Under the ERCT decision tree, when extra resources are the thing being tested, the comparison group may legitimately receive standard "business as usual" instruction without a matched substitute. The comparison pupils here received ordinary teaching, consistent with this exception, so no confounding imbalance is introduced beyond what the study design intends to test.
Final sentence: Criterion B is met because NELI's additional teaching time and materials are the explicit treatment variable being evaluated, and the comparison group's receipt of standard, business-as-usual teaching is the appropriate baseline under this exception.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific national-scale, quasi-experimental evaluation was found; prior NELI trials are precursor evidence from the same evaluation pipeline, not independent reproductions of this study.
- "This national scale-up evaluation is the final stage in the EEF's 'evaluation pipeline'" (p. 5)
Relevant Quotes:
1) "This national scale-up evaluation is the final stage in the EEF's 'evaluation pipeline' (EEF, 2023) following a pilot study (Fricke et al., 2013, funded by Nuffield Foundation), an efficacy trial (Sibieta, Kotecha, and Skipp, 2016), and an effectiveness trial (Dimova et al., 2020)." (p. 5)
2) "This is, therefore, comparable to the estimate made by this evaluation, although we estimated the effect size to be slightly lower (0.297)." (p. 42), referring to the earlier effectiveness trial's broadly similar effect size.
3) "Further long-term data analysis building on the effectiveness trial has recently been undertaken (Groom, Brown and Lymperis, 2023) but there is the opportunity for long-term impacts of the scale-up to be understood." (p. 44)
Detailed Analysis:
The prior pilot, efficacy, and effectiveness trials of NELI predate this scale-up evaluation and used a different design (individually/cluster randomised controlled trials with smaller, more controlled samples), so they are earlier stages of the same evaluation pipeline rather than independent replications of this specific national-scale, quasi-experimental FRD study.
Internet search update (2026): a search for independent replications of this specific study identified one related, more recent publication: Thompson, D.M., Nelson, N., Hulme, C., Snowling, M.J. and Hogan, T.P. (2025) "From research to practice: The Nuffield early language intervention's (NELI's) journey across the Atlantic" (Emerald, edited volume chapter). Its abstract describes it as "a case study" of the "adaptation and implementation of the Nuffield Early Language Intervention (NELI) in US schools", integrated into a Montana school district's multi-tiered support system. This is an implementation case study, not a controlled replication of this scale-up's FRD design or findings, and it is co-authored by two of NELI's original developers (Hulme and Snowling), so it would not satisfy the independence requirement of criterion R even if it were a replication. No other independent replication of this specific wave two national scale-up evaluation, by a different research team, was identified in available sources.
Final sentence: Criterion R is not met because no independent replication of this specific scale-up evaluation was identified; earlier NELI trials are precursor studies rather than reproductions of this study, and the one related 2025 publication found is a non-independent implementation case study rather than a replication.
-
A
All-subject Exams
- The evaluation measured only oral language outcomes via LanguageScreen; no other core subjects were assessed and no exception rationale is provided.
- "The research questions all focused on one primary outcome: pupils' oral language outcomes measured by the LanguageScreen standardised score" (p. 10)
Relevant Quotes:
1) "This evaluation was undertaken to understand the impact of NELI, when delivered at scale, on children's early language outcomes. The research questions all focused on one primary outcome: pupils' oral language outcomes measured by the LanguageScreen standardised score post-intervention." (p. 10)
2) "Secondary outcomes: The evaluation did not include any secondary outcomes." (p. 16-17)
3) "The evaluation focused on one outcome, a decision which was shaped by some of the constraints and preferences—such as timeline, lack of usable secondary data, and the preference not to collect further primary data from schools." (p. 43)
Detailed Analysis:
Although criterion E is met (LanguageScreen is a standardised assessment), criterion A additionally requires that impact be assessed across all main subjects taught, not just the subject targeted by the intervention, unless a specialised-context exception applies. Here only a single outcome (oral language, via LanguageScreen) was measured; the authors explicitly confirm there were no secondary outcomes and explain this was a deliberate scope decision, not a specialised vocational context exception. No other subjects (e.g., early mathematics/numeracy) were assessed.
Final sentence: Criterion A is not met because only one outcome domain (oral language) was assessed, with no other subjects covered and no qualifying exception.
-
G
Graduation Tracking
- There is no tracking of pupils through to graduation in this report; criterion G also fails automatically because criterion Y is not met.
- "Future research may focus on the other pupil-level outcome identified by the logic model, long-term reading comprehension" (p. 43)
Relevant Quotes:
1) "measured in the summer of 2022 (endline), around 20 weeks after the expected start of the delivery of the intervention to pupils" (p. 16), confirming the only outcome measurement point is the short-term endline.
2) "Future research may focus on the other pupil-level outcome identified by the logic model, long-term reading comprehension, and may also take the opportunity to explore the effect of the intervention on attainment in English measured by national assessments." (p. 43-44)
3) "Data from this evaluation will be deposited in the EEF's archive, linking it to NPD national assessment data (for example, KS1, KS2) thus providing the basis for subsequent analysis of the NELI wave two scale-up without the need for additional data collection." (p. 44)
Detailed Analysis:
Per the ERCT rules, since criterion Y (Year Duration) is not met, criterion G is automatically not met. Independently, the report itself confirms there is no follow-up beyond the ~20-week endline measurement within this document; long-term tracking through to graduation is described only as a possibility for future research using archived, linked administrative data, not as work that has been completed and reported here.
Internet search update (2026): a search for subsequent papers by these authors tracking the wave two scale-up cohort to graduation found none. The only relevant material located was this report's own prospective statement about future EEF archive linkage to KS1/KS2 data, which is a plan rather than completed graduation-tracking evidence. No follow-up publication reporting outcomes for this specific cohort was identified.
Final sentence: Criterion G is not met, both because criterion Y is not met and because no tracking through to graduation was conducted or reported in this evaluation, nor found in a search for subsequent publications.
-
P
Pre-Registered
- The study's quasi-experimental design was registered with OSF in February 2023, after data collection had already been completed in September 2022, so pre-registration did not precede data collection.
- "This QED was registered on 2 February 2023 with OSF Registries." (p. 11)
Relevant Quotes:
1) "This QED was registered on 2 February 2023 with OSF Registries." (Registration DOI: https://doi.org/10.17605/OSF.IO/5M8JF) (p. 11)
2) "Schools were invited to take part in the evaluation in June 2022 and data collection was completed by September that year." (p. 5)
3) Table 5 (Timeline): "November–December 2022: Finalise and publish study plan." and "December 2022–March 2023: Data linking with NPD. Complete preliminary analysis. Main analysis." (p. 24-25), showing both the study plan and formal registration postdate data collection.
Detailed Analysis:
Criterion P requires that the full study protocol be pre-registered before data collection begins. Here, school recruitment, baseline testing, and endline data collection were completed by September 2022, yet the formal OSF registration occurred in February 2023 and even the study plan was only finalised in November-December 2022 — both well after data collection ended. This sequencing (data collection substantially preceding registration) directly contravenes the timing requirement of criterion P, regardless of the fact that a registration exists.
Internet search update (2026): an attempt was made to independently confirm the OSF registration timing directly on the OSF registry (DOI 10.17605/OSF.IO/5M8JF), but the registry page could not be rendered as readable text by available tools. The paper's own explicit, internally consistent statement of the registration date (2 February 2023) versus data-collection completion (September 2022) is taken as sufficient evidence for this determination.
Final sentence: Criterion P is not met because the protocol/registration postdates completion of data collection rather than preceding it.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.