Level 1 Criteria
-
C Class-level RCT
- Randomisation was conducted at the individual pupil level within each nursery/school rather than at the class or school level, and the authors themselves flag this as creating a risk of contamination between treatment and control pupils.
- "The trial was designed as a randomised controlled trial, with randomisation at the pupil level within nurseries attached to schools." (p. 11)
- Relevant Quotes: 1) "The trial was designed as a randomised controlled trial, with randomisation at the pupil level within nurseries attached to schools." (p. 11) 2) "The randomisation was undertaken independently by members of the impact evaluation team who randomly allocated pupils within each nursery to one of three groups—a control group or one of two treatment groups." (p. 11) 3) "The main downside of within-school randomisation is that spillover effects are more likely resulting from, for example, TAs applying the intervention techniques to pupils in the control group." (p. 11) 4) "Randomisation was at the child level within nurseries/schools." (p. 4, Executive summary) Detailed Analysis: The ERCT 'C' criterion requires randomisation at the class level or stronger (school level), with an exception only where the intervention is designed for personal, one-to-one tutoring. Here, randomisation was explicitly conducted at the individual pupil level within each nursery, meaning treatment and control children came from the same classes/settings. The intervention is delivered mainly in small groups of two to four children (3 x 20-30 minute sessions per week), supplemented in reception only by two 15-minute individual sessions per week; it is therefore not a purely one-to-one personal tutoring intervention that would qualify for the stated exception. Critically, the authors themselves acknowledge that this within-nursery/school, pupil-level design creates a real risk of spillover, i.e. TAs delivering intervention techniques to control pupils in the same setting, which is precisely the contamination problem criterion C is designed to prevent. Final sentence: Criterion C is not met because randomisation occurred at the individual pupil level within the same nursery/school, with no valid one-to-one tutoring exception, and the authors explicitly acknowledge the resulting contamination risk.
-
E Exam-based Assessment
- The primary and secondary outcomes were measured using established, externally-valid standardised instruments (CELF, Renfrew APT, YARC), not custom-built tests.
- "The primary outcome for this evaluation is a language skills score which is a composite score of four different externally-valid measures of language skill:" (p. 13)
- Relevant Quotes: 1) "The primary outcome for this evaluation is a language skills score which is a composite score of four different externally-valid measures of language skill: Renfrew Action Picture Test (APT)... CELF-Preschool 2 UK: Expressive Vocabulary... Listening Comprehension (based on the York Assessment of Reading Comprehension test, YARC)." (p. 13) 2) "YARC: Letter Knowledge... YARC: Early Word Reading... Spelling: children are asked to write a series of simple words." (p. 13) 3) "All these component measures are standardised, age appropriate, and were chosen by the project team to be consistent with the measures used in their Randomised Controlled Trial (Fricke et al., 2013), and the aims of the intervention to improve language and literacy skills." (p. 13) 4) "All tests were administered and scored by research assistants trained by the project team who were blind to the allocation of children to groups." (p. 13) Detailed Analysis: The outcomes are drawn from recognised, externally-published standardised psychometric instruments (the Renfrew Action Picture Test, CELF-Preschool, and YARC), each with established norms and prior use in other studies, rather than being custom-designed solely for this trial. Tests were administered by blinded, trained assessors, further supporting objectivity. Final sentence: Criterion E is met because the study used widely-used, externally-valid standardised assessments (CELF, Renfrew APT, YARC) rather than bespoke measures.
-
T Term Duration
- Outcomes were measured well over a full academic term after intervention start in both arms; since the stronger Year Duration criterion is met, Term Duration is automatically considered met.
- "Pre-test: this was conducted in April 2013 just before the start of the intervention phase... Post-test: undertaken between May and July 2014 after the end of the intervention phase." (p. 13)
- Relevant Quotes: 1) "Pre-test: this was conducted in April 2013 just before the start of the intervention phase." (p. 13) 2) "Post-test: undertaken between May and July 2014 after the end of the intervention phase." (p. 13) 3) "Pupils in one treatment group received a 30-week programme starting in the final term of nursery and continuing for the first two terms of reception year in primary school. Pupils in the second treatment group received a 20-week programme that ran during the first two terms of primary school." (p. 6) Detailed Analysis: For both the 30-week arm (starting April 2013, measured May-July 2014, over a year later) and the 20-week arm (starting September 2013, measured May-July 2014, roughly 8-9 months later), the interval from intervention start to outcome measurement clearly exceeds one academic term. As the Year Duration criterion (Y) is also met (see below), the ERCT standard states the weaker Term Duration criterion is automatically satisfied. Final sentence: Criterion T is met because the measured interval from intervention start to post-test substantially exceeds one academic term in both trial arms, and because Year Duration (Y) is separately met.
-
D Documented Control Group
- The control ("waitlist") group is extensively documented, including its business-as-usual status and detailed baseline demographic and language characteristics compared with treatment groups.
- "Children in the control group received no additional language support during the trial beyond that normally received in a business-as-usual scenario, but their language and word-level literacy skills were monitored in the same way as the two treatment groups." (p. 8)
- Relevant Quotes: 1) "Children in the control group received no additional language support during the trial beyond that normally received in a business-as-usual scenario, but their language and word-level literacy skills were monitored in the same way as the two treatment groups." (p. 8) 2) Table 3: "% female 49.2% (30-week) / 48.1% (20-week) / 48.8% (Control); Average age in months 47.4 / 47.4 / 47.4; Average language composite score -0.065 / -0.073 / -0.090." (p. 16) 3) Table 9 provides extensive baseline characteristics for the control group versus both treatment groups, including EAL status, known speech/language difficulties, FSM eligibility, SEN status, ethnicity, and pre-test scores on all language and literacy components. (pp. 25-27) 4) "A waitlist control group was chosen to address the ethical issues of identifying a group of struggling pupils and then not providing them any additional support." (p. 11) Detailed Analysis: The control group's composition, size, demographic characteristics, baseline test scores, and treatment status (business-as-usual, later offered an alternative literacy intervention) are all thoroughly documented and directly compared with the treatment arms across multiple detailed tables. Final sentence: Criterion D is met because the control group is comprehensively documented, including demographics, baseline scores, and confirmation of business-as-usual conditions.