Level 1 Criteria
-
C Class-level RCT
- Randomisation occurred at the individual-student level within schools, but the intervention was a pulled-out, small-group (3-5 students) supplemental reading program delivered by dedicated instructors outside the regular classroom, matching the personal-teaching/tutoring exception that allows student-level randomisation.
- "Randomization was carried out within schools without further stratification." (p. 455)
- Relevant Quotes: 1) "Of these students, 94 were randomized into intervention (n = 42) and comparison (n = 52) groups... Randomization was carried out within schools without further stratification." (p. 455, Spanish study) 2) "For the English study, 362 students were screened, and 108 (30%) met the criteria; of these students, 96 were randomized into intervention (n = 46) and comparison (n = 50) groups." (p. 455) 3) "Eight different intervention instructors provided the Spanish or English intervention (or both interventions) to treatment participants in small groups of 3 to 5 students for 50 minutes a day, 5 days a week, from October through May... The intervention was provided in addition to the students' core reading lessons and did not conflict with their daily scheduled reading time." (pp. 461, 464) Detailed Analysis: The quotes make clear that randomisation was performed at the level of the individual student within a school, not at the class or school level. Under a strict reading of Criterion C this would normally fail, since students in the same classroom could be split between intervention and comparison conditions, risking contamination (e.g., a teacher applying intervention techniques to comparison students, or comparison students overhearing intervention content). However, the ERCT standard's exception for personal/tutoring-style interventions applies here. The intervention was not delivered by the regular classroom teacher to a subset of students in situ; instead, dedicated external intervention instructors pulled small groups of 3 to 5 at-risk students out of the classroom for a separate 50-minute daily session "in addition to" and separate from core classroom instruction. This pull-out, small-group delivery model closely parallels one-to-one/small-group tutoring arrangements, where contamination risk from within-classroom spillover is much lower than for an in-classroom curricular change, because the comparison students do not experience the intervention instructor, materials, or session at all - they simply continue with business-as-usual classroom reading instruction. Because the intervention functionally resembles a personal/small-group tutoring program delivered outside the regular classroom rather than an in-class pedagogical change, the exception for tutoring-like interventions applies, and student-level randomisation is acceptable. Final: Criterion C is met via the tutoring/personal- teaching exception, despite randomisation being carried out at the individual-student level within schools.
-
E Exam-based Assessment
- Primary outcome measures were well-established, widely used standardized instruments (WLPB-R, CTOPP/TOPP-S, DIBELS, TOWRE), even though one supplementary experimental spelling measure was also used.
- "The Woodcock Language Proficiency Battery-Revised (WLPB-R; Woodcock, 1991...) is a well-standardized instrument with a normative sample of 6,359..." (p. 457)
- Relevant Quotes: 1) "The Comprehensive Test of Phonological Processing (CTOPP; Wagner, Torgesen, & Rashotte, 1999) is a well-validated measure of wide-ranging phonological processing abilities... The Test of Phonological Processing-Spanish TOPP-S was developed to align with the English CTOPP." (p. 456) 2) "The Woodcock Language Proficiency Battery- Revised (WLPB-R; Woodcock, 1991; Woodcock & Munoz- Sandoval, 1995) is a well-standardized instrument with a normative sample of 6,359... Internal consistency values for the subtests used... have ranged from .77 to .95." (p. 457) 3) "The Dynamic Indicators of Basic Early Literacy Skills (DIBELS; Good & Kaminski, 2002)... were administered to assess oral reading fluency." (p. 458) 4) "The word reading efficiency subtest of the Test of Word Reading Efficiency (Torgesen, Wagner, & Rashotte, 1999) measures decontextualized reading fluency." (p. 458) 5) "An experimental spelling measure developed for this study consisted of 25 words of one or two syllables..." (p. 458) Detailed Analysis: The core outcome battery used to assess decoding, comprehension, oral language, phonological awareness, and fluency consists of nationally normed, widely used standardized instruments: the WLPB-R (a large, normed battery with over 6,000 normative cases), the CTOPP/TOPP-S, DIBELS/IDEL, and the TOWRE. These are not custom-built for this study; they are established assessments used across many studies of reading development, with published psychometric properties reported in the text. Only the spelling measure was purpose-built for this study, and it functioned as one supplementary outcome among many primarily standardized measures of reading and language. Because the study's central reading and language outcomes rely on widely recognised standardized tests rather than tests custom-designed to favor the intervention, Criterion E is satisfied. Final: Criterion E is met because standardized, widely used assessments (WLPB-R, CTOPP/TOPP-S, DIBELS, TOWRE) formed the core outcome battery.
-
T Term Duration
- Because the stronger Year Duration criterion (Y) is met, the weaker Term Duration criterion is automatically considered met as well.
- "...administered in both English and Spanish to all participants... prior to the intervention (October) and following its completion (May)." (p. 456)
- Relevant Quotes: 1) "Eight different intervention instructors provided the Spanish or English intervention... from October through May." (p. 461) 2) "The same comprehensive battery of language/ literacy measures was administered in both English and Spanish to all participants by well-trained members of the research team prior to the intervention (October) and following its completion (May)." (p. 456) Detailed Analysis: Pretest occurred in October and posttest in May, an interval of roughly seven to eight months, which far exceeds the minimum one-term (3-4 month) requirement. As documented under Criterion Y below, this interval also satisfies the stronger year-long duration requirement, and per the ERCT standard, meeting Y automatically satisfies T. Final: Criterion T is met, both directly (the October-to-May interval exceeds one term) and by inheritance from the met Y criterion.
-
D Documented Control Group
- The comparison group's demographics, baseline scores, and the amount/type of any supplemental instruction they received outside the study are documented in detail across four results tables and the narrative text.
- "27 of 46 (59%) comparison group students received one or more types of reading instruction over their core instruction. The amount of time additional instruction was provided to these 27 students ranged from 7.9 to 93.2 hours..." (p. 465)
- Relevant Quotes: 1) "All of the students were Hispanic, and 44% were female (42% in the Spanish study, 46% in the English study). The mean age at pretest among the 183 students who began the study was 6.6 years (SD = 0.4)..." (p. 455) 2) Tables 1-4 report, for the comparison ("C") group in each study/language: sample size (n), pretest mean and SD, posttest mean and SD, and gain scores for every outcome measure (pp. 462-463, 472-473, 476-477, 480-481). 3) "For the students in the Spanish study with additional instruction data, 27 of 46 (59%) comparison group students received one or more types of reading instruction over their core instruction. The amount of time additional instruction was provided to these 27 students ranged from 7.9 to 93.2 hours (M = 41.2 hours, SD = 27.2, Mdn = 31.2)..." (p. 465) 4) "Among the students in the English study with additional instruction data, 28 of 40 (70%) of those in the comparison group received one or more types of supplemental reading instruction over their core instruction..." (p. 465) 5) "As expected from the random assignment to groups, there were no significant group mean differences in performance on either of the intervention screening skills... in either language in either study (all ps > .05)." (p. 468) Detailed Analysis: The paper provides extensive documentation of the comparison ("business as usual") group: overall demographic composition, pretest equivalence checks confirming comparability with the intervention group, and detailed per-measure pretest/posttest statistics in Tables 1-4. Crucially, the authors also directly measured and reported how much supplemental reading instruction comparison students received from other sources outside the study, including exact percentages and hour ranges. This level of transparency about the control condition's composition and experience goes well beyond a bare mention of a control group. Final: Criterion D is met given the thorough demographic, baseline, and supplemental-instruction documentation of the comparison group.