Level 1 Criteria
-
C Class-level RCT
- Randomization occurred at the small-group (2-3 student) level within classrooms rather than at the class level, but the intervention is a pull-out, small-group tutoring program delivered outside the regular classroom, which falls under the ERCT exception for personal/tutoring-style teaching.
- "For both cohorts, students were randomly assigned, within classroom, to small groups of two to three students... Small groups were randomly assigned to conditions within classroom..." (p. 496-497)
- Relevant Quotes: 1) "A two-cohort cluster-randomized trial was conducted to estimate effects of small-group supplemental vocabulary instruction for at-risk kindergarten English learners (ELs)." (Abstract, p. 490) 2) "For both cohorts, students were randomly assigned, within classroom, to small groups of two to three students, with group sizes based on eligible students within each classroom. Small groups were randomly assigned to conditions within classroom, with the added constraint that classrooms with an uneven number of small groups would have an 'extra' group randomly assigned to one or the other of the conditions." (p. 496-497) 3) "Children in both conditions received small-group instruction outside their classroom for 30 min per day, four days per week, for an average of 20 weeks." (p. 497) 4) "All tutors were recruited from their school communities and hired as district employees paid by the schools with funds provided by the research grant." (p. 499-500) Detailed Analysis: The unit of randomization here is the small group (dyads/triads of 2-3 students) within a classroom, not the class or school as a whole. Under a strict reading of criterion C this would fail, since contamination between conditions could occur among children in the same classroom. However, the ERCT standard includes an explicit exception: if an intervention is designed for personal teaching such as tutoring, student-level (or small-group-level) randomization is acceptable. The intervention studied here is explicitly a pull-out, small-group tutoring program: trained "tutors" (paraeducators) deliver individualized, small-group (2-3 student) lessons outside the regular classroom for 30 minutes per day. This is materially closer to one-to-one/small-group tutoring than to whole-class instruction, and the pull-out delivery outside the classroom substantially reduces opportunity for cross-condition contamination relative to within-classroom instruction. Given this is a personal/tutoring-style intervention, the exception applies. Criterion C is met because, although randomization occurred below the class level, the intervention is a pull-out small-group tutoring program that qualifies for the ERCT tutoring exception.
-
E Exam-based Assessment
- Core outcomes (receptive vocabulary, decoding, spelling) were measured with widely used, standardized norm-referenced tests, satisfying the exam-based assessment requirement even though one supplementary curriculum-based measure was researcher-developed.
- "Receptive vocabulary was measured using the Peabody Picture Vocabulary Test-IIIA (PPVT-IIIA; Dunn & Dunn, 2006)." (p. 501-502)
- Relevant Quotes: 1) "Receptive vocabulary was measured using the Peabody Picture Vocabulary Test-IIIA (PPVT-IIIA; Dunn & Dunn, 2006)... Test manual split-half reliability is 0.76 for five-year-olds and .80 for six-year-olds." (p. 501-502) 2) "Decoding was measured using the Woodcock Reading Mastery Test-Revised/Norm Referenced (WRMT-R/NU; Woodcock, 1987/1998) Word Attack subtest... Test manual split-half reliability for kindergartners is 0.94." (p. 502) 3) "Spelling was measured at pretest and posttest using the Wide Range Achievement Test-4 (WRAT-4; Wilkinson & Robertson, 2006) Spelling subtest... Test manual internal consistency reliability coefficient for kindergarteners is .94." (p. 502) 4) "Reading vocabulary was assessed using an experimenter-developed 25-item curriculum-based measure (CBM) of target word reading vocabulary." (p. 502) Detailed Analysis: Three of the four outcome measures used in this study are widely recognized, standardized, norm-referenced instruments with published reliability data: the PPVT-IIIA (receptive vocabulary), the WRMT-R/NU Word Attack subtest (decoding), and the WRAT-4 Spelling subtest (spelling). Only the reading-vocabulary measure (a curriculum-based measure of taught root words) was experimenter-developed and specific to this study's word corpus. Because the study relies substantially on multiple standard, validated psychometric instruments for its principal outcome domains (not solely a custom instrument), the exam-based assessment criterion is satisfied. Criterion E is met because the study used multiple widely recognized standardized tests (PPVT-IIIA, WRMT-R/NU, WRAT-4) alongside one supplementary custom measure.
-
T Term Duration
- Outcomes were measured about seven months after the intervention began (fall pretest to spring posttest), exceeding the minimum one-term requirement, and this is reinforced by the stronger Year Duration criterion also being met.
- "students were pretested in the fall prior to treatment and posttested in the spring just after treatment, approximately seven months apart, during their kindergarten year" (p. 501)
- Relevant Quotes: 1) "Children in both conditions received small-group instruction outside their classroom for 30 min per day, four days per week, for an average of 20 weeks." (p. 497) 2) "For each cohort, students were pretested in the fall prior to treatment and posttested in the spring just after treatment, approximately seven months apart, during their kindergarten year, and were followed up approximately seven months after posttest, in the winter of the following year." (p. 501) Detailed Analysis: The immediate outcome measurement (posttest) occurred approximately seven months after the pretest/start of the intervention period, which exceeds the minimum one-academic-term (3-4 month) requirement. In addition, the study also includes a longer-term follow-up roughly 14 months after the intervention began, which independently satisfies the stronger Year Duration criterion (see criterion Y), and per the ERCT standard, meeting Y automatically satisfies T. Criterion T is met because the measurement interval from intervention start to posttest (about seven months) exceeds one academic term, and is further reinforced by the year-long follow-up tracking.
-
D Documented Control Group
- Although this is a two-active-treatment comparison rather than a treatment-vs-no-treatment design, the comparison (IBR) condition is extensively documented, including demographics, sample sizes, and detailed procedures.
- "Table 1. Disaggregated student demographic characteristics" showing Connections and IBR groups separately by gender, SPED/ELL services, language proficiency level, and home language. (p. 498)
- Relevant Quotes: 1) "Table 1. Disaggregated student demographic characteristics" comparing Connections (n = 163) and IBR (n = 161) groups on gender, SPED services, ELL services, state language test level, and home language. (p. 498) 2) "Students assigned to the IBR condition received instruction in the same target vocabulary provided in the Connections condition. Instruction was provided in the context of reading aloud a storybook..." (p. 500) 3) "...163 students in 75 Connections small groups and 161 students in 72 IBR small groups." (p. 497), with IBR procedures described in full in Table A4, Appendix A (p. 526) Detailed Analysis: This study does not include a business-as-usual, no-treatment control group; instead it compares two active supplemental treatments (Connections vs. IBR), a limitation the authors themselves acknowledge ("difficult to interpret these gains without a no-treatment control comparison"). Treating IBR as the comparison ("control") arm relative to Connections, the paper provides thorough documentation: sample sizes, demographic breakdowns by condition (Table 1), baseline pretest scores by condition (Table 2), and a full description of the IBR procedures (Table A4). This level of detail allows readers to verify comparability of the comparison group at baseline and to understand exactly what instruction it received. Criterion D is met because the comparison group's characteristics, size, and procedures are documented in detail, even though the design lacks a no-treatment control.