Level 1 Criteria
-
C Class-level RCT
- Randomisation was at the individual-student level within classrooms, not at the class or school level, and the small-group (non-tutoring) nature of the intervention does not qualify for the personal-tutoring exception.
- "We randomly assigned individual students classified by their schools as ELs (n = 84) to either homogeneous groups (four ELs) or heterogeneous groups (two ELs and two non-ELs)."
- Relevant Quotes: 1) "In this preregistered within-teacher randomized controlled trial (n = 84), we tested the effects of grouping English learners (ELs) in homogeneous groups (all ELs) versus heterogeneous groups (ELs and non-ELs) on language, reading comprehension, and argumentative writing." (Abstract, p. 909) 2) "Students (n = 84) were individually randomly assigned to within-teacher groups, with teachers providing instruction to both the homogeneous and heterogeneous groups." (p. 912) 3) "We randomly assigned individual students classified by their schools as ELs (n = 84) to either homogeneous groups (four ELs) or heterogeneous groups (two ELs and two non-ELs)." (p. 916) 4) "After selection, participants were individually randomly assigned within self-contained fourth and fifth grade classrooms to homogeneous groups (n groups = 14; n ELs = 56) or heterogeneous groups (n groups = 14; n ELs = 28), as shown in Figure 1." (p. 922) Detailed Analysis: The unit of randomisation in this study is unambiguously the individual student. Students within the same self-contained classroom, taught by the same teacher, were individually randomly assigned to a four-EL homogeneous group or a two-EL/two-non-EL heterogeneous group. This is explicitly a within-classroom, within-teacher design; no whole classes or schools were randomised to conditions. The ERCT exception for Criterion C applies only to interventions "designed for personal teaching like tutoring," where student-level randomisation is acceptable even without class- or school-level assignment. Here, the intervention is delivered to small groups of four students at a time (either four ELs, or two ELs plus two non-ELs), not one-to-one tutoring, so this exception does not clearly apply. Because both conditions being compared are active instructional groups delivered by the same teacher in the same classroom, cross-contamination between "treatment" and "control" (in the classic sense) is less of a concern than in a typical treatment-vs-no- treatment RCT. Nonetheless, per the letter of the ERCT standard, the unit of randomisation must be the class or school (absent the tutoring exception), and here it was the individual student. Final: Criterion C is not met because randomisation occurred at the individual-student level within single classrooms, and the tutoring exception does not apply to this small-group (not one-to-one) intervention.
-
E Exam-based Assessment
- The study used ReadBasix, an externally developed and psychometrically validated standardized assessment battery, for its reading comprehension and language skills outcomes.
- "English reading comprehension was measured using the ReadBasix assessment program... developed by researchers at the Educational Testing Service (Sabatini et al., 2019)."
- Relevant Quotes: 1) "English reading comprehension was measured using the ReadBasix assessment program (previously known as the Reading Inventory and Scholastic Evaluation or RISE) developed by researchers at the Educational Testing Service (Sabatini et al., 2019)." (p. 924) 2) "ReadBasix is a web-administered reading skills componential battery designed with vertical scales for children in grades 3-12." (p. 924) 3) "The technical manual for ReadBasix reports an acceptable (IRT) marginal reliability of .753 for grade 4 and .674 for grade 5 of the norming sample (Sabatini et al., 2019)." (p. 924) 4) "The English vocabulary subtest of ReadBasix presents students with a target word and asks them to choose the most closely related word from three options... The technical manual for ReadBasix reports an IRT marginal reliability of .832 to .867 for grades 4 and 5 of the norming sample." (p. 924) Detailed Analysis: The study's key distal outcome, reading comprehension, was measured with ReadBasix, a standardized, norm-referenced battery developed by researchers at the Educational Testing Service and validated with published technical manuals and reliability statistics across grades 3-12. ReadBasix was used at both pretest and posttest for reading comprehension, and it is a widely used, externally developed standardized assessment, not a custom test built solely for this study. Other outcomes (CLAVES language skills factor, argumentative writing rubric, CALS-I) are more proximal/researcher- or field-developed instruments, but the criterion requires that a standard exam-based assessment be used somewhere in the outcome battery, which is satisfied by the ReadBasix reading comprehension and vocabulary/morphology subtests. Final: Criterion E is met because the study used ReadBasix, an externally developed, standardized, psychometrically validated assessment battery, for its reading comprehension, vocabulary, and morphology measures.
-
T Term Duration
- Outcomes were measured at the end of an approximately 12-week (about 3-month) intervention period, which approximates the standard's definition of one academic term.
- "All participating ELs received ~12 weeks (18 total hours) of CLAVES..."
- Relevant Quotes: 1) "All participating ELs received ~12 weeks (18 total hours) of CLAVES, with half randomly assigned to homogeneous groups (consisting of four ELs) and half assigned to heterogeneous groups..." (p. 912) 2) "Teachers were asked to implement CLAVES for 30 minutes a day three times a week for 12 weeks, or a total of 18 hours of instruction." (p. 922) 3) "They reported teaching a median of 31 lessons with some variability (SD = 5.20 lessons; minimum = 20; maximum = 35)." (p. 923) Detailed Analysis: The intervention was designed to run for approximately 12 weeks (about 2.75-3 months), with posttest measures administered at the conclusion of this period. The ERCT standard defines a term as "approximately 3-4 months." Twelve weeks sits at the low end of this range and is consistent with a school trimester/quarter-length term in many US districts. There is no explicit statement of a longer gap between intervention start and outcome measurement beyond the instructional period itself, but the approximately 12-week span of active tracking from intervention onset to posttest is reasonably close to the standard's "approximately 3-4 months" definition of a term. Final: Criterion T is met, on balance, because the approximately 12-week (about 3-month) span from intervention start to posttest measurement approximates the standard's definition of one academic term.
-
D Documented Control Group
- The paper provides detailed, tabulated demographic and baseline performance data for both comparison groups, confirming they were comparable at pretest.
- "As shown in Table 2, there were no statistically significant differences between the heterogeneous and homogeneous groups at pretest on any variables..."
- Relevant Quotes: 1) "Table 2. Demographics and pretest descriptive statistics, overall and by group (n = 84)" reporting gender, home-language background, reading comprehension, decoding, vocabulary, morphology, English language skills composite, and argumentative writing separately for "Homogeneous mean (SD)(n = 56)" and "Heterogeneous mean (SD) (n = 28)." (p. 921, Table 2) 2) "As shown in Table 2, there were no statistically significant differences between the heterogeneous and homogeneous groups at pretest on any variables, including gender (p = .609), home-language background (all p values > .10), reading comprehension (p = .274), word reading (p = .694), or English language skills (p = .600)..." (p. 928) 3) "Following selection, student participants included 84 fourth and fifth graders classified as ELs (41 girls; see Table 2). On family surveys, the most common home language was Mandarin or Cantonese (61%), followed by Spanish (21%)..." (p. 919) Detailed Analysis: Although this study does not have a classic untreated "control" arm (both homogeneous and heterogeneous groups receive the same CLAVES curriculum), the comparison group in each contrast (e.g., the homogeneous condition when heterogeneous is the focal comparison, or vice versa) is extensively documented. Table 2 provides detailed demographic and pretest performance data (gender, home-language background, reading comprehension, word recognition, vocabulary, morphology, English language skills composite, and argumentative writing) broken out separately by group, along with significance tests confirming baseline equivalence. This level of detail satisfies the intent of Criterion D, which requires that the comparison group's composition, baseline characteristics, and conditions be clearly documented so that comparability can be assessed. Final: Criterion D is met because the paper provides detailed demographic and baseline performance data for both comparison groups, confirming their comparability at the outset of the study.