Level 1 Criteria
-
C Class-level RCT
- Randomization occurred at the school level (30 schools), which is stronger than and therefore satisfies the class-level RCT requirement.
- "A total of 30 elementary schools (N = 2,870 students) were randomized to a treatment or control condition."
- Relevant Quotes: 1) "A total of 30 elementary schools (N = 2,870 students) were randomized to a treatment or control condition." (p. 1279, Abstract) 2) "For this longitudinal cluster (school) RCT study, 30 elementary schools in one urban school district located in the southeastern United States were recruited." (p. 1284) 3) "Schools were blocked by three levels of school size (i.e., small, medium, and large schools based on total student enrollment), two levels of prior reading proficiency ... and prior experience with MORE and then randomized to a treatment (full MORE spiral curriculum from Grades 1 to 3) or a control condition." (pp. 1284-1285) Detailed Analysis: The paper explicitly and repeatedly states that the unit of randomization was the entire school, not individual classes or students. Thirty elementary schools were recruited, blocked on size, prior reading proficiency, and prior MORE experience, and then randomized as whole units to treatment or control. School-level randomization is the stronger design required for Level 2 (criterion S), and per the ERCT standard, meeting the stronger school-level criterion automatically satisfies the weaker class-level criterion. There is no indication of within-school or within-class randomization anywhere in the methods. Criterion C is met because randomization occurred at the school level, which exceeds the class-level requirement.
-
E Exam-based Assessment
- The study used state-wide End-of-Grade standardized assessments for its primary domain-general reading and mathematics outcomes.
- "We measured students' domain-general reading comprehension and mathematics ability with posttests in Grades 3 and 4, using state-wide end-of-grade (EOG) standardized assessments, during the last week of the school year."
- Relevant Quotes: 1) "We measured students' domain-general reading comprehension and mathematics ability with posttests in Grades 3 and 4, using state-wide end-of-grade (EOG) standardized assessments, during the last week of the school year." (p. 1289) 2) "For each of the domains, we used the item response theory-scaled EOG test score to estimate the overall ability in respective domains. Internal consistencies are approximately .90 for Grades 3 and 4 across demographic subgroups (North Carolina Department of Public Instruction, 2020)." (p. 1289) 3) "We used the MAP (NWEA, 2019) reading and mathematics assessment scores (test-retest reliabilities = .79-.86) obtained in Grade 1 to control for students' baseline reading and mathematics abilities. The MAP assessment is a vertically scaled and computer-adaptive assessment..." (p. 1289) 4) By contrast, the domain-specific outcomes were researcher-developed: "We used a semantic association task ... to assess students' ability to identify semantically related words..." (p. 1288) and "We developed near-, mid-, and far-transfer domain-specific (science) reading comprehension passages and 29 multiple-choice questions..." (p. 1289) Detailed Analysis: The study's key domain-general reading and mathematics outcomes, which anchor the headline effects reported in the abstract (ES = .11-.16), were measured using statewide End-of-Grade (EOG) assessments, which are standard, government-administered, widely recognized standardized exams, not instruments created by the researchers. The Grade 1 baseline measures (MAP) are likewise a nationally normed, standardized, computer- adaptive assessment (NWEA). While the study also used custom, researcher-designed domain-specific vocabulary and reading-comprehension instruments (which are appropriate as supplementary transfer measures but would not by themselves satisfy this criterion), the presence of a genuine standardized exam-based assessment (EOG) as a primary outcome is sufficient to meet this criterion. Criterion E is met because the study used state-wide EOG standardized assessments as a primary outcome measure for reading and mathematics.
-
T Term Duration
- Outcomes were measured years after intervention start, which greatly exceeds the one-term minimum.
- "In the current study, there were nearly 36 months between the beginning and end of the full MORE spiral curriculum."
- Relevant Quotes: 1) "In the treatment condition (i.e., full spiral curriculum), students participated in content literacy lessons from Grades 1 to 3 during the school year and wide reading of thematically related informational texts in the summer following Grades 1 and 2." (p. 1279, Abstract) 2) "Treatment impacts were sustained at 14-month follow-up on Grade 4 reading comprehension (ES = .12) and mathematics achievement (ES = .16)." (p. 1279, Abstract) 3) "In the current study, there were nearly 36 months between the beginning and end of the full MORE spiral curriculum." (p. 1293) Detailed Analysis: The intervention began in Grade 1 (spring 2019) and outcomes were measured as late as spring of Grade 4 (roughly 2022), a 14-month follow-up after the Grade 3 posttest and nearly three years after the intervention began. This interval vastly exceeds the minimum one-term requirement, and it also satisfies the stronger Year Duration criterion (see Y below), which per the standard automatically satisfies this weaker Term Duration criterion. Criterion T is met because outcomes were tracked for multiple years after the intervention began, far exceeding one academic term.
-
D Documented Control Group
- The control group's size, demographics, and baseline characteristics are documented in detail (Tables 2-3) and compared to the treatment group.
- "In the control condition (i.e., partial spiral curriculum), students participated in lessons in only Grade 3."
- Relevant Quotes: 1) "In the control condition (i.e., partial spiral curriculum), students participated in lessons in only Grade 3." (p. 1279, Abstract) 2) "Table 2 compares the demographic and baseline achievement characteristics of the students in the RCT sample to students in the non-RCT sample. ... more students received individual education plan (IEP; p < .01) and attained lower reading and mathematics scores at Grade 1 baseline (ps < .05) compared to the non-RCT sample." (p. 1284) 3) "Table 3, Balance Checks for Analytic Sample of Students Remaining in the Long-Term Impact Analysis (School-Level Averages)" lists control-school n = 15 and detailed means/SDs for White, Black, Hispanic, Asian, Male, Limited English proficiency, IEP, SES bands, and baseline MAP reading/mathematics scores for control schools versus treatment schools. (p. 1287) 4) "Follow-up 1 (G3 Spring) ... CONTROL: Analyzed (n = 871), Lost to follow-up (missing demographic or outcome data) (n = 412)." (Figure 3, p. 1286) Detailed Analysis: The paper documents the control condition in substantial detail: what the control group received ("lessons in only Grade 3," i.e., the partial spiral), its sample size at each wave (1,283 randomized, 871-902 analyzed at follow-ups), and its demographic and baseline-achievement characteristics via Tables 2 and 3, including race/ ethnicity, sex, English proficiency, IEP status, SES, and baseline MAP reading/mathematics scores, directly compared to the treatment group. A formal baseline-equivalence analysis (Table 3) is also reported. This level of detail allows a reader to assess whether the control group was comparable to the treatment group at baseline. Criterion D is met because the control group's composition, size, and characteristics are thoroughly documented and compared to the treatment group.