Level 1 Criteria
-
C Class-level RCT
- The study used a two-stage clustered design in which entire remedial-math class sections (not individual students within a class) were randomly assigned to the tutoring or control condition, satisfying class-level RCT.
- "we used a two-stage clustered randomization procedure with the high-impact tutoring treatment being administered at the classroom level within schools" (p. 12)
- Relevant Quotes: 1) "To estimate causal effects, we used a two-stage clustered randomization procedure with the high-impact tutoring treatment being administered at the classroom level within schools (see Figure 1)." (p. 12) 2) "In late spring, eligible eighth grade students were randomized to remedial mathematics class sections at each high school for the following academic year. These remedial class sections were then randomly selected to be either a high-impact tutoring treatment group or a remedial mathematics class control group taught solely by a classroom teacher." (p. 13) 3) "In year 3 of the study, we included an additional treatment condition within the high-impact tutoring course sections by randomly assigning students to either 2:1 or 3:1 student-tutor group sizes." (p. 13) Detailed Analysis: The randomisation unit for treatment/control assignment was the class section, not the individual student. Students were first randomised to remedial-math sections, and those whole sections were then randomised to the tutoring or control condition, which is exactly the class-level design the criterion requires and avoids within-class contamination. (The subsequent 2:1 vs. 3:1 randomisation in year 3 is a secondary sub-assignment within the tutoring arm and does not affect the primary treatment/control randomisation unit.) Because this is also a personal/small-group tutoring intervention, the standard's tutoring exception would in any case permit even student-level assignment, but the study exceeds that bar by randomising at the class level. Criterion C is met because whole class sections, not individual students within a section, were randomly assigned to treatment or control.
-
E Exam-based Assessment
- Outcomes were measured with the NWEA MAP mathematics assessment, a widely used, externally validated, standardized computer adaptive test, not a custom instrument.
- "Developed by NWEA, the MAP assessment is a computer adaptive test that is used to evaluate mathematics ability and growth throughout the school year. More than 9,700 schools use the MAP assessment in 145 countries" (p. 14)
- Relevant Quotes: 1) "The MAP assessment was our primary outcome of interest. Developed by NWEA, the MAP assessment is a computer adaptive test that is used to evaluate mathematics ability and growth throughout the school year." (p. 14) 2) "More than 9,700 schools use the MAP assessment in 145 countries, and NWEA has a 40-year history of developing such assessments." (p. 14) 3) "The reliability evidence (test-retest reliability and marginal reliability) for the MAP is strong. For the ninth grade mathematics assessment, test-retest reliability is .0.9, and marginal reliability is 0.96 (NWEA, 2019). External alignment studies offer further evidence to support the validity of the MAP assessment (Egan & Davidson, 2017)." (p. 14) 4) "These studies have shown that 97% of MAP items are aligned with Common Core State Standards (ninth grade mathematics: r = .72)." (p. 14) Detailed Analysis: The primary outcome measure is the NWEA MAP mathematics assessment, an externally developed and independently validated, widely used standardized test administered in thousands of schools internationally, with documented reliability and alignment to curriculum standards. It was not designed by the study authors for this specific intervention. Criterion E is met because the study relied on a widely recognised, externally validated standardized exam (NWEA MAP) rather than a custom-built assessment.
-
T Term Duration
- Outcomes were measured at the end of a full academic year, far exceeding the minimum one-term interval required.
- "This assessment was administered again at the end of the academic year, enabling us to determine the effects of high-impact tutoring on students' mathematics test scores during the school year." (p. 14)
- Relevant Quotes: 1) "At the start of the academic year, all ninth grade students took the NWEA MAP assessment, providing us with baseline achievement data in mathematics. This assessment was administered again at the end of the academic year, enabling us to determine the effects of high-impact tutoring on students' mathematics test scores during the school year." (p. 14) 2) "In the pooled sample, students (n = 525) in the treatment group participated in high-impact tutoring (i.e., groups of 2:1 or 3:1) three class periods per week for an entire academic year." (p. 3) Detailed Analysis: Baseline achievement was measured at the start of ninth grade and the primary outcome at the end of the same academic year, an interval far longer than the one-term minimum required by this criterion. Because this interval also satisfies the stronger Year Duration (Y) criterion, the weaker Term Duration criterion is automatically met as well. Criterion T is met because outcomes were measured a full academic year after the intervention began, well beyond one term.
-
D Documented Control Group
- The control group is described in detail, including its size, demographic composition, baseline scores, class size, and the specific instruction it received.
- "we compared the outcomes of ninth grade students who were randomly assigned to either a remedial mathematics class providing high-impact tutoring (treatment) or to a remedial mathematics class delivered by a classroom teacher only (control)." (p. 7)
- Relevant Quotes: 1) "In this study, we compared the outcomes of ninth grade students who were randomly assigned to either a remedial mathematics class providing high-impact tutoring (treatment) or to a remedial mathematics class delivered by a classroom teacher only (control)." (p. 7) 2) "Class sections for both study conditions were capped at 22 students. In the control group, the average class size was 19 students over the 3-year period of analysis." (p. 7) 3) Table 3, "Summary Statistics," reports means, standard deviations, and ranges of baseline and end-of-year math RIT scores, GPA, gender, FRL status, race/ethnicity, ELL status, and special education status separately for the control group sample (n = 438). (p. 15) 4) Table 4, "Comparison of the Baseline Characteristics of the Treatment and Control Groups," reports baseline math scores and demographic proportions for the control group in the pooled sample and each of the 3 years. (p. 16) Detailed Analysis: The paper documents the control condition thoroughly: what instruction control students received (a remedial mathematics class led solely by a classroom teacher, without a tutor), its size (n = 438; average class size 19), and detailed baseline demographic and achievement data broken out by year and pooled sample in Tables 3 and 4. This level of detail allows readers to assess comparability of the control group to the treatment group. Criterion D is met because the control group's composition, size, baseline characteristics, and the instruction it received are clearly and quantitatively documented.