Level 1 Criteria
-
C Class-level RCT
- Randomisation was via an individual lottery rather than by class or school, but winners and losers attend different schools, so classroom-level contamination (the concern behind this criterion) does not arise.
- "Students who do not win an immersion slot are assigned to the regular instructional program in their default neighborhood schools." (p. 290S)
- Relevant Quotes: 1) "Students receive admission to immersion programs in Portland through a lottery process administered by the school district." (p. 290S) 2) "Students receive a random lottery number for each preference choice, but in practice, all immersion slots are filled in the first lottery round." (p. 290S) 3) "Students who do not win an immersion slot are assigned to the regular instructional program in their default neighborhood schools." (p. 290S) 4) "Of the 3,457 students who applied to Portland immersion lotteries during the study years, 1,946 (56.3%) were truly randomized within a binding lottery category and subcategory." (p. 291S) Detailed Analysis: The formal unit of randomisation here is the individual applicant: each student receives a random lottery number within a school/preference/subcategory stratum, and winners and losers are determined student by student, not by randomly assigning whole classes or schools to condition. Taken purely mechanically, this is student-level randomisation, which the standard's procedure treats as failing unless a valid exception (the personal-tutoring exception) applies. That named exception does not apply here. However, the substantive problem criterion C is meant to guard against is contamination: control students being taught alongside, or influenced by, treatment students in the same classroom or school. In this design, losers of the lottery do not remain in the same classroom or school as winners; they are routed to "the regular instructional program in their default neighborhood schools" - ordinarily a different building and different peer/teacher environment entirely from the immersion school the winner attends. There is essentially no possibility of a control student receiving spillover immersion instruction from a treated classmate, because in the large majority of cases they do not share a school. This is unlike within-class or within-school randomisation, where contamination is the central risk. Because the practical contamination concern that motivates criterion C is not present (treatment and control students are, in effect, enrolled in different institutions), and because the randomisation process and analytic strata are clearly and transparently described (CONSORT diagram, lottery categories/subcategories, binding-lottery restriction), I judge that the substance of criterion C is satisfied even though the paper does not describe a cluster (class/school) random assignment mechanism. This is a genuinely borderline call given the strict letter of the standard, and a stricter reading (requiring literal class/school-level randomisation with no applicable exception) would mark this as not met. Final: Criterion C is met on balance because, despite individual-level lottery randomisation, winners and losers are routed to different schools, eliminating the within-classroom contamination risk the criterion is designed to prevent.
-
E Exam-based Assessment
- Outcomes were measured using OAKS, Oregon's official state-mandated standardized test, not a custom instrument.
- "Student achievement in reading, mathematics, and science is measured by performance on the state-mandated accountability test, the Oregon Assessment of Knowledge and Skills (OAKS)." (p. 292S)
- Relevant Quotes: 1) "Student achievement in reading, mathematics, and science is measured by performance on the state-mandated accountability test, the Oregon Assessment of Knowledge and Skills (OAKS)." (p. 292S) 2) "Mathematics and reading tests are administered annually in Grades 3 through 8 and once in high school; science is tested in Grades 5 and 8. The tests are administered solely in English." (p. 292S) 3) "We standardize scores to have mean zero and standard deviation one within grade level, subject, and school year." (p. 292S) Detailed Analysis: The outcome measures are drawn entirely from Oregon's state-mandated accountability testing system (OAKS), a standardized, statewide test used for official accountability purposes rather than an instrument created specifically for this study. This is precisely the kind of widely recognised, objective, comparable assessment the E criterion requires. Final: Criterion E is met because achievement is measured using OAKS, Oregon's official state-mandated standardized accountability test.
-
T Term Duration
- Outcomes were measured multiple years after the intervention began (kindergarten through Grade 8/9), well beyond the one-term minimum.
- "Outcome data are measured through the 2013-2014 academic year, so the oldest cohort can be observed through ninth grade and the youngest through third grade." (p. 291S)
- Relevant Quotes: 1) "The study focuses on the seven cohorts of students who applied to a pre-k or kindergarten immersion slot in Portland for the fall terms of 2004 through 2010." (p. 291S) 2) "Outcome data are measured through the 2013-2014 academic year, so the oldest cohort can be observed through ninth grade and the youngest through third grade." (p. 291S) 3) "In Grade 5, lottery winners outperform their counterparts by 13% of a standard deviation, and they do so by 22% of a standard deviation in Grade 8." (p. 297S) Detailed Analysis: The intervention begins at kindergarten entry, and primary outcomes are reported across multiple subsequent grades (from Grade 3 through Grade 8/9), i.e., years after intervention start. This interval vastly exceeds the minimum of one academic term between intervention start and outcome measurement. Final: Criterion T is met because outcomes are tracked for multiple years after kindergarten entry, far exceeding a single academic term.
-
D Documented Control Group
- Table 2 gives detailed demographic, baseline, and attrition statistics for the control group, with balance tests against the treatment group.
- "The left side of Table 2 presents descriptive statistics for the randomized (binding) analytic sample, and the right side presents comparable information for the full sample." (p. 291S)
- Relevant Quotes: 1) "The left side of Table 2 presents descriptive statistics for the randomized (binding) analytic sample... Table 2 also presents the difference between groups for each variable and p values for t tests of the differences." (p. 291S-292S) 2) Table 2 reports, for "Won Slot" versus "Not Placed" (control), proportions female, Asian, Black, Hispanic, White, subsidized-meal-eligible, special needs in kindergarten, gifted in kindergarten, English learner in kindergarten, and first-language-not-English, along with unadjusted differences and strata-adjusted p values. (Table 2, p. 292S) 3) "This combination of overall and differential attrition rates lies very near the conservative threshold for meeting What Works Clearinghouse (2014) evidence standards... our models adjust for observed baseline characteristics as well as lottery strata fixed effects." (p. 291S) Detailed Analysis: The paper provides an extensive, table-based description of the control ("Not Placed") group's demographic and baseline characteristics, directly compared against the treatment ("Won Slot") group, with formal balance tests. Attrition for the control group is also explicitly quantified and discussed. This constitutes clear, quantitative documentation of the control group sufficient to assess comparability. Final: Criterion D is met because Table 2 documents the control group's demographics, baseline characteristics, and attrition in detail, alongside formal balance tests against the treatment group.