Level 1 Criteria
-
C Class-level RCT
- Randomisation was conducted at the class level (30 classes), which satisfies the class-level RCT requirement.
- "The randomization took part at the class level." (p. 3)
- Relevant Quotes: 1) "722 students (48.7% girls) from 30 sixth-grade classes from 10 Realschulen located in middle-class neighbourhoods of small and midsize towns in Baden-Wuerttemberg, Germany were randomly assigned to receive either German or German/English instruction." (p. 3) 2) "The randomization took part at the class level." (p. 3) 3) "There were 362 bilingually educated students ... and 360 monolingually educated students ..." (p. 3) Detailed Analysis: The paper explicitly states that whole classes (not individual students within a shared classroom) were the unit of randomisation: 30 intact sixth-grade classes, drawn from 10 schools, were assigned wholly to either the monolingual or the bilingual condition. This is exactly the unit of randomisation the ERCT Standard's C criterion requires, and it rules out the within-class contamination problem the criterion is designed to prevent (students in the same class receiving different treatments). There is no indication that individual students within the same class were split across conditions. Final summary: Criterion C is met because the study explicitly reports that randomisation occurred at the whole-class level across 30 classes.
-
E Exam-based Assessment
- The primary outcome was a custom-assembled "Floating and Sinking" test built largely from prior research items plus new author-written items, not a widely recognised standardised exam.
- "Altogether the test consisted of 36 items. ... four items were developed by us." (p. 4)
- Relevant Quotes: 1) "To measure students' learning gains on the topic 'Floating and Sinking', a test based on students' typical preconcepts ... was administered. Altogether the test consisted of 36 items." (p. 4) 2) "Twenty-eight items were published in Blumberg (2008); Hardy et al. (2006); Kleickmann (2008); Möller (2005); Möller et al. (2006); Stern, Möller, Hardy, and Jonen (2002). Three more items were translated from former Trends in International Mathematics and Science Studies (TIMSS ...), and four items were developed by us." (p. 4) 3) "For the IRT-based scaling of the test, a 2-parameter logistic model (Birnbaum model) was chosen ..." (p. 4) Detailed Analysis: The study's primary and only academic outcome measure is a bespoke, topic-specific instrument assembled by the research team from earlier studies' items (on the same "Floating and Sinking" topic), a handful of translated TIMSS items, and four items the authors wrote themselves, then custom-calibrated with an IRT model for this study. This is precisely the kind of researcher-constructed, intervention-aligned test the ERCT Standard's E criterion warns against: it is not a widely recognised, standardised exam (e.g., a national curriculum or state-wide test) administered as-is. Even the physics-preknowledge covariate (partly from TIMSS) is a secondary/control measure, not the primary outcome, and it too was adapted/supplemented by the authors. Final summary: Criterion E is not met because the outcome measure is a custom, author-assembled test rather than a standardised, widely recognised exam.
-
T Term Duration
- Outcomes were measured immediately after a five-lesson intervention and again only 6 weeks later, far short of one academic term.
- "The lessons were followed by a 60-min posttest and a 30-min follow-up test 6 weeks after instruction ended." (p. 3)
- Relevant Quotes: 1) "Then, the intervention was implemented through a teaching unit on the topic 'Floating and Sinking' consisting of five lessons lasting 90 min each." (p. 3) 2) "The lessons were followed by a 60-min posttest and a 30-min follow-up test 6 weeks after instruction ended." (p. 3) Detailed Analysis: The intervention itself was very brief (five 90-minute lessons), and the final, longest-delayed measurement point (the follow-up test) occurred only 6 weeks after the intervention ended. Six weeks is well short of the 3-4 month academic term the ERCT Standard's T criterion requires between intervention start and outcome measurement. No later assessment point is reported anywhere in the paper. Final summary: Criterion T is not met because the longest follow-up interval reported is 6 weeks, far below the required one-term minimum.
-
D Documented Control Group
- The comparison (monolingual) group's demographics, baseline achievement, and instructional exposure are documented in detail in Table 1 and the Procedure section.
- "There were 362 bilingually educated students (52.7% girls, ...) and 360 monolingually educated students (44.8% girls, ...)." (p. 3)
- Relevant Quotes: 1) "There were 362 bilingually educated students (52.7% girls, Mage = 11.5 years, SD = 0.61, 35.7% immigration background) and 360 monolingually educated students (44.8% girls, Mage = 11.5 years, SD = 0.57, 33.2% immigration background)." (p. 3) 2) "Table 1 provides an overview of the descriptive statistics. ... the subsamples differed significantly only on gender." (p. 5, Results 6.1) - Table 1 reports, separately for the monolingual and bilingual groups: gender, immigration background, parental education, general cognitive ability, English ability, science grade, physics preknowledge, science self-concept, interest in science, and floating-and-sinking pretest scores. 3) "The teaching material and the instructional time were held constant between groups." (p. 3) Detailed Analysis: Although this study compares two active instructional conditions rather than an intervention-vs-nothing design, the monolingually taught group functions as the comparison/control condition against which the bilingual condition is evaluated. Its size, gender split, immigration background, socioeconomic proxy (parental education), baseline cognitive ability, baseline physics knowledge, baseline English ability, motivation measures, and pretest performance are all explicitly reported and statistically compared against the treatment group in Table 1. This level of detail allows a reader to judge comparability at baseline. Final summary: Criterion D is met because the comparison group's characteristics and baseline data are thoroughly documented in Table 1 and the text.