Level 1 Criteria
-
C Class-level RCT
- Randomisation was carried out on stratified pairs of students within each classroom, with all three conditions present in the same class, so the unit of randomisation was below the class level and contamination was not prevented.
- "Pairs of students were stratified according to Pretest Part I scores on algebra prior understanding and academic language proficiency and then randomly assigned to the three treatment conditions, ensuring that the three conditions were represented in each classroom..."
- Relevant Quotes: 1) "Pairs of students were stratified according to Pretest Part I scores on algebra prior understanding and academic language proficiency and then randomly assigned to the three treatment conditions, ensuring that the three conditions were represented in each classroom and the groups were comparable in language and algebra during the 85-minute experiment session (including posttest)." (p. 16) 2) "The intervention sample consisted of N = 239 students in Grades 9-11 (ages 14 to 24 years) from 16 classes from four schools; all in their second chance units for understanding algebra." (p. 16) 3) "In the knowledge organization phase, students worked in self-chosen pairs in three treatment conditions with different worked examples and explicit instructions..." (p. 10) 4) "In a 45-minute preparation session, students completed questionnaires on background variables and Pretest Part I, and selected partners for pair work in the experiment session, as the selection of volunteer partners can promote interactive engagement (Chi et al., 2017)." (p. 16) 5) "The study was conducted in their regular mathematics classes with informed consent from students and guardians to use the data for research purposes." (p. 16) Detailed Analysis: Criterion C requires randomisation of entire classes (or schools), so that treatment and control students are not mixed within the same classroom, which would allow contamination between conditions. The paper is explicit that the randomised unit was the student pair, not the class: pairs were stratified on prior algebra understanding and academic language proficiency and then randomly assigned, and the design deliberately ensured that "the three conditions were represented in each classroom". This means students in the control condition and students in both video conditions worked simultaneously in the same room, during the same 85-minute lesson, which is exactly the arrangement the class-level requirement is designed to avoid. Although 16 classes from four schools participated, classes and schools were not the units of allocation; they merely supplied the sample. The tutoring exception does not apply. The intervention is not one-to-one personal tutoring: it is a self-learning environment delivered to self-chosen student pairs working in their regular mathematics classes, with no dedicated tutor per student. The pair setting was chosen to promote peer interactive engagement, not to deliver individualised tutoring, so the standard class-level requirement applies in full. Criterion C is not met because randomisation was performed on student pairs within classrooms rather than on whole classes or schools, and no tutoring exception applies.
-
E Exam-based Assessment
- The conceptual understanding outcome was measured with a ten-item pretest and posttest assembled and adapted by the researchers for this study, not with a recognised standardised exam.
- "The conceptual understanding pretest included ten unit-weighted items which were adopted and adapted from existing tests on students' conceptual understanding of variables and expressions, covering misconceptions such as letters as specific unknowns (Hodgen et al., 2024; Kuchemann, 1981)."
- Relevant Quotes: 1) "The conceptual understanding pretest included ten unit-weighted items which were adopted and adapted from existing tests on students' conceptual understanding of variables and expressions, covering misconceptions such as letters as specific unknowns (Hodgen et al., 2024; Kuchemann, 1981). Cognitive labs with students' think-aloud work provided deep insights into their reasoning about the items, so that the construct validity of the items could be optimized (Clark & Watson, 1995)." (p. 12) 2) "The pretest combined two items from the algebra prior understanding test (Pretest Part I in Fig. 2) administered during the preparation session and eight items from the e-scooter exploration tasks (Fig. 1, center) administered during the experiment session (Pretest Part II)." (p. 13) 3) "The conceptual understanding posttest, administered after the knowledge organization phase, consisted of ten structurally similar unit-weighted items (maximum score of 10) with changed contexts, to measure conceptual understanding rather than retention." (p. 13) 4) "For both tests, at least 20% of the data of the items were coded by two independent raters. The determined interrater agreement was substantial (k > 0.70)." (p. 13) 5) "Internal consistency was acceptable (pretest: wtot = 0.74, posttest: wtot = 0.75) and the average inter-item correlation (pretest: r = 0.21, posttest: r = 0.23) fell within the recommended range for scale homogeneity" (p. 13) 6) "Students' academic language proficiency (ALP) in German was measured by a C-Test, a widely used, economical, and valid measure that uses cloze texts without explicit mathematical language." (p. 12) Detailed Analysis: Criterion E requires that educational outcomes be measured with standardised, widely recognised exams rather than instruments built for the study, because researcher-built instruments risk being over-aligned with the intervention. Here the dependent variable - conceptual understanding of variables and expressions - was measured by a ten-item instrument that the authors assembled themselves. Individual items were "adopted and adapted from existing tests" and the authors refined them through their own cognitive labs; this is a bespoke research instrument, not a national, state-wide or otherwise recognised standardised examination. The posttest is described as consisting of structurally similar items with changed contexts, and eight of the ten pretest items are drawn directly from the e-scooter exploration task that also formed the material of the intervention itself, which is precisely the alignment concern Criterion E guards against. Reporting of interrater agreement and omega total demonstrates psychometric care, but internal consistency statistics do not make a test a standardised exam. The one genuinely standardised instrument used, the C-Test of academic language proficiency, was a background/control variable, not the educational outcome, so it cannot satisfy this criterion. Criterion E is not met because the educational outcome was measured with a researcher-constructed conceptual understanding test rather than a standardised exam.
-
T Term Duration
- The whole intervention and outcome measurement took place inside a single 85-minute session, with the posttest administered immediately after the 40-minute knowledge organization phase, far short of one academic term.
- "In the 85-minute experiment session, students worked through a self-learning environment with a two-phase macro-design (Loibl et al., 2017), consisting of identical exploration phases (20 min) across treatment conditions, three different conditions for the knowledge organization phase (40 min), and the posttest (25 min)."
- Relevant Quotes: 1) "To examine the research question, we conducted an 85-minute randomized controlled trial with N = 239 ninth to eleventh graders, with conceptual understanding (of variables as generalizers and algebraic expressions as descriptions of general relationships) as the dependent variable." (p. 9) 2) "In the 85-minute experiment session, students worked through a self-learning environment with a two-phase macro-design (Loibl et al., 2017), consisting of identical exploration phases (20 min) across treatment conditions, three different conditions for the knowledge organization phase (40 min), and the posttest (25 min)." (p. 10) 3) "After the knowledge organization phase, students continued individually with the posttest." (p. 12) 4) "Since students completed Pretest Part I in the preparation session 1-3 weeks prior to the experimental session, Posttest Items 1 and 2 (Table 1) were similar to the taxi task..." (p. 13) 5) "Focusing on one meaning of variables and expressions in a single session obviously does not allow students to develop a comprehensive understanding of algebra with its multiple interconnected concepts and meanings." (p. 25) 6) "Hence, the short-term self-learning environment should be embedded in a longer-term curriculum and evaluated for its effectiveness." (p. 25) Detailed Analysis: Criterion T requires that the primary outcome be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Short interventions are acceptable, but the follow-up tracking must extend to a term. In this study, the intervention (the 20-minute exploration phase plus the 40-minute knowledge organization phase) and the outcome measurement (the 25-minute posttest) all fell inside one 85-minute lesson. The interval from intervention start to primary outcome measurement is therefore about one hour. The only earlier contact was a 45-minute preparation session 1-3 weeks before, during which background data and Pretest Part I were collected - this precedes the intervention rather than extending follow-up after it. No delayed or retention posttest of any kind is reported, and the authors themselves characterise the study as a "single session" and "short-term self-learning environment" that ought in future to be embedded in a longer-term curriculum. Criterion T is not met because outcomes were measured immediately within the same 85-minute session in which the intervention took place, nowhere near one academic term.
-
D Documented Control Group
- The control condition is documented in detail, including its size (n=80), the exact learning activities and eight prompts it received, and a full table of baseline demographics, language proficiency and pretest scores with statistical confirmation of baseline equivalence.
- "Control condition: Student pairs received a written worked example of the exploration task and were asked to explain it to each other and revise their initial responses from the exploration phase."
- Relevant Quotes: 1) "Control condition: Student pairs received a written worked example of the exploration task and were asked to explain it to each other and revise their initial responses from the exploration phase. Afterwards, the pairs were invited again to revise their exploration solutions and to ensure that all their initial questions were answered." (p. 10) 2) "To practice the learned content, they worked on a similarly structured near-transfer task with a given table and a context similar to the exploration task ... and a far-transfer task without a given table in a very different context (expression to convert the temperature in degrees Fahrenheit for every possible degree Celsius)." (p. 10) 3) "In total, the control condition received eight prompts to support focused cognitive engagement, including the two transfer tasks and a follow-up task that required the students to summarize what they had learned." (p. 10) 4) "Table 2 Descriptive statistics of student variables in three treatment conditions ... Control condition (n = 80) ... Gender (girls in %) 62.5% ... Immigrant background (in %) 55.0% ... Multilingual background (in %) 62.5% ... Low socioeconomic status (in %) 43.8% ... Age in years (M (SD)) 16.9 (1.60) ... Academic language proficiency (M (SD)) 45.5 (10.8) ... Prior conceptual understanding (M (SD)) 3.75 (2.10) ... Post conceptual understanding (M (SD)) 3.94 (2.16)" (p. 20) 5) "The MANOVA showed no differences between conditions (with F(2, 236) = 0.646, p = .826). ANOVAs for each individual variable also revealed no significant group differences (F(2, 236) < 1.265, p > .283). Hence, the three groups started with comparable background variables and comparable prior conceptual understanding." (p. 19) 6) "The descriptive statistics in Table 2 show that, on average, students in the control condition used only 4.38 of the 8 (55%) consolidation prompts..." (p. 19) Detailed Analysis: Criterion D requires the control group to be well documented: its size, demographic composition, baseline performance and the conditions it experienced. All of these elements are present. The size is given exactly (n = 80). The treatment received by the control group is described concretely and in operational detail - a written worked example of the same e-scooter exploration task, mutual peer explanation, an invitation to revise their exploration solutions, a near-transfer water-price task, a far-transfer Fahrenheit-to-Celsius task, and a summarising follow-up task, adding up to eight consolidation prompts. Table 2 reports the control group's gender split, immigrant and multilingual background, socioeconomic status, age, academic language proficiency, engagement scores and both pretest and posttest conceptual understanding means and standard deviations, side by side with the two video conditions. The authors additionally run a MANOVA and per-variable ANOVAs demonstrating baseline equivalence across the three arms. Criterion D is met because the control group's size, demographics, baseline performance and exact instructional conditions are all explicitly documented.