Level 1 Criteria
-
C Class-level RCT
- Randomisation was performed on individual students within each classroom, not on whole classes or schools, and no tutoring exception applies.
- "students in each class were randomly assigned to either the EG (n = 151) or the CG (n = 141) by drawing face-down cards, ensuring a balanced distribution" (p. 5)
- Relevant Quotes: 1) "The study was conducted with a sample of N = 292 sixth-grade students from German Realschulen (intermediate secondary schools in the German three-track system) and students in each class were randomly assigned to either the EG (n = 151) or the CG (n = 141) by drawing face-down cards, ensuring a balanced distribution." (p. 5) 2) "In each classroom learners were randomly allocated to EG and CG through face-down card drawing." (p. 5) 3) "To test these hypotheses, a randomized controlled trial was implemented in regular sixth-grade mathematics classes. One EG (working with a digitally-enriched learning environment) and one CG (working solely paper-based) were exposed to the 'part of many wholes' subdimension of the part-whole concept in an intervention setting." (p. 4) 4) "In both conditions, (1) researchers conducted the lesson after this exploration phase, (2) results of the exploration phase were written down, and (3) individual practice tasks were administered paper based." (p. 5) Detailed Analysis: Criterion C requires randomisation at the class level (or the stronger school level), so that intervention and control groups are isolated from one another and contamination between conditions is avoided. The paper is explicit and unambiguous that allocation happened at the level of the individual student inside each classroom: "In each classroom learners were randomly allocated to EG and CG through face-down card drawing." Both conditions therefore sat in the same room, worked through the same workbook simultaneously, and shared the same teacher-led systematization and practice phases conducted by the researchers. This is precisely the design the criterion is intended to exclude, because control students can observe treatment students using tablets and the same instructor delivers both conditions, so contamination and expectancy effects cannot be ruled out. The exception in the standard applies only where the intervention is inherently personal teaching such as one-to-one tutoring. This intervention is a whole-class exploration phase embedded in regular sixth-grade mathematics lessons using a shared workbook, not tutoring, so the exception does not apply. Nothing in the paper indicates that whole classes or schools were the unit of allocation; class membership was simply the setting in which within-class randomisation occurred. Criterion C is not met because randomisation was carried out on individual students within the same classroom rather than on entire classes or schools, and no tutoring exception applies.
-
E Exam-based Assessment
- Learning outcomes were measured with researcher-built, self-piloted pre- and posttests rather than a recognised standardised examination.
- "Prior to and following the intervention, both groups were given identical paper-based knowledge tests and self-assessments-piloted with n = 43 students." (p. 6)
- Relevant Quotes: 1) "Prior to and following the intervention, both groups were given identical paper-based knowledge tests and self-assessments-piloted with n = 43 students." (p. 6) 2) "Both tests included open- and closed-ended items that were scored according to a piloted coding rubric, with a maximum of two points attainable per item." (p. 8) 3) "The pretest on relevant prior knowledge comprised seven items (Cronbach's alpha = 0.76)-that were graded as 'fully solved' (2 points), 'partially solved' (1 point), or 'not solved' (0 points)-with a maximum of 14 points available" (p. 8) 4) "The posttest on context specific knowledge after the intervention contained twelve items (Cronbach's alpha = 0.75)-scored in the same way as the pre-test-whereby a maximum of 24 points could have been achieved." (p. 8) 5) "Five items were of a conceptual nature (e.g., 'Circle 3/4 of the pizzas,' given a picture of 8 randomly distributed pizzas ...) and two of procedural aspects (e.g., 'Calculate: 3/5 of 45')." (p. 8) 6) "This assessment utilized self-reported data from the pretest ... Using standardized scales from the PISA survey study (OECD, 2013)." (p. 6) Detailed Analysis: Criterion E requires that the educational outcome be measured with a standardised, widely recognised exam that was not built for the study, so that the assessment cannot be over-aligned with the intervention content. Here the achievement measures are bespoke instruments constructed by the authors: a seven-item pretest and a twelve-item posttest, hand-scored against "a piloted coding rubric" and piloted internally with 43 students. The example items are transparently tailored to the intervention material - the posttest asks about distributing salami pizzas among six guests, mirroring the pizza-distribution exploration task used in the treatment condition. Only internal-consistency statistics (Cronbach's alpha) are reported; no external validation, norming, or national/state standing is claimed. The only standardised instrument in the study is the PISA-derived motivational scale (interest, anxiety, work ethic, self-concept), but that is a self-report trait questionnaire used as a baseline covariate, not an exam-based measure of educational achievement. Similarly, the subjective task value scale is a self-report state measure adapted by the authors. Neither can substitute for a standardised achievement exam. Criterion E is not met because the achievement outcomes were measured by custom, researcher-designed and self-piloted tests closely aligned to the intervention task rather than by a recognised standardised exam.
-
T Term Duration
- The whole trial, from pretest to posttest, took place inside a single 90-minute lesson, far short of one academic term.
- "The total intervention had a duration of 90 minutes and followed seven steps chronologically" (p. 5)
- Relevant Quotes: 1) "The total intervention had a duration of 90 minutes and followed seven steps chronologically (Fig. 4)." (p. 5) 2) "Table 1 Time schedule of the individual successive intervention phases ... 7 Questionnaire 'trait assessment' ... 8 Pretest to assess relevant prior knowledge ... 15 Intervened introduction and exploration phase ... 15 Systematization phase ... 30 Practice phase ... 7 Questionnaire 'state assessment' ... 8 Posttest to assess content-specific knowledge of the 'Part of Many Wholes'-concept" (p. 6) 3) "The exploration phase itself lasted approximately 15 minutes within a single lesson." (p. 11) 4) "Another limitation is the intervention's instructional time-a single exploration phase of 15 minutes within a single mathematics lesson of 90 minutes. While this focused and strict experimental design enabled precise measurement of immediate motivational effects, it did not capture long-term changes in attitudes or sustained learning outcomes." (p. 12) 5) "Because the posttest followed shortly after the intervention within the same lesson, students may have experienced some fatigue or reduced motivation to engage fully with a second assessment." (p. 12) Detailed Analysis: Criterion T requires that the primary outcome be measured at least one full academic term (roughly 3-4 months) after the intervention begins; short interventions are permitted, but term-long follow-up tracking is not. In this study the intervention start and the outcome measurement fall inside the same 90-minute lesson. Table 1 fixes the entire sequence: the experimental manipulation occupies 15 minutes, and the posttest is administered 8 minutes at the end of the same session, roughly one hour after the manipulation began. The authors themselves flag the absence of any longer-term measurement as a limitation and call for future longitudinal work ("This could be tested in a long-term study"). There is no delayed retention test, no follow-up at the end of the teaching unit, and no later data collection of any kind. The interval from intervention start to primary outcome measurement is therefore on the order of one hour, not one term. Criterion T is not met because the outcome was measured within the same single 90-minute lesson in which the intervention began, with no term-long follow-up.
-
D Documented Control Group
- The control group's size, its exact activity, and its baseline motivational and prior-knowledge scores are reported alongside the experimental group.
- "Table 6 Descriptive results of the pretest regarding motivational and emotional orientations regarding mathematics and relevant prior knowledge on fractions-split between the CG (n = 141) and the EG (n = 151)" (p. 10)
- Relevant Quotes: 1) "The study was conducted with a sample of N = 292 sixth-grade students from German Realschulen ... and students in each class were randomly assigned to either the EG (n = 151) or the CG (n = 141)" (p. 5) 2) "The EG worked with the digital simulation-based learning environment during the exploration phase and the CG worked with the same material in the form of a paper-based version: Both groups worked with an identically developed workbook, which only differed in the experimental manipulation" (p. 5) 3) "In CG, the above-described task was given in paper-based format, requiring students to draw the results of their equal sharing process." (p. 5) 4) "The CG worked on the same task paper based (Fig. 5, right). They were instructed to mark where they would cut the pizzas and to draw the resulting pizza slices on the plates." (p. 6) 5) "participants in the CG did not receive any feedback." (p. 6) 6) "Table 6 Descriptive results of the pretest regarding motivational and emotional orientations regarding mathematics and relevant prior knowledge on fractions-split between the CG (n = 141) and the EG (n = 151)" (p. 10) 7) "Prior knowledge 5.730 3.991 6.172 3.786 0.971 290 .333 0.114" (Table 6, p. 10) 8) "Due to the random design of the study, there were no significant differences between the EG and the CG in terms of motivational and emotional orientations regarding mathematics (i.e., motivational traits) or domain-specific prior knowledge before the intervention (Table 6)-suggesting two comparable groups." (p. 9) 9) "In accordance with the study's fully anonymous design, no demographic details, including gender and age, were collected." (p. 5) Detailed Analysis: Criterion D asks for a well-documented control group: its size, its baseline characteristics, and what it actually received during the study. All three are supplied here. The control group's size is stated exactly (n = 141, with n = 139 retained for the posttest subjective-task-value analysis per Table 7). Its condition is described concretely: the same workbook and the same pizza-distribution task in static paper-based form, marking cut lines and drawing slices, with no feedback, followed by the same researcher-led systematization and practice phases as the experimental group. Baseline comparability is documented quantitatively. Table 6 reports means and standard deviations separately for CG and EG on four PISA motivational scales (interest, anxiety, work ethic, self-concept) and on the fraction prior-knowledge pretest, together with t, df, p, and Cohen's d for each comparison, and the text confirms no significant baseline differences. This is sufficient to judge that the groups were comparable at the outset. The one gap is demographics: the authors explicitly did not collect gender or age because of the fully anonymous design. However, the standard's core requirement - baseline performance, group size, and the conditions the control group experienced - is satisfied in detail, and the population (sixth-graders in Realschulen in Baden-Wuerttemberg) is characterised at the sample level. Criterion D is met because the paper reports the control group's size, its precise paper-based condition, and its baseline motivational and prior-knowledge scores in a dedicated comparison table.