Does subjective task value mediate the effect of digital simulation-based explorations on fraction learning? A randomized-controlled trial

Maria-Martine Oppmann, Maik Beege, Sarah Isabelle Hofer, Frank Reinhold

Published:
ERCT Check Date:
DOI: 10.1007/s11858-026-01793-5
  • mathematics
  • K12
  • EU
  • EdTech app
  • mobile learning
0
  • C

    Randomisation was performed on individual students within each classroom, not on whole classes or schools, and no tutoring exception applies.

    "students in each class were randomly assigned to either the EG (n = 151) or the CG (n = 141) by drawing face-down cards, ensuring a balanced distribution" (p. 5)

  • E

    Learning outcomes were measured with researcher-built, self-piloted pre- and posttests rather than a recognised standardised examination.

    "Prior to and following the intervention, both groups were given identical paper-based knowledge tests and self-assessments-piloted with n = 43 students." (p. 6)

  • T

    The whole trial, from pretest to posttest, took place inside a single 90-minute lesson, far short of one academic term.

    "The total intervention had a duration of 90 minutes and followed seven steps chronologically" (p. 5)

  • D

    The control group's size, its exact activity, and its baseline motivational and prior-knowledge scores are reported alongside the experimental group.

    "Table 6 Descriptive results of the pretest regarding motivational and emotional orientations regarding mathematics and relevant prior knowledge on fractions-split between the CG (n = 141) and the EG (n = 151)" (p. 10)

  • S

    No schools were randomised; allocation occurred among individual students inside each participating classroom.

    "In each classroom learners were randomly allocated to EG and CG through face-down card drawing." (p. 5)

  • I

    The authors developed the digital learning environment, designed the tests, personally taught the lessons, and analysed the data, with no third-party evaluator.

    "In both conditions, (1) researchers conducted the lesson after this exploration phase" (p. 5)

  • Y

    The study ran for a single 90-minute lesson, so the year-long tracking requirement fails, and criterion T was not met either.

    "a single exploration phase of 15 minutes within a single mathematics lesson of 90 minutes" (p. 12)

  • B

    Both conditions received an identical workbook and an identical, pre-fixed time schedule, and the only difference - the digital simulation itself - is the explicit treatment variable under test.

    "Both groups worked with an identically developed workbook, which only differed in the experimental manipulation, i.e., the digitalization of the individual exploratory phase of the mathematics lesson on the 'part of many wholes' concept." (p. 5)

  • R

    Internet searching found no independent replication of this trial by any other research team; the only related studies are by the same author group.

  • A

    Only a single narrow topic - the 'part of many wholes' fraction concept - was assessed, and criterion E was not met.

    "The posttest on context specific knowledge after the intervention contained twelve items (Cronbach's alpha = 0.75)" (p. 8)

  • G

    Measurement ended minutes after the intervention, no follow-up publication tracking this cohort exists, and criterion Y was not met.

    "it did not capture long-term changes in attitudes or sustained learning outcomes" (p. 12)

  • P

    Neither the paper nor any trial registry record shows a pre-registered protocol for this study.

Abstract

Fractions are challenging but essential for mathematical learning. Educational technology may support students' acquisition of fraction concepts. This study investigates underlying cause-and-effect mechanisms, i.e., whether features, e.g., authentic simulations, have a motivating effect in learning situations-resulting in an indirect learning-promoting effect. In a randomized controlled trial with N = 292 sixth-grade students, we examined the effects of a digital, simulation-based exploration of fractions compared to a paper-based version on students' subjective task value and their learning gains related to 'part of many wholes' concept. Results showed a significant positive effect of the digital, simulation-based exploration on a composite scale of students' intrinsic value and attainment value. Mediation analysis further indicated that the digital simulation significantly enhanced students' subjective task value, which in turn significantly improved posttest achievement. Consistent with the mediation hypothesis, we found a significant indirect effect. These findings underscore the importance of integrating features into digital learning environments that enhance student subjective task value in the challenging content area of fractions.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was performed on individual students within each classroom, not on whole classes or schools, and no tutoring exception applies.
      • "students in each class were randomly assigned to either the EG (n = 151) or the CG (n = 141) by drawing face-down cards, ensuring a balanced distribution" (p. 5)
      • Relevant Quotes: 1) "The study was conducted with a sample of N = 292 sixth-grade students from German Realschulen (intermediate secondary schools in the German three-track system) and students in each class were randomly assigned to either the EG (n = 151) or the CG (n = 141) by drawing face-down cards, ensuring a balanced distribution." (p. 5) 2) "In each classroom learners were randomly allocated to EG and CG through face-down card drawing." (p. 5) 3) "To test these hypotheses, a randomized controlled trial was implemented in regular sixth-grade mathematics classes. One EG (working with a digitally-enriched learning environment) and one CG (working solely paper-based) were exposed to the 'part of many wholes' subdimension of the part-whole concept in an intervention setting." (p. 4) 4) "In both conditions, (1) researchers conducted the lesson after this exploration phase, (2) results of the exploration phase were written down, and (3) individual practice tasks were administered paper based." (p. 5) Detailed Analysis: Criterion C requires randomisation at the class level (or the stronger school level), so that intervention and control groups are isolated from one another and contamination between conditions is avoided. The paper is explicit and unambiguous that allocation happened at the level of the individual student inside each classroom: "In each classroom learners were randomly allocated to EG and CG through face-down card drawing." Both conditions therefore sat in the same room, worked through the same workbook simultaneously, and shared the same teacher-led systematization and practice phases conducted by the researchers. This is precisely the design the criterion is intended to exclude, because control students can observe treatment students using tablets and the same instructor delivers both conditions, so contamination and expectancy effects cannot be ruled out. The exception in the standard applies only where the intervention is inherently personal teaching such as one-to-one tutoring. This intervention is a whole-class exploration phase embedded in regular sixth-grade mathematics lessons using a shared workbook, not tutoring, so the exception does not apply. Nothing in the paper indicates that whole classes or schools were the unit of allocation; class membership was simply the setting in which within-class randomisation occurred. Criterion C is not met because randomisation was carried out on individual students within the same classroom rather than on entire classes or schools, and no tutoring exception applies.
    • E

      Exam-based Assessment

      • Learning outcomes were measured with researcher-built, self-piloted pre- and posttests rather than a recognised standardised examination.
      • "Prior to and following the intervention, both groups were given identical paper-based knowledge tests and self-assessments-piloted with n = 43 students." (p. 6)
      • Relevant Quotes: 1) "Prior to and following the intervention, both groups were given identical paper-based knowledge tests and self-assessments-piloted with n = 43 students." (p. 6) 2) "Both tests included open- and closed-ended items that were scored according to a piloted coding rubric, with a maximum of two points attainable per item." (p. 8) 3) "The pretest on relevant prior knowledge comprised seven items (Cronbach's alpha = 0.76)-that were graded as 'fully solved' (2 points), 'partially solved' (1 point), or 'not solved' (0 points)-with a maximum of 14 points available" (p. 8) 4) "The posttest on context specific knowledge after the intervention contained twelve items (Cronbach's alpha = 0.75)-scored in the same way as the pre-test-whereby a maximum of 24 points could have been achieved." (p. 8) 5) "Five items were of a conceptual nature (e.g., 'Circle 3/4 of the pizzas,' given a picture of 8 randomly distributed pizzas ...) and two of procedural aspects (e.g., 'Calculate: 3/5 of 45')." (p. 8) 6) "This assessment utilized self-reported data from the pretest ... Using standardized scales from the PISA survey study (OECD, 2013)." (p. 6) Detailed Analysis: Criterion E requires that the educational outcome be measured with a standardised, widely recognised exam that was not built for the study, so that the assessment cannot be over-aligned with the intervention content. Here the achievement measures are bespoke instruments constructed by the authors: a seven-item pretest and a twelve-item posttest, hand-scored against "a piloted coding rubric" and piloted internally with 43 students. The example items are transparently tailored to the intervention material - the posttest asks about distributing salami pizzas among six guests, mirroring the pizza-distribution exploration task used in the treatment condition. Only internal-consistency statistics (Cronbach's alpha) are reported; no external validation, norming, or national/state standing is claimed. The only standardised instrument in the study is the PISA-derived motivational scale (interest, anxiety, work ethic, self-concept), but that is a self-report trait questionnaire used as a baseline covariate, not an exam-based measure of educational achievement. Similarly, the subjective task value scale is a self-report state measure adapted by the authors. Neither can substitute for a standardised achievement exam. Criterion E is not met because the achievement outcomes were measured by custom, researcher-designed and self-piloted tests closely aligned to the intervention task rather than by a recognised standardised exam.
    • T

      Term Duration

      • The whole trial, from pretest to posttest, took place inside a single 90-minute lesson, far short of one academic term.
      • "The total intervention had a duration of 90 minutes and followed seven steps chronologically" (p. 5)
      • Relevant Quotes: 1) "The total intervention had a duration of 90 minutes and followed seven steps chronologically (Fig. 4)." (p. 5) 2) "Table 1 Time schedule of the individual successive intervention phases ... 7 Questionnaire 'trait assessment' ... 8 Pretest to assess relevant prior knowledge ... 15 Intervened introduction and exploration phase ... 15 Systematization phase ... 30 Practice phase ... 7 Questionnaire 'state assessment' ... 8 Posttest to assess content-specific knowledge of the 'Part of Many Wholes'-concept" (p. 6) 3) "The exploration phase itself lasted approximately 15 minutes within a single lesson." (p. 11) 4) "Another limitation is the intervention's instructional time-a single exploration phase of 15 minutes within a single mathematics lesson of 90 minutes. While this focused and strict experimental design enabled precise measurement of immediate motivational effects, it did not capture long-term changes in attitudes or sustained learning outcomes." (p. 12) 5) "Because the posttest followed shortly after the intervention within the same lesson, students may have experienced some fatigue or reduced motivation to engage fully with a second assessment." (p. 12) Detailed Analysis: Criterion T requires that the primary outcome be measured at least one full academic term (roughly 3-4 months) after the intervention begins; short interventions are permitted, but term-long follow-up tracking is not. In this study the intervention start and the outcome measurement fall inside the same 90-minute lesson. Table 1 fixes the entire sequence: the experimental manipulation occupies 15 minutes, and the posttest is administered 8 minutes at the end of the same session, roughly one hour after the manipulation began. The authors themselves flag the absence of any longer-term measurement as a limitation and call for future longitudinal work ("This could be tested in a long-term study"). There is no delayed retention test, no follow-up at the end of the teaching unit, and no later data collection of any kind. The interval from intervention start to primary outcome measurement is therefore on the order of one hour, not one term. Criterion T is not met because the outcome was measured within the same single 90-minute lesson in which the intervention began, with no term-long follow-up.
    • D

      Documented Control Group

      • The control group's size, its exact activity, and its baseline motivational and prior-knowledge scores are reported alongside the experimental group.
      • "Table 6 Descriptive results of the pretest regarding motivational and emotional orientations regarding mathematics and relevant prior knowledge on fractions-split between the CG (n = 141) and the EG (n = 151)" (p. 10)
      • Relevant Quotes: 1) "The study was conducted with a sample of N = 292 sixth-grade students from German Realschulen ... and students in each class were randomly assigned to either the EG (n = 151) or the CG (n = 141)" (p. 5) 2) "The EG worked with the digital simulation-based learning environment during the exploration phase and the CG worked with the same material in the form of a paper-based version: Both groups worked with an identically developed workbook, which only differed in the experimental manipulation" (p. 5) 3) "In CG, the above-described task was given in paper-based format, requiring students to draw the results of their equal sharing process." (p. 5) 4) "The CG worked on the same task paper based (Fig. 5, right). They were instructed to mark where they would cut the pizzas and to draw the resulting pizza slices on the plates." (p. 6) 5) "participants in the CG did not receive any feedback." (p. 6) 6) "Table 6 Descriptive results of the pretest regarding motivational and emotional orientations regarding mathematics and relevant prior knowledge on fractions-split between the CG (n = 141) and the EG (n = 151)" (p. 10) 7) "Prior knowledge 5.730 3.991 6.172 3.786 0.971 290 .333 0.114" (Table 6, p. 10) 8) "Due to the random design of the study, there were no significant differences between the EG and the CG in terms of motivational and emotional orientations regarding mathematics (i.e., motivational traits) or domain-specific prior knowledge before the intervention (Table 6)-suggesting two comparable groups." (p. 9) 9) "In accordance with the study's fully anonymous design, no demographic details, including gender and age, were collected." (p. 5) Detailed Analysis: Criterion D asks for a well-documented control group: its size, its baseline characteristics, and what it actually received during the study. All three are supplied here. The control group's size is stated exactly (n = 141, with n = 139 retained for the posttest subjective-task-value analysis per Table 7). Its condition is described concretely: the same workbook and the same pizza-distribution task in static paper-based form, marking cut lines and drawing slices, with no feedback, followed by the same researcher-led systematization and practice phases as the experimental group. Baseline comparability is documented quantitatively. Table 6 reports means and standard deviations separately for CG and EG on four PISA motivational scales (interest, anxiety, work ethic, self-concept) and on the fraction prior-knowledge pretest, together with t, df, p, and Cohen's d for each comparison, and the text confirms no significant baseline differences. This is sufficient to judge that the groups were comparable at the outset. The one gap is demographics: the authors explicitly did not collect gender or age because of the fully anonymous design. However, the standard's core requirement - baseline performance, group size, and the conditions the control group experienced - is satisfied in detail, and the population (sixth-graders in Realschulen in Baden-Wuerttemberg) is characterised at the sample level. Criterion D is met because the paper reports the control group's size, its precise paper-based condition, and its baseline motivational and prior-knowledge scores in a dedicated comparison table.
  • Level 2 Criteria

    • S

      School-level RCT

      • No schools were randomised; allocation occurred among individual students inside each participating classroom.
      • "In each classroom learners were randomly allocated to EG and CG through face-down card drawing." (p. 5)
      • Relevant Quotes: 1) "In each classroom learners were randomly allocated to EG and CG through face-down card drawing." (p. 5) 2) "students in each class were randomly assigned to either the EG (n = 151) or the CG (n = 141) by drawing face-down cards, ensuring a balanced distribution." (p. 5) 3) "Following the consent of the school principals and mathematics teachers, consent forms were distributed to students and their parents." (p. 5) 4) "As the digital simulation-based learning environment we used an electronic textbook on fractions that focuses on building up conceptual knowledge and has already been shown to be effective in authentic learning scenarios in a large cluster randomized controlled trial (Reinhold et al., 2020)." (p. 6) Detailed Analysis: Criterion S requires that the randomised unit be the school or equivalent implementing institution. Here the randomised unit is the individual pupil. Schools and teachers were recruited by consent, not allocated; within every participating classroom, students drew face-down cards to determine whether they would use the tablet version or the paper version. The paper never reports the number of participating schools or classes, nor any school-level allocation procedure, stratifica- tion, or cluster-adjusted analysis - the statistical models (independent-samples t-tests and OLS mediation regressions) treat the student as the unit of analysis. The paper does reference an earlier large cluster randomised controlled trial (Reinhold et al., 2020) that evaluated the same electronic textbook, but that is a different, prior study and cannot satisfy the criterion for the present trial. Criterion S is not met because randomisation was done among individual students within classrooms and no school-level or cluster-level allocation took place.
    • I

      Independent Conduct

      • The authors developed the digital learning environment, designed the tests, personally taught the lessons, and analysed the data, with no third-party evaluator.
      • "In both conditions, (1) researchers conducted the lesson after this exploration phase" (p. 5)
      • Relevant Quotes: 1) "In both conditions, (1) researchers conducted the lesson after this exploration phase, (2) results of the exploration phase were written down, and (3) individual practice tasks were administered paper based. The time of the individual successive intervention phases was determined in advance between the study leaders to create reliability of the results (Table 1)." (pp. 5-6) 2) "Acknowledgements Software used in this study was developed in the 'ALICE:fractions' project. We want to thank S. Hoch, B. Werner, J. Richter-Gebert and K. Reiss for their great contribution to the project." (p. 13) 3) "F. Reinhold: Conceptualization, Methodology, Software, Formal analysis, Resources, Writing-Review & Editing, Supervision, Project administration, Funding acquisition." (p. 13) 4) "M.-M. Oppmann: Conceptualization, Methodology, Validation, Formal analysis, Investigation, Data Curation, Writing - Original Draft, Visualization." (p. 13) 5) "As the digital simulation-based learning environment we used an electronic textbook on fractions ... has already been shown to be effective in authentic learning scenarios in a large cluster randomized controlled trial (Reinhold et al., 2020)." (p. 6) 6) "Reinhold, F., Hoch, S., Werner, B., Richter-Gebert, J., & Reiss, K. (2020). Learning fractions with and without educational technology: What matters for high-achieving and low-achieving students?" (p. 15) 7) "This research was supported by the Daimler and Benz Foundation through a research grant awarded to F. Reinhold (Grant No. 32-08/20)." (p. 13) 8) "Competing Interests The authors declare that they have no conflict of interest." (p. 13) Detailed Analysis: Criterion I requires that the evaluation be carried out independently of the people who designed the intervention, or that explicit third-party oversight of data collection and analysis be documented. Every link in this chain runs back to the author team. The digital environment is the ALICE:fractions electronic textbook, and the senior author F. Reinhold is credited with the "Software" contributor role and is the first author of the Reinhold et al. (2020) study that produced and previously evaluated that same textbook. The authors also built the workbook, wrote and piloted the pre- and posttests and the coding rubric, and - critically - personally delivered the instruction: "researchers conducted the lesson after this exploration phase," with the timing of each phase fixed "between the study leaders." Data curation, formal analysis, and reporting are also attributed to the authors (Oppmann: Investigation, Data Curation, Formal analysis; Reinhold: Formal analysis). No external evaluation agency, no blinded independent test administrators, and no independent oversight body are mentioned anywhere. The declared absence of a conflict of interest and the approval by the Freiburg Regional Council are ethics/administrative statements, not evidence of independent conduct. The exception patterns in the standard (provider excluded from data collection and analysis; trial led by a government body) do not apply here. Criterion I is not met because the same team designed the digital intervention and the instruments, taught the lessons, and analysed the data, with no independent evaluator involved.
    • Y

      Year Duration

      • The study ran for a single 90-minute lesson, so the year-long tracking requirement fails, and criterion T was not met either.
      • "a single exploration phase of 15 minutes within a single mathematics lesson of 90 minutes" (p. 12)
      • Relevant Quotes: 1) "The total intervention had a duration of 90 minutes and followed seven steps chronologically (Fig. 4)." (p. 5) 2) "Another limitation is the intervention's instructional time-a single exploration phase of 15 minutes within a single mathematics lesson of 90 minutes. While this focused and strict experimental design enabled precise measurement of immediate motivational effects, it did not capture long-term changes in attitudes or sustained learning outcomes." (p. 12) 3) "Future research should implement longitudinal interventions, potentially across different instructional phases (Loibl et al., 2024; Prediger et al., 2021) to evaluate the durability and transferability of the observed effects." (p. 12) 4) "Our goal during the exploration phase was to increase the subjective task value for the topic in a way that an increase in motivation for the following lessons on fractions could also be observed. This could be tested in a long-term study." (p. 12) Detailed Analysis: Criterion Y requires outcomes to be measured at least 75% of an academic year (roughly 7 months or more) after the intervention begins. The total elapsed time between the start of the experimental manipulation and the posttest here is under one hour, contained within a single 90-minute mathematics lesson. The authors explicitly acknowledge that the design "did not capture long-term changes in attitudes or sustained learning outcomes" and propose longitudinal follow-up as future work, confirming that no such tracking was carried out. In addition, the prompt-level dependency applies: since criterion T (Term Duration) is not met, criterion Y cannot be met. Criterion Y is not met because the entire trial, including outcome measurement, was completed within one 90-minute lesson, nowhere near 75% of an academic year, and criterion T also failed.
    • B

      Balanced Control Group

      • Both conditions received an identical workbook and an identical, pre-fixed time schedule, and the only difference - the digital simulation itself - is the explicit treatment variable under test.
      • "Both groups worked with an identically developed workbook, which only differed in the experimental manipulation, i.e., the digitalization of the individual exploratory phase of the mathematics lesson on the 'part of many wholes' concept." (p. 5)
      • Relevant Quotes: 1) "Both groups worked with an identically developed workbook, which only differed in the experimental manipulation, i.e., the digitalization of the individual exploratory phase of the mathematics lesson on the 'part of many wholes' concept." (p. 5) 2) "In both conditions, (1) researchers conducted the lesson after this exploration phase, (2) results of the exploration phase were written down, and (3) individual practice tasks were administered paper based. The time of the individual successive intervention phases was determined in advance between the study leaders to create reliability of the results (Table 1)." (pp. 5-6) 3) "The workbooks distributed to the EG included QR codes in the exploration phase that led to the completion of interactive tasks. In contrast, the workbooks distributed CG contained tasks that were equivalent in format to those completed by the EG but were presented in a static, paper-based format." (p. 6) 4) "The CG completed the same task using paper-based materials to facilitate the same conceptual exploration." (p. 5) 5) "Depending on whether the solution was correct or incorrect, participants in the EG received feedback (e.g., 'Correct!', 'The green plate is still empty.', or 'Too bad! You didn't give everybody the same mount. For example, the green plate has more than the red.')- participants in the CG did not receive any feedback." (p. 6) 6) "During the subsequent systematization phase, both groups were shown the distributed pizzas again so that they can both practice from the same level of knowledge paper-based." (p. 6) 7) "In this paper, we hypothesize that a digital simulation-based mathematical exploration of fractions-through its interactive, gesture-based, and contextually rich design-may enhance students' intrinsic value and attainment value compared to a paper-based control condition." (p. 1) 8) "Hypothesis 1: Simulation-based mathematical exploration leads to higher Subjective Task Value in students than comparable paper-based exploration." (p. 4) 9) "Although both groups engaged with the same mathematical content and visual representations, the control group did not perform any physical action of cutting and distributing the pizzas in the intervened exploration phase; but merely marked there solutions with pens." (p. 12) Detailed Analysis: Following the decision tree for criterion B, the first question is whether the intervention adds time or budget. Instructional time is exactly matched: Table 1 fixes a single schedule (7 + 8 + 15 + 15 + 30 + 7 + 8 minutes) applied to both conditions, and the timing was "determined in advance between the study leaders." Both groups sat the same trait questionnaire, the same pretest, the same 15-minute exploration, the same 15-minute systematization, the same 30-minute practice phase, and the same posttest, using an identically developed workbook that differed only in the exploration pages. The systematization phase explicitly re-levels both groups ("both groups were shown the distributed pizzas again so that they can both practice from the same level of knowledge"). Adult support is also equivalent: the researchers taught both conditions in the same room. What the experimental group does receive extra is the touchscreen device and the automated correctness feedback embedded in the simulation. This is not a separable, optional add-on: the interactivity, authenticity, adaptivity, and immediate feedback of the digital simulation are precisely the constructs the study is testing. The stated experimental manipulation is "Interactivity, authenticity, and adaptivity" (Fig. 2), and Hypothesis 1 frames the contrast explicitly as simulation-based versus "comparable paper-based exploration." Under the standard's integral-resource rule and Example 3 of the specification (digital tool versus teaching-as-usual), the extra digital inputs are the treatment variable, so a paper-based business-as- usual comparison is legitimate by design. The authors themselves note the residual confound that the control group performed no embodied cutting gesture, and recommend an additional manual-manipulation control arm in future work. That is a specificity limitation for attributing the effect to digitality as such, not an imbalance of educational time or budget, and it does not overturn the resource-balance judgement. Criterion B is met because instructional time, materials, and teacher input were held identical across conditions, and the only additional resources given to the experimental group - the tablet simulation and its immediate feedback - are integral to the intervention being tested against a paper-based baseline.
  • Level 3 Criteria

    • R

      Reproduced

      • Internet searching found no independent replication of this trial by any other research team; the only related studies are by the same author group.
      • Relevant Quotes: 1) "As the digital simulation-based learning environment we used an electronic textbook on fractions that focuses on building up conceptual knowledge and has already been shown to be effective in authentic learning scenarios in a large cluster randomized controlled trial (Reinhold et al., 2020)." (p. 6) 2) "However, this is contradicted by the fact that other motivational mechanisms are at work in other phases of mathematics instruction in comparable learner groups (e.g., perceived autonomy and competence support in a practice phase on equivalence of fractions, Oppmann et al., 2025)." (p. 12) 3) "The generalizability of the results may be limited by the sample context (sixth graders in Germany). Future research should investigate whether similar effects are observed across different age groups, cultures, or mathematical topics." (p. 12) Detailed Analysis: Criterion R requires that this specific study - its central experimental claim, design, and context - has been independently replicated by a different research team and published in a peer-reviewed journal. Verification searches were carried out on the internet (publisher record, DOI 10.1007/s11858-026-01793-5, author profiles, and citation searches for the article title and for replications of the mediation claim). The article is indexed as ZDM - Mathematics Education 58, 575-589 (2026), published online 21 May 2026. No paper by any independent team replicating this trial was found, and no citing replication study could be identified. No verbatim quotes from a replication can be supplied because no such publication exists at the time of this check. The two closest studies remain by the same author group and are not replications: Reinhold, F., Hoch, S., Werner, B., Richter-Gebert, J., & Reiss, K. (2020), "Learning fractions with and without educational technology: What matters for high-achieving and low-achieving students?", Learning and Instruction, 65, 101264 - an earlier trial of the same ALICE:fractions electronic textbook co-authored by the present senior author; and Oppmann, M.-M., Beege, M., & Reinhold, F. (2025), "Stimulating individual learning of the concept of fraction equivalence: How students utilize adaptive features in digital learning environments mediates their effect", Learning and Instruction, 98, 102118 - a companion 90-minute RCT by the present first and senior authors on a different instructional phase with a different mediator. Neither is independent of this team, and neither reproduces the present subjective-task-value mediation design. Criterion R is not met because no independent replication of this study by a different research team was reported in the paper or found through internet searching.
    • A

      All-subject Exams

      • Only a single narrow topic - the 'part of many wholes' fraction concept - was assessed, and criterion E was not met.
      • "The posttest on context specific knowledge after the intervention contained twelve items (Cronbach's alpha = 0.75)" (p. 8)
      • Relevant Quotes: 1) "Posttest regarding content-specific knowledge: The final part of the intervention employed a posttest to assess students' understanding of the concept of 'part of many wholes'." (p. 5) 2) "The posttest on context specific knowledge after the intervention contained twelve items (Cronbach's alpha = 0.75)-scored in the same way as the pre-test-whereby a maximum of 24 points could have been achieved." (p. 8) 3) "Seven tasks were designed to assess specific aspects of conceptual knowledge (e.g. 'Mark 1/6 of the three strips in two different ways with different colors.' ... ). The remaining five tasks focused on procedural knowledge (e.g., 'Calculate. Solve as many tasks as possible: 4/7 of 42 and 3/9 of 6.')." (p. 9) 4) "Prior to and following the intervention, both groups were given identical paper-based knowledge tests and self-assessments-piloted with n = 43 students." (p. 6) Detailed Analysis: Criterion A requires that impact be measured across all main subjects taught at that educational level, using standardised exam-based assessments. This study measures one sub-topic of one subject: the 'part of many wholes' subdimension of the part-whole fraction concept in mathematics. There is no assessment of German, English, science, or any other core sixth-grade Realschule subject, and no rationale is offered for restricting measurement, nor would the specialised-vocational exception apply to general lower-secondary schooling. Independently, the standard and the ranking instructions make criterion E a prerequisite for criterion A. Since the outcomes were captured by custom researcher-built tests rather than standardised exams, criterion E fails and criterion A therefore fails on that ground as well. Criterion A is not met because only a single narrow mathematics topic was assessed, with custom instruments, and criterion E was not met.
    • G

      Graduation Tracking

      • Measurement ended minutes after the intervention, no follow-up publication tracking this cohort exists, and criterion Y was not met.
      • "it did not capture long-term changes in attitudes or sustained learning outcomes" (p. 12)
      • Relevant Quotes: 1) "Because the posttest followed shortly after the intervention within the same lesson, students may have experienced some fatigue or reduced motivation to engage fully with a second assessment." (p. 12) 2) "Another limitation is the intervention's instructional time-a single exploration phase of 15 minutes within a single mathematics lesson of 90 minutes. While this focused and strict experimental design enabled precise measurement of immediate motivational effects, it did not capture long-term changes in attitudes or sustained learning outcomes." (p. 12) 3) "Future research should implement longitudinal interventions ... to evaluate the durability and transferability of the observed effects." (p. 12) 4) "Our goal during the exploration phase was to increase the subjective task value for the topic in a way that an increase in motivation for the following lessons on fractions could also be observed. This could be tested in a long-term study." (p. 12) 5) "In accordance with the study's fully anonymous design, no demographic details, including gender and age, were collected." (p. 5) Detailed Analysis: Criterion G requires that participants be followed through to graduation from their educational stage, in the original paper or in a subsequent publication on the same cohort. Here data collection stopped with a posttest administered minutes after the intervention inside the same lesson. The authors state plainly that the design "did not capture long-term changes in attitudes or sustained learning outcomes" and propose longitudinal designs only as future work. Moreover, the study was fully anonymous with no demographic data collected, which makes any later linkage of these participants to graduation records practically impossible. Internet searching was performed for subsequent publications by the same authors that might report longer-term or graduation tracking of this cohort (author publication profiles for M.-M. Oppmann, M. Beege, S. I. Hofer and F. Reinhold, and searches on the article title and DOI). The only closely related recent output found is Oppmann, M.-M., Beege, M., & Reinhold, F. (2025), "Stimulating individual learning of the concept of fraction equivalence: How students utilize adaptive features in digital learning environments mediates their effect", Learning and Instruction, 98, 102118, which is itself a separate 90-minute RCT on a different instructional phase rather than a follow-up of this cohort. No follow-up paper with graduation tracking was found, and no verbatim quotes evidencing graduation tracking can be supplied because no such publication exists. In addition, the specification's dependency applies: since criterion Y is not met, criterion G cannot be met. Criterion G is not met because measurement ceased within the same lesson, no follow-up publication tracking this cohort to graduation was found, and criterion Y also failed.
    • P

      Pre-Registered

      • Neither the paper nor any trial registry record shows a pre-registered protocol for this study.
      • Relevant Quotes: 1) "The sample was built only after the study was approved by the local school authority (Freiburg Regional Council, Department 7, School and Education, approval number 7-6499.2)." (p. 5) 2) "As no straightforward a-priori power calculation exists for the present mediation model, we relied on established simulation-based recommendations to evaluate whether the available sample size can be considered sufficient." (p. 9) 3) "Note that two items from the originally adapted item pool ('In today's math lesson, I was interested' and 'In today's math lesson, I was motivated to learn') were excluded because they reflected general engagement rather than the targeted SEVT facets of subjective task value." (p. 8) 4) "Competing Interests The authors declare that they have no conflict of interest." (p. 13) 5) "This research was supported by the Daimler and Benz Foundation through a research grant awarded to F. Reinhold (Grant No. 32-08/20)." (p. 13) Detailed Analysis: Criterion P requires a publicly pre-registered protocol, with hypotheses, methods, and analysis plan lodged before data collection began, and evidence of the registration and its date. The paper contains no registration identifier, no registry name (AsPredicted, OSF, ClinicalTrials.gov, DRKS, ISRCTN, AEA RCT Registry), no protocol citation, and no data- or materials-availability statement pointing to one. The only ex-ante approval mentioned is ethical/adminis- trative clearance from the Freiburg Regional Council, which is not a study pre-registration. Because no registry or identifier is named in the paper, verification was attempted through internet searching on the article title, DOI 10.1007/s11858-026-01793-5, the author names, and pre-registration terms; the publisher record contains no pre-registration or data-availability link, and no matching registry entry was located. No registration date could therefore be checked. Two further details point away from a binding pre-specified analysis plan: the authors state that no a-priori power calculation was carried out and instead justify the sample size post hoc against published simulation benchmarks, and two items were dropped from the subjective task value scale after the fact, with the mediator structure chosen following a comparison of three competing CFA models. These are legitimate analytic choices but they are the kind of researcher degrees of freedom that pre-registration is designed to constrain. Criterion P is not met because no pre-registration of the protocol, hypotheses, or analysis plan is reported in the paper or findable in any registry.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.