Can Interactive Videos Enhance Understanding of Variables as Generalizers? Randomized Controlled Trial on Engagement Effects

Stefan Korntreff, Stephan Bach, Mike Altieri, Susanne Prediger

Published:
ERCT Check Date:
DOI: 10.1007/s10758-026-09973-8
  • mathematics
  • K12
  • EU
  • EdTech platform
0
  • C

    Randomisation was carried out on stratified pairs of students within each classroom, with all three conditions present in the same class, so the unit of randomisation was below the class level and contamination was not prevented.

    "Pairs of students were stratified according to Pretest Part I scores on algebra prior understanding and academic language proficiency and then randomly assigned to the three treatment conditions, ensuring that the three conditions were represented in each classroom..."

  • E

    The conceptual understanding outcome was measured with a ten-item pretest and posttest assembled and adapted by the researchers for this study, not with a recognised standardised exam.

    "The conceptual understanding pretest included ten unit-weighted items which were adopted and adapted from existing tests on students' conceptual understanding of variables and expressions, covering misconceptions such as letters as specific unknowns (Hodgen et al., 2024; Kuchemann, 1981)."

  • T

    The whole intervention and outcome measurement took place inside a single 85-minute session, with the posttest administered immediately after the 40-minute knowledge organization phase, far short of one academic term.

    "In the 85-minute experiment session, students worked through a self-learning environment with a two-phase macro-design (Loibl et al., 2017), consisting of identical exploration phases (20 min) across treatment conditions, three different conditions for the knowledge organization phase (40 min), and the posttest (25 min)."

  • D

    The control condition is documented in detail, including its size (n=80), the exact learning activities and eight prompts it received, and a full table of baseline demographics, language proficiency and pretest scores with statistical confirmation of baseline equivalence.

    "Control condition: Student pairs received a written worked example of the exploration task and were asked to explain it to each other and revise their initial responses from the exploration phase."

  • S

    Schools were not randomised at all; the four participating schools and 16 classes merely supplied the sample, and allocation was performed on student pairs within each classroom.

    "Pairs of students were stratified according to Pretest Part I scores on algebra prior understanding and academic language proficiency and then randomly assigned to the three treatment conditions, ensuring that the three conditions were represented in each classroom..."

  • I

    The same team designed the self-learning environment and the instructional videos in their own prior design-research work, and then ran, coded and analysed the trial themselves, with no external evaluator or third-party oversight reported.

    "Hence, the self-learning environment on variables and expressions in view of this paper was developed in a qualitative design-research study (Korntreff & Prediger, 2026) to tailor these learning opportunities to students' needs."

  • Y

    The entire trial ran within a single 85-minute session with an immediate posttest, so the tracking interval falls drastically short of 75% of an academic year, and Criterion T is also not met.

    "To examine the research question, we conducted an 85-minute randomized controlled trial with N = 239 ninth to eleventh graders..."

  • B

    All three conditions received exactly the same instructional time within one 85-minute session and an active, content-matched control that worked through a written worked example of the same task plus transfer practice, so educational time and inputs were balanced.

    "In the 85-minute experiment session, students worked through a self-learning environment with a two-phase macro-design (Loibl et al., 2017), consisting of identical exploration phases (20 min) across treatment conditions, three different conditions for the knowledge organization phase (40 min), and the posttest (25 min)."

  • R

    No independent replication of this trial by a different research team is reported in the paper, and dedicated internet searching found only same-author companion publications rather than any reproduction.

  • A

    Only a narrow slice of mathematics - conceptual understanding of variables as generalizers and algebraic expressions - was assessed, with a researcher-built instrument, so neither the all-subject coverage nor the Criterion E prerequisite is satisfied.

    "Focusing on one meaning of variables and expressions in a single session obviously does not allow students to develop a comprehensive understanding of algebra with its multiple interconnected concepts and meanings."

  • G

    Measurement ended with the posttest inside the same 85-minute session, no follow-up publication by the same authors tracks the cohort to graduation, and Criterion Y is also not met.

    "After the knowledge organization phase, students continued individually with the posttest."

  • P

    The paper contains no reference to a trial registry entry, pre-registered protocol, or registration date preceding data collection, and no registry record was found online.

Abstract

Self-learning environments with instructional videos are well-established for procedural skill remediation, but their effectiveness for conceptual understanding of complex concepts has been less examined. This study examines the impact of more or less structured prompts in instructional videos on students' understanding of variables as generalizers while controlling for effects of students' individual engagement. A randomized controlled trial with N=239 students (grades 9-11) compared three self-learning environments covering the same algebraic content - understanding variables and algebraic expressions: a control group with written explanations and practice tasks, a video group with moderately structured prompts, and a video group with highly structured prompts. All groups significantly improved their understanding; however, when accounting for students' individual engagement operationalized through prompt use, both video groups achieved significantly higher effects than the control group. No significant differences emerged between the two video groups. For students with low pretest scores, the control group proved less inclusive. The findings highlight the potential of video-supported self-learning for conceptual understanding, but also indicate that further design improvements are needed to engage all students.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was carried out on stratified pairs of students within each classroom, with all three conditions present in the same class, so the unit of randomisation was below the class level and contamination was not prevented.
      • "Pairs of students were stratified according to Pretest Part I scores on algebra prior understanding and academic language proficiency and then randomly assigned to the three treatment conditions, ensuring that the three conditions were represented in each classroom..."
      • Relevant Quotes: 1) "Pairs of students were stratified according to Pretest Part I scores on algebra prior understanding and academic language proficiency and then randomly assigned to the three treatment conditions, ensuring that the three conditions were represented in each classroom and the groups were comparable in language and algebra during the 85-minute experiment session (including posttest)." (p. 16) 2) "The intervention sample consisted of N = 239 students in Grades 9-11 (ages 14 to 24 years) from 16 classes from four schools; all in their second chance units for understanding algebra." (p. 16) 3) "In the knowledge organization phase, students worked in self-chosen pairs in three treatment conditions with different worked examples and explicit instructions..." (p. 10) 4) "In a 45-minute preparation session, students completed questionnaires on background variables and Pretest Part I, and selected partners for pair work in the experiment session, as the selection of volunteer partners can promote interactive engagement (Chi et al., 2017)." (p. 16) 5) "The study was conducted in their regular mathematics classes with informed consent from students and guardians to use the data for research purposes." (p. 16) Detailed Analysis: Criterion C requires randomisation of entire classes (or schools), so that treatment and control students are not mixed within the same classroom, which would allow contamination between conditions. The paper is explicit that the randomised unit was the student pair, not the class: pairs were stratified on prior algebra understanding and academic language proficiency and then randomly assigned, and the design deliberately ensured that "the three conditions were represented in each classroom". This means students in the control condition and students in both video conditions worked simultaneously in the same room, during the same 85-minute lesson, which is exactly the arrangement the class-level requirement is designed to avoid. Although 16 classes from four schools participated, classes and schools were not the units of allocation; they merely supplied the sample. The tutoring exception does not apply. The intervention is not one-to-one personal tutoring: it is a self-learning environment delivered to self-chosen student pairs working in their regular mathematics classes, with no dedicated tutor per student. The pair setting was chosen to promote peer interactive engagement, not to deliver individualised tutoring, so the standard class-level requirement applies in full. Criterion C is not met because randomisation was performed on student pairs within classrooms rather than on whole classes or schools, and no tutoring exception applies.
    • E

      Exam-based Assessment

      • The conceptual understanding outcome was measured with a ten-item pretest and posttest assembled and adapted by the researchers for this study, not with a recognised standardised exam.
      • "The conceptual understanding pretest included ten unit-weighted items which were adopted and adapted from existing tests on students' conceptual understanding of variables and expressions, covering misconceptions such as letters as specific unknowns (Hodgen et al., 2024; Kuchemann, 1981)."
      • Relevant Quotes: 1) "The conceptual understanding pretest included ten unit-weighted items which were adopted and adapted from existing tests on students' conceptual understanding of variables and expressions, covering misconceptions such as letters as specific unknowns (Hodgen et al., 2024; Kuchemann, 1981). Cognitive labs with students' think-aloud work provided deep insights into their reasoning about the items, so that the construct validity of the items could be optimized (Clark & Watson, 1995)." (p. 12) 2) "The pretest combined two items from the algebra prior understanding test (Pretest Part I in Fig. 2) administered during the preparation session and eight items from the e-scooter exploration tasks (Fig. 1, center) administered during the experiment session (Pretest Part II)." (p. 13) 3) "The conceptual understanding posttest, administered after the knowledge organization phase, consisted of ten structurally similar unit-weighted items (maximum score of 10) with changed contexts, to measure conceptual understanding rather than retention." (p. 13) 4) "For both tests, at least 20% of the data of the items were coded by two independent raters. The determined interrater agreement was substantial (k > 0.70)." (p. 13) 5) "Internal consistency was acceptable (pretest: wtot = 0.74, posttest: wtot = 0.75) and the average inter-item correlation (pretest: r = 0.21, posttest: r = 0.23) fell within the recommended range for scale homogeneity" (p. 13) 6) "Students' academic language proficiency (ALP) in German was measured by a C-Test, a widely used, economical, and valid measure that uses cloze texts without explicit mathematical language." (p. 12) Detailed Analysis: Criterion E requires that educational outcomes be measured with standardised, widely recognised exams rather than instruments built for the study, because researcher-built instruments risk being over-aligned with the intervention. Here the dependent variable - conceptual understanding of variables and expressions - was measured by a ten-item instrument that the authors assembled themselves. Individual items were "adopted and adapted from existing tests" and the authors refined them through their own cognitive labs; this is a bespoke research instrument, not a national, state-wide or otherwise recognised standardised examination. The posttest is described as consisting of structurally similar items with changed contexts, and eight of the ten pretest items are drawn directly from the e-scooter exploration task that also formed the material of the intervention itself, which is precisely the alignment concern Criterion E guards against. Reporting of interrater agreement and omega total demonstrates psychometric care, but internal consistency statistics do not make a test a standardised exam. The one genuinely standardised instrument used, the C-Test of academic language proficiency, was a background/control variable, not the educational outcome, so it cannot satisfy this criterion. Criterion E is not met because the educational outcome was measured with a researcher-constructed conceptual understanding test rather than a standardised exam.
    • T

      Term Duration

      • The whole intervention and outcome measurement took place inside a single 85-minute session, with the posttest administered immediately after the 40-minute knowledge organization phase, far short of one academic term.
      • "In the 85-minute experiment session, students worked through a self-learning environment with a two-phase macro-design (Loibl et al., 2017), consisting of identical exploration phases (20 min) across treatment conditions, three different conditions for the knowledge organization phase (40 min), and the posttest (25 min)."
      • Relevant Quotes: 1) "To examine the research question, we conducted an 85-minute randomized controlled trial with N = 239 ninth to eleventh graders, with conceptual understanding (of variables as generalizers and algebraic expressions as descriptions of general relationships) as the dependent variable." (p. 9) 2) "In the 85-minute experiment session, students worked through a self-learning environment with a two-phase macro-design (Loibl et al., 2017), consisting of identical exploration phases (20 min) across treatment conditions, three different conditions for the knowledge organization phase (40 min), and the posttest (25 min)." (p. 10) 3) "After the knowledge organization phase, students continued individually with the posttest." (p. 12) 4) "Since students completed Pretest Part I in the preparation session 1-3 weeks prior to the experimental session, Posttest Items 1 and 2 (Table 1) were similar to the taxi task..." (p. 13) 5) "Focusing on one meaning of variables and expressions in a single session obviously does not allow students to develop a comprehensive understanding of algebra with its multiple interconnected concepts and meanings." (p. 25) 6) "Hence, the short-term self-learning environment should be embedded in a longer-term curriculum and evaluated for its effectiveness." (p. 25) Detailed Analysis: Criterion T requires that the primary outcome be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Short interventions are acceptable, but the follow-up tracking must extend to a term. In this study, the intervention (the 20-minute exploration phase plus the 40-minute knowledge organization phase) and the outcome measurement (the 25-minute posttest) all fell inside one 85-minute lesson. The interval from intervention start to primary outcome measurement is therefore about one hour. The only earlier contact was a 45-minute preparation session 1-3 weeks before, during which background data and Pretest Part I were collected - this precedes the intervention rather than extending follow-up after it. No delayed or retention posttest of any kind is reported, and the authors themselves characterise the study as a "single session" and "short-term self-learning environment" that ought in future to be embedded in a longer-term curriculum. Criterion T is not met because outcomes were measured immediately within the same 85-minute session in which the intervention took place, nowhere near one academic term.
    • D

      Documented Control Group

      • The control condition is documented in detail, including its size (n=80), the exact learning activities and eight prompts it received, and a full table of baseline demographics, language proficiency and pretest scores with statistical confirmation of baseline equivalence.
      • "Control condition: Student pairs received a written worked example of the exploration task and were asked to explain it to each other and revise their initial responses from the exploration phase."
      • Relevant Quotes: 1) "Control condition: Student pairs received a written worked example of the exploration task and were asked to explain it to each other and revise their initial responses from the exploration phase. Afterwards, the pairs were invited again to revise their exploration solutions and to ensure that all their initial questions were answered." (p. 10) 2) "To practice the learned content, they worked on a similarly structured near-transfer task with a given table and a context similar to the exploration task ... and a far-transfer task without a given table in a very different context (expression to convert the temperature in degrees Fahrenheit for every possible degree Celsius)." (p. 10) 3) "In total, the control condition received eight prompts to support focused cognitive engagement, including the two transfer tasks and a follow-up task that required the students to summarize what they had learned." (p. 10) 4) "Table 2 Descriptive statistics of student variables in three treatment conditions ... Control condition (n = 80) ... Gender (girls in %) 62.5% ... Immigrant background (in %) 55.0% ... Multilingual background (in %) 62.5% ... Low socioeconomic status (in %) 43.8% ... Age in years (M (SD)) 16.9 (1.60) ... Academic language proficiency (M (SD)) 45.5 (10.8) ... Prior conceptual understanding (M (SD)) 3.75 (2.10) ... Post conceptual understanding (M (SD)) 3.94 (2.16)" (p. 20) 5) "The MANOVA showed no differences between conditions (with F(2, 236) = 0.646, p = .826). ANOVAs for each individual variable also revealed no significant group differences (F(2, 236) < 1.265, p > .283). Hence, the three groups started with comparable background variables and comparable prior conceptual understanding." (p. 19) 6) "The descriptive statistics in Table 2 show that, on average, students in the control condition used only 4.38 of the 8 (55%) consolidation prompts..." (p. 19) Detailed Analysis: Criterion D requires the control group to be well documented: its size, demographic composition, baseline performance and the conditions it experienced. All of these elements are present. The size is given exactly (n = 80). The treatment received by the control group is described concretely and in operational detail - a written worked example of the same e-scooter exploration task, mutual peer explanation, an invitation to revise their exploration solutions, a near-transfer water-price task, a far-transfer Fahrenheit-to-Celsius task, and a summarising follow-up task, adding up to eight consolidation prompts. Table 2 reports the control group's gender split, immigrant and multilingual background, socioeconomic status, age, academic language proficiency, engagement scores and both pretest and posttest conceptual understanding means and standard deviations, side by side with the two video conditions. The authors additionally run a MANOVA and per-variable ANOVAs demonstrating baseline equivalence across the three arms. Criterion D is met because the control group's size, demographics, baseline performance and exact instructional conditions are all explicitly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Schools were not randomised at all; the four participating schools and 16 classes merely supplied the sample, and allocation was performed on student pairs within each classroom.
      • "Pairs of students were stratified according to Pretest Part I scores on algebra prior understanding and academic language proficiency and then randomly assigned to the three treatment conditions, ensuring that the three conditions were represented in each classroom..."
      • Relevant Quotes: 1) "The intervention sample consisted of N = 239 students in Grades 9-11 (ages 14 to 24 years) from 16 classes from four schools; all in their second chance units for understanding algebra." (p. 16) 2) "Pairs of students were stratified according to Pretest Part I scores on algebra prior understanding and academic language proficiency and then randomly assigned to the three treatment conditions, ensuring that the three conditions were represented in each classroom and the groups were comparable in language and algebra during the 85-minute experiment session (including posttest)." (p. 16) 3) "Table 2 Descriptive statistics of student variables in three treatment conditions ... Control condition (n = 80) ... Video condition with moderately structured prompts (n = 79) ... Video condition with highly structured prompts (n = 80)" (p. 20) Detailed Analysis: Criterion S requires that the educational institution or implementing unit - the school, centre or site - be the unit of random assignment. No such randomisation occurred here. The paper reports that four schools and 16 classes contributed participants, but allocation happened one level below the class: stratified student pairs were randomly assigned to the three conditions, and the design explicitly required all three conditions to appear inside every classroom. Consequently each of the four schools, and each of the 16 classes, contained control, MSP and HSP students simultaneously. There is no statement anywhere of schools or classes being allocated to conditions, and the resulting arm sizes (80, 79, 80) reflect balanced individual/pair allocation rather than cluster allocation. Criterion S is not met because randomisation occurred at the level of student pairs within classrooms and no school-level assignment took place.
    • I

      Independent Conduct

      • The same team designed the self-learning environment and the instructional videos in their own prior design-research work, and then ran, coded and analysed the trial themselves, with no external evaluator or third-party oversight reported.
      • "Hence, the self-learning environment on variables and expressions in view of this paper was developed in a qualitative design-research study (Korntreff & Prediger, 2026) to tailor these learning opportunities to students' needs."
      • Relevant Quotes: 1) "Hence, the self-learning environment on variables and expressions in view of this paper was developed in a qualitative design-research study (Korntreff & Prediger, 2026) to tailor these learning opportunities to students' needs." (p. 5) 2) "Following the suggestion of Kwon et al. (2011), some of our highly structured prompts include response options (e.g., drag-and-drop textboxes) with elaborative feedback and subsequent deepening prompts." (p. 8) 3) "To promote constructive (explaining the meaning) rather than active ICAP behavior (dragging textboxes), we added audio feedback for incorrect responses (e.g., 'times minute price') (following Kwon et al., 2011) and a subsequent self-explanation prompt encouraging reflection on the difference between typical non-viable and intended descriptions (see Fig. 3)." (p. 12) 4) "Our project addressed this gap by designing a self-learning environment to give students a second chance to understand this crucial part of algebra where schools had previously failed them." (p. 24) 5) "This randomized controlled trial was conducted within the project MuM-Video - Instructional videos as resource for language-responsive mathematics classrooms (financially supported by BMBF-grant 01JD2001A to S. Prediger and M. Altieri by the German Ministry of Education and Research, 2020-25)." (p. 30) 6) "We are grateful to all the teachers, student helpers and participants who made the data collection possible, and we would like to thank Eric Großart and Aylina Westermann for their help with the data analysis." (p. 30) 7) "The authors declare that they have no competing interests." (p. 30) 8) "Our non-finding regarding H2b ... inspired a deeper qualitative analysis of the quality and correctness of student responses in the knowledge organization phase (Korntreff, 2026)." (p. 26) Detailed Analysis: Criterion I requires the evaluation to be conducted independently of the people who designed the intervention, or at minimum to document third-party oversight of data collection, analysis and conclusions. The paper documents the opposite. The self-learning environment under test was developed by two of the authors in their own preceding design-research study (Korntreff & Prediger, 2026), and the paper repeatedly uses first-person designer language about the intervention's micro-design ("some of our highly structured prompts", "we added audio feedback", "Our project addressed this gap by designing a self-learning environment"). The same authors then specified the hypotheses, ran the trial in participating classrooms, coded worksheets and screencasts for consolidation engagement, and performed the regression modelling. Support is acknowledged from teachers, student helpers and two named assistants for data analysis, but these are collaborators working under the author team, not an independent evaluation agency, and there is no statement of blinding of coders to condition or of an external body overseeing analysis or conclusions. The declaration of no competing interests and the public research funding are commendable transparency measures, but the standard asks specifically about separation between intervention designer and evaluator, which is absent here. Internet verification of the author team's publication record confirms the overlap: the same first author's monograph "Enhancing conceptual understanding of variables with videos: Design research and controlled trial considering engagement and conceptual focus" (Korntreff, 2026, Springer Spektrum) presents both the design research and this controlled trial, showing that design and evaluation were carried out within one research programme. Criterion I is not met because the authors both designed the video-based intervention and conducted, coded and analysed its evaluation, with no independent evaluator reported.
    • Y

      Year Duration

      • The entire trial ran within a single 85-minute session with an immediate posttest, so the tracking interval falls drastically short of 75% of an academic year, and Criterion T is also not met.
      • "To examine the research question, we conducted an 85-minute randomized controlled trial with N = 239 ninth to eleventh graders..."
      • Relevant Quotes: 1) "To examine the research question, we conducted an 85-minute randomized controlled trial with N = 239 ninth to eleventh graders, with conceptual understanding ... as the dependent variable." (p. 9) 2) "In the 85-minute experiment session, students worked through a self-learning environment with a two-phase macro-design (Loibl et al., 2017), consisting of identical exploration phases (20 min) across treatment conditions, three different conditions for the knowledge organization phase (40 min), and the posttest (25 min)." (p. 10) 3) "After the knowledge organization phase, students continued individually with the posttest." (p. 12) 4) "In addition, longitudinal studies are needed to examine the effects of curriculum materials that repeatedly integrate conceptually focused self-learning environments." (p. 26) Detailed Analysis: Criterion Y requires outcomes to be measured at least 75% of one academic year (roughly 7 months or more) after the intervention begins. The prompt-specific rule also states that if Criterion T is not met, Criterion Y cannot be met. Criterion T is not met in this paper, which already settles the matter. On the substance, the intervention started and the primary outcome was collected within the same 85-minute lesson: 20 minutes of exploration, 40 minutes of knowledge organization, then a 25-minute posttest. The elapsed interval is about one hour rather than seven to ten months. The authors themselves call for future longitudinal studies, confirming that no year-long tracking was attempted. Criterion Y is not met because the tracking interval was a single 85-minute session and Criterion T is also not met.
    • B

      Balanced Control Group

      • All three conditions received exactly the same instructional time within one 85-minute session and an active, content-matched control that worked through a written worked example of the same task plus transfer practice, so educational time and inputs were balanced.
      • "In the 85-minute experiment session, students worked through a self-learning environment with a two-phase macro-design (Loibl et al., 2017), consisting of identical exploration phases (20 min) across treatment conditions, three different conditions for the knowledge organization phase (40 min), and the posttest (25 min)."
      • Relevant Quotes: 1) "In the 85-minute experiment session, students worked through a self-learning environment with a two-phase macro-design (Loibl et al., 2017), consisting of identical exploration phases (20 min) across treatment conditions, three different conditions for the knowledge organization phase (40 min), and the posttest (25 min)." (p. 10) 2) "Control condition: Student pairs received a written worked example of the exploration task and were asked to explain it to each other and revise their initial responses from the exploration phase." (p. 10) 3) "To practice the learned content, they worked on a similarly structured near-transfer task with a given table and a context similar to the exploration task ... and a far-transfer task without a given table in a very different context..." (p. 10) 4) "In total, the control condition received eight prompts to support focused cognitive engagement, including the two transfer tasks and a follow-up task that required the students to summarize what they had learned." (p. 10) 5) "This condition received nine prompts to support focused cognitive engagement, including follow-up tasks that required summarizing the important content and posing and answering content-related questions with their partner." (p. 11) 6) "This condition included 9 prompts to support focused cognitive engagement and 6 follow-up prompts that required to provide written responses to typical errors..." (p. 12) 7) "A randomized controlled trial with N = 239 students (grades 9-11) compared three self-learning environments covering the same algebraic content - understanding variables and algebraic expressions: a control group with written explanations and practice tasks, a video group with moderately structured prompts, and a video group with highly structured prompts." (p. 1) 8) "This study used a two-phase problem-solving before instruction design (PS-I) across all treatment conditions (Loibl et al., 2017)." (p. 25) 9) "We examined how much structure is necessary to support students' conceptual focus by comparing moderately and highly structured prompts in two video conditions, using prompts of different quality and quantity across conditions (Fig. 3)." (p. 25) Detailed Analysis: Step 1 - Determine study intent. The study compares three ways of delivering the same knowledge organization phase for the same algebraic content within the same fixed time budget. It is not a study of whether extra instructional time or extra resources help; the abstract states all three environments cover "the same algebraic content". Step 2 - Identify intervention resources. The video conditions received a 6:25-minute interactive instructional video with embedded prompts (drag-and-drop textboxes, on-screen text entry, audio feedback, deepening prompts), delivered on digital devices. Step 3 - Determine whether additional resources were provided. Time is exactly matched: every condition had the identical 20-minute exploration phase, a 40-minute knowledge organization phase, and a 25-minute posttest, totalling 85 minutes. The control condition was an active, not a passive or no-treatment, control: it received a written worked example of the very same e-scooter task, peer explanation, an opportunity to revise its exploration solutions, a near-transfer task, a far-transfer task and a summarising follow-up task. The educational substance offered to the control - worked examples plus practice, an approach the authors note is itself "known to enhance conceptual understanding (Schwartz & Martin, 2004)" - is a genuine comparable substitute for the video, not an absence of instruction. The one non-trivial asymmetry is in the number of consolidation prompts (8 in control, 9 in MSP, 15 in HSP). However, prompt count here is the treatment contrast itself - the study's entire research question is how much prompt structure supports focused cognitive engagement - and the prompts occupy the same fixed 40 minutes rather than adding time. The authors are transparent about this design feature, flag it as a limitation, and address it analytically by z-standardising engagement within each condition so that relative rather than absolute prompt use is compared. The marginal material cost (a digitally delivered video versus a printed worked example) is small and integral to the instructional medium under test. Step 4/5 - The differing prompt structure is integral to the intervention being tested rather than a separable confounding add-on, and no group received extra instructional time or a separate budget uplift. Criterion B is met because instructional time was identical across arms and the control was an active, content-equivalent worked-example-plus-practice condition, with the only difference - the number and structure of prompts - being the treatment variable itself.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this trial by a different research team is reported in the paper, and dedicated internet searching found only same-author companion publications rather than any reproduction.
      • Relevant Quotes: 1) "Hence, the self-learning environment on variables and expressions in view of this paper was developed in a qualitative design-research study (Korntreff & Prediger, 2026) to tailor these learning opportunities to students' needs." (p. 5) 2) "Our non-finding regarding H2b (no difference in posttest understanding between moderately and highly structured prompts regardless of engagement) inspired a deeper qualitative analysis of the quality and correctness of student responses in the knowledge organization phase (Korntreff, 2026)." (p. 26) 3) "Yet, only one study reviewed by Wylie and Chi (2014), the study by Kwon et al. (2011), compared two versions of structured prompts..." (p. 8) 4) "Since most studies in Rittle-Johnson et al.'s (2017) meta-analysis were not conducted in classrooms and did not include interactive videos, and few studies compared versions of structured prompts (Kwon et al., 2011), it is unclear if those findings apply to prompts in interactive video aimed at enhancing conceptual understanding in algebra." (p. 8) 5) "Future PS-I studies should compare PS-I and flipped classroom I-PS designs using instructional videos with prompts that encourage active knowledge organization to clarify the merits of this phase structure when using interactive video." (p. 25) Detailed Analysis: Criterion R requires that this specific study, or its central experimental claim in a comparable design, has been independently replicated by a different research team and published in a peer-reviewed journal. The paper reports no such replication. The companion works it cites - Korntreff and Prediger (2026) in the Journal of Mathematical Behavior and the Korntreff (2026) Springer Spektrum monograph - are by the same author team and are the design-research and qualitative-analysis siblings of this very trial, not independent reproductions. Internet verification was carried out for this criterion. Searches of the publisher record, Google-style web search, and the author group's own publication lists at TU Dortmund University returned only publications by the same team: Korntreff, Bach, Altieri and Prediger (2026) for this trial; Korntreff and Prediger (2026), "How can students' engagement with instructional videos on generalizing with algebraic expressions be scaffolded? Design research for specifying the content-specific focus of scaffolds", Journal of Mathematical Behavior 82, Article 101317; Korntreff (2026), "Enhancing conceptual understanding of variables with videos: Design research and controlled trial considering engagement and conceptual focus", Springer Spektrum; and two CERME14 conference papers (2025) by the same group. No paper by any other research team attempting to reproduce this trial's design or findings could be found; no verbatim quotes from any replication study can therefore be provided. The article was only published online in 2026, so an independent replication could not plausibly have appeared yet. Criterion R is not met because no independent replication of this trial by a different research team has been published, and none was found by internet search.
    • A

      All-subject Exams

      • Only a narrow slice of mathematics - conceptual understanding of variables as generalizers and algebraic expressions - was assessed, with a researcher-built instrument, so neither the all-subject coverage nor the Criterion E prerequisite is satisfied.
      • "Focusing on one meaning of variables and expressions in a single session obviously does not allow students to develop a comprehensive understanding of algebra with its multiple interconnected concepts and meanings."
      • Relevant Quotes: 1) "To examine the research question, we conducted an 85-minute randomized controlled trial with N = 239 ninth to eleventh graders, with conceptual understanding (of variables as generalizers and algebraic expressions as descriptions of general relationships) as the dependent variable." (p. 9) 2) "While conceptual understanding of algebra is acquired over many years of schooling ... our randomized controlled trial focuses on only two particular aspects: variables as generalizers and expressions for describing general relationships." (p. 3) 3) "The conceptual understanding pretest included ten unit-weighted items which were adopted and adapted from existing tests on students' conceptual understanding of variables and expressions..." (p. 12) 4) "Focusing on one meaning of variables and expressions in a single session obviously does not allow students to develop a comprehensive understanding of algebra with its multiple interconnected concepts and meanings." (p. 25) 5) "Students' academic language proficiency (ALP) in German was measured by a C-Test ... It tests both, vocabulary and grammar knowledge of elaborated language, in a comprehensive way (Grotjahn et al., 2002)." (p. 12) Detailed Analysis: Criterion A requires impact to be measured across all main subjects of the curriculum, using standardised exam-based assessments, and explicitly fails whenever Criterion E fails. Criterion E is not met here, so Criterion A cannot be met. Substantively, the outcome measurement is narrower still than a single subject: it covers two specific conceptual elements within elementary algebra (variables as generalizers, and expressions as descriptions of general relationships), which the authors themselves acknowledge is "just one part of understanding algebra among others". No outcomes were collected in German, science, social studies, foreign languages or any other main subject taught at Grades 9-11. The German C-Test does touch language skills, but it is a background covariate measured before the intervention to control for academic language proficiency, not an outcome assessing impact on language attainment. No specialised-intervention exception applies: this is general secondary-school mathematics remediation, not vocational or highly specialised upper-secondary training for which restricted outcome measurement could be justified. Criterion A is not met because only a narrow area of algebraic conceptual understanding was assessed, with a non-standardised instrument, and Criterion E is not met.
    • G

      Graduation Tracking

      • Measurement ended with the posttest inside the same 85-minute session, no follow-up publication by the same authors tracks the cohort to graduation, and Criterion Y is also not met.
      • "After the knowledge organization phase, students continued individually with the posttest."
      • Relevant Quotes: 1) "After the knowledge organization phase, students continued individually with the posttest." (p. 12) 2) "In the 85-minute experiment session, students worked through a self-learning environment ... consisting of identical exploration phases (20 min) across treatment conditions, three different conditions for the knowledge organization phase (40 min), and the posttest (25 min)." (p. 10) 3) "Hence, the short-term self-learning environment should be embedded in a longer-term curriculum and evaluated for its effectiveness." (p. 25) 4) "In addition, longitudinal studies are needed to examine the effects of curriculum materials that repeatedly integrate conceptually focused self-learning environments." (p. 26) Detailed Analysis: Criterion G requires participants to be followed until they graduate from the relevant educational stage, and the prompt-specific rule states that if Criterion Y is not met, Criterion G is not met either. Criterion Y is not met, which is decisive. On the substance, the study collected its final data point - the 25-minute posttest - immediately after the knowledge organization phase within the same lesson. There is no delayed posttest, no end-of-year measurement, no linkage to school records, and no reported plan to follow the Grade 9-11 cohort to the end of secondary schooling. Internet verification was carried out for this criterion, as graduation tracking can appear in later papers by the same authors. The publication lists of Stefan Korntreff and Susanne Prediger at TU Dortmund University were checked together with publisher records. The subsequent same-author outputs are Korntreff and Prediger (2026), "How can students' engagement with instructional videos on generalizing with algebraic expressions be scaffolded?", Journal of Mathematical Behavior 82, Article 101317; Korntreff (2026), "Enhancing conceptual understanding of variables with videos: Design research and controlled trial considering engagement and conceptual focus", Springer Spektrum; and CERME14 conference papers (2025) by Korntreff and Prediger and by Bach, Altieri, Korntreff and Prediger. All of these concern the design research and the qualitative analysis of responses from the same single session; none reports any longitudinal follow-up or graduation outcomes for the participating cohort, so no verbatim quote evidencing graduation tracking can be provided. Criterion G is not met because no follow-up beyond the immediate posttest exists in this paper or in any subsequent publication by the same authors, and Criterion Y is also not met.
    • P

      Pre-Registered

      • The paper contains no reference to a trial registry entry, pre-registered protocol, or registration date preceding data collection, and no registry record was found online.
      • Relevant Quotes: 1) "Data Availability The statistical data are available upon request from the corresponding author. The teaching material is available in German and English here: sima.dzlm.de/um/8-002." (p. 30) 2) "Declarations Conflict of interest The authors declare that they have no competing interests." (p. 30) 3) "This randomized controlled trial was conducted within the project MuM-Video - Instructional videos as resource for language-responsive mathematics classrooms (financially supported by BMBF-grant 01JD2001A to S. Prediger and M. Altieri by the German Ministry of Education and Research, 2020-25)." (p. 30) 4) "Thus, we conducted an a priori power analysis for a two-way mixed ANOVA with three treatment conditions as the between-subjects factor and two times of measurement (pretest and posttest) as the within-subjects factor, assuming an effect size of partial eta2 = 0.02." (p. 16) 5) "After receiving unexpected hypothesis testing results, we investigated an exploratory question to examine the role of prior understanding, a known predictor of both learning outcomes and learning gains (Kuhlmann et al., 2024)." (p. 17) Detailed Analysis: Criterion P requires a publicly pre-registered full protocol - hypotheses, methods and planned analyses - lodged before data collection began, with a verifiable registry reference and date. The paper provides none. The Declarations, Data Availability and Funding sections mention data on request, a teaching-material URL and a BMBF grant number, but no ClinicalTrials.gov, OSF, AEA, ISRCTN, DRKS or comparable registry identifier, and no registration date. The authors do report an a priori power analysis and state hypotheses H1 and H2a-b in advance of the results, which is good practice, but an a priori power calculation described inside the results paper is not a public pre-registration and cannot be verified as having preceded data collection. Internet verification was carried out for this criterion. Because the paper cites no registry identifier, searches were run against the article record and the MuM-Video project name for any linked pre-registration; no registry entry for this trial was found in any registry, so there is no registration date to compare against the start of data collection. The paper is also candid that part of the analysis was generated after seeing the results - the exploratory question EQ was formulated "after receiving unexpected hypothesis testing results" and the model selection proceeded through a sequence of nested models (Models 0 to 6b) chosen partly in light of observed p-values. This post-hoc model exploration is exactly the practice pre-registration is meant to constrain, and it is transparently flagged as exploratory, but its presence underlines the absence of a binding pre-specified analysis plan. Criterion P is not met because no pre-registration reference or date is provided anywhere in the paper and no registry record could be located online.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.