Abstract
Background: Mathematical structuring competence (MSC)—the ability to recognize, apply, and build patterns and structures—is a crucial component of mathematics education and an essential characteristic of mathematical talent. However, systematic approaches to fostering MSC, especially for talent promotion, are still lacking. Aims: We evaluated the efficacy of a newly developed intervention for enhancing talented students' MSC as a primary outcome and their arithmetic performance and mathematics-related motivational dispositions as secondary outcomes. The intervention was conducted online and involved two core components: inquiry-based learning and mathematically rich pattern tasks. Sample: We collected data from 104 talented 2nd to 5th graders (32 girls) participating in an extracurricular STEM enrichment program. Methods: We used a randomized controlled field trial with repeated measures and a treated control group to evaluate the efficacy of the intervention. Multiple linear regression analyses were calculated to test for intervention effects. Results: Students in the intervention group developed more sophisticated MSC than those in the control group (β = 0.46, p < .001). No differential intervention effects on MSC were found, indicating that all students, regardless of their prior knowledge, fluid intelligence, gender, or grade level, benefited equally from the intervention. The intervention did not significantly affect arithmetic performance or motivational dispositions. Conclusions: The study showed that talented students' MSC can be promoted by combining inquiry-based learning with mathematically rich pattern tasks.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was conducted at the individual student level (not class or school), and the intervention was a small-group, not one-to-one, design, so the criterion is not met.
- Within each block, students were first randomly assigned to the intervention or control condition. Subsequently, they were randomly and evenly assigned to one of the available courses within their group.
Relevant Quotes:
1) "A total of 104 students (32 girls) attending the Hector Children's Academy Program (HCAP) participated in the intervention." (p. 4)
2) "were randomly assigned to one of two online courses: the MSC intervention or the control condition, a computational thinking course." (pp. 4-5)
3) "Within each block, students were first randomly assigned to the intervention or control condition. Subsequently, they were randomly and evenly assigned to one of the available courses within their group." (p. 5)
4) "In total, six intervention courses and eight control courses were conducted, with course sizes ranging from four to twelve students (M = 7.36, SD = 2.06)." (p. 5)
Detailed Analysis:
Criterion C requires randomisation at the class level or stronger (school level), to avoid contamination within a single classroom. Here, the unit of randomisation was the individual student: students registered in the HCAP summer program were first randomly assigned to the intervention or control condition, and then randomly distributed across the available courses within their assigned condition. The courses were small groups (4 to 12 students) newly formed for the study, not intact classes or schools.
The exception for personal/one-to-one tutoring does not apply, because the intervention was delivered to small groups using a think-pair-share collaborative approach rather than individual tutoring. Since randomisation was performed at the student level (assigning individuals to conditions and then to groups), the class-level requirement is not satisfied.
Criterion C is not met because randomisation was carried out at the individual student level rather than at the class or school level, and no tutoring exception applies.
-
E
Exam-based Assessment
- The primary outcome (MSC) was measured with a researcher-adapted indicator exercise test rather than a standardized, widely recognized exam.
- "To measure MSC, we adapted Ehrlich's (2015) indicator exercise test to the target group and purpose of the investigation. The test was shortened by one task ... most open response fields (except four) were replaced by several multiple-choice items." (p. 6)
Relevant Quotes:
1) "To measure MSC, we adapted Ehrlich's (2015) indicator exercise test to the target group and purpose of the investigation. The test was shortened by one task so that the test could be completed in 45 min. Items were modified so that answers could be given on the computer and coded as dichotomous variables ... most open response fields (except four) were replaced by several multiple-choice items. The test was divided into two test versions and extended by parallel tasks for the measurements at pretest and posttest." (p. 6)
2) "We assessed arithmetic performance with the eight items of the scala on arithmetic word problems from the German Mathematics Test for fourth grade for the pretest (Gölitz et al., 2006) and with the 15 items of the subtest on arithmetic word problems for fifth grade for the posttest (Götz et al., 2013)." (p. 7)
3) "First, we assessed arithmetic performance using subscales of the DEMAT (Gölitz et al., 2006; Götz et al., 2013), a standardized achievement test widely used and psychometrically validated in Germany." (pp. 9-10)
Detailed Analysis:
Criterion E requires that the study's outcomes be measured with a standardized, widely recognized exam rather than an instrument specially designed for the study. The primary outcome of this trial is MSC, which the intervention was specifically designed to improve. MSC was assessed using a researcher-adapted version of Ehrlich's (2015) indicator exercise test, which the authors shortened, restructured into multiple-choice items, split into two versions, and extended with self-created parallel tasks. This is a custom/adapted research instrument, not a standardized exam. While a genuinely standardized test (the DEMAT 4 and DEMAT 5+) was used, it was applied only to the secondary outcome of arithmetic performance, on which no intervention effect was found. The core measure of the intervention's targeted effect therefore rests on a non-standardized, study-specific instrument.
Criterion E is not met because the primary outcome (MSC) was measured with a researcher-adapted, study- specific instrument rather than a standardized, widely recognized exam.
-
T
Term Duration
- The intervention lasted one week and the posttest was given immediately afterwards with no follow-up, far short of one academic term.
- Finally, we were able to observe short-term effects by administering the posttest immediately after the intervention. Because no follow-up assessment was implemented, we did not gain insights into long-term effects.
Relevant Quotes:
1) "The intervention was conducted as a 1-week online course within a broad HCAP 2023 summer break program..." (p. 4)
2) "Each course took place over 5 days." (p. 5)
3) "In the first and last online sessions of each intervention block, students in both conditions took the pretest and posttest, respectively, in a joint 2-hr testing session..." (p. 5)
4) "Finally, we were able to observe short-term effects by administering the posttest immediately after the intervention. Because no follow-up assessment was implemented, we did not gain insights into long-term effects." (p. 11)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins, allowing for term-long follow-up tracking even for short interventions. Here, the intervention lasted only one week (5 days), and the posttest was administered immediately after the intervention in the last session of the same week. The authors explicitly state that no follow-up assessment was implemented.
The interval between intervention start and outcome measurement was therefore about one week, far shorter than one academic term, and there was no term-long follow-up tracking.
Criterion T is not met because outcomes were measured immediately after a one-week intervention, with no term-long follow-up tracking.
-
D
Documented Control Group
- The control group's size, demographics, baseline and posttest performance, and the computational thinking course it received are all clearly documented in the text and tables.
- Students assigned to the control condition attended a course aimed at promoting computational thinking...
Relevant Quotes:
1) "Students assigned to the control condition attended a course aimed at promoting computational thinking..." (p. 5)
2) "Table 1. Composition of the sample." showing Control N = 51, Boys 33, Girls 18, with grade-level breakdown. (p. 4)
3) "Table 2. Descriptive statistics for all variables." reporting pretest and posttest means and SDs separately for the control group (CG) on MSC, arithmetic performance, self-concept, intrinsic value, attainment value, math grade, German grade, and fluid intelligence. (p. 7)
4) "Table 4. Baseline equivalence between the intervention and control groups." with Hedges' g effect sizes and p values for all baseline characteristics. (p. 8)
Detailed Analysis:
Criterion D requires the control group to be well-documented, including demographic information, baseline performance, and the treatment received. The paper provides detailed documentation of the control group: Table 1 reports its size and the gender and grade-level composition; Table 2 reports its baseline and posttest means and standard deviations across all outcomes and covariates; and Table 4 reports baseline equivalence statistics. The control condition's treatment (a computational thinking / Python programming course) is described in detail in Section 2.4.
This level of detail allows readers to confirm the control group's characteristics and baseline comparability and to understand exactly what the control group received.
Criterion D is met because the control group's size, demographics, baseline performance, and the treatment it received are all clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred at the individual student level rather than at the school or site level, so the school-level RCT criterion is not met.
- Within each block, students were first randomly assigned to the intervention or control condition. Subsequently, they were randomly and evenly assigned to one of the available courses within their group.
Relevant Quotes:
1) "were randomly assigned to one of two online courses: the MSC intervention or the control condition, a computational thinking course." (pp. 4-5)
2) "Within each block, students were first randomly assigned to the intervention or control condition. Subsequently, they were randomly and evenly assigned to one of the available courses within their group." (p. 5)
3) "The HCAP is an enrichment program with 69 locations in the German federal state of Baden-Württemberg and offers numerous enrichment courses, mainly focusing on STEM education..." (p. 4)
Detailed Analysis:
Criterion S requires randomisation at the school level, i.e., entire educational institutions or implementation units randomly assigned to conditions. In this study, randomisation occurred at the level of individual students who registered for the HCAP summer program, who were assigned to conditions and then to small online courses. There is no indication that any school, site, or HCAP location was the unit of randomisation; students were drawn into online courses and assigned individually.
Because the randomisation unit was the individual student rather than the school or implementation site, the school-level RCT requirement is not satisfied.
Criterion S is not met because randomisation was conducted at the individual student level, not at the school or site level.
-
I
Independent Conduct
- The intervention was developed, conducted, and analysed by the same research team (the authors), with no independent third-party evaluator, so the criterion is not met.
- The courses were developed by researchers from the fields of math education (intervention condition) and computer science education (control condition) and implemented by trained course instructors.
Relevant Quotes:
1) "The courses were developed by researchers from the fields of math education (intervention condition) and computer science education (control condition) and implemented by trained course instructors." (p. 5)
2) "Therefore, to systematically support and enhance MSC in mathematically talented second to fifth graders, we developed an online intervention based on inquiry-based learning (Dorier & Maass, 2020) and mathematically rich pattern tasks (Boaler, 2016)..." (p. 2)
3) "Judith Havemann: Writing – original draft, Visualization, Software, Resources, Project administration, Methodology, Investigation, Formal analysis, Data curation, Conceptualization." (p. 11)
4) "Declarations of interest. none." (p. 12)
Detailed Analysis:
Criterion I requires the study to be conducted independently from the authors who designed the intervention, to reduce bias in implementation and analysis. Here, the intervention (the MSC course) was developed by the same research team (math education researchers / the authors), who also designed the study, conducted the investigation, and performed the formal analysis, as documented in the CRediT statement. The control condition was likewise developed by the related research group (adapted from Kunz et al., 2023, one of the co-authors). Trained instructors delivered the courses, but the design, implementation oversight, and data analysis were carried out by the intervention developers themselves.
There is no statement of an independent third-party evaluation team conducting data collection, analysis, or formulation of conclusions; the research assistants who administered tests were part of the same project team. Therefore, the study does not meet the independence requirement.
Criterion I is not met because the same team that developed the intervention also designed, conducted, and analysed the study, with no independent third-party evaluation.
-
Y
Year Duration
- Outcomes were measured about one week after the intervention began, far short of an academic year, and criterion T was not met, so Y is not met.
- The intervention was conducted as a 1-week online course within a broad HCAP 2023 summer break program...
Relevant Quotes:
1) "The intervention was conducted as a 1-week online course within a broad HCAP 2023 summer break program..." (p. 4)
2) "Each course took place over 5 days." (p. 5)
3) "Because no follow-up assessment was implemented, we did not gain insights into long-term effects." (p. 11)
Detailed Analysis:
Criterion Y requires outcomes to be measured at least 75% of a full academic year (~9-10 months) after the intervention begins. The intervention here lasted one week, with the posttest administered immediately afterwards and no follow-up. The tracking interval was therefore about one week, far short of an academic year.
In addition, per the prompt's criterion-specific instruction, if criterion T (Term Duration) is not met, then criterion Y is not met. Criterion T is not met here, so criterion Y is automatically not met as well.
Criterion Y is not met because the study tracked outcomes for only about one week, far short of an academic year, and criterion T was not met.
-
B
Balanced Control Group
- The study used an active treated control group (a parallel computational thinking course of comparable planned instructional time and structure), so educational time and resources were broadly balanced across conditions.
- were randomly assigned to one of two online courses: the MSC intervention or the control condition, a computational thinking course.
Relevant Quotes:
1) "were randomly assigned to one of two online courses: the MSC intervention or the control condition, a computational thinking course." (pp. 4-5)
2) "It consisted of eight 75-min units covering the introduction and three MSC topics (magic squares, Josephus problem, sequences of numbers and figures) within 1 week." (p. 5)
3) "It included eight 85-min online units introducing the students to Python..." (control condition, pp. 5-6)
4) "In both conditions, course instructors received 2 hr of training from the respective course developer 1 week in advance." (p. 6)
5) "Notably, despite lower adherence in the control group, the quality of delivery was comparable across both conditions... Nevertheless, the reduced dosage and incomplete delivery in the control condition restrict the conclusions..." (p. 11)
Detailed Analysis:
Criterion B asks whether both conditions received comparable time, budget, and resources, so that the specific intervention effect can be isolated, unless additional resources are explicitly the treatment variable. This study used a treated (active) control group: rather than a no-treatment control, students in the control condition received an equivalent-dose computational thinking course (eight 85-min units) delivered in parallel over the same one-week period, while the intervention group received eight 75-min MSC units. Both conditions involved comparable online instructional time, trained instructors (each receiving 2 hr of training), and small-group structure.
Applying the decision-tree logic: extra resources (course time, materials, instructor support) were present in the intervention, but the control group received a comparable, parallel active course with broadly equivalent planned time and structure (eight units of similar length). Thus the control matches the intervention's time and resource inputs, isolating the specific MSC content. The authors note lower implementation adherence in the control condition, but the planned dosage and the quality of delivery were comparable, and the design intent was an equivalent active treated control.
Criterion B is met because the study used a treated control group (a parallel computational thinking course of comparable planned instructional time and structure), so educational time and resources were broadly balanced across conditions.
-
Level 3 Criteria
-
R
Reproduced
- The study is described as the first experimental evidence of its kind, and no independent peer-reviewed replication of this specific intervention was found.
- Taken together, our results contribute to the growing body of research on MSC by providing the first experimental field evidence on how to foster MSC in mathematically talented primary school students.
Relevant Quotes:
1) "However, to the best of our knowledge, no systematic approaches or empirical studies exist on promoting MSC in general or particularly in talented students..." (p. 2)
2) "Taken together, our results contribute to the growing body of research on MSC by providing the first experimental field evidence on how to foster MSC in mathematically talented primary school students." (p. 9)
3) "However, whereas this controlled implementation provides valuable insights into the intervention's efficacy, we do not know whether similar effects can be replicated in more varied, real-world settings." (p. 11)
Detailed Analysis:
Criterion R requires that the specific study be independently replicated by a different research team in a different context, published in a peer-reviewed journal. The authors explicitly describe their study as providing "the first experimental field evidence" on fostering MSC in mathematically talented students and note that no prior systematic intervention studies exist for this population. They themselves state that it remains unknown whether the effects can be replicated in other settings.
An external internet search (ScienceDirect, the Tübingen FDAT data repository, ResearchGate, and general web sources) for independent replications of this specific newly developed MSC online intervention (the think-pair-share inquiry-based pattern-task program for talented students) did not identify any peer-reviewed replication by a different team. The study was published in 2026 and the related PASMAP/early-algebra studies cited in the paper concern different interventions and populations and do not replicate this specific study. No replication quotes are available because no replication study was found.
Criterion R is not met because this is described as the first experimental study of its kind and no independent replication of this specific intervention was found.
-
A
All-subject Exams
- Only mathematics-domain outcomes (MSC and arithmetic) were measured; no other core school subjects were assessed with standardised exams, so the criterion is not met.
- We evaluated the efficacy of a newly developed intervention for enhancing talented students' MSC as a primary outcome and their arithmetic performance and mathematics-related motivational dispositions as secondary outcomes.
Relevant Quotes:
1) "We evaluated the efficacy of a newly developed intervention for enhancing talented students' MSC as a primary outcome and their arithmetic performance and mathematics-related motivational dispositions as secondary outcomes." (p. 1)
2) "To measure MSC, we adapted Ehrlich's (2015) indicator exercise test to the target group and purpose of the investigation." (p. 6)
3) "We assessed arithmetic performance with the eight items of the scala on arithmetic word problems from the German Mathematics Test for fourth grade for the pretest... and with the 15 items of the subtest on arithmetic word problems for fifth grade for the posttest..." (p. 7)
Detailed Analysis:
Criterion A requires that the study measure impact on all main subjects taught at that educational level, not just the intervention subject, using standardised exam-based assessments. This study measured only mathematics-domain outcomes: mathematical structuring competence (the primary outcome) and arithmetic performance, plus mathematics-related motivational dispositions. No assessment was made of other core primary-school subjects such as reading/language, science, or other curriculum areas (German and math grades were collected only as covariates, not as standardised outcome exams).
Because outcomes were confined to the mathematics domain and other main subjects were not assessed with standardised exams, the all-subject coverage requirement is not satisfied. The specialised-intervention exception (upper secondary/vocational) does not apply to this primary-school enrichment context.
Criterion A is not met because only mathematics-domain outcomes were assessed; no other core subjects were measured with standardised exams.
-
G
Graduation Tracking
- No follow-up beyond the immediate posttest was conducted, no graduation-tracking follow-up paper was found, and criterion Y was not met.
- Because no follow-up assessment was implemented, we did not gain insights into long-term effects.
Relevant Quotes:
1) "Because no follow-up assessment was implemented, we did not gain insights into long-term effects." (p. 11)
2) "Finally, we were able to observe short-term effects by administering the posttest immediately after the intervention." (p. 11)
3) "Hence, future research should adopt a more longitudinal approach to address sustainability questions." (p. 11)
Detailed Analysis:
Criterion G requires the study to follow up and track participants until their graduation from the relevant educational stage. The authors explicitly state that no follow-up assessment was implemented and that only short-term effects (posttest immediately after the one-week intervention) were observed. There is no tracking of students through to graduation.
An external internet search for subsequent or follow-up publications by the same authors (Havemann, Jaggy, Kunz, Trautwein, Paravicini) tracking this HCAP cohort to graduation did not identify any such paper; no graduation tracking quotes are available because no such follow-up study was found. In addition, per the prompt's criterion-specific instruction, if criterion Y (Year Duration) is not met, then criterion G is not met. Criterion Y is not met here, so criterion G is automatically not met as well.
Criterion G is not met because the authors explicitly conducted no follow-up beyond the immediate posttest, no graduation-tracking follow-up paper was found, and criterion Y was not met.
-
P
Pre-Registered
- The authors explicitly state that the study and analysis plan were not preregistered, so the criterion is not met.
- The study and analysis plan were not preregistered.
Relevant Quotes:
1) "The study and analysis plan were not preregistered." (p. 4)
2) "Research data and analysis codes are available at https://doi.org/10.57754/FDAT.n4w8r-z9873." (p. 4)
Detailed Analysis:
Criterion P requires that the full study protocol be pre-registered (with hypotheses, methods, and planned analyses) before data collection begins. The paper explicitly and unambiguously states, "The study and analysis plan were not preregistered." While the authors share data and analysis code in an open repository (the Tübingen FDAT repository), this is post-hoc data sharing, not pre-registration of a protocol before data collection. Because the paper explicitly states the study was not preregistered, there is no registry entry to verify and no registration date to check.
Because there is an explicit statement that the study was not preregistered, the criterion fails.
Criterion P is not met because the authors explicitly state that the study and analysis plan were not preregistered.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.