Abstract
This pre-registered randomized controlled trial investigated whether generative AI (GenAI) tools can support secondary school students' self-regulated learning (SRL) when integrated into regular classroom lessons. A total of 371 students (Grades 7-9) were randomly assigned at the individual level to one of three conditions: (a) GenAI-supported utility value reflection, (b) GenAI-supported cognitive learning strategy prompting, or (c) a control condition (standard ChatGPT use without pedagogical prompting). Students participated in six 45-minute sessions during regular physics or English lessons (one pretest, four learning sessions, one posttest). All conditions used GPT-4o-powered systems that differed only in theory-informed prompting. Utility value prompts helped preserve students' perceived utility value of the learning content, but no clear condition differences emerged for other learning outcomes. Exploratory analyses suggested meaningful GPT interactions buffered motivational declines. The study offers ideas for scalable, classroom-integrated SRL support using GenAI.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomization was performed at the individual student level within shared classes rather than at the class or school level, so the class-level RCT criterion is not met.
- "Students were randomly assigned at the individual level, and the assignments were implemented through the system's backend without disclosing them to teachers or students before the intervention."
Relevant Quotes:
1) "We conducted an RCT with a between-subjects design across two academic subjects (i.e., physics and English). Students were randomly assigned to one of three conditions: (a) GenAI-supported utility value reflection, (b) GenAI-supported cognitive learning strategy prompting, or (c) control condition (standard ChatGPT use without pedagogical prompting)." (Study Design, p. 12)
2) "Simple randomization was implemented within the IWM Author learning platform (v3.014) using a computer-generated allocation sequence. Students were randomly assigned at the individual level, and the assignments were implemented through the system's backend without disclosing them to teachers or students before the intervention." (Study Design, p. 12)
3) "No blocking or stratification was applied, and slight variations in group sizes resulted from randomization." (p. 12)
4) "The intervention was conducted simultaneously across all participating classes and conditions." (p. 14)
Detailed Analysis:
The ERCT C criterion requires randomisation at the class level (or stronger school level), unless the intervention is a personal/one-to-one tutoring intervention in which case student-level randomisation is allowed. This study explicitly randomised individual students to the three conditions, and the three conditions co-existed within the same classes simultaneously ("conducted simultaneously across all participating classes and conditions"). This is a student-level RCT within shared classrooms, which is precisely the design the C criterion is meant to exclude because of contamination risk among students in the same room. The exception for personal tutoring does not apply: although each student worked individually with a chatbot, this was a classroom-integrated group lesson with teachers present, not a one-to-one tutoring intervention designed for personal teaching, and the three conditions were deliberately mixed within the same classes.
Criterion C is not met because randomisation was at the individual student level within shared classes, with no class-level or school-level randomisation and no valid tutoring exception.
-
E
Exam-based Assessment
- Outcomes were measured with self-report scales and custom, curriculum-aligned domain-knowledge items developed by the study teachers, not a widely recognised standardised exam.
- "Students' domain-specific knowledge was assessed using three subject-specific test items ... tailored to the content covered in the interventions."
Relevant Quotes:
1) "Our primary outcome for the utility value intervention condition was students' perceived utility value of the learning content for daily life, which was assessed using a self-report scale comprising three items ... translated from Gaspard et al., 2015." (Measures, p. 16-17)
2) "Our primary outcome for the cognitive learning condition was students' self-reported use of the elaboration-based cognitive learning strategy, which was assessed with a three-item self-report scale ... translated from Klingsieck, 2018." (Measures, p. 17)
3) "Students' domain-specific knowledge was assessed using three subject-specific test items (e.g., single-choice, matching, multiple-choice, fill-in-the-blank) tailored to the content covered in the interventions." (Measures, p. 18)
4) "For each academic subject, experienced in-service teachers developed curriculum-aligned test items and instructional materials, which were subsequently reviewed, discussed, and refined ... The items were rated by the same subject teachers who had developed the tests, tasks, and instructional materials." (p. 18)
5) "Students' cognitive ability was assessed at T2 using the KFT 4-12+R cognitive ability test (Heller & Perleth, 2000)." (p. 20-21)
Detailed Analysis:
Criterion E requires a standardised, widely recognised exam-based assessment that was not custom-built for the study. The principal learning outcomes here are self-report scales (perceived utility value, self-reported strategy use, maintained interest, effort), which are not exams at all. The closest exam-based measure is the domain-specific knowledge test, but it was explicitly developed by the study teachers, tailored to the intervention content, and limited to a few items; it is a researcher/teacher-made instrument rather than a standardised national or state-wide curriculum exam. The KFT is a standardised cognitive-ability test, but it measures reasoning ability (used as a control/baseline variable), not a standardised educational achievement exam for the intervention subjects.
Criterion E is not met because the educational outcomes were measured with self-report scales and custom, teacher-developed knowledge items rather than a recognised standardised exam.
-
T
Term Duration
- The intervention and outcome measurement spanned only six 45-minute sessions over a few weeks, far shorter than one academic term.
- "The study spanned six school sessions (of 45 min each) during regular class time: one pretest session (T1), four consecutive learning sessions, and one posttest session (T2)."
Relevant Quotes:
1) "The study spanned six school sessions (of 45 min each) during regular class time: one pretest session (T1), four consecutive learning sessions, and one posttest session (T2)." (Common Design Features, p. 14)
2) "The two measurements and interventions took place in April and May 2025." (p. 12)
3) "At T2, which was scheduled directly after the intervention, the participating students were asked to provide sociodemographic information ... In addition, they were tested on their domain-specific knowledge." (Data Collection, p. 16)
4) "all outcome measures were collected immediately after the intervention, which does not allow conclusions about whether the observed effects extend beyond the immediate learning situation; future research should therefore include delayed follow-up assessments." (Limitations, p. 46)
Detailed Analysis:
Criterion T requires outcomes to be measured at least one full academic term (~3-4 months) after the intervention begins. Here the entire study consisted of six 45-minute sessions, with the posttest scheduled directly after the four learning sessions; the whole sequence took place within April-May 2025. The authors themselves note that all outcomes were collected immediately after the intervention. There is no term-long follow-up tracking from intervention start to outcome measurement.
Criterion T is not met because the interval from intervention start to outcome measurement was a few weeks of short sessions, well short of one academic term.
-
D
Documented Control Group
- The control group's size, demographics, baseline characteristics, and "standard ChatGPT" condition are documented in detail, including baseline-equivalence testing.
- "Students were randomly assigned to one of three conditions: ... (c) control condition (standard ChatGPT use without pedagogical prompting)."
Relevant Quotes:
1) "control condition (standard ChatGPT use without pedagogical prompting)." (Study Design, p. 12)
2) "Control Condition. The GPT engaged in light conversation about the task, but refrained from offering any pedagogical strategies. It acknowledged the student's input but did not encourage reflection, elaboration, or the use of strategies." (p. 13)
3) "Condition n Female Drop-outs (n) % ... Control condition 122 32.8 42 53.3 Total 371 ... Table 1 Descriptive Statistics for the Sample per Condition." (p. 12)
4) "Following the What Works Clearinghouse (WWC) Standards (2022), we examined baseline equivalence for these variables. ... Table 4 Baseline differences in potential control variables ... co (control) values reported per variable." (p. 20-21)
Detailed Analysis:
Criterion D requires detailed documentation of the control group's composition, size, baseline performance, and the conditions it experienced. The paper reports the control group's sample size (n = 122), gender composition, and dropout, gives baseline equivalence statistics for the control versus treatment arms across numerous variables (cognitive ability, conscientiousness, openness, academic self-concept, dispositional interest, growth mindset, AI literacy) in Table 4, and explicitly describes the control treatment (standard ChatGPT use without pedagogical prompting). This is comprehensive documentation.
Criterion D is met because the control group's characteristics, size, baseline equivalence, and the exact nature of its (non-pedagogical) treatment are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation was at the individual student level, not at the school level, so the school-level RCT criterion is not met.
- "Students were randomly assigned at the individual level, and the assignments were implemented through the system's backend."
Relevant Quotes:
1) "Simple randomization was implemented within the IWM Author learning platform (v3.014) using a computer-generated allocation sequence. Students were randomly assigned at the individual level." (p. 12)
2) "No blocking or stratification was applied, and slight variations in group sizes resulted from randomization." (p. 12)
3) "The intervention was conducted simultaneously across all participating classes and conditions." (p. 14)
Detailed Analysis:
Criterion S requires that entire schools (or equivalent implementing units) be the unit of randomisation. In this study the unit of randomisation was the individual student, with all three conditions running concurrently within the same classes and schools. There is no school-level (or even class-level) randomisation.
Criterion S is not met because randomisation occurred at the individual student level rather than at the school level.
-
I
Independent Conduct
- The same author team designed the GenAI intervention, developed the prompts and measures, and conducted and analysed the study, with no independent third-party evaluator.
- "All prompts were developed through an intensive prompt-engineering process involving multiple iterative feedback loops among the authors and repeated testing phases."
Relevant Quotes:
1) "All prompts were developed through an intensive prompt-engineering process involving multiple iterative feedback loops among the authors and repeated testing phases, following established best practices for effective prompting." (p. 12)
2) "We collected data in schools online using software we developed in-house, called IWM Study (v3.014), which was also used for the intervention." (Data Collection, p. 15)
3) "The materials and tasks were developed by three teachers in close collaboration with two subject-specific education experts for physics and English." (p. 14)
4) "We have no known conflicts of interest to disclose. The authors are responsible for the content of this publication." (Declarations)
Detailed Analysis:
Criterion I requires that the study be conducted independently of the team that designed the intervention, to reduce bias in implementation, measurement, and analysis. Here the authors designed the GenAI-based intervention (developed all condition prompts), built the in-house IWM Study platform used to run it, developed the outcome measures, and carried out the data collection and analysis themselves. Although teachers and subject experts helped develop materials, and scorers were blinded to condition for the open-ended/knowledge items, there is no external, independent evaluation team conducting the trial. The blinding of raters mitigates measurement bias but does not establish independent conduct of the trial as a whole.
Criterion I is not met because the intervention designers (the authors) also implemented, measured, and analysed the study, with no independent third-party evaluator.
-
Y
Year Duration
- The study lasted only six 45-minute sessions with immediate posttesting, far short of (and because T is not met, automatically failing) the year-duration requirement.
- "The study spanned six school sessions (of 45 min each) ... one pretest session (T1), four consecutive learning sessions, and one posttest session (T2)."
Relevant Quotes:
1) "The study spanned six school sessions (of 45 min each) during regular class time: one pretest session (T1), four consecutive learning sessions, and one posttest session (T2)." (p. 14)
2) "The two measurements and interventions took place in April and May 2025." (p. 12)
3) "all outcome measures were collected immediately after the intervention." (Limitations, p. 46)
Detailed Analysis:
Criterion Y requires outcome measurement at least 75% of a full academic year after the intervention begins. Per the prompt instructions, if T (Term Duration) is not met, then Y is automatically not met. T is not met here, and independently the study lasted only a few weeks of short sessions with immediate posttesting, nowhere near a full academic year.
Criterion Y is not met because T is not met and the study duration was only a few weeks, far below the year-duration threshold.
-
B
Balanced Control Group
- All three conditions received identical instructional content, time, task formats, and technical settings, with the only difference being the chatbot prompts, so time and resources were balanced across groups.
- "All sessions followed the same instructional structure across conditions, including identical instructional content, task formats, session duration, and technical settings; conditions differed only in the chatbot prompts."
Relevant Quotes:
1) "All students interacted with GPT-4o-powered systems. However, system behavior differed across conditions due to differential theory-informed prompting." (p. 12)
2) "All sessions followed the same instructional structure across conditions, including identical instructional content, task formats, session duration, and technical settings; conditions differed only in the chatbot prompts." (p. 15)
3) "control condition (standard ChatGPT use without pedagogical prompting)." (p. 12)
4) "All prompts were adapted to the academic subject (physics or English) but followed the same intervention logic." (p. 13)
Detailed Analysis:
Applying the Criterion B decision tree: the intervention does not add extra instructional time or budget relative to the control. All three conditions, including the control, received the same GPT-4o-powered chatbot, the same six 45-minute sessions, the same digital tasks, instructional input, worked examples, and technical settings. The only difference between conditions was the content of the chatbot prompts (utility-value reflection, cognitive strategy prompting, or neutral conversation). Because the control group received an equivalent active chatbot experience with the same time-on-task and resources, the groups are balanced; the manipulated prompting logic is the treatment variable and is integral to what is being tested.
Criterion B is met because all conditions received identical time, materials, and technical resources, with the only difference being the prompting content that constitutes the experimental manipulation.
-
Level 3 Criteria
-
R
Reproduced
- This specific RCT has not been independently replicated by a different research team in a peer-reviewed outlet.
Relevant Quotes:
1) "Against this backdrop, we conducted the present study to implement a theory-informed randomized controlled trial (RCT) in real classrooms, with clearly defined treatments and outcome measures." (Introduction, p. 2)
Detailed Analysis:
Criterion R requires that this specific study (its central experimental claim, design, and context) be independently replicated by a different research team and published in a peer-reviewed journal. This is a newly published 2026 RCT (Educational Psychology Review, DOI 10.1007/s10648-026-10133-8) of GenAI-supported SRL prompts; the paper presents itself as novel evidence and reports no independent replication. An internet search for independent replications of this specific CustomGPT utility-value / cognitive-strategy prompting RCT in secondary classrooms (physics/English) returned only the original article itself and thematically related but distinct GenAI-SRL studies by other teams (e.g., EFL mediation models, metacognitive-support studies); none is an independent reproduction of this particular trial. No replication was found, which is expected given the very recent (April 2026) publication.
Criterion R is not met because no independent replication of this specific study by a different team in a peer-reviewed outlet could be found.
-
A
All-subject Exams
- Outcomes covered only the two intervention subjects (physics and English) using custom items, and because criterion E is not met, criterion A automatically fails.
- "We conducted an RCT with a between-subjects design across two academic subjects (i.e., physics and English)."
Relevant Quotes:
1) "We conducted an RCT with a between-subjects design across two academic subjects (i.e., physics and English)." (p. 12)
2) "Students' domain-specific knowledge was assessed using three subject-specific test items ... tailored to the content covered in the interventions." (p. 18)
3) "raw test scores from the physics and English knowledge tests were standardized (z scores) prior to analysis." (p. 19)
Detailed Analysis:
Criterion A requires that the study measure impact on all main school subjects using standardised exam-based assessments, and it explicitly depends on criterion E being met. Criterion E is not met (custom, teacher-developed items and self-report scales rather than standardised exams), so by the prompt's rule criterion A cannot be met. Additionally, knowledge outcomes were only assessed in the two intervention subjects (physics and English), not across all core subjects of the curriculum.
Criterion A is not met because criterion E is not met and only the two intervention subjects were assessed with custom, non-standardised items.
-
G
Graduation Tracking
- Students were assessed immediately after the intervention with no follow-up to graduation, and because criterion Y is not met, criterion G automatically fails.
- "all outcome measures were collected immediately after the intervention, which does not allow conclusions about whether the observed effects extend beyond the immediate learning situation."
Relevant Quotes:
1) "At T2, which was scheduled directly after the intervention, the participating students were asked to provide sociodemographic information ... In addition, they were tested on their domain-specific knowledge." (p. 16)
2) "all outcome measures were collected immediately after the intervention ... future research should therefore include delayed follow-up assessments to examine the persistence and longer-term impact." (Limitations, p. 46)
Detailed Analysis:
Criterion G requires tracking participants until graduation from their educational stage, and it depends on criterion Y being met. Criterion Y is not met, so by the prompt's rule criterion G cannot be met. Independently, the study collected all outcomes immediately after a few weeks of sessions, with no follow-up tracking through graduation, and the authors explicitly call for future delayed follow-ups. An internet search for subsequent follow-up or graduation-tracking publications by the same author team tracking this Grades 7-9 cohort returned none (the only directly related output is the original 2026 article), consistent with the paper's own statement that delayed follow-up is left to future research.
Criterion G is not met because criterion Y is not met and the study performed no graduation tracking, measuring outcomes only immediately after the intervention, with no follow-up papers found.
-
P
Pre-Registered
- The study, including its hypotheses and data-analysis plan, was pre-registered at OSF, and data collection (T1/T2 in April-May 2025) occurred after pre-registration.
- "This study, including the data analysis approaches, was preregistered at OSF: https://doi.org/10.17605/OSF.IO/VBSQ4"
Relevant Quotes:
1) "This study, including the data analysis approaches, was preregistered at OSF: https://doi.org/10.17605/OSF.IO/VBSQ4. The full reproducible code ... and data are online on our OSF project page." (Methods/Analysis, p. 25)
2) "This study was preregistered at the Open Science Framework (OSF: https://doi.org/10.17605/OSF.IO/VBSQ4)." (Data Availability)
3) "whereas our primary hypotheses were preregistered, and all planned analyses were transparently reported, we conducted several exploratory analyses." (Limitations, p. 46)
4) "The two measurements and interventions took place in April and May 2025." (p. 12)
5) "An ethics committee approved the study and the data collection ... (date of approval: February 20, 2025)." (Declarations)
Detailed Analysis:
Criterion P requires that the study protocol, including hypotheses and planned analyses, be publicly pre-registered before data collection begins. The paper provides an explicit OSF pre-registration DOI (10.17605/OSF.IO/VBSQ4) covering hypotheses and data analysis approaches, and confirms that the primary hypotheses and all planned analyses were preregistered. The OSF DOI was checked online and resolves to the OSF project osf.io/vbsq4, confirming the registration reference is genuine. The data collection (T1 pretest and T2 posttest) took place in April-May 2025, after ethics approval (February 20, 2025), indicating the pre-registration and protocol were in place around the start of the trial rather than after the data were collected.
Criterion P is met because the study's hypotheses and analysis plan were pre-registered at OSF (DOI verified to resolve), with data collection occurring afterward in April-May 2025.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.