Abstract
This pilot randomized controlled trial compared AI-assisted self-directed learning, lecture-based instruction, and simulation-based education in teaching emergency choking management to pre-professional health students. Twenty students (n = 6-7 per group) enrolled in a two-week preparatory program were randomly assigned to one of three instructional groups: lecture, simulation, or AI-assisted self-directed learning using ChatGPT-4o. The lecture and simulation groups received roughly 60-minute faculty-led sessions, while the AI-assisted group received the standardized educational outline and was asked to self-study using ChatGPT-4o. All participants completed a post-intervention simulation scenario and a knowledge assessment on the third day of the two-week program, followed by a repeat assessment on the final day (nine days later). Students in the lecture and simulation groups demonstrated higher performance and confidence than those in the AI-assisted group. In this small pilot, AI-assisted self-directed learning appeared to underperform faculty-supported lecture and simulation formats.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was performed at the individual student level, not at the class or school level, and no personal-tutoring exception applies.
- "Participants were randomly assigned via computer-generated allocation to one of three instructional groups: lecture-based, simulation-based, or AI-assisted self-learning."
Relevant Quotes:
1) "This pilot randomized controlled trial used a parallel three-arm design [14] involving 20 pre-professional students voluntarily enrolled in a two-week intensive summer program for aspiring healthcare providers at a small university in Boston, Massachusetts." (p. 3)
2) "Participants were randomly assigned via computer-generated allocation to one of three instructional groups: lecture-based, simulation-based, or AI-assisted self-learning." (p. 3)
3) "A total of 20 pre-professional students were enrolled and randomly assigned to one of three groups (lecture n = 7, simulation n = 6, AI-assisted n = 7)." (p. 5)
Detailed Analysis:
The ERCT 'C' criterion requires random assignment of entire classes (or the stronger school unit), unless the intervention is one-to-one/personal tutoring. Here the 20 individual students were the unit of randomisation: they were "randomly assigned via computer-generated allocation" as individuals into three instructional arms within a single two-week cohort at one small university. This is student-level randomisation, which raises exactly the contamination risk the criterion guards against, since all participants were in the same program. The lecture and simulation arms are group-delivered sessions, not personal tutoring, so the tutoring exception does not apply, and no class- or school-level randomisation is described.
Criterion C is not met because randomisation was conducted at the individual student level rather than at the class or school level, with no qualifying tutoring exception.
-
E
Exam-based Assessment
- All outcome instruments were custom tools built by the study team for this study, not widely recognised standardised exams.
- "The Choking Knowledge Test was developed by content experts and was based on the course content with essay questions that intentionally included new content..."
Relevant Quotes:
1) "The Choking Simulation Performance Measure Tool was created using course content and American Heart Association guidelines [17,18] of choking management (Appendix II)." (p. 4)
2) "The Choking Knowledge Test was developed by content experts and was based on the course content with essay questions that intentionally included new content to provide space for ChatGPT searches and critical thinking (Appendix III)." (p. 5)
3) "The Holistic Critical Thinking Scoring Rubric (HCTSR) evaluates overall critical thinking quality using a four-point scale... it was adapted in this study with an added column to assess critical thinking demonstrated in AI chat transcripts (Appendix IV)." (p. 5)
Detailed Analysis:
Criterion E requires a standardised, widely recognised exam-based assessment that was not specially designed for the study. Every outcome instrument here was created by the authors/content experts specifically for this study: the performance checklist was "created using course content and American Heart Association guidelines," the knowledge test was "developed by content experts and was based on the course content," and the critical-thinking rubric was an adapted/modified HCTSR. Although the performance tool is informed by AHA guidelines and showed good internal consistency (alpha = 0.83), it is a bespoke study checklist, not a standardised national or state exam administered to the study population. No recognised standardised exam is named.
Criterion E is not met because all assessments were custom-built for the study rather than being standardised, widely recognised exams.
-
T
Term Duration
- The intervention was a single ~60-minute session and final measurement occurred only about nine days later, far short of one academic term.
- "The lecture and simulation arms each consisted of a single approximately 60-minute faculty-led session..."
Relevant Quotes:
1) "The lecture and simulation arms each consisted of a single approximately 60-minute faculty-led session delivered by the same core teaching team to ensure content consistency." (p. 3)
2) "All participants completed a post-intervention simulation scenario and a knowledge assessment on the third day of the two-week program (multiple-choice and essay), followed by a repeat assessment on the final day of the program (nine days later)." (Abstract)
3) "An exploratory focus group was conducted after testing, followed by a repeat post-test without AI access nine days later (at the end of the two-week program) to measure retention and independent knowledge performance." (p. 3)
Detailed Analysis:
Criterion T requires outcomes to be measured at least one full academic term (~3-4 months) after the intervention begins. The intervention here was a single ~60-minute session, and the entire study ran inside a two-week summer program, with the initial assessment on day three and the final follow-up assessment nine days after that. The longest interval from intervention start to outcome measurement is therefore roughly two weeks, which is far shorter than a single academic term.
Criterion T is not met because the interval from intervention to final measurement was about two weeks, well below one academic term.
-
D
Documented Control Group
- The comparison groups' demographics, sizes, and conditions are documented in Table 1 and the methods, including confirmation of what each arm received.
- "A total of 20 pre-professional students were enrolled and randomly assigned to one of three groups (lecture n = 7, simulation n = 6, AI-assisted n = 7)."
Relevant Quotes:
1) "A total of 20 pre-professional students were enrolled and randomly assigned to one of three groups (lecture n = 7, simulation n = 6, AI-assisted n = 7). Participants across groups were similar in grade and prior AI usage. The AI group, by chance of randomization, was slightly older, entirely female, more racially diverse, and had the highest proportion of English as an additional language (EAL) students." (p. 5)
2) "TABLE 1: Demographics by Groups" - reports age, grade, gender, race/ethnicity, English as primary or additional language, and prior AI use for each of the lecture, simulation, and AI arms. (p. 6)
3) "Third, no baseline knowledge or critical-thinking assessment was administered, which limits inferences about change attributable to instruction." (p. 12)
Detailed Analysis:
Criterion D requires that the control/comparison group be documented in terms of characteristics, size, and conditions. Although this is a three-arm active comparison rather than a study with a single untreated control, the comparison arms are documented in detail: Table 1 provides each arm's size and demographic profile (age, grade, gender, race/ethnicity, language background, prior AI use), and the methods clearly describe what each arm received (faculty-led lecture, simulation session, or AI self-study). A weakness is that no baseline knowledge or critical-thinking measure was collected, but the demographic and condition documentation of the comparison groups is clear and detailed.
Criterion D is met because the comparison groups' size, demographic characteristics, and conditions are clearly documented, despite the absence of a baseline performance measure.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation was at the individual student level in a single institution, not at the school level.
- "Participants were randomly assigned via computer-generated allocation to one of three instructional groups..."
Relevant Quotes:
1) "This pilot randomized controlled trial used a parallel three-arm design [14] involving 20 pre-professional students voluntarily enrolled in a two-week intensive summer program... at a small university in Boston, Massachusetts." (p. 3)
2) "Participants were randomly assigned via computer-generated allocation to one of three instructional groups: lecture-based, simulation-based, or AI-assisted self-learning." (p. 3)
Detailed Analysis:
Criterion S requires randomisation at the school level (the educational institution or implementing unit), with multiple schools/sites randomised. This study was a single-site trial at one small university, and the unit of randomisation was the individual student, not the school. There is no randomisation of schools, centres, or sites.
Criterion S is not met because the study randomised individual students within a single institution rather than randomising schools.
-
I
Independent Conduct
- The same team designed the instruments and intervention, delivered the sessions, and scored the outcomes, with no independent or third-party evaluator.
- "performance measures were scored independently by four researchers from video recordings: two researchers were observers; two were the standardized patients."
Relevant Quotes:
1) "The lecture and simulation arms each consisted of a single approximately 60-minute faculty-led session delivered by the same core teaching team to ensure content consistency." (p. 3)
2) "performance measures were scored independently by four researchers from video recordings: two researchers were observers; two were the standardized patients." (p. 5)
3) "essay questions in combination with the AI Chat threads were scored independently by two researchers using the HCTSR via consensus." (p. 5)
4) "Concept and design: Janice C. Palaganas, Alex Morton, Khaulah Jawed... Acquisition, analysis, or interpretation of data: Janice C. Palaganas, Alex Morton, Fatima Adem, Jianna Ramos, Arjun Kumar, Maria Bajwa." (p. 20)
Detailed Analysis:
Criterion I requires that the evaluation be conducted independently of the people who designed and delivered the intervention, to reduce bias in implementation, measurement, and analysis. Here the same core teaching team designed the study, built the custom instruments, delivered the lecture and simulation sessions, and then scored performance, knowledge, and critical thinking (the four researchers who scored performance and the two who scored essays are members of the study team). No external evaluation agency, third-party assessor, or blinded independent evaluator separate from the intervention team is described.
Criterion I is not met because the intervention was designed, delivered, and evaluated by the same research team without any independent or third-party oversight.
-
Y
Year Duration
- Term duration (T) is not met, and the study spanned only about two weeks, so the year-duration criterion cannot be satisfied.
- "involving 20 pre-professional students voluntarily enrolled in a two-week intensive summer program..."
Relevant Quotes:
1) "The lecture and simulation arms each consisted of a single approximately 60-minute faculty-led session..." (p. 3)
2) "...followed by a repeat assessment on the final day of the program (nine days later)." (Abstract)
3) "involving 20 pre-professional students voluntarily enrolled in a two-week intensive summer program for aspiring healthcare providers..." (p. 3)
Detailed Analysis:
Criterion Y requires that outcomes be measured at least 75% of a full academic year after the intervention begins, and per the ERCT rules Y cannot be met when the weaker Term Duration (T) criterion is not met. Here T is not met, and the whole study occurred within a two-week program, with final measurement about nine days after a single-session intervention. This is far below 75% of an academic year.
Criterion Y is not met because the study lasted only about two weeks and the prerequisite Term Duration criterion is not met.
-
B
Balanced Control Group
- All three arms received a session of comparable scheduled duration, and the differing instructional support is integral to the teaching modalities being compared, which is the study's treatment variable.
- "and was instructed to use ChatGPT-4o to learn the content during a session of comparable scheduled duration."
Relevant Quotes:
1) "Notably, the lecture and simulation groups received roughly 60-minute faculty-led sessions, while the AI-assisted group received the standardized educational outline and was asked to self-study using ChatGPT-4o." (Abstract, p. 1)
2) "The AI-assisted self-learning group was provided the same standardized educational outline (Appendix I) and individual paid ChatGPT-4o accounts [16], and was instructed to use ChatGPT-4o to learn the content during a session of comparable scheduled duration. A faculty member was in the room to provide technical and session support." (p. 3)
3) "We deliberately compared AI-assisted self-directed learning against both lecture and simulation as these three modalities represent the most common observed formats encountered by pre-professional learners... didactic transfer of content (lecture), experiential practice with feedback (simulation), and increasingly, AI-mediated self-study." (p. 2)
4) "First, by design, the AI-assisted group received considerably less instructional support than the lecture and simulation groups... In effect, the comparison may reflect unguided self-study with AI versus faculty-supported teaching as much as it reflects AI as a modality." (p. 12)
Detailed Analysis:
Criterion B compares the nature, quantity, and quality of resources (time, budget, materials, adult support) given to each condition, and asks whether any imbalance is integral to the treatment variable being tested or is a separable confound. On time, the study states that all three arms received a session of comparable scheduled duration (roughly 60 minutes), and all groups worked from the same standardized educational outline, so instructional time and core content were balanced across arms. On budget, the lecture and simulation arms received faculty instruction while the AI arm received individual paid ChatGPT-4o accounts, so each arm was resourced.
The substantive difference between arms is the form of instructional support: faculty-led delivery (lecture and simulation) versus AI-mediated self-study. This difference is not a separable add-on layered on top of a shared intervention; it is the very definition of the three modalities the study set out to compare. A lecture is inseparable from faculty delivery, and simulation is inseparable from facilitated practice and debriefing, so the presence or absence of faculty support is integral to each modality being tested, which is the study's explicit treatment variable. Under the decision tree, because session time was comparable and the resource/support differences are integral to the treatment being tested rather than a matched confound, the criterion is satisfied. The authors do caution that the AI arm received less instructional support and recommend equalizing support in future work; this asymmetry is important for interpreting effects but does not, in itself, constitute an unbalanced extra resource added to one arm on top of a shared intervention.
Criterion B is met because instructional time and core content were comparable across arms and the differing instructional support is integral to the teaching modalities being compared, which is the treatment variable, while noting the authors' acknowledged support asymmetry.
-
Level 3 Criteria
-
R
Reproduced
- This is a novel single-site pilot with no independent replication reported or found.
- "This study is among the first to integrate generative AI-based self-learning with simulation and lecture methods to assess pre-professional students' clinical and critical thinking skills."
Relevant Quotes:
1) "This study is among the first to integrate generative AI-based self-learning with simulation and lecture methods to assess pre-professional students' clinical and critical thinking skills." (p. 12)
2) "This was a single-site pilot study, and therefore, the findings should be interpreted as preliminary and hypothesis-generating." (p. 12)
Detailed Analysis:
Criterion R requires independent replication of the study by a different research team in a different context, published in a peer-reviewed journal. The authors describe this work as "among the first" of its kind and a single-site pilot; no replication is referenced. An external replication search could not be completed because the session's web-search quota was exhausted; however, the paper resolves to its publisher landing page (Cureus Journal of Computer Science, DOI 10.7759/s44389-026-00094-y) and, given its very recent publication (June 2026) and self-described novelty, no independent replication would be expected to exist yet.
Criterion R is not met because no independent replication of this study exists.
-
A
All-subject Exams
- Exam-based Assessment (E) is not met, and only a single content domain (choking management) was assessed, so the all-subject criterion fails.
- "The Choking Knowledge Test was developed by content experts and was based on the course content..."
Relevant Quotes:
1) "The Choking Simulation Performance Measure Tool was created using course content and American Heart Association guidelines [17,18] of choking management (Appendix II)." (p. 4)
2) "The Choking Knowledge Test was developed by content experts and was based on the course content..." (p. 5)
Detailed Analysis:
Criterion A requires that impact be measured across all main subjects using standardised exam-based assessments, and it explicitly cannot be met if Criterion E (Exam-based Assessment) is not met. Here E is not met because all instruments were custom-built. In addition, the study assessed only a single narrow content domain (emergency choking management), not the range of core subjects taught to these students.
Criterion A is not met because the prerequisite Exam-based Assessment criterion fails and only one narrow content area was measured.
-
G
Graduation Tracking
- Year Duration (Y) is not met and there was no follow-up to graduation, with tracking ending nine days after the intervention.
- "...followed by a repeat post-test without AI access nine days later (at the end of the two-week program)..."
Relevant Quotes:
1) "An exploratory focus group was conducted after testing, followed by a repeat post-test without AI access nine days later (at the end of the two-week program) to measure retention and independent knowledge performance." (p. 3)
2) "Third, no baseline knowledge or critical-thinking assessment was administered, which limits inferences about change attributable to instruction." (p. 12)
Detailed Analysis:
Criterion G requires tracking participants through to graduation from their educational stage, and per the ERCT rules it cannot be met when Year Duration (Y) is not met. Y is not met here, and the study's final measurement was a repeat post-test nine days after the intervention, at the end of the two-week program. No long-term or graduation follow-up is described or planned, and no subsequent follow-up publication by the same authors tracking this cohort could be identified (the session's web-search quota was exhausted, but the paper itself describes no planned graduation tracking).
Criterion G is not met because there was no graduation-level follow-up and the prerequisite Year Duration criterion is not met.
-
P
Pre-Registered
- No pre-registration of the study protocol is reported; only an IRB exemption is noted, and the authors recommend prospective registration for future work.
- "This study was deemed exempt by the Mass General Brigham (MGB) Institutional Review Board (Protocol ID 2025P001575)."
Relevant Quotes:
1) "This study was deemed exempt by the Mass General Brigham (MGB) Institutional Review Board (Protocol ID 2025P001575)." (p. 3)
2) "Future studies should include baseline assessment, larger and more diverse cohorts, comparable instructional support across arms, explicit AI-literacy and prompt-design scaffolding, and prospective registration." (p. 12)
Detailed Analysis:
Criterion P requires that the full study protocol, including hypotheses and planned analyses, be pre-registered on a public registry before data collection begins. The only registration-like item is an institutional review board exemption with a protocol ID (2025P001575), which is an internal ethics approval number, not a public pre-registration of the study protocol and analysis plan on a trial registry. No ClinicalTrials.gov, ISRCTN, OSF, or equivalent registration is mentioned anywhere in the paper, and the publisher landing page lists no such registration. Moreover, the authors explicitly list "prospective registration" as a step for future studies, indicating this pilot was not prospectively registered.
Criterion P is not met because no public pre-registration of the protocol is reported, only an IRB exemption.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.