Abstract
Background This study aims to investigate the effects of using patient scenarios generated by artificial intelligence in rehabilitation education through different training models (artificial intelligence-based, internet-supported+traditional, and traditional only) on students' digital competence, clinical self-efficacy, and attitudes towards artificial intelligence. Methods Ninety volunteer students were included in the study and divided into three groups using block randomisation: (1) artificial intelligence-supported (2), internet-supported + traditional method, and (3) traditional method only. An 8-week training programme was conducted for each scenario, consisting of weekly 90-minute sessions that alternated between assessment and treatment applications. The Digital Competence Self-Assessment Scale, Clinical Self-Efficacy Scale, and Artificial Intelligence Attitude Scale were administered before and after the intervention. One-way ANOVA or Kruskal-Wallis tests were used for between-group comparisons, and paired t-tests were used for within-group changes (alpha = 0.05). Results In the intra-group analyses, a significant increase was observed in clinical self-efficacy and artificial intelligence attitude scores in all groups (p<.05). Digital competence increased in the AI-supported and internet-supported+traditional groups (p<.05). Intergroup comparisons revealed significant differences in digital competence and AI attitude scores (p<.05). The increase in clinical self-efficacy scores was not significant at the intergroup level (p>.05). Conclusion Artificial intelligence-based scenario training increases the level of digital competence in rehabilitation students and develops positive attitudes towards artificial intelligence. The findings indicate that this method can be integrated into educational programmes to strengthen clinical training processes. Further studies with larger samples, longer-term interventions, and objective performance measures are recommended to understand the effects on clinical skills.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was carried out at the individual student level (block randomisation, block size 3) within a single cohort, not at the class or school level, and the intervention is not one-to-one tutoring.
- "After providing informed consent, participants were randomly allocated into three groups using a computer-generated simple block randomisation method (block size=3)..."
Relevant Quotes:
1) "This study was designed as a three-group randomized controlled educational intervention study and conducted among second-year physiotherapy associate degree students." (p. 2)
2) "After providing informed consent, participants were randomly allocated into three groups using a computer-generated simple block randomisation method (block size=3): Group 1 (artificial intelligence), Group 2 (internet+traditional methods), and Group 3 (traditional methods only)." (p. 2)
3) "Ninety volunteer students were included in the study and divided into three groups using block randomisation..." (Abstract, p. 1)
4) "The study was conducted with second-year students enrolled in the Therapy and Rehabilitation Department (physiotherapy associate degree programme) at a local university health services vocational school." (p. 3)
Detailed Analysis:
The ERCT 'C' criterion requires randomisation at the class level (or stronger school level), unless the intervention is personal/one-to-one tutoring, in which case student-level randomisation is acceptable. Here the participants were 90 individual volunteer students drawn from a single second-year associate-degree cohort and randomly assigned individually ("participants were randomly allocated into three groups using ... block randomisation"). The unit of randomisation is the individual student, not a class or a school. All three groups were drawn from the same student population at the same institution, creating exactly the within-cohort contamination risk the criterion is designed to prevent. The intervention is a group-based scenario-analysis course rather than personal one-to-one tutoring, so the tutoring exception does not apply.
Criterion C is not met because randomisation was performed at the individual student level within a single cohort rather than at the class or school level, and no tutoring exception applies.
-
E
Exam-based Assessment
- Outcomes were measured only with self-report Likert scales (digital competence, clinical self-efficacy, AI attitude), not with any standardised exam-based assessment of educational achievement.
- "The Digital Competence Self-Assessment Scale, Clinical Self-Efficacy Scale, and Artificial Intelligence Attitude Scale were administered before and after the intervention."
Relevant Quotes:
1) "The Digital Competence Self-Assessment Scale, Clinical Self-Efficacy Scale, and Artificial Intelligence Attitude Scale were administered before and after the intervention." (Abstract, p. 1)
2) "Artificial Intelligence Attitude Scale ... The scale consists of 13 items covering three sub-dimensions and is rated on a 5-point Likert scale, with higher scores indicating a more positive attitude [13]." (p. 4)
3) "Clinical Self-Efficacy Scale ... The scale consists of 10 items and is rated on a 5-point Likert scale, with higher scores indicating greater self-confidence [14]." (p. 4)
4) "Digital Competence Self-Assessment Scale (DYÖ) ... The scale consists of 21 items covering five sub-dimensions and is rated on a 5-point Likert scale, with higher scores indicating higher levels of digital competence [15]." (p. 4)
5) "In addition, the outcomes were assessed using self-report scales, which may reflect participants' perceived competencies rather than objective performance. Clinical performance was not directly evaluated using objective performance-based assessments..." (p. 6)
Detailed Analysis:
The 'E' criterion requires the use of standardised, widely-recognised exam-based assessments of educational outcomes, not instruments designed or assembled for the study and not subjective self-report tools. All three outcome measures in this study are validated psychometric self-report Likert scales measuring perceptions (digital competence self-assessment, clinical self-efficacy perception, and attitudes toward AI). They are not exam-based achievement assessments, and the authors explicitly acknowledge that "outcomes were assessed using self-report scales, which may reflect participants' perceived competencies rather than objective performance," with no objective performance-based assessment used. There is no standardised national/state exam or comparable achievement test of subject knowledge.
Criterion E is not met because the study relied solely on self-report Likert-scale instruments rather than any standardised exam-based assessment of educational achievement.
-
T
Term Duration
- The intervention lasted only eight weekly 90-minute sessions and outcomes were measured immediately at the end, far short of a full academic term of follow-up from intervention start.
- "The intervention lasted eight weeks, with one session per week (approximately 90 min per session)."
Relevant Quotes:
1) "An 8-week training programme was conducted for each scenario, consisting of weekly 90-minute sessions that alternated between assessment and treatment applications." (Abstract, p. 1)
2) "The intervention lasted eight weeks, with one session per week (approximately 90 min per session). During this period, participants analysed patient scenarios in one week and developed treatment programmes for the same case in the following week." (p. 3)
3) "At the end of the intervention, all baseline scales were re-administered in the same order." (p. 3)
4) "First, the relatively short duration of the intervention may have limited the development of observable differences in clinical skills." (p. 6)
Detailed Analysis:
The 'T' criterion requires that primary outcomes be measured at least one full academic term (~3-4 months) after the intervention begins. Here the intervention ran for eight weeks (eight 90-minute sessions, roughly 2 months) and the outcome scales were re-administered "at the end of the intervention," i.e. immediately after the final session. There was no term-length follow-up beyond the end of the eight-week programme; the authors themselves describe the duration as "relatively short" and recommend "longer-term interventions" and "follow-up periods." The interval from intervention start to measurement is about eight weeks, shorter than one academic term.
Criterion T is not met because the eight-week intervention and immediate post-test fall short of one full academic term of tracking from intervention start.
-
D
Documented Control Group
- The traditional-methods control group is documented with size (n=30), demographics, baseline scores, and a clear description of the business-as-usual condition it received.
- "participants in Group 3 used only traditional learning methods."
Relevant Quotes:
1) "Group 3 (traditional methods only). The randomisation process was performed by an independent researcher who was not involved in the study." (p. 2)
2) "participants in Group 3 used only traditional learning methods." (p. 3)
3) "Table 3 Demographic characteristics of the participants ... Traditional Training ... Age 20(19-21) ... Female 20 (66.6) ... Male 10 (33.4) ... Prior Simulation Training ... No 30 (100.0)." (p. 4)
4) "Traditional Training (n=30)" with pre-test scores: "Digital Competence 32.00 (29-33) ... Clinical Self-Efficacy 31.67±0.96 ... AI Attitude 54.50 (52.00-56.75)" (Table 4, p. 5)
Detailed Analysis:
The 'D' criterion requires the control group to be well-documented, including size, demographics, baseline performance, and the conditions it received. Group 3 serves as the comparison/control condition (traditional methods only). The paper documents its size (n=30), demographic characteristics (age, sex distribution, daily computer/tablet usage, prior simulation training) in Table 3, and baseline (pre-test) scores on all three outcome measures in Table 4. The condition is clearly described: students used "only traditional learning methods." This provides sufficient documentation to assess comparability at baseline.
Criterion D is met because the control (traditional-methods) group's size, demographics, baseline scores, and treatment condition are clearly documented in the text and Tables 3-4.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation was at the individual student level within a single institution; no schools (or comparable institutional units) were randomised.
- "participants were randomly allocated into three groups using a computer-generated simple block randomisation method (block size=3)..."
Relevant Quotes:
1) "After providing informed consent, participants were randomly allocated into three groups using a computer-generated simple block randomisation method (block size=3)..." (p. 2)
2) "The study was conducted with second-year students enrolled in the Therapy and Rehabilitation Department (physiotherapy associate degree programme) at a local university health services vocational school." (p. 3)
3) "the participants were recruited from a single institution using a convenience sampling approach, which may limit the generalisability of the results." (p. 6)
Detailed Analysis:
The 'S' criterion requires randomisation at the school (or comparable institutional/site) level. In this study the unit of randomisation was the individual student, and all participants came from a single department at one vocational school. No multiple schools, sites, or institutions were randomly assigned; the study is explicitly single-institution.
Criterion S is not met because randomisation occurred at the individual student level within one institution, with no school- or site-level randomisation.
-
I
Independent Conduct
- The same two authors designed the scenarios, ran the intervention, collected and analysed the data; only the randomisation step was delegated, with no independent external evaluator for conduct or analysis.
- "RY contributed to study design, data collection, data analysis, and manuscript writing. OSN contributed to scenario validation, methodological supervision, and critical manuscript revision."
Relevant Quotes:
1) "The randomisation process was performed by an independent researcher who was not involved in the study." (p. 2)
2) "Participants were then presented with patient scenarios generated using an artificial intelligence programme (ChatGPT) and reviewed by two academics specialising in physiotherapy and rehabilitation." (p. 2)
3) "RY contributed to study design, data collection, data analysis, and manuscript writing. OSN contributed to scenario validation, methodological supervision, and critical manuscript revision. Both authors read and approved the final manuscript." (p. 7)
4) "Clinical performance was not directly evaluated using objective performance-based assessments, and blinding was not applied." (p. 6)
Detailed Analysis:
The 'I' criterion requires that the evaluation be conducted independently of those who designed the intervention, to reduce bias in implementation, data collection, analysis, and reporting. Here the two authors designed the study, validated the AI-generated scenarios, collected the data, performed the analysis, and wrote the manuscript. The only independent element is the randomisation step, performed by "an independent researcher who was not involved in the study" - but this person did not conduct the data collection or analysis, and blinding was not applied. The core design, delivery, data collection, and analysis were carried out by the intervention designers themselves, so there is no independent third-party evaluation.
Criterion I is not met because the same authors who designed and delivered the intervention also collected and analysed the data, with independence limited to the randomisation step only.
-
Y
Year Duration
- Criterion T is not met (eight-week intervention with immediate post-test), and the study covers nowhere near 75% of an academic year, so Y automatically fails.
- "The intervention lasted eight weeks, with one session per week (approximately 90 min per session)."
Relevant Quotes:
1) "An 8-week training programme was conducted ... consisting of weekly 90-minute sessions..." (Abstract, p. 1)
2) "The intervention lasted eight weeks, with one session per week (approximately 90 min per session)." (p. 3)
3) "At the end of the intervention, all baseline scales were re-administered in the same order." (p. 3)
4) "Future research should include ... longer intervention and follow-up periods..." (p. 7)
Detailed Analysis:
The 'Y' criterion requires outcomes to be measured at least 75% of one academic year (~9-10 months) after the intervention begins. Per the criterion-specific instruction, if T (Term Duration) is not met then Y is not met. T is not met here, and independently the eight-week duration with immediate post-test is far below 75% of an academic year.
Criterion Y is not met because the intervention and follow-up spanned only eight weeks, well short of an academic year, and T is also not met.
-
B
Balanced Control Group
- All three groups received the same eight-week, weekly 90-minute scenario-based course with equivalent time on task; the only difference was the information source (AI vs internet vs traditional materials), which is the treatment variable itself, so educational time and resources are balanced.
- "The intervention lasted eight weeks, with one session per week (approximately 90 min per session)."
Relevant Quotes:
1) "An 8-week training programme was conducted for each scenario, consisting of weekly 90-minute sessions that alternated between assessment and treatment applications." (Abstract, p. 1)
2) "Group-specific learning approaches were defined as follows: participants in Group 1 used only artificial intelligence tools (ChatGPT); participants in Group 2 used a combination of academic resources (e.g., Google Scholar, PubMed) and traditional materials (e.g., books, lecture notes); and participants in Group 3 used only traditional learning methods." (p. 3)
3) "During this period, participants analysed patient scenarios in one week and developed treatment programmes for the same case in the following week." (p. 3)
4) "In the artificial intelligence group, students used ChatGPT as an informational support tool to search for case-related information. However, clinical reasoning and treatment programme development were performed independently by the participants based on their own knowledge." (p. 3)
Detailed Analysis:
The 'B' criterion asks whether intervention and control groups receive comparable educational time and resources, unless the additional resource is itself the explicit treatment variable. Here all three groups followed the identical eight-week schedule of weekly 90-minute sessions analysing the same set of patient scenarios (cervical disc herniation, patellofemoral pain, carpal tunnel, total hip replacement) and developing treatment plans. The instructional time on task is the same across groups. The only systematic difference is the source of information used to research the cases: AI tool (ChatGPT) vs internet/academic resources plus traditional materials vs traditional methods only. This information source is precisely the treatment variable the study is testing (comparative effectiveness of AI-based vs internet-supported vs traditional learning). Applying the decision tree: extra time/budget is essentially equal across arms (same sessions, same duration), and to the extent the AI tool is an "extra resource," it is the integral treatment variable being tested against a business-as-usual traditional condition.
Criterion B is met because all groups received equivalent instructional time and the same scenario-based curriculum, with the differing information source being the explicit treatment variable rather than an unbalanced add-on.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study by a different research team in a peer-reviewed journal is reported or identifiable; the paper is a single original study.
Relevant Quotes:
1) "Therefore, the present study aims to investigate the comparative effectiveness of AI-supported clinical case analysis in rehabilitation education..." (p. 2)
2) "Future studies including larger samples, longer intervention periods, and objective performance-based assessments are needed to better evaluate the potential impact of AI-based education on clinical skills." (p. 6)
Detailed Analysis:
The 'R' criterion requires that the specific study be independently replicated by a different research team in a different context, published in a peer-reviewed journal. This paper presents a single original study (published April 2026). It cites related literature on AI in education and digital competence, but none of those are replications of this specific AI-scenario rehabilitation-education trial; rather, the authors call for future studies. An internet search for replications found a separate, different study ("Effects of artificial intelligence based physiotherapy educational approach in developing clinical reasoning skills: a randomized controlled trial", BMC Medical Education 2025, doi 10.1186/s12909-025-07926-w), but it is a distinct trial by a different author team with different outcome measures and does not replicate this specific study. No external replication of this particular study could be identified.
Criterion R is not met because there is no evidence of an independent replication of this specific study by another research team in a peer-reviewed outlet.
-
A
All-subject Exams
- Criterion E is not met, and outcomes covered only a single narrow domain (rehabilitation/clinical scenarios) via self-report scales, not standardised exams across all main subjects, so A fails.
- "The Digital Competence Self-Assessment Scale, Clinical Self-Efficacy Scale, and Artificial Intelligence Attitude Scale were administered before and after the intervention."
Relevant Quotes:
1) "The Digital Competence Self-Assessment Scale, Clinical Self-Efficacy Scale, and Artificial Intelligence Attitude Scale were administered before and after the intervention." (Abstract, p. 1)
2) "The scenarios focused on common musculoskeletal conditions frequently encountered in clinical practice and were designed to require assessment, clinical reasoning, and treatment planning based on anatomical and rehabilitation principles." (p. 2)
Detailed Analysis:
The 'A' criterion requires measuring impact across all main subjects using standardised exam-based assessments, and per the criterion-specific instruction, if E is not met then A is not met. E is not met here (only self-report scales were used). Furthermore, the outcomes are confined to a single specialised domain (rehabilitation clinical scenarios) and three self-report perception scales, with no standardised exams in multiple core subjects.
Criterion A is not met because criterion E fails and the study assessed only a single narrow domain via self-report scales rather than standardised exams across all main subjects.
-
G
Graduation Tracking
- Criterion Y is not met, and outcomes were measured immediately after an eight-week programme with no follow-up to graduation, so G fails.
- "At the end of the intervention, all baseline scales were re-administered in the same order."
Relevant Quotes:
1) "At the end of the intervention, all baseline scales were re-administered in the same order." (p. 3)
2) "Future research should include more diverse samples, longer intervention and follow-up periods, and objective performance-based measures..." (p. 7)
Detailed Analysis:
The 'G' criterion requires tracking participants until graduation, and per the criterion-specific instruction, if Y is not met then G is not met. Y is not met here. The study measured outcomes only immediately at the end of the eight-week intervention, with no follow-up through graduation and no subsequent tracking publications; the authors explicitly note the absence of follow-up periods. An internet search found no follow-up or graduation-tracking publications by the same authors for this cohort.
Criterion G is not met because Y fails and there was no tracking of participants beyond the immediate post-intervention assessment, let alone to graduation.
-
P
Pre-Registered
- The authors explicitly state the study was not registered in any clinical trial registry, so no pre-registered protocol exists.
- "This study was not registered in a clinical trial registry."
Relevant Quotes:
1) "In order to carry out this study, an application was submitted to the Scientific Research and Publication Ethics Committee of a local university, and ethical approval was obtained on 12 June 2025 (Reference No: E-18457941-050.99- 181729). All stages of the study were conducted in accordance with the Declaration of Helsinki. This study was not registered in a clinical trial registry." (p. 3)
Detailed Analysis:
The 'P' criterion requires the full study protocol to be pre-registered (with hypotheses, methods, and planned analyses) before data collection begins, in a recognised registry. The paper explicitly states, "This study was not registered in a clinical trial registry." There is therefore no registry entry to verify, and the only prospective approval was institutional ethics approval, which is not a public pre-registration of the study protocol/analysis plan.
Criterion P is not met because the authors explicitly confirm the study was not registered in any clinical trial registry.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.