Abstract
Background: This study aimed to explore the efficacy of a teaching model integrating artificial intelligence-assisted problem-based learning (PBL) with case-based learning (CBL) in standardized clinical teaching of postherpetic neuralgia (PHN) for medical interns. Methods: A total of 120 interns at the Affiliated Hospital of Yunnan University were randomly divided into an observation group and a control group (n=60 each). The control group received conventional lecture-based teaching, while the observation group was taught using the AI-assisted PBL-CBL combined model. After four weeks of teaching, the educational outcomes were evaluated through theoretical examinations, clinical skill assessments, Mini-Clinical Evaluation Exercise (Mini-CEX) scores, and a teaching satisfaction questionnaire. Results: After four weeks, the theoretical score of the observation group was 87.35+/-5.66, and the clinical practice score was 90.12+/-3.47, both significantly higher than those of the control group (84.67+/-4.05 and 88.23+/-3.94, respectively; P<0.01). The Mini-CEX evaluation showed that students in the observation group had significantly improved clinical practice skills (P<0.001), whereas no significant differences were observed between the two groups in humanistic qualities, professionalism, or organization/efficiency (P>0.05). In addition, the observation group reported higher teaching satisfaction than the control group. Conclusions: The AI-assisted PBL-CBL teaching model significantly improves medical interns' theoretical knowledge and clinical skills related to PHN, enhances their clinical reasoning ability, and increases teaching satisfaction.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was at the individual intern level (random number table), not at the class or school level, and the intervention is not one-to-one tutoring.
- Participants were randomly assigned at a 1:1 ratio using a random number table, with allocation concealment and assessor blinding implemented to control bias.
Relevant Quotes:
1) "A total of 120 interns at the Affiliated Hospital of Yunnan University were randomly divided into an observation group and a control group (n=60 each)." (Abstract, p. 2)
2) "Participants were randomly assigned at a 1:1 ratio using a random number table, with allocation concealment and assessor blinding implemented to control bias. Based on the study design, a total sample size of 120 cases was enrolled, with 60 cases in both the control group and the observation group." (Methods, p. 4)
3) "From March to May 2025, 120 students were randomly assigned to observation group (n=60) or control group (n=60)." (Results, p. 8)
4) "Medical interns were divided into 10 groups of six students each for discussion and learning." (p. 5)
Detailed Analysis:
Criterion C requires randomisation at the class level (or the stronger school level), or, as an exception, allows student-level randomisation only when the intervention is personal one-to-one tutoring. Here the unit of randomisation was the individual intern: 120 interns were individually allocated 1:1 by a random number table into the observation and control groups. There is no indication that intact classes were the unit of assignment; instead individuals were split, and within each condition interns were further sub-divided into small groups. Both conditions were drawn from the same single site (one hospital), so the two groups were not isolated intact classes, creating potential for contamination. The intervention is a group-based teaching model (AI-assisted PBL-CBL with group discussion and simulation), not one-to-one tutoring, so the tutoring exception does not apply.
Criterion C is not met because randomisation was performed at the individual student level rather than at the class or school level, and no tutoring exception applies.
-
E
Exam-based Assessment
- The outcome measures were instructor-developed theoretical and clinical-skill tests and a Mini-CEX rating tool specific to PHN, not widely recognised standardised exams.
- "The test paper jointly developed by the two instructors is as per the fundamental theories of PHN, standardized diagnostic and treatment process and rehabilitation management."
Relevant Quotes:
1) "The theoretical test is a closed book test. The test paper jointly developed by the two instructors is as per the fundamental theories of PHN, standardized diagnostic and treatment process and rehabilitation management. The total score was 100 marks." (p. 7)
2) "The assessment of clinical skills used simulations... The scoring is done by three dermatology specialists using a standard rubric with a maximum score of 100." (p. 7)
3) "A mini-clinical evaluation exercise (Mini-CEX) was used to assess students' abilities in history taking, physical examination, and clinical reasoning... The validated Mini-CEX tool employs a 9-point rating scale across seven domains..." (p. 7)
4) "The teaching satisfaction questionnaire was created by the author to determine the participants' level of teaching satisfaction (Cronbach's alpha value was 0.823)..." (p. 7)
Detailed Analysis:
Criterion E requires that outcomes be measured with standardised, widely recognised exam-based assessments that were not specially designed for the study (e.g., national curriculum or state-wide examinations). In this paper the theoretical examination was "jointly developed by the two instructors" specifically for the PHN module, the clinical-skill assessment was a simulation scored on a study "standard rubric," and the satisfaction instrument was "created by the author." These are all custom instruments built for this study.
The primary outcome, the Mini-CEX, is described as a "validated" workplace-based rating tool. However, it is a clinician-scored observational rating instrument applied to a PHN encounter rather than a standardised, externally administered subject examination of the kind the standard requires, and its content here was study-specific. No national or otherwise widely recognised standardised exam was used.
Criterion E is not met because all outcome measures were custom instructor-developed tests and rating tools rather than widely recognised standardised exams.
-
T
Term Duration
- The intervention lasted only four weeks with outcomes measured immediately, far short of the one-term requirement, and no term-length follow-up was conducted.
- After four weeks of teaching, the educational outcomes were evaluated through theoretical examinations, clinical skill assessments, Mini-Clinical Evaluation Exercise (Mini-CEX) scores, and a teaching satisfaction questionnaire.
Relevant Quotes:
1) "After four weeks of teaching, the educational outcomes were evaluated through theoretical examinations, clinical skill assessments, Mini-Clinical Evaluation Exercise (Mini-CEX) scores, and a teaching satisfaction questionnaire." (Abstract, p. 2)
2) "Every module was handled by the same instructor over a duration of four weeks." (p. 5)
3) "Second, the short duration of intervention failed to reflect the sustained effects and long-term value of the teaching model." (Limitations, p. 14)
4) "In addition, only short-term outcomes were analyzed in the present study; therefore, long-term efficacy and impacts still require prolonged follow-up observation." (p. 15)
Detailed Analysis:
Criterion T requires outcomes to be measured at least one full academic term (approximately 3-4 months) after the intervention begins, with term-long follow-up tracking accepted even for naturally short interventions. Here the teaching intervention lasted four weeks, and all outcomes (theoretical exam, clinical skills, Mini-CEX, satisfaction) were measured immediately after those four weeks. There is no follow-up extending to one term or beyond; the authors themselves acknowledge the "short duration of intervention" and that only "short-term outcomes" were analysed. Four weeks with immediate measurement is well short of a full term.
Criterion T is not met because the interval from intervention start to outcome measurement was only about four weeks, far shorter than one academic term, with no term-length follow-up.
-
D
Documented Control Group
- The control group's size, demographics, baseline course grades, and lecture-based conditions are clearly documented in Table 1 and the text.
- Table 1. General baseline characteristics ... Control Group (n=60) 42/18 ... 22.19+/-0.66 ... 89.10+/-3.34.
Relevant Quotes:
1) "The control group received conventional lecture-based teaching, while the observation group was taught using the AI-assisted PBL-CBL combined model." (Abstract, p. 2)
2) "The control group was taught using lecture-based instruction. The instructors delivered systematic lectures on the theoretical knowledge of PHN using PowerPoint presentations... Medical interns were divided into 10 groups of six students each for discussion and learning." (p. 5)
3) "Table 1. General baseline characteristics ... Control Group (n=60) 42/18 ... 22.19+/-0.66 ... 89.10+/-3.34." (Results, p. 8)
4) "There was no significant difference in the mean grades of basic medical courses between the two groups." (p. 8)
Detailed Analysis:
Criterion D requires the control group to be well-documented, including size, demographics, baseline performance, and the conditions it received. The paper documents the control group thoroughly: it has a defined size (n=60), reported gender split (42 male / 18 female), mean age (22.19+/-0.66 years), and baseline mean grade in basic medical courses (89.10+/-3.34) in Table 1, which also shows the two groups did not differ significantly at baseline. The treatment received by the control group (conventional lecture-based instruction with PowerPoint and small-group discussion) is clearly described.
Criterion D is met because the control group's size, demographics, baseline performance, and instructional conditions are all clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- This was a single-center trial randomising individual interns, with no randomisation at the school or institutional-unit level.
- This study adopted a single-center randomized controlled trial (RCT) design in strict accordance with the CONSORT guidelines.
Relevant Quotes:
1) "This study adopted a single-center randomized controlled trial (RCT) design in strict accordance with the CONSORT guidelines." (Methods, p. 4)
2) "Participants were randomly assigned at a 1:1 ratio using a random number table..." (Methods, p. 4)
3) "First, the single center design with small sample size may lead to selection bias and limit the extrapolation of research results to a wider clinical environment." (Limitations, p. 14)
Detailed Analysis:
Criterion S requires randomisation at the school level, i.e., entire schools or equivalent implementing units being randomly assigned. This study is explicitly a single-center trial conducted within one hospital, with randomisation performed at the individual intern level via a random number table. Only one institution was involved, so there was no randomisation of schools or comparable units.
Criterion S is not met because the study was a single-center trial randomising individual students, with no school-level (or equivalent unit-level) randomisation.
-
I
Independent Conduct
- The same authors designed the intervention, delivered the teaching, and analysed the data, so the study was not conducted independently of the intervention designers despite assessor blinding.
- Yang F and He LP jointly proposed the research direction and designed the study content. Yang F completed the clinical teaching work. ... Yang F conducted the statistical analysis.
Relevant Quotes:
1) "Yang F and He LP jointly proposed the research direction and designed the study content. Yang F completed the clinical teaching work. Wang JM performed data collation. Yang F conducted the statistical analysis. Yang F and Wang JM drafted the manuscript." (Authors' contributions, p. 16)
2) "Participants were randomly assigned at a 1:1 ratio using a random number table, with allocation concealment and assessor blinding implemented to control bias." (p. 4)
3) "All assessments were conducted by the same independent evaluator to ensure inter-rater reliability." (p. 7)
Detailed Analysis:
Criterion I requires the study to be conducted independently from the authors who designed the intervention, to reduce bias in implementation and analysis. Here the same small author team designed the study and the AI-assisted PBL-CBL teaching model ("jointly proposed the research direction and designed the study content"), delivered the teaching ("Yang F completed the clinical teaching work"), and performed the analysis ("Yang F conducted the statistical analysis"). Although the study reports allocation concealment, assessor blinding, and a single "independent evaluator" for the Mini-CEX (measures of assessment independence), the overall evaluation was not conducted by a third party independent of the intervention designers; the designers themselves implemented and analysed the trial. This does not meet the requirement for independent conduct by an external evaluator.
Criterion I is not met because the same team that designed the intervention also delivered the teaching and analysed the data, with no independent third-party conduct of the evaluation.
-
Y
Year Duration
- The study lasted only four weeks with immediate outcome measurement, far short of an academic year, and Term Duration (T) is also not met.
- Every module was handled by the same instructor over a duration of four weeks.
Relevant Quotes:
1) "Every module was handled by the same instructor over a duration of four weeks." (p. 5)
2) "After four weeks of teaching, the educational outcomes were evaluated..." (Abstract, p. 2)
3) "In addition, only short-term outcomes were analyzed in the present study; therefore, long-term efficacy and impacts still require prolonged follow-up observation." (p. 15)
Detailed Analysis:
Criterion Y requires outcomes to be measured at least 75% of one full academic year (roughly 9-10 months) after the intervention begins. This study ran for four weeks with outcomes measured immediately afterward, which is far below the year-duration threshold. Moreover, under the standard, if the weaker Term Duration (T) criterion is not met, then Y cannot be met; here T is not met.
Criterion Y is not met because the study spanned only about four weeks with no year-long tracking, and the weaker Term Duration criterion is also not met.
-
B
Balanced Control Group
- Both groups received the same content over the same four-week period from the same instructor, and the AI-assisted PBL-CBL method is the integral treatment variable, so resource allocation is balanced.
- The two teaching modules on PHN focus on its core cognition, clinical evaluation, treatment plan and rehabilitation management. Every module was handled by the same instructor over a duration of four weeks.
Relevant Quotes:
1) "The two teaching modules on PHN focus on its core cognition, clinical evaluation, treatment plan and rehabilitation management. Every module was handled by the same instructor over a duration of four weeks." (p. 5)
2) "The control group was taught using lecture-based instruction. The instructors delivered systematic lectures on the theoretical knowledge of PHN using PowerPoint presentations. They also provided explanations using typical case images. Medical interns were divided into 10 groups of six students each for discussion and learning." (p. 5)
3) "The observation group adopted a teaching model that combines AI-assisted PBL with CBL, utilizing Doubao AI (V3.8...) as the core AI tool throughout the process." (p. 5)
4) "Both student cohorts receive the same simulated patients presenting with PHN..." (p. 7)
Detailed Analysis:
Criterion B asks whether the intervention and control groups receive balanced time, budget, and materials, unless the additional resource is itself the explicit treatment variable. Comparing the two conditions: both groups covered the same PHN teaching modules (core cognition, clinical evaluation, treatment plan, rehabilitation management), were taught by the same instructor, over the same four-week duration, and both used small-group work. The instructional time and content scope therefore appear balanced. The principal difference is the pedagogical approach itself - the AI-assisted PBL-CBL model (Doubao AI) versus conventional lectures. The AI tool is not a separable add-on of extra instructional hours or budget; it is the integral core of the very intervention being tested ("utilizing Doubao AI ... as the core AI tool throughout the process"). Under the decision logic, where the extra resource is integral to the treatment being evaluated and delivered within equivalent instructional time, the business-as-usual lecture control is the appropriate comparator and the criterion is satisfied.
Criterion B is met because both groups received the same content over the same four-week period from the same instructor, and the AI-assisted PBL-CBL approach is the integral treatment variable rather than an unmatched extra allocation of time or budget.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific PHN AI-assisted PBL-CBL trial exists; the cited studies are related but distinct interventions and contexts.
- In comparison with previous studies on artificial intelligence-assisted teaching, Gizem Ergezen Sahin and colleagues reported higher Mini-CEX scores in the AI-PBL group than in the conventional PBL group, though no statistically significant difference was identified between the two cohorts (21).
Relevant Quotes:
1) "In comparison with previous studies on artificial intelligence-assisted teaching, Gizem Ergezen Sahin and colleagues reported higher Mini-CEX scores in the AI-PBL group than in the conventional PBL group, though no statistically significant difference was identified between the two cohorts (21)." (Discussion, p. 12)
2) "In another relevant study conducted by Zeng Hui and team, the PBL-ChatGPT group presented remarkable improvements in medical interview skills, clinical judgement and overall clinical competence relative to the lecture-based learning group." (pp. 12-13)
3) "Future research will expand the sample size, conduct multicenter prospective studies, extend the follow-up period..." (p. 15)
Detailed Analysis:
Criterion R requires that this specific study be independently replicated by a different research team in a different context, published in a peer-reviewed journal. The paper cites other AI-assisted teaching studies (Ergezen Sahin et al., ref 21, a physiotherapy RCT; Hui et al., ref 17, a ChatGPT-assisted PBL study in general clinical medical education), but these are different interventions in different contexts, not independent replications of this specific AI-assisted PBL-CBL PHN trial. An external internet search (Springer/BMC, PubMed/PMC, and general web) as of the review date found no independent reproduction of this particular PHN teaching study by any other author team; the study is an in-press 2026 single-center trial with no identified replication.
Criterion R is not met because no independent replication of this specific AI-assisted PBL-CBL PHN study has been conducted; the cited works are related but distinct interventions and contexts.
-
A
All-subject Exams
- Outcomes covered only the single topic of post-herpetic neuralgia, not all main subjects across the curriculum.
- This study aimed to explore the efficacy of a teaching model integrating artificial intelligence-assisted problem-based learning (PBL) with case-based learning (CBL) in standardized clinical teaching of postherpetic neuralgia (PHN) for medical interns.
Relevant Quotes:
1) "This study aimed to explore the efficacy of a teaching model integrating artificial intelligence-assisted problem-based learning (PBL) with case-based learning (CBL) in standardized clinical teaching of postherpetic neuralgia (PHN) for medical interns." (Abstract, p. 2)
2) "The theoretical test ... is as per the fundamental theories of PHN, standardized diagnostic and treatment process and rehabilitation management." (p. 7)
3) "A mini-clinical evaluation exercise (Mini-CEX) was used to assess students' abilities in history taking, physical examination, and clinical reasoning." (p. 7)
Detailed Analysis:
Criterion A requires that the study measure impact across all main subjects of the curriculum using standardised exam-based assessments, to detect unintended effects on non-target subjects. This study is confined to a single, narrow clinical topic - post-herpetic neuralgia (PHN) - and all outcomes (theoretical exam, clinical skills, Mini-CEX) assess only PHN-related knowledge and clinical competence. No other subjects across the medical curriculum were assessed, and no justification is given that would frame this as a permissible specialised exception broad enough to satisfy the all-subject requirement.
Criterion A is not met because outcomes were measured only for the single PHN topic rather than across all main subjects of the curriculum.
-
G
Graduation Tracking
- Outcomes were measured immediately after four weeks with no follow-up to graduation, and Year Duration (Y) is also not met.
- In addition, only short-term outcomes were analyzed in the present study; therefore, long-term efficacy and impacts still require prolonged follow-up observation.
Relevant Quotes:
1) "After four weeks of teaching, the educational outcomes were evaluated..." (Abstract, p. 2)
2) "In addition, only short-term outcomes were analyzed in the present study; therefore, long-term efficacy and impacts still require prolonged follow-up observation." (p. 15)
3) "Second, the short duration of intervention failed to reflect the sustained effects and long-term value of the teaching model." (p. 14)
Detailed Analysis:
Criterion G requires participants to be tracked until graduation to assess long-term impact. This study measured outcomes immediately after the four-week intervention, with no follow-up and explicit acknowledgement that only short-term outcomes were analysed and that long-term follow-up is left for future work. An external internet search for any subsequent follow-up publication by the same authors tracking this intern cohort to graduation found none. Additionally, under the standard, if Year Duration (Y) is not met then G cannot be met; here Y is not met.
Criterion G is not met because outcomes were assessed immediately after a four-week intervention with no tracking to graduation, and the Year Duration criterion is also not met.
-
P
Pre-Registered
- No trial pre-registration (registry, ID, or date) is reported or locatable externally; ethics approval and CONSORT adherence do not satisfy the pre-registration requirement.
Relevant Quotes:
1) "This study adopted a single-center randomized controlled trial (RCT) design in strict accordance with the CONSORT guidelines." (Methods, p. 4)
2) "The study was approved by the medical ethics committee of the Affiliated Hospital of Yunnan University (ethics approval number: KJK#030-25-5R1), and written informed consent was obtained from all participants." (p. 4)
3) "This study was conducted in accordance with the Declaration of Helsinki and approved by the Medical Ethics Committee of Yunnan University Affiliated Hospital. The approval number is KJK#030-25-5R1." (p. 15)
Detailed Analysis:
Criterion P requires that the full study protocol be pre-registered in a trial registry before data collection begins, including hypotheses, methods, and planned analyses, with a verifiable registration date. The paper reports CONSORT adherence and ethics-committee approval, but ethics approval and CONSORT compliance are not the same as prospective public pre-registration. There is no mention of any trial registry (e.g., ClinicalTrials.gov, ISRCTN, ChiCTR), no registration number, and no registration date. An external internet search of trial registries and the web for any pre-registration of this study returned no matching registration record, confirming that the protocol was not publicly pre-registered before data collection.
Criterion P is not met because the paper provides no trial pre-registration reference, ID, or date, and none could be located externally; ethics approval and CONSORT adherence do not satisfy pre-registration.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.