AI-assisted case-based learning and flipped classroom to improve clinical decision-making: a randomized controlled trial in reproductive medicine

Qi Wang, Cong Hu, Yiyang Li, Ting Zhang, Songling Zhang

Published:
ERCT Check Date:
DOI: 10.1080/10872981.2026.2670047
  • higher education
  • China
  • flipped classroom
  • blended learning
  • EdTech platform
  • mobile learning
0
  • C

    Randomisation was performed at the individual student level (by lottery), not at the class or school level, and no one-to-one tutoring exception applies.

    "Fifty resident physicians who met the inclusion criteria were randomly assigned to two groups via lottery, with 25 participants in each group (n = 25 each)."

  • E

    The study relied on custom, study-specific instruments (a bespoke ART knowledge test, plus study-administered Mini-CEX and OSCE), not a widely recognised standardised exam.

    "Post-course assessment included: Theoretical Knowledge Test (maximum score: 100 points) covering key elements of ART theory."

  • T

    Outcomes were assessed only one to two weeks after a short six-session course, well short of the required term-long tracking from intervention start.

    "To evaluate knowledge retention, a post-course examination was administered two weeks after the completion of the teaching intervention."

  • D

    The control group (n = 25, traditional lecture) is clearly documented with baseline demographics and performance in Table 1.

    "Baseline characteristics, including age, pre-course grade point average (GPA), pre-course test scores, total duration of internship, and sex were compared between the two groups (Table 1)."

  • S

    This is a single-centre trial with individual student-level randomisation, so no school-level randomisation occurred.

    "Participants were randomly allocated by lot drawing into either an AI assistied CBL + FC group or a control group."

  • I

    The same author team designed, delivered, and analysed the intervention; only outcome scoring was delegated to blinded examiners, which does not constitute independent conduct.

    "In contrast to conventional AI-assisted instructional models, our approach implements a dynamic adaptation to individual learning trajectories."

  • Y

    Term Duration (T) is not met and follow-up was only one to two weeks, far short of a full academic year, so Y is not met.

    "The relatively short follow-up period also poses challenges in assessing the long-term impact of AI-assisted learning on clinical decision-making."

  • B

    Instructional time was comparable across groups and the additional AI resources are the explicit treatment variable being tested, so the balance requirement is satisfied.

    "The total cumulative study time remained comparable between the two cohorts (Supplementary Table S2)."

  • R

    The study is a novel single-centre trial with no independent replication by another research team, and none was found in an external search.

    "In addition, the study design was a single-centre trial with a limited sample size ... Multi-centre, longitudinal trials with larger sample sizes are needed to further substantiate our findings."

  • A

    Criterion E is not met and outcomes were limited to the single domain of reproductive medicine, so the all-subject standardised-exam requirement fails.

    "Post-course assessment included: Theoretical Knowledge Test (maximum score: 100 points) covering key elements of ART theory."

  • G

    Year Duration (Y) is not met and participants were not tracked to graduation, only assessed one to two weeks post-course, so G is not met.

    "The relatively short follow-up period also poses challenges in assessing the long-term impact of AI-assisted learning on clinical decision-making."

  • P

    No public pre-registration on a trial registry is reported; the protocol is only available on request and an ethics approval number is not pre-registration.

Abstract

Background: Efficient training of reproductive medicine clinicians is critical in the context of declining global fertility and increasing infertility. Traditional lecture-based instruction often fails to sufficiently develop clinical decision-making skills within limited residency rotations. Innovative strategies that integrate artificial intelligence (AI) with case-based learning (CBL) and flipped classroom (FC) formats may enhance clinical reasoning, however rigorous evidence in reproductive medicine education remains limited. Methods: We conducted a randomized controlled trial involving 50 obstetrics and gynecology residents at the First Hospital of Jilin University. Participants were randomly assigned to an AI-assisted CBL+FC group or a traditional lecture control group. The AI-assisted CBL+FC group completed pre-class interactive case work with virtual standardized patients on the DoctorU platform and case analyses on the Superstar Learning platform, followed by interactive in-class discussions. Primary outcomes included post-course theoretical knowledge tests, Mini-Clinical Evaluation Exercise (Mini-CEX), and Objective Structured Clinical Examination (OSCE) scores. Secondary outcomes assessed learner motivation, clinical thinking, self-directed learning, and perceived course effectiveness using a 5-point Likert scale. Results: Baseline characteristics were comparable between groups. After the intervention, the AI-assisted CBL+FC group achieved significantly higher theoretical test scores than the control group. The AI-assisted CBL+FC group also demonstrated superior overall clinical competence in Mini-CEX assessments and higher OSCE total scores. Participants in the AI-assisted CBL+FC group reported greater improvements in learning motivation, clinical reasoning, self-directed learning, and perceived course effectiveness. Conclusions: The AI-assisted CBL+FC instructional model significantly enhances theoretical knowledge, clinical decision-making skills, and learner engagement among reproductive medicine residents. This blended learning model offers an efficacious and generalizable methodology for training practitioners to address the evolving clinical requirements within contemporary fertility care.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was performed at the individual student level (by lottery), not at the class or school level, and no one-to-one tutoring exception applies.
      • "Fifty resident physicians who met the inclusion criteria were randomly assigned to two groups via lottery, with 25 participants in each group (n = 25 each)."
      • Relevant Quotes: 1) "Participants were randomly allocated by lot drawing into either an AI assistied CBL + FC group or a control group." (p. 2) 2) "Fifty resident physicians who met the inclusion criteria were randomly assigned to two groups via lottery, with 25 participants in each group (n = 25 each)." (p. 4) 3) "We conducted a randomized controlled trial involving 50 obstetrics and gynecology residents at the First Hospital of Jilin University. Participants were randomly assigned to an AI-assisted CBL+FC group or a traditional lecture control group." (Abstract) Detailed Analysis: Criterion C requires that randomisation be conducted at the class level (or the stronger school level), assigning entire intact classes rather than individual students within a single class, to prevent contamination between treatment and control learners. In this study, randomisation was performed at the individual student level: 50 residents were individually allocated "by lot drawing" / "via lottery" into the two arms. The unit of randomisation was the individual resident, not a class or school. The exception for personal/one-to-one tutoring does not apply here, because the intervention is a group-based classroom teaching model (pre-class work plus in-class small-group discussions and lectures), not one-to-one tutoring. Because individual residents from the same residency cohort were randomised into the two arms and then presumably taught in group settings, there is a clear risk of cross-group contamination, which is exactly what class-level randomisation is meant to avoid. Criterion C is not met because randomisation was conducted at the individual student level rather than at the class or school level, and no tutoring exception applies.
    • E

      Exam-based Assessment

      • The study relied on custom, study-specific instruments (a bespoke ART knowledge test, plus study-administered Mini-CEX and OSCE), not a widely recognised standardised exam.
      • "Post-course assessment included: Theoretical Knowledge Test (maximum score: 100 points) covering key elements of ART theory."
      • Relevant Quotes: 1) "Post-course assessment included: Theoretical Knowledge Test (maximum score: 100 points) covering key elements of ART theory. Comprehensive Skills Assessment-combining learning progress analysis from Superstar Learning, the Mini-CEX assessmen and an OSCE format." (p. 4) 2) "The Mini-CEX scoring system uses a 9-point scale, with seven criteria: medical interviewing skills, physical examination skills, humanistic qualities/professionalism, clinical judgement, counselling skills, organisation efficiency an overall clinical competence." (p. 3) 3) "The OSCE assessment encompassed three domains: medical history collection, physical examination, and report analysis. Each item has a maximum score of 20 points, with a total possible score of 60 points." (p. 3) Detailed Analysis: Criterion E requires that outcomes be measured with a standardised, widely recognised exam-based assessment that was not specially designed for the study (e.g., a national curriculum exam or a state-wide standardised achievement test). Here, the primary outcome measures are a Theoretical Knowledge Test "covering key elements of ART theory" that was constructed for the course, a Mini-CEX rated by the study's own examiners on a stable infertile patient case, and an OSCE with three study-defined stations scored out of 60. Mini-CEX and OSCE are recognised assessment formats in medical education, but they are not fixed, externally administered standardised examinations; the specific cases, content, and scoring here were assembled by the study team for this trial. The theoretical test is a custom instrument. None of these constitutes a widely recognised standardised exam in the sense required by the ERCT standard. Criterion E is not met because the assessments were custom-built study-specific instruments (a bespoke ART knowledge test plus study-administered Mini-CEX and OSCE), not a standardised, widely recognised external exam.
    • T

      Term Duration

      • Outcomes were assessed only one to two weeks after a short six-session course, well short of the required term-long tracking from intervention start.
      • "To evaluate knowledge retention, a post-course examination was administered two weeks after the completion of the teaching intervention."
      • Relevant Quotes: 1) "This randomised controlled trial was conducted between September 2025 and December 2025, in the Department of Obstetrics and gynaecologist, First Hospital of Jilin University." (p. 2) 2) "Total instructional time was six sessions for each group." (p. 3) 3) "To evaluate knowledge retention, a post-course examination was administered two weeks after the completion of the teaching intervention." (p. 4) 4) "Objective clinical performance was further assessed via an OSCE conducted within one week post-course." (p. 5) Detailed Analysis: Criterion T requires that primary outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins, i.e., term-long tracking from intervention start to measurement. Although the trial window spans September to December 2025, the actual intervention consisted of only six instructional sessions, and the primary outcome measurements were taken shortly after the intervention: the theoretical knowledge test "two weeks after the completion of the teaching intervention," the OSCE "within one week post-course," and the Mini-CEX immediately after the knowledge test. There is no evidence that outcomes were tracked for at least a full term from the intervention start; the follow-up interval is on the order of one to two weeks after a short course. Criterion T is not met because the outcomes were measured only one to two weeks after a short six-session intervention, far short of the required term-long (~3-4 month) tracking from intervention start.
    • D

      Documented Control Group

      • The control group (n = 25, traditional lecture) is clearly documented with baseline demographics and performance in Table 1.
      • "Baseline characteristics, including age, pre-course grade point average (GPA), pre-course test scores, total duration of internship, and sex were compared between the two groups (Table 1)."
      • Relevant Quotes: 1) "Participants in the control group were assigned individual pre-class preparation, which was followed by a traditional didactic lecture where the instructor provided detailed, pre-scripted explanations." (p. 3) 2) "Fifty resident physicians who met the inclusion criteria were randomly assigned to two groups via lottery, with 25 participants in each group (n = 25 each)." (p. 4) 3) "Baseline characteristics, including age, pre-course grade point average (GPA), pre-course test scores, total duration of internship, and sex were compared between the two groups (Table 1). There was no statistically significant differences (P > 0.05), confirming baseline comparability." (p. 4) 4) "Table 1. Comparison of baseline characteristics between the two groups. Control group ... Age (years) 24.60 +/- 1.12 ... Pre-course GPA 2.94 +/- 0.20 ... Pre-course test 63.80 +/- 7.54 ... Duration of internship 1.54 +/- 0.89 ... Male(0.12) Female(0.88)." (p. 5) Detailed Analysis: Criterion D requires that the control group be well-documented, including its size, demographic and baseline characteristics, and the conditions/treatment it received. The paper clearly documents the control arm: it had n = 25 participants, received individual pre-class preparation followed by a traditional didactic lecture (the "business as usual" comparator), and Table 1 provides the control group's baseline age, GPA, pre-course test scores, internship duration, and sex distribution alongside the intervention group. This level of detail allows readers to assess baseline comparability and the nature of the control condition. Criterion D is met because the control group's size, baseline demographic and performance characteristics, and instructional condition are explicitly documented in the text and Table 1.
  • Level 2 Criteria

    • S

      School-level RCT

      • This is a single-centre trial with individual student-level randomisation, so no school-level randomisation occurred.
      • "Participants were randomly allocated by lot drawing into either an AI assistied CBL + FC group or a control group."
      • Relevant Quotes: 1) "This randomised controlled trial was conducted between September 2025 and December 2025, in the Department of Obstetrics and gynaecologist, First Hospital of Jilin University." (p. 2) 2) "Participants were randomly allocated by lot drawing into either an AI assistied CBL + FC group or a control group." (p. 2) 3) "Fifty resident physicians who met the inclusion criteria were randomly assigned to two groups via lottery, with 25 participants in each group (n = 25 each)." (p. 4) Detailed Analysis: Criterion S requires randomisation at the school level (the institution or implementing unit), with multiple schools/sites randomly assigned. This study was conducted at a single institution (the First Hospital of Jilin University) and randomised individual residents into two groups. There was no randomisation of schools, sites, or institutions; it is a single-centre, student-level RCT. Criterion S is not met because randomisation occurred at the individual student level within a single institution, not at the school/site level.
    • I

      Independent Conduct

      • The same author team designed, delivered, and analysed the intervention; only outcome scoring was delegated to blinded examiners, which does not constitute independent conduct.
      • "In contrast to conventional AI-assisted instructional models, our approach implements a dynamic adaptation to individual learning trajectories."
      • Relevant Quotes: 1) "In contrast to conventional AI-assisted instructional models, our approach implements a dynamic adaptation to individual learning trajectories. By integrating with FC framework, we facilitate targeted pedagogical interventions through intensive teacher-student discourse." (p. 2) 2) "All sessions were taught by the same instructor, using the same syllabus covering assisted reproductive medicine topics." (p. 3) 3) "Two independent examiners, who were not involved in the teaching process, were responsible for scoring the Mini-CEX and OSCE assessments. To ensure objectivity, the two examiners were blinded to the students' group assignments and received prior training to standardise the scoring criteria." (p. 4) 4) "CRediT: Qi Wang: Conceptualization, Data curation, Investigation, Software, Writing - original draft; Cong Hu: Conceptualization, Data curation, Formal analysis, Methodology..." (p. 8) Detailed Analysis: Criterion I requires that the study be conducted independently from the authors who designed the intervention, so as to reduce bias in implementation, measurement, analysis, and reporting. In this study, the authors themselves designed the AI-assisted CBL+FC model ("our approach"), delivered the teaching, and performed the data curation, formal analysis, and writing, as shown by the CRediT statement. The only element of independence is that two examiners "not involved in the teaching process" and "blinded to the students' group assignments" scored the Mini-CEX and OSCE. While blinded outcome scoring is a good methodological feature, it is limited to the marking of two outcome measures; it does not amount to independent, third-party conduct of the trial as a whole, because the same team designed the intervention, implemented it, and analysed the results. Criterion I is not met because the intervention was designed, delivered, and analysed by the same author team, with only the scoring of two assessments delegated to blinded examiners rather than independent conduct of the overall study.
    • Y

      Year Duration

      • Term Duration (T) is not met and follow-up was only one to two weeks, far short of a full academic year, so Y is not met.
      • "The relatively short follow-up period also poses challenges in assessing the long-term impact of AI-assisted learning on clinical decision-making."
      • Relevant Quotes: 1) "Total instructional time was six sessions for each group." (p. 3) 2) "To evaluate knowledge retention, a post-course examination was administered two weeks after the completion of the teaching intervention." (p. 4) 3) "The relatively short follow-up period also poses challenges in assessing the long-term impact of AI-assisted learning on clinical decision-making. Multi-centre, longitudinal trials with larger sample sizes are needed to further substantiate our findings." (p. 8) Detailed Analysis: Criterion Y requires that outcomes be measured for at least 75% of a full academic year (roughly 9-10 months) after the intervention begins. Per the specification, if the weaker Term Duration criterion (T) is not met, then Year Duration (Y) cannot be met. Here T is not met, and independently the study only involved six sessions with outcome measurement one to two weeks afterwards; the authors themselves acknowledge a "relatively short follow-up period." This is far short of a full academic year. Criterion Y is not met because criterion T is not met and the study tracked outcomes for only one to two weeks after a short intervention, nowhere near a full academic year.
    • B

      Balanced Control Group

      • Instructional time was comparable across groups and the additional AI resources are the explicit treatment variable being tested, so the balance requirement is satisfied.
      • "The total cumulative study time remained comparable between the two cohorts (Supplementary Table S2)."
      • Relevant Quotes: 1) "All sessions were taught by the same instructor, using the same syllabus covering assisted reproductive medicine topics ... Total instructional time was six sessions for each group." (p. 3) 2) "Participants in the control group were assigned individual pre-class preparation, which was followed by a traditional didactic lecture ... participants in the AI assisted CBL + FC group were instructed to interact with virtual SPs ... as part of their pre-class work and case analysis tasks completed via Superstar Learning." (p. 3) 3) "Although the AI-assisted CBL + FC model required more intensive pre-class preparation compared to the traditional model, which was offset by a significant reduction in post-class review time. The total cumulative study time remained comparable between the two cohorts (Supplementary Table S2)." (p. 7) 4) "This study systematically integrated an AI-assisted CBL approach with the FC model to examine its utility in strengthening clinical decision-making competencies in assisted reproductive medicine." (p. 7) Detailed Analysis: Criterion B compares the nature, quantity, and quality of resources (time, budget, materials) given to the intervention and control groups, unless the additional resource is itself the explicit treatment variable being tested. Both groups received the same number of sessions ("six sessions for each group"), were taught by the same instructor with the same syllabus, and the paper reports that "the total cumulative study time remained comparable between the two cohorts," so instructional time is balanced. The intervention group additionally used the AI platforms (DoctorU virtual standardised patients and Superstar Learning). However, these AI tools are not a separable, confounding add-on; they are precisely the treatment variable under investigation - the study's stated purpose is "to examine [the] utility [of the AI-assisted CBL+FC model] in strengthening clinical decision-making." Applying the decision tree: extra resources are present, but time is comparable and the AI technology is integral to and is the explicit treatment variable being tested against a business-as-usual lecture control. Therefore the criterion is satisfied. Criterion B is met because instructional time was comparable between groups and the additional AI resources constitute the explicit treatment variable being tested, not a separable confounding add-on.
  • Level 3 Criteria

    • R

      Reproduced

      • The study is a novel single-centre trial with no independent replication by another research team, and none was found in an external search.
      • "In addition, the study design was a single-centre trial with a limited sample size ... Multi-centre, longitudinal trials with larger sample sizes are needed to further substantiate our findings."
      • Relevant Quotes: 1) "However, few rigorous empirical studies have systematically examined AI-assisted CBL + FC approaches, particularly in relation to the development of interdisciplinary clinical decision-making competence among reproductive medicine trainers." (p. 2) 2) "In addition, the study design was a single-centre trial with a limited sample size, which may limit the generalisability of our findings ... Multi-centre, longitudinal trials with larger sample sizes are needed to further substantiate our findings." (p. 8) Detailed Analysis: Criterion R requires that this specific study be independently replicated by a different research team in a different context, published in a peer-reviewed journal. The paper is a single, novel single-centre trial published in May 2026; the authors explicitly note that "few rigorous empirical studies have systematically examined AI-assisted CBL + FC approaches" and call for further multi-centre trials, indicating no existing independent replication of this specific intervention and design. An external internet search (PubMed, Taylor & Francis, and general web sources) was carried out and did not identify any independent peer-reviewed replication of this particular AI-assisted CBL+FC reproductive-medicine RCT (using the DoctorU and Superstar Learning platforms) by a different research team. No replication quotes could be found. Criterion R is not met because there is no independent replication of this specific study by a different research team; it is presented as a novel single-centre trial and no replication was found in an external search.
    • A

      All-subject Exams

      • Criterion E is not met and outcomes were limited to the single domain of reproductive medicine, so the all-subject standardised-exam requirement fails.
      • "Post-course assessment included: Theoretical Knowledge Test (maximum score: 100 points) covering key elements of ART theory."
      • Relevant Quotes: 1) "Post-course assessment included: Theoretical Knowledge Test (maximum score: 100 points) covering key elements of ART theory." (p. 4) 2) "The OSCE assessment encompassed three domains: medical history collection, physical examination, and report analysis." (p. 3) 3) "This study systematically integrated an AI-assisted CBL approach with the FC model to examine its utility in strengthening clinical decision-making competencies in assisted reproductive medicine." (p. 7) Detailed Analysis: Criterion A requires that the study measure impact across all main subjects using standardised exam-based assessments, and it explicitly depends on criterion E: if E is not met, A cannot be met. Here criterion E is not met (the assessments are custom, study-specific instruments rather than standardised exams). Furthermore, all outcomes are confined to a single specialised domain (assisted reproductive medicine / ART), with no assessment of other subjects. Both grounds independently preclude satisfying this criterion. Criterion A is not met because criterion E is not met and outcomes were measured only within the single specialised domain of reproductive medicine, not across main subjects using standardised exams.
    • G

      Graduation Tracking

      • Year Duration (Y) is not met and participants were not tracked to graduation, only assessed one to two weeks post-course, so G is not met.
      • "The relatively short follow-up period also poses challenges in assessing the long-term impact of AI-assisted learning on clinical decision-making."
      • Relevant Quotes: 1) "To evaluate knowledge retention, a post-course examination was administered two weeks after the completion of the teaching intervention." (p. 4) 2) "The relatively short follow-up period also poses challenges in assessing the long-term impact of AI-assisted learning on clinical decision-making." (p. 8) Detailed Analysis: Criterion G requires that participants be tracked through to graduation from their educational stage to assess long-term impact, and per the specification it depends on criterion Y: if Y is not met, G cannot be met. Here Y is not met. Additionally, the study measured outcomes only one to two weeks after the intervention and explicitly acknowledges a "relatively short follow-up period," with no tracking of residents through to completion of their residency/graduation. An external search for subsequent follow-up publications by the same author team (Wang, Hu, Li, Zhang, Zhang) that track this cohort to graduation returned no such papers; none could be found, and no graduation-tracking quotes are available. Criterion G is not met because criterion Y is not met and there was no follow-up tracking of participants through to graduation, only a one-to-two-week post-course assessment; no follow-up publications were found.
    • P

      Pre-Registered

      • No public pre-registration on a trial registry is reported; the protocol is only available on request and an ethics approval number is not pre-registration.
      • Relevant Quotes: 1) "The datasets during and/or analysed during the current study are available from the corresponding author on reasonable request. The trial protocol and statistical analysis plan are also available upon request." (p. 8) 2) "This study was approved by the Ethics Committee of the First Hospital of Jilin University [approval number 25K579-001]." (p. 8) Detailed Analysis: Criterion P requires that the full study protocol be publicly pre-registered on a recognised registry before data collection began, including a registry link/identifier and a registration date preceding data collection. The paper states only that the trial protocol and statistical analysis plan are "available upon request" from the corresponding author, and it reports ethics committee approval (approval number 25K579-001). Neither constitutes public pre-registration: there is no reference to a trial registry (e.g., ClinicalTrials.gov, ISRCTN, or the Chinese Clinical Trial Registry), no registration identifier, and no registration date that can be verified as preceding data collection. Because the paper cites no registry identifier, there is no registry entry to verify; an ethics approval number is not a pre-registration. Criterion P is not met because the paper provides no public pre-registration on a recognised registry (no registry ID or pre-data-collection registration date); the protocol is only available on request.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.