Effectiveness of a generative AI-powered digital tutor integrated with a knowledge graph in anatomy education for nursing students: a randomized controlled trial

Can Zhao, Jianzhong Zhu, Jianhui Liu, Wentao Zhao, Yin Pang

Published:
ERCT Check Date:
DOI: 10.1186/s12909-026-09469-0
  • science
  • higher education
  • China
  • blended learning
  • EdTech platform
0
  • C

    Randomisation was performed on intact classes rather than on individual students, explicitly to avoid within-class contamination, satisfying the class-level RCT requirement.

    "To minimize potential contamination between students within the same classroom environment, group allocation was performed at the class level. Each class was assigned to one of the three teaching models according to the generated random sequence." (p. 3)

  • E

    All outcomes relied on locally administered course examinations and questionnaires developed specifically for this study, with no widely recognised standardised exam used.

    "A Blended Teaching Effectiveness Questionnaire was specifically developed for this study based on the study objectives." (p. 4)

  • T

    The intervention spanned a full 16-week semester with outcomes measured at its end and again one month later, exceeding the one-term requirement.

    "All three groups followed the same 16-week teaching schedule with four class hours per week." (p. 3)

  • D

    The control arm's size, demographics, baseline pre-test scores and exact instructional conditions are all documented in the text, Table 1, Table 2 and the CONSORT diagram.

    "In the Traditional Teaching Group, teaching relied on conventional lecture-based instruction using textbooks, slides, and physical anatomical models without digital tutoring support." (p. 3)

  • S

    Randomisation was at the class level within a single medical college, so no school- or institution-level randomisation occurred.

    "First, the study was conducted at a single medical college, which may limit the generalizability of the results to other institutions or educational contexts." (p. 9)

  • I

    The same Cangzhou Medical College team designed the teaching model, delivered it, developed the measures and analysed the data, with independence reported only for the allocation procedure.

    "Authors' contributions CZ Conceptualization, Methodology, Writing - Original Draft. JZ Data Curation. JL Formal Analysis. WZ Investigation. YP Writing - Review & Editing, Supervision." (p. 9)

  • Y

    Tracking spanned only a 16-week semester plus a one-month follow-up, roughly five months, which is well short of 75% of an academic year.

    "Second, the intervention was implemented over a single semester, and long-term learning outcomes beyond the follow-up test were not evaluated." (p. 9)

  • B

    Contact time and instructional structure were identical across arms with active control activities in every instructional link, and the only added resource, the AI digital tutor and knowledge graph, is the explicit treatment variable being tested.

    "All groups followed the same 16-week course structure based on Gagne's Nine Events of Instruction, differing only in the teaching tools and tutoring strategies used." (Abstract, p. 1)

  • R

    No independent replication was found in the paper or by internet search; the study is framed as novel, calls for future validation, and was published only two months before this check.

    "Future multicenter studies are needed to validate the effectiveness of the proposed teaching model across different educational settings." (p. 9)

  • A

    Outcomes were confined to the single intervention course with no measurement of other curriculum subjects, and criterion E was not met, which independently precludes criterion A.

    "All students were undertaking the course Human Anatomy and Histology & Embryology for the first time as part of their nursing curriculum." (p. 3)

  • G

    Tracking stopped one month after the final examination for first-year students years away from graduation, no follow-up publication was found by internet search, and criterion Y was not met.

    "Second, the intervention was implemented over a single semester, and long-term learning outcomes beyond the follow-up test were not evaluated." (p. 9)

  • P

    The paper reports only institutional ethics approval and contains no trial registry identifier, link or registration date, and no registration record was found by internet search of Europe PMC or clinical trial registries.

Abstract

Background: Human Anatomy and Histology & Embryology are foundational courses in nursing education but are often challenging for students due to their complex spatial structures, dense knowledge points, and high cognitive load. Recent advances in generative artificial intelligence (AI) and knowledge graph technologies provide new opportunities for enhancing medical education. This study aimed to evaluate the effectiveness of a blended teaching model integrating a generative AI-powered digital tutor with a knowledge graph in improving learning outcomes among nursing students. Methods: A three-arm parallel randomized controlled trial was conducted among 362 first-year nursing students at Cangzhou Medical College, of whom 301 completed the study. Participants were randomly assigned to an AI-Enhanced Group, a Blended Teaching Group, or a Traditional Teaching Group. All groups followed the same 16-week course structure based on Gagne's Nine Events of Instruction, differing only in the teaching tools and tutoring strategies used. Learning outcomes were evaluated using module examination scores, final comprehensive scores, knowledge graph comprehension scores, knowledge retention rates, case-based inference accuracy, and multidimensional questionnaires. Statistical analyses were conducted using one-way ANOVA with post-hoc comparisons. Results: Significant differences were observed among the three groups across multiple learning outcomes. The AI-Enhanced Group achieved higher module scores (82.5 +/- 7.3) compared with the Blended Group (77.6 +/- 8.1) and Traditional Teaching Group (74.1 +/- 8.5) (F = 28.64, p < 0.001). Final comprehensive scores were also significantly higher in the AI-Enhanced Group (84.2 +/- 6.8) (F = 32.16, p < 0.001). Knowledge graph comprehension scores showed a large effect size (F = 103.72, p < 0.001). Furthermore, the AI-Enhanced Group demonstrated higher knowledge retention rates (82.3%) and case-based inference accuracy (91.2%) than the comparison groups. Questionnaire results indicated stronger recognition of digital tutors, higher acceptance of blended learning, and improved autonomous learning ability. Conclusions: The blended teaching model integrating a generative AI-powered digital tutor with a knowledge graph significantly improved nursing students' academic performance, knowledge retention, and clinical reasoning ability.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was performed on intact classes rather than on individual students, explicitly to avoid within-class contamination, satisfying the class-level RCT requirement.
      • "To minimize potential contamination between students within the same classroom environment, group allocation was performed at the class level. Each class was assigned to one of the three teaching models according to the generated random sequence." (p. 3)
      • Relevant Quotes: 1) "This study employed a three-arm parallel randomized controlled trial to evaluate the effectiveness of an AI-enhanced teaching model integrating a generative AI-powered digital tutor with a knowledge graph in anatomy and histology & embryology education for nursing students." (p. 3) 2) "Randomization was conducted using a computer-generated randomization sequence. To minimize potential contamination between students within the same classroom environment, group allocation was performed at the class level. Each class was assigned to one of the three teaching models according to the generated random sequence." (p. 3) 3) "Participants were recruited from six classes of first-year nursing students enrolled in the 2025 cohort at Cangzhou Medical College, China." (p. 3) 4) "The allocation procedure was conducted by a researcher who was not involved in the teaching intervention or outcome evaluation." (p. 3) Detailed Analysis: Criterion C requires that randomisation be carried out at the class level (or stronger), so that treatment and control conditions are properly isolated and contamination between students in the same room is avoided. The paper states explicitly that the unit of allocation was the intact class: "group allocation was performed at the class level. Each class was assigned to one of the three teaching models according to the generated random sequence." Six intact classes of first-year nursing students were the randomised units, and the paper gives the exact rationale the ERCT standard cares about, namely minimising "potential contamination between students within the same classroom environment." The randomisation mechanism is described (a computer-generated random sequence) and the allocation was performed by a researcher not involved in teaching or outcome evaluation, which supports proper implementation. Sample size is reported (362 enrolled, 301 analysed: n = 100 / 100 / 101). One weakness is that with only six classes the number of randomised clusters is very small, the paper describes no stratification or clustering adjustment in the analysis, and the CONSORT diagram in Fig. 1 labels the post-exclusion sample of 301 as "Randomized (n = 301)", which sits awkwardly with the stated class-level allocation. Nonetheless the unit of randomisation as described in the Methods clearly satisfies the class-level requirement. Criterion C is met because entire classes, not individual students within a class, were randomly assigned to the three teaching conditions.
    • E

      Exam-based Assessment

      • All outcomes relied on locally administered course examinations and questionnaires developed specifically for this study, with no widely recognised standardised exam used.
      • "A Blended Teaching Effectiveness Questionnaire was specifically developed for this study based on the study objectives." (p. 4)
      • Relevant Quotes: 1) "Module scores Module examinations were conducted every four weeks during the semester, resulting in three module tests in total. Each test had a full score of 100 points. The average score across the three module examinations was calculated as the module score." (p. 4) 2) "Final comprehensive score The final comprehensive score was calculated using a weighted evaluation system that included a theoretical examination (60%), practical laboratory assessment (30%), and process evaluation of learning activities (10%). The total score ranged from 0 to 100." (p. 4) 3) "Knowledge graph comprehension score Students' understanding of the knowledge graph learning system was assessed using a Likert-scale questionnaire (1 = completely not understood, 5 = completely understood). The instrument included 20 items across five dimensions: knowledge structure recognition, graph navigation ability, concept association, application ability, and perceived learning support." (p. 4) 4) "Case-based inference accuracy Ten clinical case-based questions were designed to evaluate students' ability to apply anatomical knowledge to clinical scenarios." (p. 4) 5) "A Blended Teaching Effectiveness Questionnaire was specifically developed for this study based on the study objectives." (p. 4) 6) "The questionnaire used in this study has not been previously published elsewhere." (p. 4) 7) "Knowledge retention rate To assess long-term knowledge retention, a follow-up test was conducted one month after the final examination." (p. 4) Detailed Analysis: Criterion E requires that outcomes be measured with standardised, widely recognised exams that were not created for the purposes of the study, so that the assessment is not aligned to the intervention in a way that inflates apparent effectiveness. Every outcome instrument in this trial is internal to the course or purpose-built by the authors. The academic outcomes are in-house module examinations administered every four weeks and a final comprehensive score that is a locally defined weighted composite of a theoretical exam (60%), a practical laboratory assessment (30%) and a "process evaluation of learning activities" (10%). No name of any national, provincial or otherwise externally validated examination is given, and no reliability or validity evidence external to this study is reported for the achievement measures. The remaining outcomes are even further from standardised exams: the knowledge graph comprehension score is a 20-item Likert questionnaire, the case-based inference measure consists of ten clinical questions that were "designed" for this study, and the Blended Teaching Effectiveness Questionnaire was "specifically developed for this study" and "has not been previously published elsewhere." The internal-consistency figures the authors report (Cronbach's alpha values of 0.892, 0.826, 0.863 and 0.845) are reliability statistics computed within this sample and do not make the instruments standardised in the sense the criterion requires. The knowledge retention rate is a ratio of a follow-up test score to the study's own final examination score, so it inherits the non-standardised character of those instruments. There is an additional concern specific to this trial: the knowledge graph comprehension instrument is explicitly about understanding of the knowledge graph learning system that forms part of the intervention itself, which is exactly the kind of intervention-aligned custom measure that criterion E is designed to exclude. Criterion E is not met because all outcomes were measured with in-house course examinations and researcher-developed questionnaires rather than with any recognised standardised exam.
    • T

      Term Duration

      • The intervention spanned a full 16-week semester with outcomes measured at its end and again one month later, exceeding the one-term requirement.
      • "All three groups followed the same 16-week teaching schedule with four class hours per week." (p. 3)
      • Relevant Quotes: 1) "All three groups followed the same 16-week teaching schedule with four class hours per week." (p. 3) 2) "Module examinations were conducted every four weeks during the semester, resulting in three module tests in total." (p. 4) 3) "Knowledge retention rate To assess long-term knowledge retention, a follow-up test was conducted one month after the final examination." (p. 4) 4) "Before the teaching intervention, a pre-test was administered to assess students' baseline course-related knowledge in Human Anatomy and Histology & Embryology." (p. 4) 5) "Second, the intervention was implemented over a single semester, and long-term learning outcomes beyond the follow-up test were not evaluated." (p. 9) Detailed Analysis: Criterion T requires at least one full academic term (roughly 3-4 months) to elapse between the start of the intervention and measurement of the primary outcomes. The intervention here ran for a full 16-week teaching schedule, which the authors themselves describe as "a single semester." The primary academic outcomes (the final comprehensive score, and the third module examination) are measured at the end of that 16-week semester, giving an interval from intervention start to primary measurement of approximately four months. The knowledge retention follow-up test extends measurement by a further month beyond the final examination, so the longest start-to-measurement interval is roughly 20 weeks, about five months. Sixteen weeks corresponds to a standard full semester in the Chinese higher education calendar, and comfortably exceeds the 3-4 month term threshold set by the standard. Baseline anchoring is also clear, since a pre-test was administered before the teaching intervention began. Criterion T is met because outcomes were measured at the end of a full 16-week semester, and again one month later, which is at least one full academic term after the intervention began.
    • D

      Documented Control Group

      • The control arm's size, demographics, baseline pre-test scores and exact instructional conditions are all documented in the text, Table 1, Table 2 and the CONSORT diagram.
      • "In the Traditional Teaching Group, teaching relied on conventional lecture-based instruction using textbooks, slides, and physical anatomical models without digital tutoring support." (p. 3)
      • Relevant Quotes: 1) "In the Traditional Teaching Group, teaching relied on conventional lecture-based instruction using textbooks, slides, and physical anatomical models without digital tutoring support." (p. 3) 2) "The participants were randomly assigned to three groups: the AI-Enhanced Group (n=100), the Blended Teaching Group (n=100), and the Traditional Teaching Group (n=101)." (p. 5) 3) "Baseline characteristics of the participants, including the pre-intervention pre-test score and self-assessment of learning foundation, are presented in Table 2. No statistically significant differences were observed among the three groups in terms of gender distribution (chi2 = 0.126, p=0.94), age (F=0.582, p=0.56), pre-test scores (F=0.143, p=0.87), or self-assessment of learning foundation (F = 0.472, p=0.62). These results indicate that the three groups were comparable at baseline." (p. 5) 4) "Table 2 Comparison of Baseline Data of the Three Groups... Gender (Male/Female, n) 9/91 10/90 9/92... Age (Years, x+/-s) 18.2+/-0.5 18.3+/-0.4 18.2+/-0.6... Pre-test Score (Points, x+/-s) 62.3+/-8.5 61.8+/-9.1 62.1+/-8.8... Self-Assessment of Learning Foundation (x+/-s) 2.86+/-0.5 2.81+/-0.5 2.79+/-0.5" (Table 2, p. 5) 5) "Table 1 Comparison of Teaching Intervention Programs Among the Three Groups... Traditional Teaching Group: Textbook reading+after-class exercise preview; PPT lecture+model demonstration; Offline Q&A+exercise book practice; Paper-based test correction+in-class comments" (Table 1, p. 4) 6) "Follow-up Lost to follow-up (n = 0) Analyzed (n = 101)" (Fig. 1 CONSORT diagram, p. 5) Detailed Analysis: Criterion D requires the control group to be well-documented, including demographic information, baseline performance and the conditions or treatment it received. The paper documents the Traditional Teaching Group thoroughly on all three dimensions. Its size is stated (n = 101 allocated, n = 101 analysed, with zero lost to follow-up per the CONSORT diagram in Fig. 1). Demographics and baseline performance are given in Table 2, which reports gender split (9 male / 92 female), age (18.2 +/- 0.6 years), baseline pre-test score (62.1 +/- 8.8) and baseline self-assessment of learning foundation (2.79 +/- 0.5) separately for the control arm, together with formal tests confirming baseline comparability across arms. The conditions the control group experienced are also described in narrative form and, component by component, in Table 1: conventional lecture-based instruction with textbooks, slides and physical anatomical models, textbook preview and after-class exercise books, offline Q&A, and paper-based test correction with in-class comments. It is explicitly stated that the control group received no digital tutoring support, confirming that no part of the intervention leaked into the control condition. Criterion D is met because the control group's size, demographics, baseline scores and instructional conditions are all reported in detail.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was at the class level within a single medical college, so no school- or institution-level randomisation occurred.
      • "First, the study was conducted at a single medical college, which may limit the generalizability of the results to other institutions or educational contexts." (p. 9)
      • Relevant Quotes: 1) "Randomization was conducted using a computer-generated randomization sequence. To minimize potential contamination between students within the same classroom environment, group allocation was performed at the class level." (p. 3) 2) "Participants were recruited from six classes of first-year nursing students enrolled in the 2025 cohort at Cangzhou Medical College, China." (p. 3) 3) "First, the study was conducted at a single medical college, which may limit the generalizability of the results to other institutions or educational contexts. Future multicenter studies are needed to validate the effectiveness of the proposed teaching model across different educational settings." (p. 9) Detailed Analysis: Criterion S requires randomisation among schools, that is, among the educational institutions or implementing units delivering the intervention, rather than among classes within a single institution. The paper is unambiguous that the unit of randomisation was the class, not the institution: "group allocation was performed at the class level." All six randomised classes were drawn from a single institution, Cangzhou Medical College, so there was only one implementing unit in the entire trial and no institution-level variation could be randomised over. The authors themselves acknowledge this as a design limitation, noting that "the study was conducted at a single medical college" and calling for "future multicenter studies." Since a single-site trial cannot by construction randomise at the school level, criterion S cannot be satisfied here. Criterion S is not met because randomisation occurred among classes within one medical college rather than among schools or institutions.
    • I

      Independent Conduct

      • The same Cangzhou Medical College team designed the teaching model, delivered it, developed the measures and analysed the data, with independence reported only for the allocation procedure.
      • "Authors' contributions CZ Conceptualization, Methodology, Writing - Original Draft. JZ Data Curation. JL Formal Analysis. WZ Investigation. YP Writing - Review & Editing, Supervision." (p. 9)
      • Relevant Quotes: 1) "To address these gaps, the present study proposes an AI-enhanced blended teaching model that integrates a generative AI-powered digital tutor with a knowledge graph-based learning environment, guided by Gagne's Nine Events of Instruction." (p. 2) 2) "The allocation procedure was conducted by a researcher who was not involved in the teaching intervention or outcome evaluation." (p. 3) 3) "The authors would like to thank the nursing students who participated in this study and the faculty members of Cangzhou Medical College for their support in the teaching intervention and data collection." (Acknowledgements, p. 9) 4) "Authors' contributions CZ Conceptualization, Methodology, Writing - Original Draft. JZ Data Curation. JL Formal Analysis. WZ Investigation. YP Writing - Review & Editing, Supervision. All authors read and approved the final manuscript." (p. 9) 5) "Author details 1 Cangzhou Medical College, No. 39, Jiuhe West Road, Yunhe District, Cangzhou, Hebei Province 061001, China" (p. 10) 6) "Fourth, this study evaluated the educational application of the AI-enhanced teaching model but did not independently validate the technical development process of the generative AI-powered digital tutor or systematically audit the factual accuracy of each AI-generated response." (p. 9) 7) "This research received no external funding." (p. 9) Detailed Analysis: Criterion I requires the trial to be conducted independently of those who designed the intervention, with a clear statement identifying who ran the study and their relationship to the intervention designers, or evidence of third-party oversight. Here the same team both proposed the instructional model and evaluated it. The Background states that "the present study proposes an AI-enhanced blended teaching model," and the author contribution statement assigns conceptualisation, methodology, investigation, data curation and formal analysis all to the five co-authors, every one of whom is affiliated to Cangzhou Medical College, the site where the trial was run. The acknowledgements thank Cangzhou Medical College faculty for support "in the teaching intervention and data collection," which places delivery and measurement inside the same institution as the designers. The one element of separation reported is narrow: "The allocation procedure was conducted by a researcher who was not involved in the teaching intervention or outcome evaluation." This is a useful safeguard against allocation bias, but it covers only the randomisation step. There is no external evaluation agency, no blinded independent assessors for the module and final examinations, and no statement of third-party oversight of the outcome measurement or analysis. Indeed the outcome instruments themselves were developed by the authors, so the same team designed the intervention, built the measures, taught the classes, marked the assessments and ran the analysis. This is materially different from the worked exceptions in the standard, where an external body or a non-developer government agency led the evaluation. The authors also concede in the limitations that they "did not independently validate the technical development process of the generative AI-powered digital tutor," further indicating the absence of independent scrutiny. The funding statement records no external funder and there are no declared competing interests, but absence of a commercial sponsor does not establish evaluator independence from the intervention designers. Criterion I is not met because the intervention designers, the teaching implementers and the outcome evaluators were the same institutional team, with independence documented only for the allocation step.
    • Y

      Year Duration

      • Tracking spanned only a 16-week semester plus a one-month follow-up, roughly five months, which is well short of 75% of an academic year.
      • "Second, the intervention was implemented over a single semester, and long-term learning outcomes beyond the follow-up test were not evaluated." (p. 9)
      • Relevant Quotes: 1) "All three groups followed the same 16-week teaching schedule with four class hours per week." (p. 3) 2) "Knowledge retention rate To assess long-term knowledge retention, a follow-up test was conducted one month after the final examination." (p. 4) 3) "Module examinations were conducted every four weeks during the semester, resulting in three module tests in total." (p. 4) 4) "Second, the intervention was implemented over a single semester, and long-term learning outcomes beyond the follow-up test were not evaluated. Longitudinal studies may provide further insights into the sustained impact of AI-enhanced teaching on students' clinical competence." (p. 9) Detailed Analysis: Criterion Y requires outcomes to be measured at least 75% of one full academic year after the intervention begins. With a typical academic year of roughly 9-10 months, the threshold is approximately 7 to 7.5 months of tracking from intervention start. The intervention ran for 16 weeks, about four months, and the last measurement was a knowledge retention follow-up test administered one month after the final examination. The maximum interval from intervention start to final measurement is therefore approximately 20 weeks, or about five months. That is roughly half of a standard academic year and clearly falls short of the 75% threshold; the shortfall is far larger than the minor technical under-run the standard tolerates. The authors describe the study duration in exactly these terms, stating as a limitation that "the intervention was implemented over a single semester, and long-term learning outcomes beyond the follow-up test were not evaluated," and calling for longitudinal studies. There is no indication anywhere in the paper of tracking into a second semester or across a full academic year. Criterion Y is not met because the total tracking interval of roughly five months is well below 75% of an academic year.
    • B

      Balanced Control Group

      • Contact time and instructional structure were identical across arms with active control activities in every instructional link, and the only added resource, the AI digital tutor and knowledge graph, is the explicit treatment variable being tested.
      • "All groups followed the same 16-week course structure based on Gagne's Nine Events of Instruction, differing only in the teaching tools and tutoring strategies used." (Abstract, p. 1)
      • Relevant Quotes: 1) "All three groups followed the same 16-week teaching schedule with four class hours per week. The instructional process was designed according to Gagne's Nine Events of Instruction, including gaining attention, informing learners of objectives, stimulating recall of prior learning, presenting the content, providing learning guidance, eliciting performance, providing feedback, assessing performance, and enhancing retention and transfer." (p. 3) 2) "The three groups differed primarily in the teaching tools and tutoring strategies employed." (p. 3) 3) "All groups followed the same 16-week course structure based on Gagne's Nine Events of Instruction, differing only in the teaching tools and tutoring strategies used." (Abstract, p. 1) 4) "In the AI-Enhanced Group, students used a generative AI-powered digital tutor integrated with a knowledge graph-based learning environment. In this study, the digital tutor was used as a course-support tool embedded in the learning environment. Its functions were limited to course-related learning support, including personalized preview tasks before class, real-time question answering during class, and adaptive feedback and customized review plans after class based on students' learning performance within the Human Anatomy and Histology & Embryology curriculum." (p. 3) 5) "In the Blended Teaching Group, the same knowledge graph learning resources were provided, but tutoring and feedback were delivered primarily by instructors through conventional blended teaching approaches." (p. 3) 6) "In the Traditional Teaching Group, teaching relied on conventional lecture-based instruction using textbooks, slides, and physical anatomical models without digital tutoring support." (p. 3) 7) "Table 1 Comparison of Teaching Intervention Programs Among the Three Groups: Pre-class Preview - Digital tutor pushes knowledge graph+personalized tasks / Teacher pushes knowledge graph+unified preview tasks / Textbook reading+after-class exercise preview; In-class Teaching - Graph interactive demonstration+AI real-time Q&A / Graph demonstration+teacher Q&A / PPT lecture+model demonstration; After-class Tutoring - Digital tutor's wrong question analysis+customized review plan / Teacher's online Q&A+unified review materials / Offline Q&A+exercise book practice; Assessment and Feedback - AI-generated competency diagnosis report+dynamic task adjustment / Teacher's manual correction+phased feedback / Paper-based test correction+in-class comments" (Table 1, p. 4) Detailed Analysis: Criterion B asks whether the intervention and control conditions received comparable time, budget and materials, unless the additional resources are themselves the treatment variable being tested. Working through the decision tree. Are extra resources present? Instructional time is explicitly matched: all three arms "followed the same 16-week teaching schedule with four class hours per week," under an identical Gagne-based instructional sequence, and the arms are stated to differ "only in the teaching tools and tutoring strategies used." Table 1 confirms that each of the four instructional links (pre-class preview, in-class teaching, after-class tutoring, assessment and feedback) is present in all three arms, with an active substitute in every cell: where the AI arm has digital tutor preview tasks the control has textbook reading and exercise preview; where the AI arm has AI real-time Q&A the control has PPT lecture and model demonstration plus offline Q&A; where the AI arm has an AI competency diagnosis report the control has paper-based test correction and in-class comments. So the control condition is a genuinely active one with a comparable structure and matched contact time, not an empty no-treatment arm. There is nonetheless a real resource difference: the AI-Enhanced arm gains access to the generative AI digital tutor and knowledge graph platform, and the Blended arm gains the knowledge graph resources without the AI tutor. These are technology and tutoring resources the Traditional arm does not have. The question is whether they are integral to the treatment being tested. They plainly are: the study's entire object is to "evaluate the effectiveness of a blended teaching model integrating a generative AI-powered digital tutor with a knowledge graph in improving learning outcomes." The digital tutor and the knowledge graph are not separable add-ons layered onto a curriculum change; they are the treatment variable itself, and the three-arm design is deliberately constructed to decompose them (AI tutor plus graph, versus graph with teacher tutoring, versus neither). The tutor's role is also explicitly bounded to course-related support, so it does not smuggle in extra curriculum content. This matches the standard's exception, and it parallels the worked examples in which added devices and support integral to the intervention being tested against business-as-usual do not break criterion B. Criterion B is met because instructional time was identical across all three arms (16 weeks, four class hours per week, same Gagne-based structure) with active teacher-delivered substitutes in the control condition, and the only resource difference, the AI digital tutor and knowledge graph platform, is the explicit treatment variable under test.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication was found in the paper or by internet search; the study is framed as novel, calls for future validation, and was published only two months before this check.
      • "Future multicenter studies are needed to validate the effectiveness of the proposed teaching model across different educational settings." (p. 9)
      • Relevant Quotes: 1) "Research exploring the integrated application of generative AI tutoring and knowledge graph-supported learning environments remains limited, particularly in foundational medical courses such as anatomy and histology & embryology." (p. 2) 2) "Third, few studies have investigated whether combining AI-powered digital tutoring with knowledge graph-based knowledge organization can improve higher-order learning outcomes such as knowledge retention and clinical reasoning ability." (p. 2) 3) "First, the study was conducted at a single medical college, which may limit the generalizability of the results to other institutions or educational contexts. Future multicenter studies are needed to validate the effectiveness of the proposed teaching model across different educational settings." (p. 9) 4) "Our findings extend this evidence by demonstrating the effectiveness of an integrated AI-tutoring and knowledge-graph framework in a randomized controlled trial setting." (p. 8) 5) "Received: 11 March 2026 / Accepted: 14 May 2026" (p. 10) 6) "Published online: 22 May 2026" (p. 10) Detailed Analysis: Criterion R requires that this study, or its central experimental claim, have been independently replicated by a different research team in a different context and published in a peer-reviewed journal. The paper reports no replication of its own design. Instead it repeatedly frames the intervention as novel and unreplicated: research on integrated generative AI tutoring plus knowledge graph learning "remains limited," "few studies have investigated" the combination, and the authors present their contribution as extending existing evidence "by demonstrating the effectiveness of an integrated AI-tutoring and knowledge-graph framework in a randomized controlled trial setting." The limitations section calls for "future multicenter studies... to validate the effectiveness of the proposed teaching model," which is an explicit statement that such validation has not yet occurred. Internet searching was carried out for this verification on 21 July 2026, covering Europe PMC (the record for this article is PMID 42174555, PMCID PMC13295220), Springer Nature Link, and general web search for replications of a generative AI digital tutor combined with a knowledge graph in anatomy or nursing education. No paper by any other research team replicating this trial was found. The nearest related works located are different studies with different designs, for example "Improving medical students' learning absorption in a knowledge graph based blended learning course" (BMC Medical Education, 2025) and "Clicking one dot opens a whole new world: a qualitative study on using knowledge graphs in surgical nursing education" (BMC Medical Education, 2025); neither reproduces this bundled AI-tutor plus knowledge-graph three-arm RCT in nursing anatomy, and the second is qualitative rather than a trial. No verbatim quotes from a replication study can be provided because no replication study was found. The related work the paper itself cites, such as the AI tutoring surgical-skills trial of Fazlollahi et al. and various ChatGPT and virtual-reality studies in medical education, concerns different interventions in different domains and does not reproduce this specific model. No independent replication could plausibly exist in any case given the timeline: the paper was received 11 March 2026 and published online 22 May 2026, only two months before this assessment, and the underlying cohort is the 2025 intake at a single college. Criterion R is not met because no independent replication of this trial by a different research team has been published, and none was found by internet search.
    • A

      All-subject Exams

      • Outcomes were confined to the single intervention course with no measurement of other curriculum subjects, and criterion E was not met, which independently precludes criterion A.
      • "All students were undertaking the course Human Anatomy and Histology & Embryology for the first time as part of their nursing curriculum." (p. 3)
      • Relevant Quotes: 1) "All students were undertaking the course Human Anatomy and Histology & Embryology for the first time as part of their nursing curriculum." (p. 3) 2) "Before the teaching intervention, a pre-test was administered to assess students' baseline course-related knowledge in Human Anatomy and Histology & Embryology." (p. 4) 3) "The final comprehensive score was calculated using a weighted evaluation system that included a theoretical examination (60%), practical laboratory assessment (30%), and process evaluation of learning activities (10%)." (p. 4) 4) "A Blended Teaching Effectiveness Questionnaire was specifically developed for this study based on the study objectives." (p. 4) 5) "Ten clinical case-based questions were designed to evaluate students' ability to apply anatomical knowledge to clinical scenarios." (p. 4) Detailed Analysis: Criterion A requires that impact be measured across all main subjects taught at that educational level, using standardised exam-based assessments, so that gains in the target subject are not achieved at the expense of other subjects. Two independent grounds defeat this criterion here. First, the prerequisite fails. The standard and the instructions state that if criterion E is not met then criterion A cannot be met. Criterion E was judged not met because all outcomes were in-house course examinations and researcher-developed questionnaires rather than recognised standardised exams. Second, on subject coverage, every reported outcome lies inside the single course under intervention. The pre-test, the three module examinations, the final comprehensive score, the knowledge graph comprehension score, the knowledge retention rate and the case-based inference questions all concern Human Anatomy and Histology & Embryology. First-year nursing students in this programme would concurrently be studying other core subjects such as physiology, biochemistry, nursing fundamentals and general education courses, yet no performance data are reported for any of them, so any spillover or displacement effect on the wider curriculum is invisible. The paper also offers no rationale for restricting measurement to the intervention subject that would engage the specialised or vocational exception. Criterion A is not met because criterion E was not satisfied and because outcomes were confined to the single intervention course with no assessment of other main subjects.
    • G

      Graduation Tracking

      • Tracking stopped one month after the final examination for first-year students years away from graduation, no follow-up publication was found by internet search, and criterion Y was not met.
      • "Second, the intervention was implemented over a single semester, and long-term learning outcomes beyond the follow-up test were not evaluated." (p. 9)
      • Relevant Quotes: 1) "Knowledge retention rate To assess long-term knowledge retention, a follow-up test was conducted one month after the final examination." (p. 4) 2) "Second, the intervention was implemented over a single semester, and long-term learning outcomes beyond the follow-up test were not evaluated. Longitudinal studies may provide further insights into the sustained impact of AI-enhanced teaching on students' clinical competence." (p. 9) 3) "Participants were recruited from six classes of first-year nursing students enrolled in the 2025 cohort at Cangzhou Medical College, China." (p. 3) 4) "Follow-up Lost to follow-up (n = 0) Analyzed (n = 100)" (Fig. 1 CONSORT diagram, p. 5) Detailed Analysis: Criterion G requires that participants be tracked through to graduation from their educational stage, so that long-term impact can be assessed. This is ruled out on two grounds. First, by the prerequisite rule: criterion G cannot be met when criterion Y is not met, and criterion Y failed because tracking covered only about five months. Second, on the substance, measurement stopped at a follow-up test one month after the final examination. The participants were first-year nursing students from the 2025 cohort, so they had at least two further years of their programme ahead of them at the point the study ended; the trial finished long before any of them could graduate. The authors state plainly that "long-term learning outcomes beyond the follow-up test were not evaluated" and identify longitudinal follow-up as future work rather than something already undertaken. For this verification, searches were performed on 21 July 2026 across Europe PMC, Springer Nature Link and general web search for subsequent publications by Can Zhao, Jianzhong Zhu, Jianhui Liu, Wentao Zhao or Yin Pang of Cangzhou Medical College that track this 2025 nursing cohort to graduation. No such follow-up publication was found, so no verbatim quotes from a follow-up paper can be provided. This is unsurprising, since the paper itself was published only in May 2026, two months before this assessment, and the cohort enrolled in 2025. Criterion G is not met because measurement ceased one month after the end-of-semester examination, with first-year students far from graduation, no follow-up tracking published, and criterion Y not met.
    • P

      Pre-Registered

      • The paper reports only institutional ethics approval and contains no trial registry identifier, link or registration date, and no registration record was found by internet search of Europe PMC or clinical trial registries.
      • Relevant Quotes: 1) "The study protocol was reviewed and approved by the Ethics Committee of Cangzhou Medical College. All participants provided informed consent before participation." (p. 3) 2) "The study protocol was approved by the Ethics Committee of Cangzhou Medical College (Approval No. CZMC-2025-06-20). All procedures involving human participants were conducted in accordance with the ethical standards of the institutional research committee and the principles of the Declaration of Helsinki." (Declarations, p. 9) 3) "The datasets generated and/or analyzed during the current study are available from the corresponding author on reasonable request." (p. 9) 4) "The online version contains supplementary material available at https://doi.org/10.1186/s12909-026-09469-0. Supplementary Material 1." (p. 9) 5) "An English-language version of the questionnaire is provided as Supplementary File 1." (p. 4) 6) "This research received no external funding." (p. 9) Detailed Analysis: Criterion P requires that the full study protocol, including hypotheses, methods and planned analyses, be registered on a public registry before data collection begins, with a verifiable link or identifier and a registration date preceding data collection. The paper contains no trial registration statement anywhere. There is no registry name, no registration identifier and no registration date. The only protocol reference is to institutional ethics approval from the Ethics Committee of Cangzhou Medical College under approval number CZMC-2025-06-20. Ethics approval is a necessary institutional safeguard but is not pre-registration: it does not place the hypotheses, outcome definitions or analysis plan in the public domain before data collection, and so provides no protection against selective outcome reporting. For this verification, the Europe PMC record for the article (PMID 42174555, PMCID PMC13295220) was retrieved on 21 July 2026 and contains no trial registration or clinical trial number field. Searches of the Chinese Clinical Trial Registry and general web searches for a registration linked to Cangzhou Medical College, this author team, or this AI digital tutor and knowledge graph intervention returned no matching registry entry. Accordingly no registration record exists whose date could be compared against the start of data collection. Notably this is a randomised controlled trial published in a medical education journal that ordinarily expects prospective registration, yet no registry entry is cited and no explanation for its absence is given. The supplementary material is described only as containing an English-language version of the study questionnaire, not a registered protocol. The absence is material for this trial because the outcome set is large and heterogeneous (module scores, final comprehensive score, knowledge graph comprehension, retention rate, case inference accuracy and five questionnaire indicators) with no primary outcome designated, exactly the situation in which a public pre-specified analysis plan would matter most. Criterion P is not met because the paper reports no trial pre-registration, citing only institutional ethics approval, and no registry entry was located by internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.