Level 1 Criteria
-
C Class-level RCT
- Randomisation was performed on intact classes rather than on individual students, explicitly to avoid within-class contamination, satisfying the class-level RCT requirement.
- "To minimize potential contamination between students within the same classroom environment, group allocation was performed at the class level. Each class was assigned to one of the three teaching models according to the generated random sequence." (p. 3)
- Relevant Quotes: 1) "This study employed a three-arm parallel randomized controlled trial to evaluate the effectiveness of an AI-enhanced teaching model integrating a generative AI-powered digital tutor with a knowledge graph in anatomy and histology & embryology education for nursing students." (p. 3) 2) "Randomization was conducted using a computer-generated randomization sequence. To minimize potential contamination between students within the same classroom environment, group allocation was performed at the class level. Each class was assigned to one of the three teaching models according to the generated random sequence." (p. 3) 3) "Participants were recruited from six classes of first-year nursing students enrolled in the 2025 cohort at Cangzhou Medical College, China." (p. 3) 4) "The allocation procedure was conducted by a researcher who was not involved in the teaching intervention or outcome evaluation." (p. 3) Detailed Analysis: Criterion C requires that randomisation be carried out at the class level (or stronger), so that treatment and control conditions are properly isolated and contamination between students in the same room is avoided. The paper states explicitly that the unit of allocation was the intact class: "group allocation was performed at the class level. Each class was assigned to one of the three teaching models according to the generated random sequence." Six intact classes of first-year nursing students were the randomised units, and the paper gives the exact rationale the ERCT standard cares about, namely minimising "potential contamination between students within the same classroom environment." The randomisation mechanism is described (a computer-generated random sequence) and the allocation was performed by a researcher not involved in teaching or outcome evaluation, which supports proper implementation. Sample size is reported (362 enrolled, 301 analysed: n = 100 / 100 / 101). One weakness is that with only six classes the number of randomised clusters is very small, the paper describes no stratification or clustering adjustment in the analysis, and the CONSORT diagram in Fig. 1 labels the post-exclusion sample of 301 as "Randomized (n = 301)", which sits awkwardly with the stated class-level allocation. Nonetheless the unit of randomisation as described in the Methods clearly satisfies the class-level requirement. Criterion C is met because entire classes, not individual students within a class, were randomly assigned to the three teaching conditions.
-
E Exam-based Assessment
- All outcomes relied on locally administered course examinations and questionnaires developed specifically for this study, with no widely recognised standardised exam used.
- "A Blended Teaching Effectiveness Questionnaire was specifically developed for this study based on the study objectives." (p. 4)
- Relevant Quotes: 1) "Module scores Module examinations were conducted every four weeks during the semester, resulting in three module tests in total. Each test had a full score of 100 points. The average score across the three module examinations was calculated as the module score." (p. 4) 2) "Final comprehensive score The final comprehensive score was calculated using a weighted evaluation system that included a theoretical examination (60%), practical laboratory assessment (30%), and process evaluation of learning activities (10%). The total score ranged from 0 to 100." (p. 4) 3) "Knowledge graph comprehension score Students' understanding of the knowledge graph learning system was assessed using a Likert-scale questionnaire (1 = completely not understood, 5 = completely understood). The instrument included 20 items across five dimensions: knowledge structure recognition, graph navigation ability, concept association, application ability, and perceived learning support." (p. 4) 4) "Case-based inference accuracy Ten clinical case-based questions were designed to evaluate students' ability to apply anatomical knowledge to clinical scenarios." (p. 4) 5) "A Blended Teaching Effectiveness Questionnaire was specifically developed for this study based on the study objectives." (p. 4) 6) "The questionnaire used in this study has not been previously published elsewhere." (p. 4) 7) "Knowledge retention rate To assess long-term knowledge retention, a follow-up test was conducted one month after the final examination." (p. 4) Detailed Analysis: Criterion E requires that outcomes be measured with standardised, widely recognised exams that were not created for the purposes of the study, so that the assessment is not aligned to the intervention in a way that inflates apparent effectiveness. Every outcome instrument in this trial is internal to the course or purpose-built by the authors. The academic outcomes are in-house module examinations administered every four weeks and a final comprehensive score that is a locally defined weighted composite of a theoretical exam (60%), a practical laboratory assessment (30%) and a "process evaluation of learning activities" (10%). No name of any national, provincial or otherwise externally validated examination is given, and no reliability or validity evidence external to this study is reported for the achievement measures. The remaining outcomes are even further from standardised exams: the knowledge graph comprehension score is a 20-item Likert questionnaire, the case-based inference measure consists of ten clinical questions that were "designed" for this study, and the Blended Teaching Effectiveness Questionnaire was "specifically developed for this study" and "has not been previously published elsewhere." The internal-consistency figures the authors report (Cronbach's alpha values of 0.892, 0.826, 0.863 and 0.845) are reliability statistics computed within this sample and do not make the instruments standardised in the sense the criterion requires. The knowledge retention rate is a ratio of a follow-up test score to the study's own final examination score, so it inherits the non-standardised character of those instruments. There is an additional concern specific to this trial: the knowledge graph comprehension instrument is explicitly about understanding of the knowledge graph learning system that forms part of the intervention itself, which is exactly the kind of intervention-aligned custom measure that criterion E is designed to exclude. Criterion E is not met because all outcomes were measured with in-house course examinations and researcher-developed questionnaires rather than with any recognised standardised exam.
-
T Term Duration
- The intervention spanned a full 16-week semester with outcomes measured at its end and again one month later, exceeding the one-term requirement.
- "All three groups followed the same 16-week teaching schedule with four class hours per week." (p. 3)
- Relevant Quotes: 1) "All three groups followed the same 16-week teaching schedule with four class hours per week." (p. 3) 2) "Module examinations were conducted every four weeks during the semester, resulting in three module tests in total." (p. 4) 3) "Knowledge retention rate To assess long-term knowledge retention, a follow-up test was conducted one month after the final examination." (p. 4) 4) "Before the teaching intervention, a pre-test was administered to assess students' baseline course-related knowledge in Human Anatomy and Histology & Embryology." (p. 4) 5) "Second, the intervention was implemented over a single semester, and long-term learning outcomes beyond the follow-up test were not evaluated." (p. 9) Detailed Analysis: Criterion T requires at least one full academic term (roughly 3-4 months) to elapse between the start of the intervention and measurement of the primary outcomes. The intervention here ran for a full 16-week teaching schedule, which the authors themselves describe as "a single semester." The primary academic outcomes (the final comprehensive score, and the third module examination) are measured at the end of that 16-week semester, giving an interval from intervention start to primary measurement of approximately four months. The knowledge retention follow-up test extends measurement by a further month beyond the final examination, so the longest start-to-measurement interval is roughly 20 weeks, about five months. Sixteen weeks corresponds to a standard full semester in the Chinese higher education calendar, and comfortably exceeds the 3-4 month term threshold set by the standard. Baseline anchoring is also clear, since a pre-test was administered before the teaching intervention began. Criterion T is met because outcomes were measured at the end of a full 16-week semester, and again one month later, which is at least one full academic term after the intervention began.
-
D Documented Control Group
- The control arm's size, demographics, baseline pre-test scores and exact instructional conditions are all documented in the text, Table 1, Table 2 and the CONSORT diagram.
- "In the Traditional Teaching Group, teaching relied on conventional lecture-based instruction using textbooks, slides, and physical anatomical models without digital tutoring support." (p. 3)
- Relevant Quotes: 1) "In the Traditional Teaching Group, teaching relied on conventional lecture-based instruction using textbooks, slides, and physical anatomical models without digital tutoring support." (p. 3) 2) "The participants were randomly assigned to three groups: the AI-Enhanced Group (n=100), the Blended Teaching Group (n=100), and the Traditional Teaching Group (n=101)." (p. 5) 3) "Baseline characteristics of the participants, including the pre-intervention pre-test score and self-assessment of learning foundation, are presented in Table 2. No statistically significant differences were observed among the three groups in terms of gender distribution (chi2 = 0.126, p=0.94), age (F=0.582, p=0.56), pre-test scores (F=0.143, p=0.87), or self-assessment of learning foundation (F = 0.472, p=0.62). These results indicate that the three groups were comparable at baseline." (p. 5) 4) "Table 2 Comparison of Baseline Data of the Three Groups... Gender (Male/Female, n) 9/91 10/90 9/92... Age (Years, x+/-s) 18.2+/-0.5 18.3+/-0.4 18.2+/-0.6... Pre-test Score (Points, x+/-s) 62.3+/-8.5 61.8+/-9.1 62.1+/-8.8... Self-Assessment of Learning Foundation (x+/-s) 2.86+/-0.5 2.81+/-0.5 2.79+/-0.5" (Table 2, p. 5) 5) "Table 1 Comparison of Teaching Intervention Programs Among the Three Groups... Traditional Teaching Group: Textbook reading+after-class exercise preview; PPT lecture+model demonstration; Offline Q&A+exercise book practice; Paper-based test correction+in-class comments" (Table 1, p. 4) 6) "Follow-up Lost to follow-up (n = 0) Analyzed (n = 101)" (Fig. 1 CONSORT diagram, p. 5) Detailed Analysis: Criterion D requires the control group to be well-documented, including demographic information, baseline performance and the conditions or treatment it received. The paper documents the Traditional Teaching Group thoroughly on all three dimensions. Its size is stated (n = 101 allocated, n = 101 analysed, with zero lost to follow-up per the CONSORT diagram in Fig. 1). Demographics and baseline performance are given in Table 2, which reports gender split (9 male / 92 female), age (18.2 +/- 0.6 years), baseline pre-test score (62.1 +/- 8.8) and baseline self-assessment of learning foundation (2.79 +/- 0.5) separately for the control arm, together with formal tests confirming baseline comparability across arms. The conditions the control group experienced are also described in narrative form and, component by component, in Table 1: conventional lecture-based instruction with textbooks, slides and physical anatomical models, textbook preview and after-class exercise books, offline Q&A, and paper-based test correction with in-class comments. It is explicitly stated that the control group received no digital tutoring support, confirming that no part of the intervention leaked into the control condition. Criterion D is met because the control group's size, demographics, baseline scores and instructional conditions are all reported in detail.