Randomized controlled trial evaluating mixed reality technology for training scrub nurses in anterior cervical spine surgery

Wenqing Yang, Tingting Ma

Published:
ERCT Check Date:
DOI: 10.3389/fmed.2026.1818354
  • adult education
  • China
  • EdTech app
0
  • C

    Randomisation was at the individual nurse level within one operating room department, not at class or school level, and no tutoring exception applies.

    "The 60 enrolled participants were randomly allocated in a 1:1 ratio to either the MR training group or the Traditional training group."

  • E

    All outcome measures were custom instruments built or adapted by the research team for this study, not recognised standardised exams.

    "The test was developed by two senior spine surgeons and one senior nursing educator, reviewed by an external panel of three surgical education experts, and pilot-tested with 10 operating room nurses"

  • T

    Primary outcomes were measured just one week after a single 4-hour training session, far short of one academic term.

    "Assessments were conducted 1 week after training completion by blinded assessors."

  • D

    The control group's size, demographics, professional experience and the exact content of the training it received are all documented in detail.

    "Participants received standard training over 4 h, consisting of: A 2-h lecture from a senior spine surgery nurse educator and an orthopedic surgeon ... A 1-h session observing two standardized ACSS procedure videos ... A 1-h hands-on session for physical identification and handling of the actual surgical instrument sets"

  • S

    The trial was conducted in a single operating room department with individual nurses randomised, so no school- or site-level randomisation occurred.

    "Second, the single-center design and modest sample size (n = 60) limit generalizability; multi-center replication is needed."

  • I

    The same authors developed the MR module and the outcome instruments and ran and analysed the trial themselves, with only partial assessor blinding as a safeguard.

    "This study addresses this gap by developing and implementing a targeted MR training module for ACSS scrub nurse education. We conducted a randomized controlled trial to rigorously compare the efficacy of this innovative approach"

  • Y

    Outcomes were tracked for only one week, nowhere near 75% of an academic year, and criterion T was not met.

    "Fifth, we did not assess long-term skill retention beyond 1 week; it remains unclear whether the observed advantages persist at 3, 6, or 12 months."

  • B

    Both arms received an identical 4 hours of structured training including a shared 2-hour lecture, and the only extra resource, the MR hardware and module, is the treatment variable itself.

    "This group received the same 2-h foundational lecture as the control group. Subsequently, they undertook a 2-h immersive training session"

  • R

    The study presents itself as the first RCT in this area and internet searching found no independent reproduction of this specific trial.

    "However, to our knowledge, no randomized controlled trial has systematically evaluated MR-based training specifically for scrub nurses in anterior cervical spine surgery."

  • A

    Criterion E is not met and outcomes covered only the single narrow domain of ACSS scrub nursing, so the all-subject requirement fails.

    "Primary outcomes were scores on theoretical knowledge, practical performance, and emergency response assessments."

  • G

    Tracking stopped one week after training, no follow-up publication by the same authors was found, and criterion Y was not met.

    "First, the transfer of skills from simulation to live operating rooms remains to be formally evaluated."

  • P

    The paper states the trial was not prospectively registered and no registry entry could be located in any trial registry.

    "At the time of study initiation, prospective registration was not mandated by institutional policy for non-clinical educational studies."

Abstract

Background: Mastery of instrument sequencing and spatial anatomy is crucial for scrub nurses in anterior cervical spine surgery (ACSS), yet conventional training often fails to provide immersive, three-dimensional practice. Aim: This randomized controlled trial evaluated a novel Mixed Reality (MR) training module against standard training for ACSS scrub nurse education. Methods: Sixty nurses were randomly assigned to an MR training group (n = 30) or a traditional training group (n = 30). The intervention group completed a structured MR curriculum using Microsoft HoloLens 2, featuring interactive 3D anatomy, virtual instrument handling, and emergency scenario simulation. The control group received standard lectures and video-based instruction. Primary outcomes were scores on theoretical knowledge, practical performance, and emergency response assessments. Secondary outcomes included cognitive load (NASA-TLX) and training satisfaction. Results: The MR group achieved significantly higher post-training scores in theoretical knowledge (92.47 vs. 86.33, p < 0.001), practical performance (94.50 vs. 88.97, p < 0.001), and emergency response (91.73 vs. 85.40, p < 0.001). Participants in the MR group also reported a significantly lower overall cognitive load (NASA-TLX total: 28.60 vs. 36.43, p < 0.001) and markedly higher satisfaction across all measured domains (p < 0.001). Conclusion: Immersive MR training is more effective than traditional methods in developing essential knowledge, skills, and situational preparedness for ACSS. It facilitates learning by reducing cognitive burden and increasing engagement, presenting a valuable tool for advancing surgical nursing education.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was at the individual nurse level within one operating room department, not at class or school level, and no tutoring exception applies.
      • "The 60 enrolled participants were randomly allocated in a 1:1 ratio to either the MR training group or the Traditional training group."
      • Relevant Quotes: 1) "Sixty nurses were randomly assigned to an MR training group (n = 30) or a traditional training group (n = 30)." (p. 1, Abstract) 2) "The 60 enrolled participants were randomly allocated in a 1:1 ratio to either the MR training group or the Traditional training group. Allocation was performed by an independent statistician using a computer-generated random number sequence, concealed in sequentially numbered, opaque envelopes." (p. 2, Section 2.2) 3) "From June to December 2023, a parallel-group, randomized controlled trial was conducted within the operating room department of our institution." (p. 2, Section 2.1) 4) "Eligible participants were registered nurses with at least 1 year of general operating room experience but no prior independent scrub nurse role in ACSS." (p. 2, Section 2.1) Detailed Analysis: Criterion C requires randomisation at the class level or stronger (school/site level), so that intact teaching groups rather than individual learners within one group are allocated to conditions. The exception permitting individual-level randomisation applies only when the intervention is inherently one-to-one personal teaching or tutoring. Here the unit of randomisation is unambiguously the individual nurse: 60 individual participants recruited from a single operating room department were allocated 1:1 by a computer-generated sequence. There is no mention of classes, cohorts, wards, units, or sites as clusters. All participants were drawn from the same department, so treated and control nurses continued to work alongside one another, creating exactly the contamination risk that criterion C is designed to prevent. The tutoring exception does not apply. The intervention is a group training curriculum: both arms began with a shared "same 2-h foundational lecture," and the MR arm then received a structured module. This is classroom- style/simulation-based group instruction rather than personalised one-to-one tutoring, so the exception cannot rescue individual-level allocation. Criterion C is not met because randomisation was performed on individual nurses within a single department rather than on classes or schools, and no one-to-one tutoring exception applies.
    • E

      Exam-based Assessment

      • All outcome measures were custom instruments built or adapted by the research team for this study, not recognised standardised exams.
      • "The test was developed by two senior spine surgeons and one senior nursing educator, reviewed by an external panel of three surgical education experts, and pilot-tested with 10 operating room nurses"
      • Relevant Quotes: 1) "A 100-point multiple-choice and short-answer test covering anatomy, instrumentation, procedural sequence, and complication management. The test was developed by two senior spine surgeons and one senior nursing educator, reviewed by an external panel of three surgical education experts, and pilot-tested with 10 operating room nurses (not included in the study) to establish face validity and appropriate difficulty." (p. 3, Section 2.5.1) 2) "Performance was rated by two blinded expert nurses using a 10-item checklist scoring instrument preparation, passing accuracy, timeliness, and aseptic technique (0-100). The checklist was developed based on a review of established scrub nurse competency frameworks [citations] and refined through iterative discussion with three senior spine surgery nurses" (p. 3, Section 2.5.2) 3) "Response was scored on a 5-item scale (0-100) assessing problem identification, communication, and solution correctness. This scale was adapted from the validated Non-Technical Skills for Surgeons (NOTSS) framework and modified for scrub nurse roles." (p. 3, Section 2.5.3) 4) "Assessed via a 10-item, 5-point Likert scale questionnaire covering interest, comprehensibility, perceived skill gain, method appeal, and overall satisfaction (total score 10-50). The questionnaire was constructed based on the Simulation Effectiveness Tool- Modified (SET-M) and adapted for the MR context." (p. 3, Section 2.5.5) Detailed Analysis: Criterion E requires that outcomes be measured with standardised, widely recognised exams that were not created for the purpose of the study. Every primary outcome instrument in this trial was purpose-built by the investigating institution. The theoretical knowledge test was "developed by two senior spine surgeons and one senior nursing educator" specifically for this study. The practical performance measure is a bespoke 10-item checklist "developed based on a review of established scrub nurse competency frameworks" and "refined through iterative discussion" with local senior nurses. The emergency response measure is a 5-item scale "adapted from" NOTSS and "modified for scrub nurse roles." The authors report reasonable psychometric work (external expert review, pilot testing, ICC = 0.89, Cronbach's alpha = 0.89 and 0.92). However, local validation of a custom instrument is not the same as using a recognised standardised examination. Adapting and modifying an existing framework (NOTSS, SET-M) likewise produces a study-specific instrument rather than a standardised exam administered in its validated form. No national, professional-body, licensure, or other externally administered standardised examination was used anywhere in the study. Criterion E is not met because all outcome assessments were custom instruments developed or modified by the research team for this study rather than recognised standardised exams.
    • T

      Term Duration

      • Primary outcomes were measured just one week after a single 4-hour training session, far short of one academic term.
      • "Assessments were conducted 1 week after training completion by blinded assessors."
      • Relevant Quotes: 1) "Participants received standard training over 4 h" (p. 2, Section 2.4.1) 2) "This group received the same 2-h foundational lecture as the control group. Subsequently, they undertook a 2-h immersive training session using a custom-developed "ACSS MR Training Module" deployed on the Microsoft HoloLens 2 platform." (p. 2, Section 2.4.2) 3) "Assessments were conducted 1 week after training completion by blinded assessors." (p. 3, Section 2.5) 4) "Measured immediately post-training using the NASA Task Load Index (NASA-TLX)" (p. 3, Section 2.5.4) 5) "Fifth, we did not assess long-term skill retention beyond 1 week; it remains unclear whether the observed advantages persist at 3, 6, or 12 months." (p. 5, Section 4.3) Detailed Analysis: Criterion T requires that primary outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Short interventions are permitted, but the follow-up tracking interval must still reach a term. In this trial the intervention was a single 4-hour training block, and the primary outcomes (theoretical knowledge, practical performance, emergency response) were collected "1 week after training completion." The secondary cognitive load measure was taken immediately post-training. The interval from intervention start to primary outcome measurement is therefore approximately one week, roughly an order of magnitude short of one academic term. The authors themselves acknowledge this explicitly as a limitation, stating they "did not assess long-term skill retention beyond 1 week" and that persistence at 3, 6, or 12 months is unknown. The study period runs June to December 2023, but that is the recruitment/enrolment window across successive participants, not the follow-up duration for any individual learner. Criterion T is not met because outcomes were measured only one week after a single-session intervention, far short of the required one-term follow-up interval.
    • D

      Documented Control Group

      • The control group's size, demographics, professional experience and the exact content of the training it received are all documented in detail.
      • "Participants received standard training over 4 h, consisting of: A 2-h lecture from a senior spine surgery nurse educator and an orthopedic surgeon ... A 1-h session observing two standardized ACSS procedure videos ... A 1-h hands-on session for physical identification and handling of the actual surgical instrument sets"
      • Relevant Quotes: 1) "Sixty nurses were randomly assigned to an MR training group (n = 30) or a traditional training group (n = 30). ... The control group received standard lectures and video-based instruction." (p. 1, Abstract) 2) "Participants received standard training over 4 h, consisting of: A 2-h lecture from a senior spine surgery nurse educator and an orthopedic surgeon, covering ACSS anatomy, procedural steps, instrument names/functions, and key nursing considerations. A 1-h session observing two standardized ACSS procedure videos. A 1-h hands-on session for physical identification and handling of the actual surgical instrument sets on a sterile table." (p. 2, Section 2.4.1) 3) "Baseline demographic and professional characteristics were well-balanced between the two groups (Table 1), confirming successful randomization." (p. 3, Section 3) 4) "TABLE 1 Baseline characteristics of participants. ... Age (years, mean +/- SD) 28.43 +/- 3.67 / 29.10 +/- 4.12 ... Gender ... Male 4 (13.3) / 3 (10.0) ... Female 26 (86.7) / 27 (90.0) ... Education level ... Associate degree 5 (16.7) / 4 (13.3) ... Bachelor's degree 23 (76.7) / 24 (80.0) ... Master's degree 2 (6.6) / 2 (6.7) ... OR experience (years, Mean +/- SD) 5.37 +/- 2.45 / 5.83 +/- 2.89" (p. 4, Table 1) 5) "Eligible participants were registered nurses with at least 1 year of general operating room experience but no prior independent scrub nurse role in ACSS. Nurses with a history of vestibular disorders, epilepsy, or who had participated in similar training within the preceding 3 months were excluded." (p. 2, Section 2.1) 6) "All 60 participants completed the study protocol with no dropouts." (p. 3, Section 3) Detailed Analysis: Criterion D requires detailed documentation of the control group: its size, demographic composition, baseline characteristics, and the conditions/treatment it received. This paper documents the control arm thoroughly. Its size is stated (n = 30) with no attrition. Table 1 reports control-group age, gender distribution, education level breakdown, and years of operating room experience, each with a formal comparison against the intervention arm and p-values, allowing readers to verify baseline comparability. Eligibility and exclusion criteria applying to both arms are stated. Crucially, what the control group actually received is specified in detail rather than left as a vague "business as usual": a named 4-hour package comprising a 2-hour lecture by a nurse educator and orthopaedic surgeon, a 1-hour standardised video observation session, and a 1-hour hands-on instrument handling session. This level of specificity about control-group conditions exceeds what many trials provide. The only minor gap is that no baseline (pre-training) knowledge or skill scores are reported, but the criterion is satisfied by the demographic, professional-experience, size, and treatment documentation supplied. Criterion D is met because the control group's size, demographic and professional baseline characteristics, and the exact content and duration of the training it received are all clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • The trial was conducted in a single operating room department with individual nurses randomised, so no school- or site-level randomisation occurred.
      • "Second, the single-center design and modest sample size (n = 60) limit generalizability; multi-center replication is needed."
      • Relevant Quotes: 1) "From June to December 2023, a parallel-group, randomized controlled trial was conducted within the operating room department of our institution." (p. 2, Section 2.1) 2) "The 60 enrolled participants were randomly allocated in a 1:1 ratio to either the MR training group or the Traditional training group." (p. 2, Section 2.2) 3) "Second, the single-center design and modest sample size (n = 60) limit generalizability; multi-center replication is needed." (p. 5, Section 4.3) Detailed Analysis: Criterion S requires randomisation among schools, that is among the institutions or sites implementing the intervention (hospitals, centres, clubs, preschools, etc.). Here the entire trial took place inside a single operating room department of one hospital, and the authors explicitly describe the trial as "single-center." With only one site there is by construction no site-level randomisation; the randomised units were the 60 individual nurses. The authors themselves flag that "multi-center replication is needed," confirming that no institution-level allocation occurred in this study. Criterion S is not met because the trial was single-centre with individual participants, not institutions or sites, as the unit of randomisation.
    • I

      Independent Conduct

      • The same authors developed the MR module and the outcome instruments and ran and analysed the trial themselves, with only partial assessor blinding as a safeguard.
      • "This study addresses this gap by developing and implementing a targeted MR training module for ACSS scrub nurse education. We conducted a randomized controlled trial to rigorously compare the efficacy of this innovative approach"
      • Relevant Quotes: 1) "This study addresses this gap by developing and implementing a targeted MR training module for ACSS scrub nurse education. We conducted a randomized controlled trial to rigorously compare the efficacy of this innovative approach against established traditional training methods." (p. 2, Section 1) 2) "The training module was developed in collaboration with the institutional simulation center." (p. 2, Section 2.4.2) 3) "The test was developed by two senior spine surgeons and one senior nursing educator" (p. 3, Section 2.5.1) 4) "Allocation was performed by an independent statistician using a computer-generated random number sequence" (p. 2, Section 2.2) 5) "Due to the nature of the intervention, participants and trainers could not be blinded to group assignment. However, outcome assessors for the practical and emergency response assessments were blinded to group allocation." (p. 2, Section 2.2) 6) "WY: Conceptualization, Investigation, Project administration, Resources, Writing - original draft, Writing - review & editing. TM: Methodology, Project administration, Resources, Supervision, Validation, Writing - original draft, Writing - review & editing." (p. 6, Author contributions) 7) "The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest." (p. 6, Conflict of interest) Detailed Analysis: Criterion I requires that the evaluation be conducted independently of those who designed the intervention, either through a genuinely external evaluator or through explicit third-party oversight of data collection, analysis, and conclusions. The paper states plainly that the authors themselves developed the intervention: "This study addresses this gap by developing and implementing a targeted MR training module," with the module built "in collaboration with the institutional simulation center" of the same hospital. The same team also authored the outcome instruments, ran the trial in their own department, performed the analysis, and wrote the conclusions, per the author contribution statement (Conceptualization, Methodology, Investigation, Supervision, Validation). There is no external evaluation agency, no independent evaluator, and no statement limiting the developers' role in data collection or interpretation. Two partial safeguards are present: an "independent statistician" generated the allocation sequence, and outcome assessors for the practical and emergency response assessments were blinded to allocation. These reduce allocation and measurement bias and are methodologically commendable, but they are narrow. The independent statistician's stated role is limited to allocation, not analysis or conclusions; the blinded assessors scored instruments the authors themselves designed; and the theoretical knowledge test was not covered by the blinding statement. The absence of a commercial conflict of interest is not the same as independence from the intervention's designers. Criterion I is not met because the authors designed the MR module, the outcome instruments, and the trial, and analysed and interpreted the results themselves, with only partial blinding rather than genuine third-party evaluation.
    • Y

      Year Duration

      • Outcomes were tracked for only one week, nowhere near 75% of an academic year, and criterion T was not met.
      • "Fifth, we did not assess long-term skill retention beyond 1 week; it remains unclear whether the observed advantages persist at 3, 6, or 12 months."
      • Relevant Quotes: 1) "This group received the same 2-h foundational lecture as the control group. Subsequently, they undertook a 2-h immersive training session using a custom-developed "ACSS MR Training Module"" (p. 2, Section 2.4.2) 2) "Assessments were conducted 1 week after training completion by blinded assessors." (p. 3, Section 2.5) 3) "Fifth, we did not assess long-term skill retention beyond 1 week; it remains unclear whether the observed advantages persist at 3, 6, or 12 months. Longitudinal follow-up is essential to evaluate the durability of training effects" (p. 5, Section 4.3) Detailed Analysis: Criterion Y requires outcomes to be measured at least 75% of a full academic year (roughly 7 or more months) after the intervention begins. Per the ranking instructions, if criterion T is not met then criterion Y cannot be met. Criterion T is not met here: the only follow-up interval was one week. Independently of that dependency, the evidence directly fails Y as well. The intervention was a single 4-hour session and the sole outcome measurement point was one week later, which is approximately 2% of an academic year. The authors explicitly disclaim any longer-term assessment, noting that retention at 3, 6, or 12 months remains unknown and that longitudinal follow-up is still needed. Criterion Y is not met because the tracking interval was one week rather than at least 75% of an academic year, and criterion T was likewise not met.
    • B

      Balanced Control Group

      • Both arms received an identical 4 hours of structured training including a shared 2-hour lecture, and the only extra resource, the MR hardware and module, is the treatment variable itself.
      • "This group received the same 2-h foundational lecture as the control group. Subsequently, they undertook a 2-h immersive training session"
      • Relevant Quotes: 1) "Participants received standard training over 4 h, consisting of: A 2-h lecture from a senior spine surgery nurse educator and an orthopedic surgeon ... A 1-h session observing two standardized ACSS procedure videos. A 1-h hands-on session for physical identification and handling of the actual surgical instrument sets on a sterile table." (p. 2, Section 2.4.1) 2) "This group received the same 2-h foundational lecture as the control group. Subsequently, they undertook a 2-h immersive training session using a custom-developed "ACSS MR Training Module" deployed on the Microsoft HoloLens 2 platform." (p. 2, Section 2.4.2) 3) "This randomized controlled trial evaluated a novel Mixed Reality (MR) training module against standard training for ACSS scrub nurse education." (p. 1, Abstract) 4) "The initial investment for MR hardware (e.g., Microsoft HoloLens 2) and custom software development presents a practical barrier to widespread adoption." (p. 5, Section 4.4) 5) "Third, the novelty of the MR intervention may have introduced a Hawthorne effect ... Future studies should consider a sham technology control group (e.g., using HoloLens 2 to view 2D videos) to isolate the specific effect of MR content." (p. 5, Section 4.3) Detailed Analysis: Criterion B asks whether the control condition provides a comparable substitute for the intervention's inputs in time, budget, and materials, unless the additional resource is itself the explicit treatment variable. Applying the decision tree: first, are extra resources present? On instructional time, no. Both arms received exactly 4 hours in total, and both began with the identical 2-hour foundational lecture. The control arm's remaining 2 hours were an active, structured programme (1 hour of standardised video observation plus 1 hour of hands-on handling of the real instrument sets on a sterile table), matching the MR arm's 2-hour immersive module hour for hour. This is a genuine active control with equivalent educator-led engagement, not a do-nothing comparison, and the hands-on instrument session is arguably a resource-comparable substitute for the virtual instrument handling. Contact time and instructional structure are therefore balanced. There remains a budget asymmetry: the MR arm used Microsoft HoloLens 2 headsets and a custom-developed software module, which the authors acknowledge carries substantial "initial investment for MR hardware ... and custom software development." However, that hardware and software are precisely the treatment variable. The study's stated purpose is to evaluate "a novel Mixed Reality (MR) training module against standard training," so the MR technology is integral to what is being tested, not a separable add-on that could have been given to controls. Under the standard's exception, and by analogy with the DPL and PAX-GBG worked examples, such integral technology costs do not break balance. One caveat worth recording: the authors themselves note that without a sham-technology control the novelty of the headset cannot be separated from the MR content. That is a threat to attributing the effect to MR content specifically, but it is a mechanism-isolation issue rather than a time or resource imbalance, and criterion B concerns the latter. Criterion B is met because both arms received an identical 4 hours of structured, educator-led training including a shared 2-hour lecture, and the only extra resource (the HoloLens 2 hardware and custom MR module) is the explicit treatment variable under test.
  • Level 3 Criteria

    • R

      Reproduced

      • The study presents itself as the first RCT in this area and internet searching found no independent reproduction of this specific trial.
      • "However, to our knowledge, no randomized controlled trial has systematically evaluated MR-based training specifically for scrub nurses in anterior cervical spine surgery."
      • Relevant Quotes: 1) "However, to our knowledge, no randomized controlled trial has systematically evaluated MR-based training specifically for scrub nurses in anterior cervical spine surgery." (p. 2, Section 1) 2) "Second, the single-center design and modest sample size (n = 60) limit generalizability; multi-center replication is needed." (p. 5, Section 4.3) 3) "Addressing these limitations through multi-center trials, longitudinal designs, and cost-effectiveness analyses will be critical to define the role of MR training in surgical nursing education (19)." (p. 5, Section 4.3) 4) "For example, some studies have demonstrated the feasibility of MR for instrument recognition training in orthopedic settings (5), and positive outcomes for situational awareness training using head-mounted displays have also been reported (8)." (p. 2, Section 1) Detailed Analysis: Criterion R requires that this specific study have been independently replicated by a different research team in a different context and published in a peer-reviewed journal. The paper positions itself as the first of its kind, stating that "no randomized controlled trial has systematically evaluated MR-based training specifically for scrub nurses in anterior cervical spine surgery." A first-of-its-kind study cannot by definition already have been replicated. The authors further call for replication as future work, explicitly noting that "multi-center replication is needed." Internet verification (July 2026) was carried out across the publisher record, general web search, and clinical trial registries. Searches for the article title, the DOI 10.3389/fmed.2026.1818354, the author names, and for any MR/HoloLens scrub nurse ACSS training replication returned only the original Frontiers in Medicine article itself plus thematically adjacent but distinct work by other teams. These related studies include San Martin-Rodriguez et al. (2019), "Augmented reality for training operating room scrub nurses," Med Educ 53:514-5; Turso-Finnich et al. (2023), "Virtual reality head-mounted displays in medical education: a systematic review," Simul Healthc 18:42-50; and Zabaleta et al. (2024), "Clinical trial on nurse training through virtual reality simulation of an operating room: assessing satisfaction and outcomes," Cir Esp 102:469-76. None of these reproduces this specific ACSS MR curriculum, its custom outcome instruments, or its design; several predate it. Under the standard's worked example on criterion R, thematically related trials in other contexts do not constitute independent reproduction of the particular study. No verbatim replication quotes could be located because no replication study exists. Given a publication date of 13 May 2026, roughly two months before this check, no subsequent independent replication could plausibly have been published and peer reviewed. Criterion R is not met because the study describes itself as the first RCT in this area, calls for replication as future work, and internet searching found no independent replication of this specific trial.
    • A

      All-subject Exams

      • Criterion E is not met and outcomes covered only the single narrow domain of ACSS scrub nursing, so the all-subject requirement fails.
      • "Primary outcomes were scores on theoretical knowledge, practical performance, and emergency response assessments."
      • Relevant Quotes: 1) "Primary outcomes were scores on theoretical knowledge, practical performance, and emergency response assessments. Secondary outcomes included cognitive load (NASA-TLX) and training satisfaction." (p. 1, Abstract) 2) "A 100-point multiple-choice and short-answer test covering anatomy, instrumentation, procedural sequence, and complication management." (p. 3, Section 2.5.1) 3) "In a simulated OR setting, participants performed as the scrub nurse for a standardized ACSS procedure on a spine model." (p. 3, Section 2.5.2) 4) "Response was scored on a 5-item scale (0-100) assessing problem identification, communication, and solution correctness." (p. 3, Section 2.5.3) Detailed Analysis: Criterion A requires that impact be measured across all main subjects of the curriculum using standardised exam-based assessments, and it explicitly depends on criterion E: if E is not met, A cannot be met. Criterion E is not met here, since every outcome instrument was custom-built or modified by the research team. That alone is decisive. Substantively, the outcomes are also confined to a single narrow domain: anterior cervical spine surgery scrub nursing, covering ACSS anatomy, instrumentation, procedural sequence, and intraoperative emergency response. No broader nursing curriculum subjects, and certainly no general academic subjects, were assessed. One could argue the specialised-vocational exception might apply to a highly specialised procedural training module, but the exception in the standard still requires the assessments themselves to be standardised exams, which they are not. Criterion A is not met because criterion E fails and because outcomes were confined to the single specialised domain of ACSS scrub nursing.
    • G

      Graduation Tracking

      • Tracking stopped one week after training, no follow-up publication by the same authors was found, and criterion Y was not met.
      • "First, the transfer of skills from simulation to live operating rooms remains to be formally evaluated."
      • Relevant Quotes: 1) "Assessments were conducted 1 week after training completion by blinded assessors." (p. 3, Section 2.5) 2) "Fifth, we did not assess long-term skill retention beyond 1 week; it remains unclear whether the observed advantages persist at 3, 6, or 12 months. Longitudinal follow-up is essential to evaluate the durability of training effects and to determine the appropriate frequency of refresher training." (p. 5, Section 4.3) 3) "First, the transfer of skills from simulation to live operating rooms remains to be formally evaluated." (p. 5, Section 4.3) 4) "In this controlled simulation-based setting, immersive MR training demonstrated superior short-term outcomes compared to traditional methods" (p. 6, Section 5) Detailed Analysis: Criterion G requires that participants be tracked through to graduation from their educational stage, and per the ranking instructions it cannot be met if criterion Y is not met. Criterion Y is not met, so G fails on that dependency alone. On the substance, the participants are already qualified registered nurses in continuing professional development, so there is no graduation endpoint of the conventional sort; the nearest analogue would be long-term tracking into independent ACSS scrub practice. The study did no such tracking. Measurement stopped one week after training, the authors describe their results as "short-term outcomes," and they explicitly state that retention beyond one week was not assessed and that transfer to live operating rooms "remains to be formally evaluated." Internet verification (July 2026) searched for subsequent publications by Wenqing Yang and Tingting Ma of the Department of Operating Room, The Affiliated Suqian Hospital of Xuzhou Medical University, that might report longer-term or graduation-equivalent tracking of this cohort. No such follow-up paper, conference abstract, or preprint was located in any available source; the only retrievable output by this team on this trial is the Frontiers in Medicine article itself. No verbatim follow-up quotes can therefore be supplied, because no follow-up publication was found. Criterion G is not met because tracking stopped one week after training with no long-term or graduation- equivalent follow-up, no follow-up publication exists, and criterion Y was not met.
    • P

      Pre-Registered

      • The paper states the trial was not prospectively registered and no registry entry could be located in any trial registry.
      • "At the time of study initiation, prospective registration was not mandated by institutional policy for non-clinical educational studies."
      • Relevant Quotes: 1) "This study was designed as an educational intervention trial without patient-related outcomes. At the time of study initiation, prospective registration was not mandated by institutional policy for non-clinical educational studies. To ensure transparency, the full study protocol, raw data, and ethical documentation are available from the corresponding author upon reasonable request." (p. 2, Section 2.3) 2) "The study protocol received approval from the Ethics Review Committee of Suqian Hospital Affiliated to Xuzhou Medical University (Approval No: IRB-S-2023012)." (p. 2, Section 2.1) 3) "The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation." (p. 6, Data availability statement) Detailed Analysis: Criterion P requires the full protocol, including hypotheses, methods, and planned analyses, to be publicly pre-registered before data collection began, with a verifiable registry reference and date. The paper is admirably candid that this did not happen. Section 2.3 states that "prospective registration was not mandated by institutional policy for non-clinical educational studies," and offers protocol availability "from the corresponding author upon reasonable request" as a substitute. No registry platform is named, no registration identifier is given, and no registration date is provided. Internet verification (July 2026) confirmed this. The publisher record at frontiersin.org displays no trial registration number, and registry searches for the trial (including by ethics approval number IRB-S-2023012, by the intervention name, and by the author names) returned no matching entry in ChiCTR or ClinicalTrials.gov. There is therefore no registration record whose date could be compared against the June 2023 study start. Institutional ethics approval (IRB-S-2023012) obtained before the trial is a governance safeguard, not a public pre-registration: an IRB submission is not a public registry entry and does not lock in hypotheses and analysis plans in a way third parties can inspect. Likewise, availability of a protocol on request after publication provides no timestamped guarantee that the analysis plan predated data collection, which is the precise protection criterion P exists to give. Criterion P is not met because the study was explicitly not prospectively registered and no registry entry could be found in any trial registry.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.