Level 1 Criteria
-
C Class-level RCT
- Randomisation was at the individual nurse level within one operating room department, not at class or school level, and no tutoring exception applies.
- "The 60 enrolled participants were randomly allocated in a 1:1 ratio to either the MR training group or the Traditional training group."
- Relevant Quotes: 1) "Sixty nurses were randomly assigned to an MR training group (n = 30) or a traditional training group (n = 30)." (p. 1, Abstract) 2) "The 60 enrolled participants were randomly allocated in a 1:1 ratio to either the MR training group or the Traditional training group. Allocation was performed by an independent statistician using a computer-generated random number sequence, concealed in sequentially numbered, opaque envelopes." (p. 2, Section 2.2) 3) "From June to December 2023, a parallel-group, randomized controlled trial was conducted within the operating room department of our institution." (p. 2, Section 2.1) 4) "Eligible participants were registered nurses with at least 1 year of general operating room experience but no prior independent scrub nurse role in ACSS." (p. 2, Section 2.1) Detailed Analysis: Criterion C requires randomisation at the class level or stronger (school/site level), so that intact teaching groups rather than individual learners within one group are allocated to conditions. The exception permitting individual-level randomisation applies only when the intervention is inherently one-to-one personal teaching or tutoring. Here the unit of randomisation is unambiguously the individual nurse: 60 individual participants recruited from a single operating room department were allocated 1:1 by a computer-generated sequence. There is no mention of classes, cohorts, wards, units, or sites as clusters. All participants were drawn from the same department, so treated and control nurses continued to work alongside one another, creating exactly the contamination risk that criterion C is designed to prevent. The tutoring exception does not apply. The intervention is a group training curriculum: both arms began with a shared "same 2-h foundational lecture," and the MR arm then received a structured module. This is classroom- style/simulation-based group instruction rather than personalised one-to-one tutoring, so the exception cannot rescue individual-level allocation. Criterion C is not met because randomisation was performed on individual nurses within a single department rather than on classes or schools, and no one-to-one tutoring exception applies.
-
E Exam-based Assessment
- All outcome measures were custom instruments built or adapted by the research team for this study, not recognised standardised exams.
- "The test was developed by two senior spine surgeons and one senior nursing educator, reviewed by an external panel of three surgical education experts, and pilot-tested with 10 operating room nurses"
- Relevant Quotes: 1) "A 100-point multiple-choice and short-answer test covering anatomy, instrumentation, procedural sequence, and complication management. The test was developed by two senior spine surgeons and one senior nursing educator, reviewed by an external panel of three surgical education experts, and pilot-tested with 10 operating room nurses (not included in the study) to establish face validity and appropriate difficulty." (p. 3, Section 2.5.1) 2) "Performance was rated by two blinded expert nurses using a 10-item checklist scoring instrument preparation, passing accuracy, timeliness, and aseptic technique (0-100). The checklist was developed based on a review of established scrub nurse competency frameworks [citations] and refined through iterative discussion with three senior spine surgery nurses" (p. 3, Section 2.5.2) 3) "Response was scored on a 5-item scale (0-100) assessing problem identification, communication, and solution correctness. This scale was adapted from the validated Non-Technical Skills for Surgeons (NOTSS) framework and modified for scrub nurse roles." (p. 3, Section 2.5.3) 4) "Assessed via a 10-item, 5-point Likert scale questionnaire covering interest, comprehensibility, perceived skill gain, method appeal, and overall satisfaction (total score 10-50). The questionnaire was constructed based on the Simulation Effectiveness Tool- Modified (SET-M) and adapted for the MR context." (p. 3, Section 2.5.5) Detailed Analysis: Criterion E requires that outcomes be measured with standardised, widely recognised exams that were not created for the purpose of the study. Every primary outcome instrument in this trial was purpose-built by the investigating institution. The theoretical knowledge test was "developed by two senior spine surgeons and one senior nursing educator" specifically for this study. The practical performance measure is a bespoke 10-item checklist "developed based on a review of established scrub nurse competency frameworks" and "refined through iterative discussion" with local senior nurses. The emergency response measure is a 5-item scale "adapted from" NOTSS and "modified for scrub nurse roles." The authors report reasonable psychometric work (external expert review, pilot testing, ICC = 0.89, Cronbach's alpha = 0.89 and 0.92). However, local validation of a custom instrument is not the same as using a recognised standardised examination. Adapting and modifying an existing framework (NOTSS, SET-M) likewise produces a study-specific instrument rather than a standardised exam administered in its validated form. No national, professional-body, licensure, or other externally administered standardised examination was used anywhere in the study. Criterion E is not met because all outcome assessments were custom instruments developed or modified by the research team for this study rather than recognised standardised exams.
-
T Term Duration
- Primary outcomes were measured just one week after a single 4-hour training session, far short of one academic term.
- "Assessments were conducted 1 week after training completion by blinded assessors."
- Relevant Quotes: 1) "Participants received standard training over 4 h" (p. 2, Section 2.4.1) 2) "This group received the same 2-h foundational lecture as the control group. Subsequently, they undertook a 2-h immersive training session using a custom-developed "ACSS MR Training Module" deployed on the Microsoft HoloLens 2 platform." (p. 2, Section 2.4.2) 3) "Assessments were conducted 1 week after training completion by blinded assessors." (p. 3, Section 2.5) 4) "Measured immediately post-training using the NASA Task Load Index (NASA-TLX)" (p. 3, Section 2.5.4) 5) "Fifth, we did not assess long-term skill retention beyond 1 week; it remains unclear whether the observed advantages persist at 3, 6, or 12 months." (p. 5, Section 4.3) Detailed Analysis: Criterion T requires that primary outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Short interventions are permitted, but the follow-up tracking interval must still reach a term. In this trial the intervention was a single 4-hour training block, and the primary outcomes (theoretical knowledge, practical performance, emergency response) were collected "1 week after training completion." The secondary cognitive load measure was taken immediately post-training. The interval from intervention start to primary outcome measurement is therefore approximately one week, roughly an order of magnitude short of one academic term. The authors themselves acknowledge this explicitly as a limitation, stating they "did not assess long-term skill retention beyond 1 week" and that persistence at 3, 6, or 12 months is unknown. The study period runs June to December 2023, but that is the recruitment/enrolment window across successive participants, not the follow-up duration for any individual learner. Criterion T is not met because outcomes were measured only one week after a single-session intervention, far short of the required one-term follow-up interval.
-
D Documented Control Group
- The control group's size, demographics, professional experience and the exact content of the training it received are all documented in detail.
- "Participants received standard training over 4 h, consisting of: A 2-h lecture from a senior spine surgery nurse educator and an orthopedic surgeon ... A 1-h session observing two standardized ACSS procedure videos ... A 1-h hands-on session for physical identification and handling of the actual surgical instrument sets"
- Relevant Quotes: 1) "Sixty nurses were randomly assigned to an MR training group (n = 30) or a traditional training group (n = 30). ... The control group received standard lectures and video-based instruction." (p. 1, Abstract) 2) "Participants received standard training over 4 h, consisting of: A 2-h lecture from a senior spine surgery nurse educator and an orthopedic surgeon, covering ACSS anatomy, procedural steps, instrument names/functions, and key nursing considerations. A 1-h session observing two standardized ACSS procedure videos. A 1-h hands-on session for physical identification and handling of the actual surgical instrument sets on a sterile table." (p. 2, Section 2.4.1) 3) "Baseline demographic and professional characteristics were well-balanced between the two groups (Table 1), confirming successful randomization." (p. 3, Section 3) 4) "TABLE 1 Baseline characteristics of participants. ... Age (years, mean +/- SD) 28.43 +/- 3.67 / 29.10 +/- 4.12 ... Gender ... Male 4 (13.3) / 3 (10.0) ... Female 26 (86.7) / 27 (90.0) ... Education level ... Associate degree 5 (16.7) / 4 (13.3) ... Bachelor's degree 23 (76.7) / 24 (80.0) ... Master's degree 2 (6.6) / 2 (6.7) ... OR experience (years, Mean +/- SD) 5.37 +/- 2.45 / 5.83 +/- 2.89" (p. 4, Table 1) 5) "Eligible participants were registered nurses with at least 1 year of general operating room experience but no prior independent scrub nurse role in ACSS. Nurses with a history of vestibular disorders, epilepsy, or who had participated in similar training within the preceding 3 months were excluded." (p. 2, Section 2.1) 6) "All 60 participants completed the study protocol with no dropouts." (p. 3, Section 3) Detailed Analysis: Criterion D requires detailed documentation of the control group: its size, demographic composition, baseline characteristics, and the conditions/treatment it received. This paper documents the control arm thoroughly. Its size is stated (n = 30) with no attrition. Table 1 reports control-group age, gender distribution, education level breakdown, and years of operating room experience, each with a formal comparison against the intervention arm and p-values, allowing readers to verify baseline comparability. Eligibility and exclusion criteria applying to both arms are stated. Crucially, what the control group actually received is specified in detail rather than left as a vague "business as usual": a named 4-hour package comprising a 2-hour lecture by a nurse educator and orthopaedic surgeon, a 1-hour standardised video observation session, and a 1-hour hands-on instrument handling session. This level of specificity about control-group conditions exceeds what many trials provide. The only minor gap is that no baseline (pre-training) knowledge or skill scores are reported, but the criterion is satisfied by the demographic, professional-experience, size, and treatment documentation supplied. Criterion D is met because the control group's size, demographic and professional baseline characteristics, and the exact content and duration of the training it received are all clearly documented.