Urinary Catheterization Training for Nursing Students Using Traditional Instruction, Simulation, and Augmented Reality: A Randomized Controlled Trial

Daniela Dunca, Cristian Valentin Toma, Didina-Catalina Barbalata, Romina-Marina Sima, Nina Rosu, George Andrei Popescu, Ovidiu-Catalin Nechita, Daniel Liviu Badescu, Viorel Jinga

Published:
ERCT Check Date:
DOI: 10.3390/app16105068
  • higher education
  • EU
0
  • C

    Randomisation was at the individual student level within a single cohort, not at the class or school level, and the one-to-one tutoring exception does not clearly apply.

    "Eligible participants were divided into one of the three study groups ... using a computer-generated random sequence provided by an investigator not involved in the workshop's delivery."

  • E

    All outcome instruments were custom-developed by the authors for this study and not formally validated, so no recognised standardised exam was used.

    "Assessment questionnaires and checklists were developed specifically for this study through an internal consensus involving four faculty members with expertise in urology and nursing."

  • T

    The intervention and its outcome measurement occurred within a single same-day workshop, far short of the required one-term interval.

    "the entire protocol (lecture, randomization, practice, and assessment) was conducted consecutively within a single scheduled clinical skills workshop"

  • D

    The control group's size, baseline scores, and conditions received are documented in the text and Table 1, satisfying the documentation requirement.

    "Group C proceeded directly to independent practice using a task-trainer, without feedback from the instructors."

  • S

    Randomisation was at the individual student level within a single faculty, so no school-level randomisation occurred.

    "Eligible participants were divided into one of the three study groups ... using a computer-generated random sequence."

  • I

    The same team designed, delivered, and analysed the trial, with only internal blinding and no independent third-party evaluator.

    "Author Contributions: Conceptualization, D.D., C.V.T. and V.J.; methodology, D.D., C.V.T., N.R. and R.-M.S.; ... investigation, D.D., N.R., O.-C.N. and D.L.B."

  • Y

    Outcomes were measured immediately after a one-day workshop with no year-long tracking, and criterion T is not met.

    "assessment was conducted at a single post-intervention time point, without evaluating long-term retention."

  • B

    Practice time and the shared lecture/video were balanced across arms, and the AR technology and guided practice are integral treatment variables tested against a business-as-usual control.

    "All participants were given a fixed duration of 20 min for practice."

  • R

    No independent replication of this specific trial exists (confirmed by external search); the authors explicitly call for future replication.

    "replicate these findings across institutions with varying resources and curricula"

  • A

    Only a single specialised skill was assessed with custom instruments, and criterion E (the prerequisite) is not met.

    "The primary outcome was catheterization skills in male and female patients, and the secondary outcomes were knowledge and confidence."

  • G

    There was no follow-up beyond immediate assessment and no graduation tracking (no follow-up paper found), and criterion Y is not met.

    "assessment was conducted at a single post-intervention time point, without evaluating long-term retention."

  • P

    The paper reports ethics approval and CONSORT reporting but provides no public pre-registration of the protocol before data collection.

Abstract

(1) Background: Augmented reality (AR) simulation may accelerate psychomotor skill acquisition in clinical education, but comparative evidence is scarce. This three-arm randomized controlled trial compared AR simulation, basic task-trainer simulation, and lecture-based instruction for urinary catheterization training. We hypothesized that AR would be associated with higher performance compared to the other two methods. (2) Methods: Primary outcomes included male and female catheterization skills assessed with checklists. Secondary outcomes included knowledge and confidence. One-way ANOVA with Tukey's HSD post hoc tests constituted the primary analysis; Kruskal-Wallis tests and Bayesian ANOVA provided convergent evidence. (3) Results: A total of 176 trainees were assigned to AR simulation, basic simulator, or lecture-only control groups (N = 60, 58, and 58, respectively). AR simulation was associated with higher skill scores than both basic simulation and the control for male catheterization (p < 0.001) and female catheterization (p < 0.001), with large effect sizes (Cohen's d = 3.11, AR vs. control). Knowledge showed no group difference (p = 0.11). Confidence moderately favored AR. Bayesian analysis supported a high probability of AR outperforming Control. (4) Conclusions: AR simulation training was associated with superior catheterization skills compared to both basic simulation and lecture-based instruction, with large effect sizes observed across both frequentist and Bayesian analyses. Knowledge was consistent across groups, suggesting a possible ceiling effect.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was at the individual student level within a single cohort, not at the class or school level, and the one-to-one tutoring exception does not clearly apply.
      • "Eligible participants were divided into one of the three study groups ... using a computer-generated random sequence provided by an investigator not involved in the workshop's delivery."
      • Relevant Quotes: 1) "Eligible participants were divided into one of the three study groups (group A = AR simulation; group B = basic simulation; group C = control) using a computer-generated random sequence provided by an investigator not involved in the workshop's delivery." (p. 3) 2) "Randomized (n = 176) 1 : 1 : 1 allocation, permuted blocks of 6" (Figure 1, CONSORT flow diagram, p. 4) 3) "All participants first attended a lecture adapted according to the European Association of Urology Nurses 2024 Guidelines ... A video demonstration followed the lecture." (p. 3) 4) "After the demonstrations, the trainees in Group A practiced UC, receiving feedback or guidance from the instructor ... All participants were given a fixed duration of 20 min for practice." (p. 3) Detailed Analysis: Criterion C requires that randomisation be conducted at the class level or stronger (school level), to prevent cross-group contamination. The quotes make clear that the unit of randomisation was the individual student: eligible participants were individually allocated to one of three groups via a computer-generated random sequence with 1:1:1 allocation in permuted blocks of 6. All participants belonged to a single first-year cohort at one faculty and were trained within a shared clinical skills workshop, so entire classes or schools were not the randomised unit. The tutoring/personal-teaching exception is not clearly applicable: although each trainee practised individually, this is a group workshop with a common lecture and video followed by modality-specific practice, not a one-to-one tutoring programme, and students of different arms were present in the same setting, leaving a plausible contamination risk. Criterion C is not met because randomisation was performed at the individual student level within a single cohort rather than at the class or school level, and the tutoring exception does not clearly apply.
    • E

      Exam-based Assessment

      • All outcome instruments were custom-developed by the authors for this study and not formally validated, so no recognised standardised exam was used.
      • "Assessment questionnaires and checklists were developed specifically for this study through an internal consensus involving four faculty members with expertise in urology and nursing."
      • Relevant Quotes: 1) "The knowledge assessment included twenty single-choice questions covering the concepts addressed in the lecture. The test was developed by the study investigators based on the EAUN 2024 Guidelines [30]." (p. 3) 2) "Trainees' performance was assessed by an evaluator blinded to group allocation, using two standardized checklists: one for catheterization of the male patient and one for the female patient." (pp. 3-4) 3) "Assessment questionnaires and checklists were developed specifically for this study through an internal consensus involving four faculty members with expertise in urology and nursing." (p. 4) 4) "Formal psychometric validation of the instruments was not performed." (p. 4) Detailed Analysis: Criterion E requires a standardised, widely recognised exam-based assessment that was not specially designed for the study. Here every outcome instrument (the 20-item knowledge test, the male and female skills checklists, and the confidence scale) was created by the study investigators/faculty specifically for this trial. The paper states the checklists and questionnaires were "developed specifically for this study" and that "formal psychometric validation of the instruments was not performed." The word "standardized" in the paper refers to a fixed checklist format for scoring, not to a nationally or externally recognised standardised exam. Criterion E is not met because all assessments were custom instruments developed by the authors for this study rather than recognised standardised exams.
    • T

      Term Duration

      • The intervention and its outcome measurement occurred within a single same-day workshop, far short of the required one-term interval.
      • "the entire protocol (lecture, randomization, practice, and assessment) was conducted consecutively within a single scheduled clinical skills workshop"
      • Relevant Quotes: 1) "All 176 randomized participants completed the study, as the entire protocol (lecture, randomization, practice, and assessment) was conducted consecutively within a single scheduled clinical skills workshop, eliminating the risk of loss to follow-up." (p. 5) 2) "All participants were given a fixed duration of 20 min for practice." (p. 3) 3) "Participants were evaluated at two time points: before training and after simulation training." (p. 3) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. In this trial the whole protocol, including the lecture, 20 minutes of practice, and the post-training assessment, was carried out consecutively within a single scheduled workshop. The outcome measurement occurred immediately after training on the same day, so the interval between intervention start and measurement was a matter of hours, far short of one term. No term-long follow-up or tracking was conducted. Criterion T is not met because outcomes were measured immediately after a single-session workshop, not at least one academic term after the intervention began.
    • D

      Documented Control Group

      • The control group's size, baseline scores, and conditions received are documented in the text and Table 1, satisfying the documentation requirement.
      • "Group C proceeded directly to independent practice using a task-trainer, without feedback from the instructors."
      • Relevant Quotes: 1) "A total of 176 trainees were assigned to AR simulation, basic simulator, or lecture-only control groups (N = 60, 58, and 58, respectively)." (Abstract) 2) "Group C proceeded directly to independent practice using a task-trainer, without feedback from the instructors." (p. 3) 3) "Baseline characteristics are summarized in Table 1. Groups were balanced at baseline. Knowledge scores did not differ significantly across groups (p = 0.66), nor did confidence scores (p = 0.97)." (p. 5) 4) Table 1 reports for the Control group (N = 58) baseline Knowledge 16.44 +/- 2.29 and Confidence 3.91 +/- 0.33. (p. 5) 5) "The eligibility criteria required participants to be enrolled in the first year of study at the Faculty of Midwifery and Nursing ... and to have no previous experience in performing urinary catheterization." (p. 3) Detailed Analysis: Criterion D requires that the control group be well documented, including its size, baseline performance, and the conditions/treatment it received. The paper clearly identifies the control group (Group C, N = 58), reports its baseline knowledge and confidence scores in Table 1, describes the eligibility profile shared by all arms (first-year midwifery/nursing students with no prior catheterization experience), and specifies what the control received: the common lecture and video followed by independent task-trainer practice without instructor feedback. This provides a comparable baseline and clear documentation of control conditions, even though demographic detail beyond baseline scores is limited. Criterion D is met because the control group's size, baseline outcome data, eligibility profile, and the conditions it received are documented in the text and Table 1.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was at the individual student level within a single faculty, so no school-level randomisation occurred.
      • "Eligible participants were divided into one of the three study groups ... using a computer-generated random sequence."
      • Relevant Quotes: 1) "Eligible participants were divided into one of the three study groups ... using a computer-generated random sequence provided by an investigator not involved in the workshop's delivery." (p. 3) 2) "The study was conducted at the Faculty of Midwifery and Nursing, Carol Davila University of Medicine and Pharmacy, Bucharest, Romania." (p. 3) 3) "Randomized (n = 176) 1 : 1 : 1 allocation, permuted blocks of 6" (Figure 1, CONSORT flow diagram, p. 4) Detailed Analysis: Criterion S requires randomisation at the school level, i.e., entire educational institutions or implementing units randomised to conditions. This trial was conducted at a single faculty within one university, and randomisation was performed at the individual student level. No schools, sites, or institutions were randomised; there was only one institution involved. Criterion S is not met because randomisation occurred at the individual student level within a single institution, not at the school/institution level.
    • I

      Independent Conduct

      • The same team designed, delivered, and analysed the trial, with only internal blinding and no independent third-party evaluator.
      • "Author Contributions: Conceptualization, D.D., C.V.T. and V.J.; methodology, D.D., C.V.T., N.R. and R.-M.S.; ... investigation, D.D., N.R., O.-C.N. and D.L.B."
      • Relevant Quotes: 1) "Author Contributions: Conceptualization, D.D., C.V.T. and V.J.; methodology, D.D., C.V.T., N.R. and R.-M.S.; ... formal analysis, D.-C.B., R.-M.S. and G.A.P.; investigation, D.D., N.R., O.-C.N. and D.L.B." (p. 10) 2) "Eligible participants were divided into one of the three study groups ... using a computer-generated random sequence provided by an investigator not involved in the workshop's delivery." (p. 3) 3) "Trainees' performance was assessed by an evaluator blinded to group allocation." (pp. 3-4) 4) "Assessment questionnaires and checklists were developed specifically for this study through an internal consensus involving four faculty members." (p. 4) Detailed Analysis: Criterion I requires that the study be conducted independently of the team that designed the intervention, typically via a third-party evaluator, to reduce bias. Here the same author group designed the training protocol, built the assessment instruments, delivered the workshop, and analysed the data (as shown in the author contributions covering conceptualization, methodology, investigation, and formal analysis). Two internal measures reduce some bias, namely an investigator not involved in delivery generating the random sequence and an evaluator blinded to allocation, but these are within-team safeguards rather than independent conduct by an external body. There is no external evaluation agency or third-party oversight of the study. Criterion I is not met because the same research team designed, delivered, and evaluated the intervention, with only internal blinding rather than independent third-party conduct.
    • Y

      Year Duration

      • Outcomes were measured immediately after a one-day workshop with no year-long tracking, and criterion T is not met.
      • "assessment was conducted at a single post-intervention time point, without evaluating long-term retention."
      • Relevant Quotes: 1) "the entire protocol (lecture, randomization, practice, and assessment) was conducted consecutively within a single scheduled clinical skills workshop" (p. 5) 2) "All participants were given a fixed duration of 20 min for practice." (p. 3) 3) "assessment was conducted at a single post-intervention time point, without evaluating long-term retention." (p. 9) Detailed Analysis: Criterion Y requires that outcomes be measured at least 75% of a full academic year after the intervention begins. This trial ran entirely within a single workshop with immediate post-training assessment and no follow-up. The authors explicitly acknowledge assessment at "a single post-intervention time point, without evaluating long-term retention." Additionally, because the weaker Term Duration criterion (T) is not met, the stronger Year Duration criterion cannot be met. Criterion Y is not met because outcomes were measured immediately after a single-session workshop, far short of a full academic year (and T is not met).
    • B

      Balanced Control Group

      • Practice time and the shared lecture/video were balanced across arms, and the AR technology and guided practice are integral treatment variables tested against a business-as-usual control.
      • "All participants were given a fixed duration of 20 min for practice."
      • Relevant Quotes: 1) "All participants first attended a lecture adapted according to the European Association of Urology Nurses 2024 Guidelines ... A video demonstration followed the lecture." (p. 3) 2) "Groups A and B received a modality-specific demonstration of catheterization on their corresponding simulator; Group C proceeded directly to independent practice using a task-trainer, without feedback from the instructors." (p. 3) 3) "All participants were given a fixed duration of 20 min for practice." (p. 3) 4) "This three-arm randomized controlled trial compared AR simulation, basic task-trainer simulation, and lecture-based instruction for urinary catheterization training." (Abstract) 5) "differences in instructional guidance between the groups could introduce confounding factors when comparing the effectiveness of guided AR and basic simulation training with the unguided control." (p. 9) Detailed Analysis: Criterion B asks whether the control condition offers a comparable substitute for the intervention's inputs (time, budget, materials), unless the additional resource is itself the treatment variable. On time, the groups are balanced: all arms received the same lecture and video, and all had the same fixed 20 minutes of hands-on practice on a task-trainer. The differences between arms are the AR technology (extra budget/materials) provided to Group A, the basic simulator demonstration for Group B, and instructor guidance/feedback given to Groups A and B but not to the unguided control. These additional resources are integral to the intervention being tested: the study's explicit purpose is to compare AR simulation and basic task-trainer simulation training against lecture-based/unguided instruction, so the technology and the guided-practice component are the treatment variables themselves rather than separable add-ons. The control receiving the business-as-usual lecture plus unguided practice is the intended comparison baseline. The authors note the guidance difference as a potential confounder, but under the ERCT logic the guided simulation package is the modality under test, and practice time was equalised across all groups. Criterion B is met because instructional time (shared lecture, video, and a fixed 20-minute practice) was balanced across arms, and the AR technology and guided-practice differences are integral to the instructional modalities being tested against a business-as-usual control.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific trial exists (confirmed by external search); the authors explicitly call for future replication.
      • "replicate these findings across institutions with varying resources and curricula"
      • Relevant Quotes: 1) "the current literature lacks a reliable comparison between traditional lecture-based teaching, task-trainer simulation, and AR-based training within a rigorous randomized controlled trial (RCT) design." (p. 2) 2) "Long-term skill retention, cross-institutional generalizability, and transfer to real-patient outcomes require further investigation." (p. 9) 3) "replicate these findings across institutions with varying resources and curricula" (p. 10) Detailed Analysis: Criterion R requires independent replication of the specific study by a different research team, published in a peer-reviewed journal. The paper positions itself as filling a gap where no such three-arm RCT comparison exists, and it explicitly calls for future replication across institutions, indicating that replication has not yet occurred. An external internet search (conducted 2026-07-10) found no independent replication of this specific AR three-arm trial. Related studies exist but test different technologies and/or designs and are by different teams, so none is a reproduction of this particular study, e.g.: Chang, C.-L. (2022), "Effect of Immersive Virtual Reality on Post-Baccalaureate Nursing Students' In-Dwelling Urinary Catheter Skill and Learning Satisfaction" (immersive VR); Yoon, H. (2024), "Effects of Immersive Straight Catheterization Virtual Reality Simulation on Skills, Confidence, and Flow State in Nursing Students" (immersive VR); and Schoeb et al. (2020), "Mixed Reality for Teaching Catheter Placement to Medical Students: A Randomized Single-Blinded, Prospective Trial" (mixed reality). These are distinct interventions, not replications of this AR task-trainer trial. Criterion R is not met because there is no independent replication of this specific trial; the authors themselves call for future replication, and no reproducing study was located online.
    • A

      All-subject Exams

      • Only a single specialised skill was assessed with custom instruments, and criterion E (the prerequisite) is not met.
      • "The primary outcome was catheterization skills in male and female patients, and the secondary outcomes were knowledge and confidence."
      • Relevant Quotes: 1) "Primary outcomes included male and female catheterization skills assessed with checklists. Secondary outcomes included knowledge and confidence." (Abstract) 2) "The primary outcome was catheterization skills in male and female patients, and the secondary outcomes were knowledge and confidence." (p. 4) 3) "Assessment questionnaires and checklists were developed specifically for this study." (p. 4) Detailed Analysis: Criterion A requires that all main subjects be assessed using standardised exam-based assessments, and it depends on criterion E as a prerequisite. This study measured only outcomes specific to a single procedural skill (urinary catheterization), namely skills, knowledge, and confidence, with no assessment of broader subject areas. Moreover, because criterion E is not met (all instruments were custom-built and unvalidated rather than standardised exams), criterion A cannot be met. Criterion A is not met because outcomes were confined to a single specialised procedural skill using custom instruments, and criterion E (its prerequisite) is not met.
    • G

      Graduation Tracking

      • There was no follow-up beyond immediate assessment and no graduation tracking (no follow-up paper found), and criterion Y is not met.
      • "assessment was conducted at a single post-intervention time point, without evaluating long-term retention."
      • Relevant Quotes: 1) "assessment was conducted at a single post-intervention time point, without evaluating long-term retention. Therefore, it remains unknown whether group differences persist over time" (p. 9) 2) "Future research should examine long-term retention at 1, 3, and 6 months." (p. 10) 3) "the entire protocol ... was conducted consecutively within a single scheduled clinical skills workshop" (p. 5) Detailed Analysis: Criterion G requires that participants be tracked through to graduation to assess long-term impact, and it depends on criterion Y. This study assessed outcomes at a single time point immediately after the workshop, with no follow-up of any kind, and the authors explicitly note the absence of long-term retention data and recommend future follow-up at 1, 3, and 6 months. There is no graduation tracking. Additionally, because criterion Y is not met, criterion G cannot be met. An external internet search (conducted 2026-07-10) for subsequent publications by the same author team tracking this cohort to graduation found no such follow-up papers; the study was published in May 2026 and no graduation- or retention-tracking companion paper could be located. Criterion G is not met because there was no follow-up beyond immediate post-training assessment, let alone tracking to graduation, no follow-up publication was found, and criterion Y is not met.
    • P

      Pre-Registered

      • The paper reports ethics approval and CONSORT reporting but provides no public pre-registration of the protocol before data collection.
      • Relevant Quotes: 1) "The sample size was determined using G*Power 3.1 [31] for one-way ANOVA ... This calculation yielded a minimum of 156 participants, 52 per group." (p. 3) 2) "The CONSORT 2025 [32] flow diagram is provided in Figure 1." (p. 4) 3) "The study was conducted in accordance with the Declaration of Helsinki and approved by the Institutional Ethics Committee of CAROL DAVILA UNIVERSITY OF MEDICINE AND PHARMACY (protocol code 17465) on 28 June 2024." (p. 10) Detailed Analysis: Criterion P requires a publicly pre-registered protocol, with hypotheses and analysis plan registered before data collection began, and evidence of a registry link and date. The paper reports an a priori power calculation, follows CONSORT 2025 reporting, and cites institutional ethics approval, but it provides no trial registration identifier or link (e.g., ClinicalTrials.gov, ISRCTN) and no statement that the protocol was pre-registered on a public registry before data collection. An external search (conducted 2026-07-10) did not locate any public trial-registry record for this specific trial. Ethics-committee approval is not equivalent to public pre-registration of a protocol. Criterion P is not met because the paper contains no reference to a public pre-registration of the study protocol before data collection, and none was found online.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.