Virtual reality-based training improves thoracentesis skills in medical interns: a randomized controlled trial

Shaoyu Han, Zhen Tian, Bingjun Cui, Lang Wu, Zhixiang Chen

Published:
ERCT Check Date:
DOI: 10.1186/s12909-026-09801-8
  • higher education
  • China
  • EdTech platform
0
  • C

    Randomization was performed at the individual student level within a single department, not at the class or school level, and the intervention is simulation skills practice rather than one-to-one tutoring.

    "Participants were randomly assigned (1:1) to either the VR training group or the control group using a computer-generated random number table."

  • E

    Outcomes were measured with a custom 300-point scoring rubric developed by the study team, not a widely recognized standardized exam.

    "The scoring rubric was developed by three respiratory medicine faculty members based on institutional thoracentesis competency checklists."

  • T

    The whole program lasted only four weeks with final outcomes measured at week 4, far short of a full academic term.

    "All participants completed a 4-week training program."

  • D

    The control group is well documented, including size, baseline demographics, and the traditional mannequin training it received.

    "Control group: Participants practiced on a traditional thoracentesis mannequin in groups of five under the supervision of the same instructor, for 1 hour per day, 4 days per week."

  • S

    This was a single-center trial randomizing individual interns, not schools or institutions.

    "This was a single-center, parallel-group, randomized controlled trial."

  • I

    The same authors designed the VR intervention and conducted the trial; blinding of the assessor does not make the conduct independent of the intervention designers.

    "A custom VR thoracentesis simulator was developed for this study."

  • Y

    The program lasted only four weeks, far short of a full academic year, and criterion T is already not met.

    "All participants completed a 4-week training program."

  • B

    The VR group practiced individually with unlimited repetitions while the control group shared one mannequin in groups of five, an unmatched difference in practice intensity that the authors themselves call the most significant threat to internal validity rather than the treatment variable.

    "the training intensity was not equivalent between groups. The VR group practiced independently with unlimited repetitions, while the control group practiced in groups of five on a single mannequin. This is the most significant threat to internal validity."

  • R

    No independent replication of this specific trial is reported or found; it is described as a preliminary study needing replication.

    "these findings require replication in larger, multi-center trials."

  • A

    Only a single procedural skill (thoracentesis) was assessed with a custom rubric, and since criterion E is not met, A cannot be met.

    "The primary outcome was the final thoracentesis procedural score at week 4."

  • G

    Follow-up ended at week 4 with no tracking to graduation, and criterion Y is already not met.

    "Third, we did not assess long-term skill retention (e.g., at 3 or 6 months)."

  • P

    The authors explicitly state the trial was not prospectively registered.

    "This study was not prospectively registered as it was a small-scale educational intervention. The authors recognize this as a limitation."

Abstract

Background: Thoracentesis is an essential clinical procedure, but its teaching is often limited by patient safety concerns and insufficient opportunities for repeated practice. Virtual reality (VR) offers immersive, repeatable simulation-based training; however, its effectiveness for thoracentesis has not been rigorously evaluated. Methods: In this randomized controlled trial, 20 medical interns were randomly assigned to either a VR-based training group (n=10) or a traditional training control group (n=10). The VR group received theoretical instruction plus VR simulation practice (1 hour/day, 4 days/week for 3 weeks), while the control group received the same theoretical instruction plus traditional mannequin-based practice. All participants completed a 4-week training program. Outcomes were assessed at baseline, week 2, week 3, and week 4 using a standardized 300-point scoring rubric. Results: At the final assessment, 80% (8/10) of the VR group achieved scores >=270 points (excellent), compared to only 10% (1/10) of the control group. Conclusion: VR-based training may improve thoracentesis procedural skills after an initial adaptation period, though these findings are preliminary due to the small sample size.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomization was performed at the individual student level within a single department, not at the class or school level, and the intervention is simulation skills practice rather than one-to-one tutoring.
      • "Participants were randomly assigned (1:1) to either the VR training group or the control group using a computer-generated random number table."
      • Relevant Quotes: 1) "In this randomized controlled trial, 20 medical interns were randomly assigned to either a VR-based training group (n=10) or a traditional training control group (n=10)." (Abstract) 2) "Participants were randomly assigned (1:1) to either the VR training group or the control group using a computer-generated random number table. The randomization sequence was generated by a research assistant not involved in participant training or assessment." (Section 2.2) 3) "This was a single-center, parallel-group, randomized controlled trial conducted between March 2023 and March 2024. Twenty medical interns rotating through the Department of Respiratory Medicine ... were enrolled." (Section 2.1) Detailed Analysis: Criterion C requires randomisation at the class level (or stronger school level), so that treatment and control groups are properly isolated to avoid contamination. Here, randomisation was carried out at the individual student (intern) level within a single department at one hospital: the 20 interns were assigned one-by-one to VR or control using a random number table. This is student-level, not class-level or school-level, randomisation, and all participants shared the same clinical rotation environment, so contamination between individuals cannot be excluded. The exception for personal/one-to-one tutoring does not clearly apply: although VR practice is performed individually, this is self-directed simulation skill practice rather than one-to-one teaching by a tutor, and the control group actually trained in shared groups of five. Therefore the tutoring exception is not a good fit. Criterion C is not met because randomisation was at the individual student level within a single department rather than at the class or school level, and the study is not a one-to-one tutoring intervention.
    • E

      Exam-based Assessment

      • Outcomes were measured with a custom 300-point scoring rubric developed by the study team, not a widely recognized standardized exam.
      • "The scoring rubric was developed by three respiratory medicine faculty members based on institutional thoracentesis competency checklists."
      • Relevant Quotes: 1) "Outcomes were assessed at baseline, week 2, week 3, and week 4 using a standardized 300-point scoring rubric (preoperative, intraoperative, and postoperative components)." (Abstract) 2) "The assessment used a standardized 300-point scoring rubric divided into three domains: Preoperative preparation (100 points) ... Intraoperative operation (100 points) ... Postoperative operation (100 points)." (Section 2.5) 3) "The scoring rubric was developed by three respiratory medicine faculty members based on institutional thoracentesis competency checklists. Content validity was assessed by two external experts ... who rated each item as 'essential' or 'useful' (agreement: 92%)." (Section 2.5) 4) "The cutoff of >=270 points for 'excellent' was defined post-hoc based on the distribution of final scores, as no published standard exists." (Section 3.2) Detailed Analysis: Criterion E requires a standardised, widely recognised exam-based assessment that was not specially designed for the study. Although the authors label their rubric "standardized," it was in fact created by three faculty members from institutional competency checklists specifically for this study, with a post-hoc "excellent" cutoff because "no published standard exists." Internal content validation by two experts and a pilot inter-rater reliability check do not make it a recognised standardised exam; it is a custom-made instrument aligned to the intervention content. Criterion E is not met because the outcome measure was a study-specific rubric rather than a recognised, external standardised exam.
    • T

      Term Duration

      • The whole program lasted only four weeks with final outcomes measured at week 4, far short of a full academic term.
      • "All participants completed a 4-week training program."
      • Relevant Quotes: 1) "All participants completed a 4-week training program." (Abstract and Section 2.4) 2) "The VR group received theoretical instruction plus VR simulation practice (1 hour/day, 4 days/week for 3 weeks)." (Abstract) 3) "Baseline assessment occurred at week 1 before any practical training. Practical training occurred during weeks 2, 3, and 4." (Section 2.4) 4) "The primary outcome was the final thoracentesis procedural score at week 4." (Abstract) 5) "Third, we did not assess long-term skill retention (e.g., at 3 or 6 months)." (Section 4.3) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. In this study the intervention started with practical training in week 2 and the primary outcome was measured at week 4, meaning the interval from intervention start to measurement was only about two to four weeks, and the entire program was four weeks. The authors explicitly note they did not assess retention at 3 or 6 months. Criterion T is not met because the interval from intervention start to outcome measurement was only a few weeks, well short of one full academic term.
    • D

      Documented Control Group

      • The control group is well documented, including size, baseline demographics, and the traditional mannequin training it received.
      • "Control group: Participants practiced on a traditional thoracentesis mannequin in groups of five under the supervision of the same instructor, for 1 hour per day, 4 days per week."
      • Relevant Quotes: 1) "20 medical interns were randomly assigned to either a VR-based training group (n=10) or a traditional training control group (n=10)." (Abstract) 2) "Control group: Participants practiced on a traditional thoracentesis mannequin in groups of five under the supervision of the same instructor, for 1 hour per day, 4 days per week. The instructor demonstrated, observed, and corrected errors." (Section 2.4) 3) "Table 1. Baseline characteristics of participants ... Sex (Male/Female) 6 / 4 ... Age (years, mean +/- SD) 19.10 +/- 1.10 ... Internship weeks (mean +/- SD) 7.70 +/- 1.70 ... Education (Junior college/Undergraduate) 6 / 4." (Section 3.1) 4) "Baseline (week 1): No significant difference between groups (VR: 185.4 +/- 12.3, control: 188.2 +/- 11.6; p = 0.612)." (Section 3.2) Detailed Analysis: Criterion D requires clear documentation of the control group's size, characteristics, and conditions. Table 1 provides the control group's sex, age, internship weeks, and educational level (n=10), the text describes the exact training the control group received (traditional mannequin practice in groups of five, matched schedule), and baseline scores are reported for both groups. This gives a clear picture of who the control group was and what it received. Criterion D is met because the control group's composition, baseline characteristics, and treatment conditions are described in detail.
  • Level 2 Criteria

    • S

      School-level RCT

      • This was a single-center trial randomizing individual interns, not schools or institutions.
      • "This was a single-center, parallel-group, randomized controlled trial."
      • Relevant Quotes: 1) "This was a single-center, parallel-group, randomized controlled trial conducted between March 2023 and March 2024." (Section 2.1) 2) "Participants were randomly assigned (1:1) to either the VR training group or the control group using a computer-generated random number table." (Section 2.2) Detailed Analysis: Criterion S requires randomisation at the level of whole schools or institutional units. This study was explicitly single-center and randomised individual interns within one department, so there is no school-level or site-level randomisation. Criterion S is not met because randomisation was at the individual student level within a single center, not at the school/institution level.
    • I

      Independent Conduct

      • The same authors designed the VR intervention and conducted the trial; blinding of the assessor does not make the conduct independent of the intervention designers.
      • "A custom VR thoracentesis simulator was developed for this study."
      • Relevant Quotes: 1) "A custom VR thoracentesis simulator was developed for this study." (Section 2.3) 2) "Authors' Contributions: Conceptualization: SH, ZT, ZC; Methodology: SH, ZT, BC; Software: BC, LW ... Formal analysis: SH, ZT; Investigation: SH, ZT, BC ... Supervision: ZC." (Declarations) 3) "The instructor who delivered theoretical lectures and supervised practice was not blinded due to the nature of the intervention. However, the assessor who conducted all final skill evaluations was blinded to group allocation." (Section 2.2) 4) "All assessments were performed by the same senior instructor, who was blinded to group allocation." (Section 2.5) Detailed Analysis: Criterion I requires that the evaluation be conducted independently from the team that designed the intervention. Here the same authors developed the custom VR simulator, designed the study, delivered/supervised the training, and performed the formal analysis. While a research assistant handled randomisation and the final skill assessor was blinded to allocation, the assessor was still an internal senior instructor, and there is no external or third-party evaluation team independent of the intervention designers. Blinding reduces detection bias but does not establish independent conduct. Criterion I is not met because the intervention was designed, delivered, and analysed by the same research team, with no independent external evaluator.
    • Y

      Year Duration

      • The program lasted only four weeks, far short of a full academic year, and criterion T is already not met.
      • "All participants completed a 4-week training program."
      • Relevant Quotes: 1) "All participants completed a 4-week training program." (Section 2.4) 2) "The primary outcome was the final thoracentesis procedural score at week 4." (Abstract) 3) "Third, we did not assess long-term skill retention (e.g., at 3 or 6 months)." (Section 4.3) Detailed Analysis: Criterion Y requires tracking of outcomes over at least about 75% of a full academic year. The entire program was four weeks and the final outcome was at week 4, with no longer-term follow-up. Additionally, per the standard, if criterion T (term duration) is not met then Y cannot be met, and T is not met here. Criterion Y is not met because the study spanned only four weeks, nowhere near a full academic year.
    • B

      Balanced Control Group

      • The VR group practiced individually with unlimited repetitions while the control group shared one mannequin in groups of five, an unmatched difference in practice intensity that the authors themselves call the most significant threat to internal validity rather than the treatment variable.
      • "the training intensity was not equivalent between groups. The VR group practiced independently with unlimited repetitions, while the control group practiced in groups of five on a single mannequin. This is the most significant threat to internal validity."
      • Relevant Quotes: 1) "VR group: Each participant practiced independently on the VR simulator for 1 hour per day, 4 days per week ... Participants could repeat procedures unlimited times." (Section 2.4) 2) "Control group: Participants practiced on a traditional thoracentesis mannequin in groups of five under the supervision of the same instructor, for 1 hour per day, 4 days per week." (Section 2.4) 3) "It is important to note that the VR group practiced independently with unlimited repetitions, while the control group practiced in groups of five, sharing one mannequin. This difference in training intensity--rather than VR per se--may have contributed to the observed differences." (Section 4.1) 4) "Second, the training intensity was not equivalent between groups ... This is the most significant threat to internal validity. We cannot determine whether the superior performance of the VR group was due to VR technology itself or simply to more individualized, repetitive practice." (Section 4.3) Detailed Analysis: Criterion B compares the time, budget, and materials given to intervention and control groups and asks whether the control condition offers a comparable substitute, unless the extra resource is explicitly the treatment variable. Both groups had the same nominal schedule (1 hour/day, 4 days/week), but the VR group each used their own simulator with unlimited individual repetitions, whereas the control shared a single mannequin among five learners, giving the VR group substantially more individual hands-on practice per person. Applying the decision tree: extra resources (individual device access and unlimited individual practice) are present, and the difference is not negligible -- the authors call it "the most significant threat to internal validity." The key question is whether this extra individualized practice is framed as the explicit treatment variable or is a separable confound. The authors explicitly attribute the difference to "training intensity-- rather than VR per se," i.e. they treat the individualized, repetitive practice as a confound that could have been balanced (e.g., by giving control learners individual mannequins), not as the intended treatment variable. The control group did not receive matched individual practice. Criterion B is not met because the VR group received substantially more individualized, unlimited practice than the control group, this imbalance was not matched for the control, and the authors themselves frame it as a confounding difference in training intensity rather than the treatment variable being tested.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific trial is reported or found; it is described as a preliminary study needing replication.
      • "these findings require replication in larger, multi-center trials."
      • Relevant Quotes: 1) "these findings require replication in larger, multi-center trials." (Section 4) 2) "though these findings are preliminary due to the small sample size." (Conclusion) 3) "few studies have specifically examined thoracentesis." (Section 4.2) Detailed Analysis: Criterion R requires that this specific study be independently replicated by a different team in a peer-reviewed outlet. The paper itself calls for replication and notes that few studies have examined VR for thoracentesis. The cited literature concerns VR in other procedures/disciplines (endoscopy, arthroscopy, nursing, prosthodontics), not independent replications of this thoracentesis VR trial. This is a 2026 in-press preliminary study (n=20); an internet check via the DOI landing page was performed and no independent replication of this specific study was identified (WebSearch quota was exhausted, so verification relied on the paper text and the DOI landing page). Criterion R is not met because there is no independent replication of this specific VR thoracentesis trial.
    • A

      All-subject Exams

      • Only a single procedural skill (thoracentesis) was assessed with a custom rubric, and since criterion E is not met, A cannot be met.
      • "The primary outcome was the final thoracentesis procedural score at week 4."
      • Relevant Quotes: 1) "The primary outcome was the final thoracentesis procedural score at week 4." (Abstract) 2) "The assessment used a standardized 300-point scoring rubric divided into three domains: Preoperative preparation ... Intraoperative operation ... Postoperative operation." (Section 2.5) Detailed Analysis: Criterion A requires that all main subjects be measured using standardised exam-based assessments, and it explicitly depends on criterion E being met. This study measured only thoracentesis procedural skill (three sub-domains of one procedure) using a custom rubric, with no assessment of other subjects, and criterion E is not met. Criterion A is not met because only one procedural skill was assessed, using a non-standardised custom rubric, and criterion E is not satisfied.
    • G

      Graduation Tracking

      • Follow-up ended at week 4 with no tracking to graduation, and criterion Y is already not met.
      • "Third, we did not assess long-term skill retention (e.g., at 3 or 6 months)."
      • Relevant Quotes: 1) "Third, we did not assess long-term skill retention (e.g., at 3 or 6 months)." (Section 4.3) 2) "Fifth, we did not measure transfer of skills to real patients--the ultimate validity test." (Section 4.3) 3) "The primary outcome was the final thoracentesis procedural score at week 4." (Abstract) Detailed Analysis: Criterion G requires tracking participants through to graduation, and it depends on criterion Y being met. This study measured outcomes only through week 4, explicitly did not assess retention at 3-6 months or transfer to patients, and did not track interns to graduation. No subsequent graduation-tracking paper by the same authors was identified (WebSearch quota was exhausted; the paper text and DOI landing page were checked). Criterion Y is also not met. Criterion G is not met because there was no graduation or long-term tracking and criterion Y is not satisfied.
    • P

      Pre-Registered

      • The authors explicitly state the trial was not prospectively registered.
      • "This study was not prospectively registered as it was a small-scale educational intervention. The authors recognize this as a limitation."
      • Relevant Quotes: 1) "Clinical trial number: This study was not prospectively registered as it was a small-scale educational intervention. The authors recognize this as a limitation." (Declarations) Detailed Analysis: Criterion P requires that the full study protocol be pre-registered before data collection begins, with a registry reference and date. The authors explicitly state the study was not prospectively registered and acknowledge this as a limitation, so there is no pre-registration and no registry entry to verify. Criterion P is not met because the study was explicitly not prospectively registered.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.