Abstract
Aim: This study aimed to evaluate the preliminary effectiveness of a ChatGPT-integrated educational session on nursing students' knowledge acquisition and retention regarding endotracheal suctioning, and to explore their perspectives and learning experiences. Methods: A pilot pre-test/post-test parallel-group randomized controlled trial with an embedded qualitative component was conducted with 32 undergraduate nursing students who participated voluntarily. Participants were randomly assigned to either a ChatGPT-integrated education group or a traditional education group. Knowledge was assessed using a structured test administered before, immediately after, and six weeks following the intervention. Quantitative data were analyzed using appropriate non-parametric statistical tests, while qualitative data were analyzed thematically following Braun and Clarke's approach. Results: While no significant difference was observed between groups in immediate post-test knowledge scores, the ChatGPT-integrated group demonstrated significantly higher knowledge retention over time at the six-week follow-up (p < 0.05). Qualitative findings revealed three main themes: learning supportive factors, student gains, and suggestions for improvement. Conclusions: The findings suggest that ChatGPT-integrated education has the potential to enhance knowledge retention over time, although it may not significantly impact immediate learning outcomes. Given small sample size, these result should be considered preliminary.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was at the individual student level within a single cohort, not at the class or school level, and the intervention is group teaching rather than tutoring.
- "At the beginning of the study, 34 students were randomly assigned in a 1:1 ratio to the intervention (n = 17) and control (n = 17) groups."
Relevant Quotes:
1) "Nursing students were randomly assigned to either the intervention group, which received a ChatGPT-integrated educational session, or the control group, which received a traditional education session, using a computer-generated randomization list." (p. 9)
2) "Participants were first stratified by year of study (third- and fourth-year), and then by grade point average (GPA <=2.0 and >2.0). Random assignment was subsequently conducted within each stratum using the Research Randomizer tool (https://www.randomizer.org/)." (p. 11)
3) "At the beginning of the study, 34 students were randomly assigned in a 1:1 ratio to the intervention (n = 17) and control (n = 17) groups." (p. 11)
4) "To minimize contamination between groups, the control and intervention sessions were conducted in consecutive weeks." (p. 9)
Detailed Analysis:
The ERCT 'C' criterion requires randomisation at the class level or stronger (school level), unless the intervention is one-to-one tutoring/personal teaching, in which case student-level randomisation is accepted. Here, the unit of randomisation was the individual student: 34 students drawn from a single university cohort were stratified by year and GPA and then randomly assigned one-by-one to the intervention or control condition. This is student-level randomisation within a single cohort, not assignment of entire classes or schools.
The intervention is a group classroom education session on endotracheal suctioning delivered to nursing students, not a one-to-one tutoring or personal-teaching intervention, so the tutoring exception does not apply. The authors themselves note a contamination risk (they scheduled the two group sessions in consecutive weeks and asked students not to share content), which is precisely the risk that class-level randomisation is meant to avoid.
Criterion C is not met because randomisation was performed at the individual student level within a single cohort rather than at the class or school level, and no tutoring exception applies.
-
E
Exam-based Assessment
- Outcomes were measured with a topic-specific research questionnaire (the SKQ), not a widely recognised standardised exam.
- "The Suctioning Knowledge Questionnaire (SKQ): The SKQ was developed in Turkish by Ozden and Yilmaz in 2021 for use with nursing students."
Relevant Quotes:
1) "The Suctioning Knowledge Questionnaire (SKQ): The SKQ was developed in Turkish by Ozden and Yilmaz in 2021 for use with nursing students (Yilmaz & Ozden, 2021)." (p. 12)
2) "The questionnaire includes 20 items that assess knowledge of endotracheal suctioning procedures across three phases: before, during, and after the intervention. Each item is multiple-choice with five response options. Correct responses are awarded five points... resulting in a total possible score ranging from 0 to 100." (p. 12)
3) "The Content Validity Index of the questionnaire was 0.98. Reliability analysis using the split-half method yielded a Spearman-Brown coefficient of 0.781, indicating a high level of internal consistency." (p. 12)
Detailed Analysis:
The ERCT 'E' criterion requires a standardised, widely-recognised exam-based assessment (e.g. a state-wide or national curriculum exam), not an instrument designed for research purposes. The SKQ is a 20-item research questionnaire on endotracheal suctioning developed by Ozden and Yilmaz in 2021 for use with nursing students. Although it was not created solely for this particular study and has reported content validity and reliability, it is a topic-specific research questionnaire rather than a standardised, widely recognised examination such as a national or curriculum exam.
The outcome instrument therefore functions as a custom/specialised knowledge questionnaire aligned to the intervention topic, which is the kind of researcher-type measure the criterion excludes.
Criterion E is not met because the outcome was measured with a topic-specific research questionnaire (the SKQ) rather than a standardised, widely recognised exam-based assessment.
-
T
Term Duration
- The longest follow-up was six weeks, shorter than the one-term minimum required by the criterion.
- "To evaluate knowledge retention, the SKQ was re-administered to all participants six weeks after the intervention."
Relevant Quotes:
1) "Each educational session lasted for two 45-minute class periods." (p. 13)
2) "Knowledge was assessed using a structured test administered before, immediately after, and six weeks following the intervention." (Abstract, p. 3)
3) "To evaluate knowledge retention, the SKQ was re-administered to all participants six weeks after the intervention." (p. 15)
Detailed Analysis:
The ERCT 'T' criterion requires that the primary outcome be measured at least one full academic term (roughly 3-4 months) after the intervention begins; short interventions are allowed provided term-long follow-up tracking is present. In this study the intervention was a single session of two 45-minute class periods, and outcomes were measured at baseline, immediately after the session, and again six weeks later as the retention measure.
The longest interval from intervention start to outcome measurement is six weeks, which is shorter than one full academic term (~3-4 months). No follow-up covering at least a full term is reported.
Criterion T is not met because the longest interval from intervention start to outcome measurement was six weeks, which is shorter than one full academic term.
-
D
Documented Control Group
- The control group's size, demographics, baseline scores, and condition are documented in detail in Tables 1 and 2.
- "No significant differences were found between the groups in terms of age, gender, prior ChatGPT use, or experiences with endotracheal suctioning (p > 0.05)."
Relevant Quotes:
1) "A total of 32 students completed the study, with 17 in the intervention group and 15 in the control group (Table 1)." (p. 17)
2) "Table 1. Descriptive characteristics of the participants (n = 32)." Control group (n = 15): Age 22.40 +/- 1.121; Gender Woman 14/93.3, Man 1/6.7; uses ChatGPT to prepare Yes 14/93.3; prior education using ChatGPT Yes 3/20; experienced suctioning Yes 7/46.7. (p. 18)
3) "No significant differences were found between the groups in terms of age, gender, prior ChatGPT use, or experiences with endotracheal suctioning (p > 0.05)." (p. 18)
4) "the control group, which received a traditional education session" (p. 9); "Both groups received the same core educational content, differing only in the instructional approach." (p. 15)
Detailed Analysis:
The ERCT 'D' criterion requires the control group to be well documented, including demographics, size, baseline characteristics, and the conditions/treatment it received. The paper reports the control group size (n = 15), full demographic and baseline data in Table 1 (age, gender, prior ChatGPT use, prior suctioning experience, perceived competence), baseline knowledge scores (Table 2 pre-test median), and confirms baseline equivalence between groups. It also clearly describes what the control group received (a traditional education session with the same core content).
Criterion D is met because the control group's size, demographics, baseline characteristics, and the instruction it received are all clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation was at the individual student level in a single university, not at the school/institution level.
- "At the beginning of the study, 34 students were randomly assigned in a 1:1 ratio to the intervention (n = 17) and control (n = 17) groups."
Relevant Quotes:
1) "The population of the study consisted of 44 fourth-year and 48 third-year nursing students enrolled in the spring semester of the 2024-2025 academic year at the Department of Nursing of a private university in Turkey." (pp. 10-12)
2) "At the beginning of the study, 34 students were randomly assigned in a 1:1 ratio to the intervention (n = 17) and control (n = 17) groups." (p. 11)
3) "Random assignment was subsequently conducted within each stratum using the Research Randomizer tool." (p. 11)
Detailed Analysis:
The ERCT 'S' criterion requires randomisation at the school level (the educational institution or unit implementing the intervention), with multiple schools or sites randomised. This study was conducted at a single private university, and randomisation was performed at the individual student level within one cohort. No schools, sites, or institutions were randomised; only individual students were assigned.
Criterion S is not met because the study randomised individual students within a single university rather than randomising schools or institutional units.
-
I
Independent Conduct
- The same team designed, delivered, and analysed the intervention, with no independent or external evaluator.
- "Due to the nature of the educational intervention, blinding of participants and the instructor was not feasible."
Relevant Quotes:
1) "E.S. contributed to supervision, conceptualization, investigation, methodology, and writing (original draft and review & editing). A.B.P. contributed to conceptualization, investigation, methodology, resources, and writing... S.B.K. contributed to conceptualization, investigation, methodology, resources, and writing." (p. 35)
2) "all educational sessions were conducted by the same instructor using standardized teaching materials." (p. 12)
3) "The random allocation sequence was generated by the researchers." (p. 11)
4) "Due to the nature of the educational intervention, blinding of participants and the instructor was not feasible." (p. 12)
Detailed Analysis:
The ERCT 'I' criterion requires that the evaluation be conducted independently from those who designed the intervention, e.g. by an external evaluation team, to reduce bias in delivery, measurement, and analysis. In this study the same authors conceptualized and designed the ChatGPT-integrated educational intervention, generated the randomisation sequence, delivered the sessions (same instructor), collected the data, and performed the analysis. There is no statement of any third-party or external evaluator, and blinding was not feasible.
Criterion I is not met because the same research team designed, delivered, and evaluated the intervention with no independent or third-party evaluation.
-
Y
Year Duration
- Follow-up was only six weeks, far short of a year, and criterion T was not met.
- "To evaluate knowledge retention, the SKQ was re-administered to all participants six weeks after the intervention."
Relevant Quotes:
1) "Each educational session lasted for two 45-minute class periods." (p. 13)
2) "To evaluate knowledge retention, the SKQ was re-administered to all participants six weeks after the intervention." (p. 15)
3) "Given that the present study was conducted with a small sample and retention was assessed only once, further research is required." (p. 31)
Detailed Analysis:
The ERCT 'Y' criterion requires that outcomes be measured at least 75% of a full academic year (~9-10 months) after the intervention begins. The longest follow-up here was six weeks, far short of a full academic year. In addition, per the standard, if criterion T (Term Duration) is not met then criterion Y cannot be met, and T is not met here.
Criterion Y is not met because outcomes were measured only six weeks after the intervention, far short of a full academic year, and the weaker Term Duration criterion was also not met.
-
B
Balanced Control Group
- Both groups received equal time, materials, and hands-on practice, differing only in the instructional method being tested, so resources were balanced.
- "Both groups received the same core educational content, differing only in the instructional approach."
Relevant Quotes:
1) "Both groups received the same core educational content, differing only in the instructional approach." (p. 15)
2) "Each educational session lasted for two 45-minute class periods." (p. 13)
3) "As in the control group, at the end of the educational session the students in the intervention group also observed a demonstration of the endotracheal suctioning procedure and were given the opportunity to practice the procedure themselves." (p. 15)
4) "all educational sessions were conducted by the same instructor using standardized teaching materials." (p. 12)
Detailed Analysis:
Criterion B compares the time, budget, and materials given to the intervention and control conditions and asks whether the control provides a comparable substitute, unless additional resources are explicitly the treatment variable. Here both groups received the same core educational content, the same session length (two 45-minute class periods), the same instructor, and the same mannequin demonstration and hands-on practice. The only difference was the instructional approach: ChatGPT-integrated inquiry (intervention) versus traditional PowerPoint/discussion teaching (control).
Applying the decision tree: the intervention did not add extra instructional time or a materially different budget compared with the control (equal session length and equal practice opportunities), so no extra time/budget resource is present that would need to be matched. The distinguishing element is the teaching method itself, which is integral to what is being tested. Both groups received equivalent educational engagement.
Criterion B is met because both groups received the same core content, equal instructional time, the same instructor, and the same hands-on practice, with the difference being only the instructional method being tested.
-
Level 3 Criteria
-
R
Reproduced
- The study is described as the first of its kind and, after internet searching, no independent peer-reviewed replication was found.
- "To the best of our knowledge, no study has yet investigated the impact of a ChatGPT-integrated endotracheal suctioning session... on nursing students' knowledge and experiences."
Relevant Quotes:
1) "To the best of our knowledge, no study has yet investigated the impact of a ChatGPT-integrated endotracheal suctioning session... on nursing students' knowledge and experiences." (p. 8)
2) "Given the pilot nature and the small sample size of this study, these conclusions should be considered preliminary and interpreted with caution. Future studies should focus on... conducting rigorous randomized controlled trials." (p. 34)
Detailed Analysis:
The ERCT 'R' criterion requires that this specific study (or its central experimental claim) has been independently replicated by a different research team in a different context and published in a peer-reviewed journal. The authors explicitly frame the study as the first of its kind on ChatGPT-integrated endotracheal suctioning training and call for future trials. Related ChatGPT-in-nursing RCTs by other teams exist (e.g. Arkan, Dalli & Varol, 2025, a single-blind RCT on problem-solving skills and attitudes), but these are not replications of this specific endotracheal suctioning trial. Internet searches of Google Scholar, PubMed, and general web sources in July 2026 identified no independent replication of this particular study by other authors; no such paper (title, authors, year) could be found and no verbatim replication quotes are available.
Criterion R is not met because there is no independent peer-reviewed replication of this specific study, which the authors themselves describe as the first of its kind.
-
A
All-subject Exams
- Only one specialised topic was assessed and criterion E was not met, so all-subject assessment fails.
- "The questionnaire includes 20 items that assess knowledge of endotracheal suctioning procedures."
Relevant Quotes:
1) "The aim of this pilot study was to evaluate the preliminary effect of a ChatGPT-integrated educational session on nursing students' knowledge acquisition and retention regarding endotracheal suctioning." (pp. 8-9)
2) "The questionnaire includes 20 items that assess knowledge of endotracheal suctioning procedures." (p. 12)
Detailed Analysis:
The ERCT 'A' criterion requires that impact be measured across all main subjects using standardised exam-based assessments, and it explicitly depends on criterion E: if E is not met, A cannot be met. Here the only outcome measured was knowledge of a single specialised clinical topic (endotracheal suctioning), assessed with the SKQ research questionnaire. No other subjects were assessed, and criterion E (standardised exam-based assessment) was not met.
Criterion A is not met because only a single specialised topic was assessed and the prerequisite criterion E was not met.
-
G
Graduation Tracking
- Tracking ended at a single six-week follow-up with no graduation tracking found after internet searching, and criterion Y was not met.
- "Given that the present study was conducted with a small sample and retention was assessed only once, further research is required."
Relevant Quotes:
1) "To evaluate knowledge retention, the SKQ was re-administered to all participants six weeks after the intervention." (p. 15)
2) "Given that the present study was conducted with a small sample and retention was assessed only once, further research is required." (p. 31)
Detailed Analysis:
The ERCT 'G' criterion requires that participants be followed up and tracked until their graduation to assess long-term impact, and it depends on criterion Y: if Y is not met, G cannot be met. Here the study tracked students only to a single six-week follow-up, with no tracking through to graduation. Internet searches in July 2026 for subsequent or follow-up publications by the same authors (Sezgunsay, Polat, Kilicer) tracking this cohort to graduation returned no such papers; no graduation-tracking publication could be found and no verbatim quotes are available. Criterion Y was also not met.
Criterion G is not met because tracking ended at six weeks with no graduation follow-up, no follow-up publication was found on internet searching, and the prerequisite criterion Y was not met.
-
P
Pre-Registered
- The trial was registered retrospectively, after data collection (registry start 4 March 2025 vs first posted 31 July 2025), so it was not pre-registered.
- "The Clinical Trial Number NCT07096518 has been retrospectively registered at ClinicalTrials.gov... on 7 August 2025."
Relevant Quotes:
1) "The Clinical Trial Number NCT07096518 has been retrospectively registered at ClinicalTrials.gov (https://clinicaltrials.gov/study/NCT07096518) on 7 August 2025, in accordance with ethical and clinical research protocols." (pp. 9-10)
2) "The population of the study consisted of 44 fourth-year and 48 third-year nursing students enrolled in the spring semester of the 2024-2025 academic year." (p. 10)
3) "Ethical approval for this study was granted by the Izmir University of Economics Health Sciences Research Ethics Committee on 22 April 2025." (p. 17)
Detailed Analysis:
The ERCT 'P' criterion requires that the full study protocol be pre-registered before data collection began. The paper explicitly states the trial was "retrospectively registered" at ClinicalTrials.gov. A direct check of the ClinicalTrials.gov record for NCT07096518 confirms this: the study start date was 4 March 2025 and the primary completion date was 10 June 2025, whereas the record was first submitted on 24 July 2025 and first posted on 31 July 2025. Registration therefore occurred months after data collection had begun and after it was completed. The authors' own use of the word "retrospectively" is consistent with the registry timestamps.
Criterion P is not met because the trial was registered retrospectively (registry first posted 31 July 2025; the paper cites 7 August 2025), well after data collection had already taken place.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.