Abstract
Purpose: Training clinical reasoning skills remains a critical challenge in speech-language pathology education. This study aimed to evaluate the effectiveness of two cloud-based, self-directed instructional modules—diagnostic report viewing and reasoning demonstration—in enhancing students' reasoning skills for speech sound disorders (SSDs) assessment. Methods: In a randomized between-group design, 51 undergraduate students in Taiwan were assigned to one of two video-based training modules in Mandarin. The control group viewed clinicians' presentations of diagnostic reports of SSD cases, while the experimental group watched the same reports accompanied by explicit demonstrations of clinical reasoning. Training effectiveness was evaluated using a Script Concordance Test (SCT) and self-report questionnaires on viewing experience. Results: Across all participants, SCT scores significantly improved from pre- to post-training (t = 2.82, p = 0.007), with no significant pre-training difference between groups. The experimental group showed significant within-group gains in diagnostic reasoning (z = -2.731, p = 0.006, r = 0.54) and total SCT scores (z = -2.623, p = 0.009, r = 0.51), though between-group comparisons of gain scores did not reach statistical significance (p = 0.31). Conclusions: Both instructional modules improved students' clinical reasoning for SSD assessment. The demonstration-based module showed indications of additional benefits in supporting reflective clinical reasoning, offering a scalable instructional framework for speech-language pathology education.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomization was performed at the individual student level within a single department cohort, not at the class or school level, and the intervention is not one-to-one tutoring.
- "Undergraduate students in speech-language pathology from three academic years were recruited and randomly assigned to either a control group or an experimental group" (p. 5)
Relevant Quotes:
1) "In a randomized between-group design, 51 undergraduate students in Taiwan were assigned to one of two video-based training modules in Mandarin." (Abstract, p. 1)
2) "Undergraduate students in speech-language pathology from three academic years were recruited and randomly assigned to either a control group or an experimental group (see Participants section for detailed demographic and academic profiles)." (p. 5)
3) "Students were randomly assigned to either an experimental or control group." (p. 5)
4) "A total of 51 undergraduate students from the Department of Audiology and Speech-Language Pathology at Mackay Medical University participated in this study." (p. 5)
Detailed Analysis:
Criterion C requires randomisation at the class level (or stronger school level), to prevent contamination between treatment and control groups. The exception is interventions designed for personal/one-to-one teaching, in which case student-level randomisation is acceptable.
Here, individual students drawn from a single department across three academic years were randomised to the two conditions. This is student-level randomisation, with both groups drawn from the same cohort/classes. There is no indication that whole classes or schools were the unit of assignment. The intervention is a self-directed, cloud-based video training programme delivered to individual students, not a one-to-one tutoring intervention in the sense intended by the exception (the exception concerns personalised teaching such as tutoring, whereas this is a standardised video module pushed to many students). Because randomisation occurred at the individual student level within shared classes, there is a real risk of contamination, and the class/school-level requirement is not satisfied.
Criterion C is not met because randomisation was conducted at the individual student level within a single department cohort rather than at the class or school level.
-
E
Exam-based Assessment
- Outcomes were measured with a self-developed Script Concordance Test (SCT-SSDs) created and validated by the same authors, not a widely recognised standardised exam.
- "Training effectiveness was evaluated using a self-developed SCT for SSDs assessment (SCT-SSDs) (Chan and Yeh 2025)" (p. 5)
Relevant Quotes:
1) "Training effectiveness was evaluated using a Script Concordance Test (SCT) and self-report questionnaires on viewing experience." (Abstract, p. 1)
2) "Training effectiveness was evaluated using a self-developed SCT for SSDs assessment (SCT-SSDs) (Chan and Yeh 2025), along with a self-reported viewing experience questionnaire." (p. 5)
3) "The first phase involved the construction and validation of the SCT tailored to SSDs assessment (SCT-SSDs) (Chan and Yeh 2025)." (p. 5)
4) "This study adopted a previously validated instrument, the SCT for Clinical Reasoning and Decision-Making in the Assessment of SSDs (Chan and Yeh 2025), to assess students' diagnostic reasoning abilities." (p. 8)
Detailed Analysis:
Criterion E requires the use of a standardised, widely recognised exam-based assessment (e.g., national curriculum exams, state-wide standardised achievement tests), not an instrument specially designed by the researchers for the study.
The primary outcome here is the SCT-SSDs, a Script Concordance Test that was constructed and validated by the same authors (Chan and Yeh 2025) in a prior phase of this research programme. While the SCT methodology is used across healthcare disciplines and the instrument underwent internal validation, this specific tool was self-developed and tailored to SSDs assessment by the study team. It is not a national or external standardised examination administered independently of the study. The scoring is based on the study's own expert panel consensus. This is precisely the situation the criterion warns against: a researcher-created measure aligned to the study content.
Criterion E is not met because the outcome measure was a self-developed Script Concordance Test created and validated by the same authors rather than a recognised standardised exam.
-
T
Term Duration
- The training lasted only 14 days (with at most a 19-day window) and the post-test was administered immediately after training, far short of one full academic term.
- "engaged in the online training program following a 14-day viewing schedule ... and then completed a group-based SCT post-test immediately after training" (p. 5)
Relevant Quotes:
1) "Within one month following the spring semester midterm examinations, participants completed a group-based SCT pre-test, then engaged in the online training program following a 14-day viewing schedule." (p. 5)
2) "They were encouraged to pace their viewing flexibly, typically completing one video segment every one to two days, and then completed a group-based SCT post-test immediately after training." (p. 5)
3) "To support self-directed learning and accommodate individual pacing differences, the online training platform remained accessible for 15 calendar days ... If participants fell behind the viewing schedule, a brief extension was permitted within a predefined maximum window of 19 total days." (p. 5)
4) "the training's effectiveness was evaluated solely through immediate post-tests without further follow-up assessments." (p. 15)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins, allowing for term-long follow-up even if the intervention itself is short.
In this study, the intervention spanned a 14-day viewing schedule (maximum 19 days), and the SCT post-test was administered "immediately after training." The authors explicitly acknowledge there was no follow-up beyond the immediate post-test. The interval from intervention start to outcome measurement is therefore roughly two to three weeks, far less than one academic term.
Criterion T is not met because outcomes were measured immediately after a 14-day intervention, well short of the one-term minimum tracking interval.
-
D
Documented Control Group
- The control group is well documented, with its composition, academic-level distribution, baseline SCT scores, prior course performance, and exact "report-only" condition all described.
- "The control group viewed videos presenting comprehensive diagnostic reports ... Control Group (n = 25)" (pp. 5-6)
Relevant Quotes:
1) "The control group viewed clinicians' presentations of diagnostic reports of SSD cases, while the experimental group watched the same reports accompanied by explicit demonstrations of clinical reasoning." (Abstract, p. 1)
2) "The control group viewed videos presenting comprehensive diagnostic reports, whereas the experimental group viewed videos that included explicit demonstrations of the clinical reasoning processes underlying each diagnostic report." (p. 5)
3) "Control Group (n = 25)" and Table 1 "Participant profile by group and academic level," giving control group N by academic level (Sophomore 11, Junior 9, Senior 5), sex distribution, and prior course performance (80.33 +/- 5.19, 79.82 +/- 4.41, 82.79 +/- 2.34). (Figure 1, Table 1, p. 6)
4) "Table 2 ... Descriptive statistics (mean +/- SD) for utility, interpretation, diagnosis and total SCT scores by group and assessment time," with control group pre-test Total 55.60 +/- 4.56. (Table 2, p. 10)
5) "The control group viewed the on-screen text of full diagnostic reports of the nine cases, including case history, behavioural observations, oral motor examination, speech sample analyses, summary of findings, diagnostic impressions, and recommendations." (p. 7)
Detailed Analysis:
Criterion D requires the control group to be well documented, including demographic information, baseline performance, and any treatment received.
The paper documents the control group thoroughly: its size (n = 25), its distribution across academic levels, sex distribution, and prior SSD-course grades are tabulated in Table 1. Baseline (pre-test) and post-test SCT scores for the control group across all subdomains are given in Table 2. Baseline equivalence between groups was statistically verified (chi-square on academic level; Mann-Whitney U on pre-training SCT scores). The exact condition the control group received (viewing the nine comprehensive diagnostic-report videos without the reasoning-demonstration overlays) is clearly specified.
Criterion D is met because the control group's size, demographics, baseline performance, and treatment condition are clearly and quantitatively documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation was at the individual student level within a single university department; no schools (or equivalent units) were randomised.
- "Students were randomly assigned to either an experimental or control group." (p. 5)
Relevant Quotes:
1) "A total of 51 undergraduate students from the Department of Audiology and Speech-Language Pathology at Mackay Medical University participated in this study." (p. 5)
2) "Students were randomly assigned to either an experimental or control group." (p. 5)
3) "participants were recruited from a single university, and the relatively small sample size may restrict the generalizability of the findings." (p. 15)
Detailed Analysis:
Criterion S requires randomisation at the school level (i.e., the educational institution or unit implementing the intervention), with multiple such units randomised.
In this study, the unit of randomisation was the individual student, and all participants came from a single department at one university. There is only one institution involved, and students within it were assigned individually to conditions. No school-level (or equivalent site-level) randomisation occurred.
Criterion S is not met because randomisation was at the individual student level within a single university department, not at the school/site level.
-
I
Independent Conduct
- The same authors who designed the training modules and the outcome instrument also conducted the trial and analysed the data; no independent third-party evaluation is described.
- "The training modules were collaboratively developed by three SLPs ... Under the guidance of the supervising SLP (the corresponding author)" (p. 5)
Relevant Quotes:
1) "The training modules were collaboratively developed by three SLPs: one senior SLP with over 20 years of experience in clinical education and supervision ... Under the guidance of the supervising SLP (the corresponding author), the junior SLP initially adapted nine real pediatric cases into diagnostic reports for instructional use." (p. 5)
2) "The first phase involved the construction and validation of the SCT tailored to SSDs assessment (SCT-SSDs) (Chan and Yeh 2025)." (p. 5)
3) "Preliminary questionnaires developed by the corresponding author and administered to sophomores in 2022 SSD courses..." (p. 4)
4) "To ensure consistency and content validity, each of the nine clinical reasoning modules developed for the experimental group was reviewed by an independent SLP who was not involved in the development process." (p. 8)
5) "an independent certified SLP who was not involved in the study independently categorized all responses" (p. 9)
Detailed Analysis:
Criterion I requires the study to be conducted independently from the authors who designed the intervention, to reduce bias in implementation, measurement, and analysis.
Here, the two authors (Yeh and Chan) designed the intervention modules, developed and validated the SCT outcome instrument, ran the trial, and analysed the data. The corresponding author guided module development and built the assessment tool. The only "independent" parties mentioned are an SLP who reviewed the reasoning modules for content validity, and an SLP who independently coded open-ended feedback for inter-rater reliability. These are limited review/reliability roles; they did not constitute independent conduct, delivery, or analysis of the trial. The data collection and analysis were performed by the intervention designers themselves.
Criterion I is not met because the same team that designed the intervention and the assessment tool also conducted the trial and analysed the results, with no independent evaluator running the study.
-
Y
Year Duration
- Because the term-duration criterion (T) is not met, and the study tracked only ~14 days with an immediate post-test, the year-duration criterion is also not met.
- "engaged in the online training program following a 14-day viewing schedule" (p. 5)
Relevant Quotes:
1) "participants completed a group-based SCT pre-test, then engaged in the online training program following a 14-day viewing schedule ... and then completed a group-based SCT post-test immediately after training." (p. 5)
2) "the training's effectiveness was evaluated solely through immediate post-tests without further follow-up assessments. Consequently, the current findings should be interpreted as evidence of initial skill acquisition rather than long-term reasoning competence." (p. 15)
Detailed Analysis:
Criterion Y requires outcomes to be measured at least 75% of one full academic year after the intervention begins. Per the criteria-specific instruction, if T (Term Duration) is not met, then Y is automatically not met.
T is not met here (a 14-day intervention with an immediate post-test). Independently, the total tracking interval of roughly two to three weeks is nowhere near 75% of an academic year, and the authors explicitly note the absence of any follow-up beyond the immediate post-test.
Criterion Y is not met because T is not met and outcomes were measured only ~14 days after the intervention began, far short of a full academic year.
-
B
Balanced Control Group
- The control group received a comparable active intervention (report-only videos of the same nine cases), and the reasoning-demonstration content was the explicit treatment variable being tested against this matched baseline.
- "Note: The core diagnostic report content is identical for both groups. Highlights, key messages, active thinking prompts and clinical reasoning flowcharts were provided exclusively to the experimental group." (Figure 2 note, p. 7)
Relevant Quotes:
1) "The control group viewed clinicians' presentations of diagnostic reports of SSD cases, while the experimental group watched the same reports accompanied by explicit demonstrations of clinical reasoning." (Abstract, p. 1)
2) "In contrast, the experimental group viewed the same materials supplemented with explicit clinical reasoning demonstrations and metacognitive cues." (p. 7)
3) "Note: The core diagnostic report content is identical for both groups. Highlights, key messages, active thinking prompts and clinical reasoning flowcharts were provided exclusively to the experimental group." (Figure 2 note, p. 7)
4) "In the report-only module (the control group), the length ranged from 412 to 655 seconds (M = 507.67, SD = 10.49), whereas in the demonstration-based module (the experimental group), the video length ranged from 612 to 908 seconds (M = 747.22, SD = 15.16)." (p. 8)
5) "The demonstration-based videos were significantly longer (Mdn = 747.22 seconds) than the report-only videos (Mdn = 507.67 seconds)." (p. 10)
Detailed Analysis:
Criterion B asks whether the control condition offers a comparable substitute for the intervention's inputs, unless the additional resource is itself the explicit treatment variable. Here the control group was not a passive "business as usual" group: it received an active intervention consisting of the same nine comprehensive diagnostic report videos, with identical core content, and completed the same nine viewing-experience questionnaires. The only difference between conditions is the addition of explicit clinical-reasoning demonstrations, key messages, active-thinking prompts, and flowcharts in the experimental group.
There is a modest time difference: demonstration videos were longer (mean ~747 s) than report-only videos (mean ~508 s), about four extra minutes per case across nine cases. However, this additional reasoning-demonstration content is precisely the treatment variable under investigation: the study's hypothesis is that "the demonstration-based module is hypothesized to offer additional benefits ... relative to the diagnostic report module." The extra time is intrinsic to and inseparable from the reasoning-demonstration intervention being tested, and both groups share an identical, substantial educational baseline (the diagnostic report videos). This matches the standard's allowance that when the additional resource is integral to the treatment being tested, the control may receive the business-as-usual (here, report-only) level.
Criterion B is met because the control group received a comparable active intervention (the identical diagnostic report videos), and the additional reasoning-demonstration content—the explicit treatment variable—is integral to the intervention being tested rather than a separable confounding resource.
-
Level 3 Criteria
-
R
Reproduced
- This specific trial has not been independently replicated by a different research team in a peer-reviewed journal; internet searching found no such replication.
Relevant Quotes:
1) "This randomized controlled trial compared two self-directed, cloud-based instructional modules for SSD assessment." (What this paper adds, p. 1)
2) "no prior study has systematically applied this approach in the context of speech-language pathology education." (p. 4)
3) "Research on instructional strategies designed to strengthen clinical reasoning processes specific to SSDs diagnosis remains sparse." (p. 4)
Detailed Analysis:
Criterion R requires that the specific study (or its central experimental claim) be independently replicated by a different research team in a different context and published in a peer-reviewed journal.
This is a novel study; the authors state that no prior study has applied this self-directed, demonstration-based approach in speech-language pathology education, and that research in this area is sparse. Internet searches (Wiley Online Library, PubMed, Google Scholar) for independent replications of this specific cloud-based SSD-reasoning module trial returned no results. The only closely related work is the authors' own companion validation paper, "Trends in the Acquisition of Clinical Reasoning in the Assessment of Speech Sound Disorders Using the Script Concordance Test" (Chan and Yeh 2025, IJLCD 60(5): e70105), which is by the same authors and is the prior phase that built the SCT-SSDs instrument; it is not an independent reproduction. While related illness-script instruction has been tested elsewhere (e.g., Moghadami et al. 2021 in medical students), those are different interventions and populations, not replications of this SSD-specific cloud-based module trial. No independent reproduction of this particular study by a different research team was identified.
Criterion R is not met because there is no independent replication of this specific trial by a different research team published in a peer-reviewed journal.
-
A
All-subject Exams
- Only a single domain (clinical reasoning for SSD assessment) was assessed via a self-developed SCT, not all main subjects with standardised exams; criterion E is also not met.
- "Training effectiveness was evaluated using a self-developed SCT for SSDs assessment (SCT-SSDs)" (p. 5)
Relevant Quotes:
1) "This study aimed to evaluate the effectiveness of two cloud-based, self-directed instructional modules ... in enhancing students' reasoning skills for speech sound disorders (SSDs) assessment." (Abstract, p. 1)
2) "Training effectiveness was evaluated using a self-developed SCT for SSDs assessment (SCT-SSDs) (Chan and Yeh 2025), along with a self-reported viewing experience questionnaire." (p. 5)
Detailed Analysis:
Criterion A requires impact to be measured across all main subjects using standardised exam-based assessments, and it is explicitly conditional on criterion E: if E is not met, A is not met.
Criterion E is not met (the outcome was a self-developed SCT, not a standardised exam), so A automatically fails. Separately, the study measured outcomes only in the single, specialised domain of clinical reasoning for SSD assessment; no other subjects were assessed.
Criterion A is not met because criterion E is not met and the study assessed only a single specialised domain rather than all main subjects via standardised exams.
-
G
Graduation Tracking
- Students were tested only immediately after the 14-day training, with no follow-up to graduation; criterion Y is also not met, and no follow-up publication tracking these students to graduation was found.
- "the training's effectiveness was evaluated solely through immediate post-tests without further follow-up assessments" (p. 15)
Relevant Quotes:
1) "and then completed a group-based SCT post-test immediately after training." (p. 5)
2) "due to constraints in the instructional timeline, the training's effectiveness was evaluated solely through immediate post-tests without further follow-up assessments." (p. 15)
3) "it is strongly recommended that future studies incorporate delayed assessments—such as administering the SCT one to two months post-training—to evaluate the sustainability of these learning outcomes." (p. 15)
Detailed Analysis:
Criterion G requires tracking participants until graduation to assess long-term impact, and it is explicitly conditional on criterion Y: if Y is not met, G is not met.
Criterion Y is not met, so G automatically fails. Independently, the study used only an immediate post-test, and the authors explicitly note the absence of any follow-up. Internet searches for subsequent publications by the same authors tracking this student cohort to graduation found none; the only related author publication is the prior SCT-validation paper (Chan and Yeh 2025), which does not track these students to graduation.
Criterion G is not met because criterion Y is not met and there was no follow-up beyond the immediate post-test, let alone tracking to graduation, in this or any subsequent publication.
-
P
Pre-Registered
- The paper reports ethics-committee approval but provides no pre-registration of the study protocol on a trial registry before data collection; no registration was found online.
- "This study was approved by the Research Ethics Committee in March 2023 (NTU-REC No. 202105ES012)." (p. 5)
Relevant Quotes:
1) "This study was approved by the Research Ethics Committee in March 2023 (NTU-REC No. 202105ES012)." (p. 5, and Ethics Statement, p. 15)
2) "The pre-test data were previously reported in a separate manuscript (Chan and Yeh 2025). The present study includes new post-test data." (p. 5)
3) "The participants of this study did not give written consent for their data to be shared publicly, so due to the sensitive nature of the research supporting data is not available." (Data Availability Statement, p. 15)
Detailed Analysis:
Criterion P requires that the full study protocol (hypotheses, methods, planned analyses) be pre-registered on a registry before data collection began, with a verifiable date.
The paper reports an ethics committee approval number (NTU-REC No. 202105ES012, March 2023) but provides no reference to a clinical-trials or study registry (e.g., ClinicalTrials.gov, ISRCTN, OSF), no registration ID for a pre-registered protocol, and no pre-registration date. Internet searches for a pre-registered protocol associated with this trial or the NTU-REC number returned no registry entry. Ethics approval is not equivalent to public pre-registration of the analysis plan.
Criterion P is not met because the paper provides no evidence of a pre-registered study protocol on a registry prior to data collection, and none was found through internet searching.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.