Abstract
This study investigates how the combined use of automatic speech recognition (ASR) and automated writing evaluation (AWE) influences Chinese English as a foreign language (EFL) learners' speaking anxiety and speaking performance. Adopting a mixed-methods experimental design, the study compares an experimental group (EG) that utilizes ASR-AWE tools with a control group (CG) that follows traditional speaking instruction. The results revealed significant reductions in speaking anxiety, especially in learners with low self-confidence, fear of making mistakes and fear of being laughed at. In terms of competence, the EG demonstrated statistically significant improvements in pronunciation, grammatical accuracy and interactive communication. Qualitative findings indicated that learners valued ASR for its real-time feedback on pronunciation and AWE for improving sentence structure and grammatical refinement. The integration of ASR and AWE supported autonomous, iterative speaking practice, providing targeted linguistic input and facilitating learner engagement. These findings offer practical implications for incorporating ASR-AWE technologies into oral communication curricula in EFL contexts, particularly in promoting student confidence and fluency in technology-supported environments.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was conducted at the individual student level within one course at one university, not at the class or school level, and no tutoring exception applies.
- "After selecting the final 64 participants, they were randomly assigned to EG (N = 32) and CG (N = 32) using a random number generator." (p. 7)
Relevant Quotes:
1) "All 103 first-year students enrolled in the compulsory Audio-Visual-Oral Business English course at a public university were invited to participate in the study." (p. 6)
2) "After selecting the final 64 participants, they were randomly assigned to EG (N = 32) and CG (N = 32) using a random number generator. This ensured fairness and minimized potential selection bias." (p. 7)
Detailed Analysis:
All participants were drawn from a single compulsory course at one university, and the 64 selected participants were individually randomised, via a random number generator, into EG and CG. There is no mention anywhere in the methodology of whole classes, sections, or schools being assigned to conditions; the unit of randomisation is explicitly the individual student. This creates a risk of contamination, since EG and CG students remain classmates in the same course throughout the 14-week intervention. The intervention (technology-assisted supplementary speaking practice versus traditional classroom speaking activities) is not a one-to-one personal tutoring arrangement, so the tutoring exception for student-level RCTs does not apply.
Criterion C is not met because randomisation occurred at the individual student level within a single course rather than at the class or school level, and no exception applies.
-
E
Exam-based Assessment
- The study used the Cambridge Business English Speaking Test, a widely recognised standardised assessment scored by accredited external examiners.
- "The speaking test used in this study is the Cambridge Business English Speaking Test, which evaluates pronunciation, fluency, accuracy and discourse competence." (p. 8)
Relevant Quotes:
1) "The speaking test used in this study is the Cambridge Business English Speaking Test, which evaluates pronunciation, fluency, accuracy and discourse competence." (p. 8)
2) "This test was selected because it aligns well with the study's focus on speaking competence... it is widely recognized for its reliability and validity in evaluating business-related spoken English skills, making it suitable for the participants' academic and professional contexts." (p. 8)
3) "Scoring was conducted by three certified examiners accredited by the Cambridge Business English Certificate (BEC) system. Each candidate's performance was rated based on the official Cambridge English: Business Vantage Speaking Assessment Scale..." (p. 8)
Detailed Analysis:
The primary speaking-competence outcome was measured with the Cambridge Business English Speaking Test, an internationally recognised, externally validated standardised assessment with published band descriptors, not a custom instrument devised solely for this study. Scoring was performed by accredited, certified examiners with reported high interrater reliability (ICC = 0.786).
Criterion E is met because the study used a widely recognised, standardised, externally scored exam-based assessment rather than a study-specific instrument.
-
T
Term Duration
- Outcomes were measured after a 14-week (semester-long) intervention, exceeding the minimum one-term requirement.
- "The instructional procedure of the study consists of three main phases: the pretest phase, the 14-week intervention phase and the posttest phase." (p. 9)
Relevant Quotes:
1) "The instructional procedure of the study consists of three main phases: the pretest phase, the 14-week intervention phase and the posttest phase." (p. 9)
2) "...this study seeks to (1) examine the effect of integrated ASR and AWE tools on students' speaking anxiety and (2) assess their impact on speaking performance over a semester-long intervention." (p. 6)
Detailed Analysis:
The intervention ran for 14 weeks (described by the authors as a semester-long intervention), with a baseline pretest at week 1 and the primary outcome measured at posttest in week 14. Fourteen weeks comfortably exceeds the minimum of one academic term (roughly 3-4 months).
Criterion T is met because the interval between intervention start and outcome measurement (14 weeks) exceeds one academic term.
-
D
Documented Control Group
- The control group's size, demographics, baseline scores and instructional condition are explicitly documented and compared with the treatment group.
- "There were no significant differences between EG and CG in speaking anxiety and speaking competence at baseline (p > 0.05), as determined by independent samples t-tests, ensuring equivalence between groups." (p. 7)
Relevant Quotes:
1) "There were no significant differences between EG and CG in speaking anxiety and speaking competence at baseline (p > 0.05), as determined by independent samples t-tests, ensuring equivalence between groups." (p. 7)
2) "Regarding gender distribution, both groups were predominantly female, with EG comprising 26 females and 6 males, while CG included 27 females and 5 males. The mean age in both groups was 18.98 years, ensuring age homogeneity." (p. 7)
3) "CG follows a traditional speaking curriculum." (p. 9); "CG participates in traditional speaking activities without ASR or AWE tools." (p. 9)
Detailed Analysis:
The control group's size (N = 32), gender composition, age, and baseline speaking-anxiety and speaking- competence scores are explicitly reported and statistically confirmed to be equivalent to the experimental group at baseline. The paper also describes what the control group actually experienced during the intervention (a traditional speaking curriculum without ASR/AWE tools), giving a clear picture of both its composition and its treatment condition.
Criterion D is met because the control group's demographics, baseline characteristics and instructional condition are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred among individual students within a single university course, not among schools.
- "After selecting the final 64 participants, they were randomly assigned to EG (N = 32) and CG (N = 32) using a random number generator." (p. 7)
Relevant Quotes:
1) "All 103 first-year students enrolled in the compulsory Audio-Visual-Oral Business English course at a public university were invited to participate..." (p. 6)
2) "After selecting the final 64 participants, they were randomly assigned to EG (N = 32) and CG (N = 32) using a random number generator." (p. 7)
Detailed Analysis:
As with criterion C, randomisation occurred among individual students drawn from a single course at one university, not among multiple schools, sites, or even multiple classes. No cluster- or school-level assignment is described anywhere in the methodology.
Criterion S is not met because randomisation was conducted at the individual student level within a single university course, not at the school level.
-
I
Independent Conduct
- Aside from independent exam scoring, there is no explicit statement that the trial's design, data collection, or analysis were conducted independently of the researchers.
- "No potential conflict of interest was reported by the authors." (Disclosure statement, p. 26)
Relevant Quotes:
1) "Scoring was conducted by three certified examiners accredited by the Cambridge Business English Certificate (BEC) system." (p. 8)
2) "No potential conflict of interest was reported by the authors." (Disclosure statement, p. 26)
3) No statement anywhere in the methods, acknowledgments, or disclosure sections identifies an external, independent research or evaluation team overseeing recruitment, randomisation, data collection, or analysis.
Detailed Analysis:
The only independent element described is the scoring of the speaking test, performed by accredited Cambridge examiners with reported interrater reliability. Criterion I, however, concerns independence of the overall trial conduct (recruitment, randomisation, data collection, qualitative coding and statistical analysis) from those with a stake in the intervention, not merely test scoring. The same author team appears to have designed the study, recruited and randomised participants, collected the reflective journals and interviews, and performed all quantitative and thematic analyses themselves. The generic disclosure statement does not amount to an explicit statement of independent conduct as required by the standard.
Criterion I is not met because, apart from third-party exam scoring, no explicit evidence shows that the trial's design, implementation, data collection, or analysis were conducted independently of the researchers.
-
Y
Year Duration
- The 14-week (roughly one-semester) tracking period is far shorter than 75% of a full academic year.
- "The instructional procedure of the study consists of three main phases: the pretest phase, the 14-week intervention phase and the posttest phase." (p. 9)
Relevant Quotes:
1) "The instructional procedure of the study consists of three main phases: the pretest phase, the 14-week intervention phase and the posttest phase." (p. 9)
2) "...assess their impact on speaking performance over a semester-long intervention." (p. 6)
Detailed Analysis:
The tracking interval from intervention start to final measurement was 14 weeks, roughly one semester. An academic year is generally understood as approximately 9-10 months (around 36-40 weeks), of which 75% would be roughly 27-30 weeks. Fourteen weeks falls well short of this threshold, and the paper offers no rationale for treating a single semester as equivalent to at least 75% of a full academic year in this context.
Criterion Y is not met because the 14-week tracking period is far shorter than 75% of a full academic year.
-
B
Balanced Control Group
- The extra ASR/AWE resources given to the experimental group are the explicit treatment variable being tested, so the control group's business-as-usual condition is an acceptable comparator.
- "The intervention phase involves EG using NetEase Youdao Dictionary with ASR and AWE for speaking practice, while CG follows a traditional speaking curriculum." (p. 9)
Relevant Quotes:
1) "...this study aims to investigate how the integration of ASR and AWE technologies can reduce speaking anxiety and improve speaking competence among Chinese EFL learners." (p. 3)
2) "The intervention phase involves EG using NetEase Youdao Dictionary with ASR and AWE for speaking practice, while CG follows a traditional speaking curriculum." (p. 9)
3) "EG engages in ASR-based pronunciation feedback and AWE for grammar and syntax improvement, while CG participates in traditional speaking activities without ASR or AWE tools." (p. 9)
Detailed Analysis:
Applying the criterion B decision logic: extra resources are present, since EG receives structured access to an ASR/AWE mobile application not given to CG. However, these additional resources are precisely the treatment variable under investigation - the study's explicit, stated purpose from the introduction onward is to test whether integrating ASR and AWE technologies (versus traditional instruction) reduces speaking anxiety and improves speaking competence. Both groups otherwise followed the same 14-week course syllabus and unit progression (Figure 1), with CG receiving "traditional speaking activities" as the business-as-usual comparator. Because the additional resource is explicitly and transparently the object of the study rather than an incidental add-on, the control group's business-as-usual condition satisfies the standard's exception for resources integral to the treatment being tested.
Criterion B is met because the additional ASR/AWE resources are the explicit treatment variable, tested against a business-as-usual control following the same core curriculum.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study was found or reported, confirmed via internet search.
- "...the sample size, though sufficient for statistical analysis, was limited to first-year business English students at a single university, which may affect generalizability." (p. 25-26)
Relevant Quotes:
1) "First, the sample size, though sufficient for statistical analysis, was limited to first-year business English students at a single university, which may affect generalizability. Future research could extend the investigation to diverse EFL learner populations across different proficiency levels and academic disciplines." (p. 25-26)
2) No reference anywhere in the paper to a prior or subsequent independent replication of this specific combined ASR-AWE intervention by a different research team.
Detailed Analysis:
The paper situates its findings alongside related prior studies on ASR or AWE individually (e.g., Ngo et al., Hsu, Fu et al., Shadiev et al.), but these examine those tools separately and are not independent replications of this specific combined ASR-AWE intervention with Chinese EFL learners in this course design. The authors explicitly flag generalizability as a limitation and call for future studies across different populations, confirming no independent replication currently exists.
An internet search (general web search plus checks of citing-article listings for this DOI) for independent replications of this specific combined ASR-AWE speaking intervention among Chinese EFL learners did not locate any peer-reviewed study by a different research team reproducing this design. The only closely related work found is a companion paper by the same author team, Li, W., Mohamad, M., & You, H. W. (2025), "Exploring the effects of using automatic speech recognition on EFL university students with high speaking anxiety," International Journal of Information and Education Technology, 15(1), 187-194, which tests ASR alone (no AWE) on a separate sample and is not an independent replication. Given the article was only published online on 17 September 2025, an independent peer-reviewed replication would in any case be very unlikely to exist yet.
Criterion R is not met because no independent replication of this specific study was found or reported, and none was located via internet search.
-
A
All-subject Exams
- Only English speaking outcomes were assessed, with no other subjects measured or justification given for this narrow scope.
- "The speaking test used in this study is the Cambridge Business English Speaking Test, which evaluates pronunciation, fluency, accuracy and discourse competence." (p. 8)
Relevant Quotes:
1) "The speaking test used in this study is the Cambridge Business English Speaking Test, which evaluates pronunciation, fluency, accuracy and discourse competence." (p. 8)
2) "2. Assess the effectiveness of ASR and AWE in improving learners' speaking competence." (p. 3)
3) No mention anywhere of assessment in any subject or skill domain other than English speaking.
Detailed Analysis:
The study's entire assessed outcome domain is English speaking competence (via the Cambridge Business English Speaking Test) and speaking anxiety (via a questionnaire); no other subjects or skill domains (e.g., mathematics, science, or broader literacy) were assessed. Although the intervention is embedded in a specialised Business English speaking course, the paper does not provide the explicit rationale required by the standard's exception for why measuring only speaking outcomes is sufficient; it simply focuses on the course's speaking objectives without discussing potential effects on, or the irrelevance of, other subjects.
Criterion A is not met because only English speaking outcomes were assessed, with no other subjects measured or explicit justification provided for the narrow scope.
-
G
Graduation Tracking
- Measurement stopped at the immediate posttest, with no tracking toward course completion or graduation, and criterion Y is not met.
- "In the posttest phase, all participants complete the posttest SAQ and a final speaking test." (p. 9)
Relevant Quotes:
1) "In the posttest phase, all participants complete the posttest SAQ and a final speaking test." (p. 9)
2) "...the long-term effects of these interventions remain unclear. Longitudinal studies are needed to examine sustained language development over extended periods." (p. 26)
Detailed Analysis:
Data collection concluded at the week-14 posttest, immediately after the intervention ended. The authors explicitly identify the lack of long-term follow-up as a limitation and call for future longitudinal research, confirming that no tracking toward course completion or graduation was conducted or planned.
An internet search for follow-up publications by Wenyi Li, Maslawati Mohamad, and Huay Woon You tracking this same 64-participant cohort toward graduation found none. The search did identify a related companion paper by the same authors, Li, W., Mohamad, M., & You, H. W. (2025), "Exploring the effects of using automatic speech recognition on EFL university students with high speaking anxiety," International Journal of Information and Education Technology, 15(1), 187-194, but this study uses a separate sample of 98 first-year students, tests ASR alone, and contains no graduation-tracking data for the cohort in the present paper. Per the standard's dependency rule, since criterion Y (Year Duration) is not met, criterion G cannot be met regardless of any graduation-tracking evidence.
Criterion G is not met because measurement stopped at the immediate posttest with no tracking toward graduation, no qualifying follow-up publication was found, and criterion Y is not met.
-
P
Pre-Registered
- No pre-registration of the study protocol is mentioned anywhere in the paper, and none was found via internet search.
Relevant Quotes:
No quotes referencing a trial registry, registration ID, or pre-specified protocol published before data collection were found anywhere in the paper, including the methodology, ethics, and disclosure sections.
Detailed Analysis:
The paper contains no reference to any registration platform (e.g., ClinicalTrials.gov, OSF Registries, or an equivalent), nor any registration date preceding data collection. Ethical approval from the university's research ethics committee is mentioned, but this is distinct from prospective registration of hypotheses, methods, and analysis plans.
An internet search (OSF Registries, AsPredicted, and general web search using the authors' names and study title) found no matching pre-registration record for this study.
Criterion P is not met because no pre-registration of the study protocol is mentioned anywhere in the paper, and none was located via internet search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.