Abstract
Introduction: The growing availability of generative artificial intelligence (AI) tools has created new opportunities to enhance language learning, particularly speaking practice. However, studies on the effectiveness of AI voice-chat applications in enhancing English as a Foreign Language (EFL) learners' speaking skills and autonomy remain limited. Methods: This mixed-methods quasi-experimental study investigated the effects of Microsoft 365 Copilot voice chat on speaking fluency, interactional competence, and learner autonomy of 52 university English as a Foreign Language (EFL) learners over eight weeks. The experimental group (n = 26) engaged in voice-based speaking tasks with Copilot, while the control group (n = 26) completed parallel peer-to-peer speaking activities. Results: Quantitative analyses revealed that the experimental group demonstrated greater improvements in speech rate, mean length of run, and pause reduction, as well as larger improvements in turn-taking, repair strategies, and topic management. Learner autonomy scores increased significantly only for the experimental group [t(25) = 4.12, p < .001, d = 0.82, mean difference = 0.90, 95% CI [0.45, 1.35]], with 71% of participants voluntarily exceeding the required practice time. Qualitative reflective responses with eight experimental participants indicated reduced speaking anxiety, perceived usefulness of immediate feedback, and increased motivation. Discussion: These findings suggest that AI voice chat, when integrated into learners' existing academic tools, may support multidimensional speaking development and autonomous practice in EFL contexts.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- The study is a self-described quasi-experimental design using two intact classes assigned partly on scheduling and administrative grounds, with contradictory statements about randomisation, so a properly implemented class-level RCT is not established.
- "One class was designated as the experimental group and the other as the control group based on scheduling and administrative considerations, as random assignment of individual students was not feasible within the institutional context." (p. 5)
Relevant Quotes:
1) "This mixed-methods quasi-experimental study investigated the effects of Microsoft 365 Copilot voice chat on speaking fluency, interactional competence, and learner autonomy of 52 university English as a Foreign Language (EFL) learners over eight weeks." (Abstract, p. 1)
2) "The study employed a mixed-methods quasi-experimental research design." (p. 4)
3) "The quasi-experimental design using intact classes assigned to experimental and control groups aimed to examine the instructional effects of AI-mediated speaking practice while maintaining ecological validity." (p. 4)
4) "One class was designated as the experimental group and the other as the control group based on scheduling and administrative considerations, as random assignment of individual students was not feasible within the institutional context." (p. 5)
5) "Because random assignment of individual students was not feasible, the two classes were randomly assigned to the study conditions: the experimental group (AI-mediated speaking practice, n = 26) and the control group (peer-to-peer speaking practice, n = 26)." (p. 5)
6) "First, the quasi-experimental design used intact classes, limiting causal conclusions despite pre-test equivalence and statistical controls." (p. 12)
Detailed Analysis:
Criterion C requires a randomised controlled trial with randomisation clearly described and properly implemented at the class level or stronger. The paper repeatedly labels itself quasi-experimental (abstract, Section 3.1, and the limitations section), and the two descriptions of group allocation directly contradict each other: Section 3.1 states that one class was "designated" to each condition "based on scheduling and administrative considerations", while Section 3.2 claims the two classes "were randomly assigned". No randomisation procedure (e.g., method of random allocation, who performed it) is described anywhere. With only two intact clusters and the authors themselves conceding in the limitations that the intact-class quasi-experimental design limits causal conclusions, the study cannot be credited as a properly implemented class-level RCT. The intervention is individual AI voice-chat practice within regular class sessions, not one-to-one tutoring by a teacher, so the tutoring exception does not apply; and even if it did, no random assignment of individual students took place.
Criterion C is not met because allocation of the two intact classes is described inconsistently and the authors themselves characterise the study as quasi-experimental rather than a properly randomised trial.
-
E
Exam-based Assessment
- Outcomes were measured with a researcher-designed 2-minute speaking task, a researcher-adapted rubric, and a researcher-developed questionnaire rather than any widely recognised standardised exam.
- "All speaking performances were audio-recorded, transcribed, anonymized, and evaluated using a researcher-developed scoring protocol..." (p. 7)
Relevant Quotes:
1) "A speaking fluency test was administered as both a pre-test and post-test. Participants completed a structured 2-minute individual speaking task on familiar discussion topics aligned with the themes practiced during the intervention." (p. 5)
2) "Learners' interactional competence was assessed using an analytic rubric adapted from previous research on interactional competence and technology-mediated communication (Plonsky and Ziegler, 2016; Ziegler, 2016)." (p. 6)
3) "Learner autonomy in speaking practice was measured using a researcher-developed questionnaire informed by Benson, 2011 conceptualization of learner autonomy..." (p. 6)
4) "All speaking performances were audio-recorded, transcribed, anonymized, and evaluated using a researcher-developed scoring protocol that included standardized transcription procedures, scoring guidelines, and operational definitions for fluency and interactional competence measures." (p. 7)
5) "Based on the university's English Language Institute placement test, all participants were classified at the B1-B2 proficiency levels." (p. 5)
Detailed Analysis:
Criterion E requires that outcomes be measured with standard, widely recognised standardised exams rather than instruments built for the study. All three outcome measures here were created or adapted by the researchers: a bespoke 2-minute speaking task scored with temporal fluency metrics, an analytic interactional-competence rubric adapted by the authors, and an explicitly "researcher-developed" learner autonomy questionnaire. The only standardised instrument mentioned, the university placement test, was used solely to screen participants at B1-B2 level, not to measure outcomes. No national or internationally recognised speaking exam (e.g., IELTS, TOEFL speaking) was used for the pre/post assessment, and the speaking topics were "aligned with the themes practiced during the intervention", raising exactly the intervention-alignment concern the criterion is designed to prevent.
Criterion E is not met because all outcome measures were custom, researcher-developed instruments rather than standardised exams.
-
T
Term Duration
- The intervention lasted eight weeks with post-tests immediately afterwards in Week 9, which is roughly two months and falls short of a full academic term of 3-4 months.
- "Data collection spanned ten weeks. Week 1 involved orientation, pre-testing... Weeks 2-9 constituted the eight-week treatment phase... At the end of Week 9, all participants completed post-tests." (p. 7)
Relevant Quotes:
1) "The instructional treatment was conducted over an eight-week period as part of the participants' regular speaking coursework." (p. 5)
2) "Data collection spanned ten weeks. Week 1 involved orientation, pre-testing (speaking fluency, interactional competence, learner autonomy), and explanation of study procedures. Weeks 2-9 constituted the eight-week treatment phase, during which the experimental group practiced with Microsoft 365 Copilot voice chat while the control group completed parallel peer-to-peer activities. At the end of Week 9, all participants completed post-tests." (p. 7)
3) "Third, the intervention period was relatively short, focusing on immediate rather than long-term changes." (p. 12)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. Here the treatment began in Week 2 and post-tests were completed at the end of Week 9, giving an interval of only about eight weeks (roughly two months) from intervention start to outcome measurement. There was no delayed follow-up measurement; the authors themselves flag the "relatively short" intervention period as a limitation. Eight weeks is clearly below the 3-4 month term threshold.
Criterion T is not met because the interval from intervention start to outcome measurement was only about eight weeks, shorter than one academic term.
-
D
Documented Control Group
- The control group (n = 26) is clearly documented with demographics, baseline comparisons, and a full description of the parallel peer-to-peer condition it received.
- "Participants in the control group completed parallel peer-to-peer speaking activities covering the same topics, task structures, and approximate time-on-task conditions." (p. 5)
Relevant Quotes:
1) "The experimental group (n = 26) engaged in voice-based speaking tasks with Copilot, while the control group (n = 26) completed parallel peer-to-peer speaking activities." (Abstract, p. 1)
2) "Pre-test comparisons revealed no statistically significant differences between the groups on any study variable (p > .05), and pre-test scores were included as covariates in the primary analyses." (p. 5)
3) "Background information collected prior to the study included age, gender, years of English study, prior experience with AI tools, and learners' general exposure to English and digital technologies outside the classroom. No substantial differences were observed between the two groups on these background characteristics prior to the intervention. These data are presented in Table 1." (p. 5)
4) "No statistically significant differences were observed between groups on any demographic characteristic (p > .05 for all comparisons), supporting the comparability of the two groups prior to the intervention." (Table 1 note, p. 6)
5) "Participants in the control group completed parallel peer-to-peer speaking activities covering the same topics, task structures, and approximate time-on-task conditions. Both groups received the same instructional materials and were taught by the same instructor throughout the study period." (p. 5)
Detailed Analysis:
Criterion D requires detailed documentation of the control group: who they are, baseline characteristics, and what they received. The paper reports the control group's size (n = 26), demographic characteristics in Table 1 (age, gender, years of English study, prior AI experience) with statistical comparisons to the experimental group, baseline pre-test equivalence on all study variables, and a clear description of the control condition (peer-to-peer speaking activities on the same topics with the same materials and instructor). Pre-test and post-test means and standard deviations for the control group are reported in Tables 2-8. This constitutes adequate documentation for comparison.
Criterion D is met because the control group's size, demographics, baseline performance, and received condition are all clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- The study involved two intact classes within a single university; no schools or institutions were randomised.
- "A convenience sampling approach was used involving two pre-existing intact classes taught by the researcher." (p. 5)
Relevant Quotes:
1) "The participants were 52 undergraduate students enrolled in a university-level EFL speaking course at a public university in Saudi Arabia." (p. 5)
2) "A convenience sampling approach was used involving two pre-existing intact classes taught by the researcher." (p. 5)
3) "Second, a modest sample size and a single educational setting restrict generalizability." (p. 12)
Detailed Analysis:
Criterion S requires randomisation among schools or equivalent institutional units. This study took place at a single public university in Saudi Arabia and involved only two intact classes within that one institution. No schools, campuses, or sites were randomised; at most, the two classes within one university were allocated to conditions (and even that allocation is inconsistently described, see criterion C). A single-site two-class design cannot satisfy school-level randomisation.
Criterion S is not met because the study was conducted with two classes at a single university and no school-level randomisation occurred.
-
I
Independent Conduct
- The authors designed, delivered, and analysed the intervention themselves, with the researcher also serving as the course instructor and no external evaluation team.
- "Because the researcher also served as the course instructor, participation was voluntary, and written informed consent was obtained prior to data collection." (p. 5)
Relevant Quotes:
1) "A convenience sampling approach was used involving two pre-existing intact classes taught by the researcher." (p. 5)
2) "Because the researcher also served as the course instructor, participation was voluntary, and written informed consent was obtained prior to data collection." (p. 5)
3) "Both groups received the same instructional materials and were taught by the same instructor throughout the study period." (p. 5)
4) "Two independent EFL instructors with prior experience in speaking assessment were trained to use the scoring rubrics and independently evaluated all speaking performances." (p. 6)
5) "AAb: Investigation, Writing - review & editing, Conceptualization, Data curation, Validation, Supervision, Methodology. AK: Formal analysis, Methodology, Validation, Conceptualization, Supervision, Writing - review & editing..." (Author contributions, p. 13)
Detailed Analysis:
Criterion I requires that the study be conducted independently of those who designed the intervention. Here the research team designed the treatment, and the researcher personally taught both intact classes and administered the intervention. The author contribution statement confirms the authors performed conceptualisation, investigation, data curation, and formal analysis themselves. The only independent element is the use of two external EFL instructors as raters of the recorded speaking performances; this is a blinding measure for scoring, not independent conduct of the trial, since data collection, implementation, and analysis all remained with the author team. No external evaluation agency or third-party oversight is described.
Criterion I is not met because the same team (including the instructor-researcher) designed, delivered, and analysed the intervention with no independent evaluators beyond two hired raters.
-
Y
Year Duration
- Tracking lasted only about ten weeks in total, far below 75% of an academic year, and criterion T is already unmet.
- "The instructional treatment was conducted over an eight-week period as part of the participants' regular speaking coursework." (p. 5)
Relevant Quotes:
1) "The instructional treatment was conducted over an eight-week period as part of the participants' regular speaking coursework." (p. 5)
2) "Data collection spanned ten weeks." (p. 7)
3) "Third, the intervention period was relatively short, focusing on immediate rather than long-term changes." (p. 12)
Detailed Analysis:
Criterion Y requires outcome measurement at least 75% of an academic year (roughly 7+ months) after the intervention begins. The entire study, including orientation, pre-testing, eight weeks of treatment, post-testing, and reflective responses, spanned only ten weeks within a single academic semester. Since criterion T (one term) is not met, criterion Y cannot be met either, per the standard's dependency rule.
Criterion Y is not met because the total tracking period of about ten weeks is far short of 75% of an academic year.
-
B
Balanced Control Group
- The control group completed parallel peer-to-peer speaking activities with the same topics, task structures, materials, instructor, and approximate time-on-task, so educational inputs were balanced and the AI tool itself was the treatment variable.
- "Participants in the control group completed parallel peer-to-peer speaking activities covering the same topics, task structures, and approximate time-on-task conditions. Both groups received the same instructional materials and were taught by the same instructor throughout the study period." (p. 5)
Relevant Quotes:
1) "Students completed approximately 5-10 min of AI-mediated speaking practice per session, four times per week, resulting in an estimated 160-320 min across the intervention period." (p. 5)
2) "Participants in the control group completed parallel peer-to-peer speaking activities covering the same topics, task structures, and approximate time-on-task conditions. Both groups received the same instructional materials and were taught by the same instructor throughout the study period." (p. 5)
3) "Both groups completed communicative speaking tasks aligned with the course objectives and weekly instructional themes." (p. 5)
4) "The activities were completed individually during scheduled class sessions under instructor supervision using university-supported Microsoft 365 accounts. Students used personal mobile devices and earphones to minimize classroom disruption." (p. 5)
5) "As such, learners in the experimental group performed speaking tasks while interacting with AI through Microsoft 365 Copilot voice chat, whereas those in the control condition completed identical speaking tasks with their peers." (p. 5)
Detailed Analysis:
Applying the criterion B decision tree: the experimental group's extra input was access to the Copilot voice-chat tool, delivered through existing university-supported Microsoft 365 accounts and students' own devices, so no meaningful additional budget was involved. Time was explicitly balanced: the control group completed "identical speaking tasks with their peers" covering "the same topics, task structures, and approximate time-on-task conditions", with the same materials and instructor. This is an active control matching the intervention's educational time. The only systematic difference between conditions was the interlocutor (AI vs. peer), which is precisely the treatment variable being tested and is integral to the intervention as defined. A caveat is that experimental participants could voluntarily practise beyond the required time (71% did), but this self-directed extra practice was itself an outcome of interest (learner autonomy), not a resource allocated by the researchers.
Criterion B is met because the control condition provided matched time, tasks, materials, and instruction, with the AI conversational partner being the integral treatment variable.
-
Level 3 Criteria
-
R
Reproduced
- The paper was published in July 2026 and no independent peer-reviewed replication of this specific study exists or is referenced.
Relevant Quotes:
1) "However, studies on the effectiveness of AI voice-chat applications in enhancing English as a Foreign Language (EFL) learners' speaking skills and autonomy remain limited." (Abstract, p. 1)
2) "Future researchers may also choose to include learners of different proficiency levels and in different contexts to determine how generalizable the results are." (p. 13)
3) "This result builds on previous research by Tai and Chen (2024), who found that daily AI-based speaking training increased speech rate and reduced pauses, and shows that these benefits can be achieved through an AI tool within learners' existing academic environment rather than through a separate app." (p. 11)
Detailed Analysis:
Criterion R requires independent replication of this specific study by a different research team, published in a peer-reviewed journal. The paper was published on 24 July 2026, only three days before this verification check, so no replication could plausibly exist yet. The paper positions itself as addressing a gap ("studies... remain limited") and calls for future research in other contexts, confirming its novelty. An internet search for the paper's title, authors, and DOI (10.3389/feduc.2026.1851080) found only the original Frontiers in Education publication itself, with no citing replication studies. Related prior work exists (e.g., Tai and Chen, 2024 on AI chatbots; Do, 2025 on Copilot and utterance fluency), but those are earlier, different studies with different designs and populations, not replications of this Copilot voice-chat trial. No independent replication of this specific study was found.
Criterion R is not met because no independent peer-reviewed replication of this newly published study exists.
-
A
All-subject Exams
- Only EFL speaking outcomes were measured with custom instruments; no other core subjects were assessed and criterion E is not met, so criterion A automatically fails.
- "This mixed-methods quasi-experimental study investigated the effects of Microsoft 365 Copilot voice chat on speaking fluency, interactional competence, and learner autonomy..." (Abstract, p. 1)
Relevant Quotes:
1) "This mixed-methods quasi-experimental study investigated the effects of Microsoft 365 Copilot voice chat on speaking fluency, interactional competence, and learner autonomy of 52 university English as a Foreign Language (EFL) learners over eight weeks." (Abstract, p. 1)
2) "Data were collected using five instruments designed to assess quantitative and qualitative aspects of learners' speaking development." (p. 5)
Detailed Analysis:
Criterion A requires standardised exam-based assessment across all main subjects taught at the educational level. This study measured only English speaking fluency, interactional competence, and learner autonomy; no other academic subjects were assessed. Moreover, criterion E is a prerequisite for criterion A, and E is not met because all instruments were researcher-developed rather than standardised exams. While one could argue a specialised focus for a university EFL speaking course, the paper offers no standardised assessment even of the target domain, so no exception applies.
Criterion A is not met because criterion E fails and only custom measures of one subject area were used.
-
G
Graduation Tracking
- Measurement stopped at the Week 9 post-test with no follow-up or graduation tracking, and prerequisite criterion Y is not met.
- "Third, the intervention period was relatively short, focusing on immediate rather than long-term changes." (p. 12)
Relevant Quotes:
1) "At the end of Week 9, all participants completed post-tests." (p. 7)
2) "Third, the intervention period was relatively short, focusing on immediate rather than long-term changes." (p. 12)
3) "Future research may examine extending the study's timeframe to examine whether AI has long-term impacts on learner development." (p. 13)
Detailed Analysis:
Criterion G requires tracking participants until graduation from their educational stage. Measurement here ended immediately after the eight-week intervention, at the end of Week 9, with reflective responses in Week 10. The authors explicitly acknowledge the focus on "immediate rather than long-term changes" and propose extended timeframes only as future research. A search for subsequent papers by this author team (Abdelrady, Khalil, Ali El Deen, Alnofal, Mohamed) found no follow-up publication tracking this same cohort of 52 EFL learners toward graduation; given the paper was published on 24 July 2026, only days before this check, no such follow-up could plausibly exist yet. Additionally, the standard's dependency rule states that if criterion Y is not met, criterion G cannot be met.
Criterion G is not met because tracking ended at the immediate post-test with no follow-up towards graduation, and no subsequent tracking publication by the authors was found.
-
P
Pre-Registered
- The paper contains no pre-registration statement or registry ID, and no external registration record was found.
Relevant Quotes:
1) "Ethical approval was obtained from the institution prior to data collection." (p. 8)
2) "The studies involving humans were approved by Ethical Approval Committee at North Private College of Nursing, KSA." (Ethics statement, p. 13)
3) "The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation." (Data availability statement, p. 13)
Detailed Analysis:
Criterion P requires that the full study protocol be pre-registered on a public registry before data collection began, with a verifiable registration reference and date. The paper reports institutional ethical approval and a data availability statement, but nowhere mentions any registry (e.g., ClinicalTrials.gov, OSF, AsPredicted, AEA registry), registration ID, or registration date. An internet search for the paper's title, DOI, and author names combined with "preregistration," "OSF," and "AsPredicted" likewise found no pre-registration record for this study. Ethical approval is not a substitute for pre-registration of hypotheses, methods, and analysis plans.
Criterion P is not met because no pre-registration of the study protocol is mentioned in the paper or discoverable externally.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.