Abstract
Pronunciation plays a pivotal role in oral fluency and overall communicative effectiveness in English as a Foreign Language (EFL) learning. However, pronunciation instruction remains a persistent challenge in higher education contexts, particularly where learners have limited opportunities for individualized feedback and sustained oral practice. Responding to recent advances in artificial intelligence (AI), this study investigated the effectiveness of the AI-powered ELSA Speak application in enhancing EFL learners' pronunciation development. Adopting a longitudinal experimental design, the study involved 60 university-level EFL students, equally divided into experimental and control groups. Learners' pronunciation performance was assessed at three time points using the CAFIS analytic framework. Repeated-measures analyses revealed that students using ELSA Speak demonstrated significantly greater improvement across all CAFIS dimensions compared to the control group. Accuracy showed the most immediate gains, followed by gradual but consistent development in intonation and fluency over time. These findings underscore the value of AI-assisted pronunciation instruction and highlight the importance of adopting multidimensional assessment frameworks to capture nuanced developmental changes. The study contributes empirical evidence from a Saudi higher education context and offers pedagogical insights into the integration of AI-driven tools for pronunciation development in EFL classrooms.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- The study is a self-described quasi-experimental non-equivalent group design with individual students, not classes or schools, allocated to conditions, so class-level RCT is not met.
- "Following a non-equivalent group design, 60 second-year English program students from a public university of Saudi Arabia were selected through a convenience sampling and assigned to two groups (experimental: n=30; control: n=30) with random distribution prior to the course onset." (p. 313)
Relevant Quotes:
1) "This study adopted a longitudinal quasi-experimental design to investigate the role of the ELSA Speak application in enhancing EFL students' pronunciation over a 12-week intervention." (p. 313)
2) "Following a non-equivalent group design, 60 second-year English program students from a public university of Saudi Arabia were selected through a convenience sampling and assigned to two groups (experimental: n=30; control: n=30) with random distribution prior to the course onset." (p. 313)
3) "Both groups received identical 90-minute weekly in-class pronunciation instruction... led by the same trained EFL instructor. The only variable difference was the mode of 30-minute daily out-of-class practice." (p. 314)
Detailed Analysis:
Criterion C requires a randomised controlled trial with randomisation at the class level or stronger (school level), with a clearly described randomisation process. This paper explicitly labels itself a "longitudinal quasi-experimental design" following a "non-equivalent group design" — a design that by definition lacks true random assignment. Although the same sentence loosely mentions "random distribution prior to the course onset," participants were recruited by convenience sampling from one university programme and allocated as individual students into two groups, not as intact classes or schools. No description of any class-level or school-level randomisation procedure is provided. The intervention is self-paced app practice supplementing shared whole-class instruction, not one-to-one personal tutoring, so the tutoring exception that would permit student-level randomisation does not apply. Both groups were taught by the same instructor in the same institution, which is exactly the contamination-prone setup class-level randomisation is meant to avoid.
Criterion C is not met because allocation was at the individual student level within a single university programme under a self-described quasi-experimental, non-equivalent group design, with no class- or school-level randomisation and no applicable tutoring exception.
-
E
Exam-based Assessment
- Outcomes were rated with the CAFIS research rubric on study-designed speaking tasks rather than any widely recognised standardised exam, so criterion E is not met.
- "Students' pronunciation performance of both groups was assessed at the beginning (week 1), mid (week 6), and end (week 12) using the CAFIS' criteria prescribed by Derwing and Munro (2005)." (p. 314)
Relevant Quotes:
1) "Learners' pronunciation performance was assessed at three time points using the CAFIS analytic framework." (Abstract, p. 311)
2) "Students' pronunciation performance of both groups was assessed at the beginning (week 1), mid (week 6), and end (week 12) using the CAFIS' criteria prescribed by Derwing and Munro (2005). From 1 (very poor) to 9 (outstanding or native-like), this scale offers a comprehensive range of rankings, focusing on three core dimensions, accuracy, intonation, and fluency." (p. 314)
3) "At each measurement point, participants from both groups undertook a standardized speaking task that included three distinct oral tasks aimed at eliciting comparable oral production for CAFIS rating." (p. 314)
4) "Three trained EFL instructors (with ≥5 years of pronunciation teaching experience) served as raters for the audio-recorded speech samples." (p. 314)
Detailed Analysis:
Criterion E requires outcomes to be measured with widely recognised standardised exams, not researcher-devised instruments. Here, outcomes were speech samples from three study-designed oral tasks (reading aloud of a selected 100-word passage, picture description, free speaking) scored by the researchers' trained raters on the CAFIS 1-9 rating scale adapted from Derwing and Munro's (2005) research rubric. Although the tasks are called "standardized" in the sense of being uniform across participants, CAFIS is an academic rating rubric, not a recognised standardised examination such as IELTS, TOEFL, or a national curriculum exam. The assessment battery was assembled specifically for this study, which is precisely the custom-assessment pattern the criterion excludes.
Criterion E is not met because outcomes were measured with a study-specific speaking task battery scored on a research rating rubric (CAFIS), not a widely recognised standardised exam.
-
T
Term Duration
- Outcomes were measured 12 weeks (about three months) after the intervention began, which just reaches the one-term threshold, so criterion T is met.
- "This study adopted a longitudinal quasi-experimental design to investigate the role of the ELSA Speak application in enhancing EFL students' pronunciation over a 12-week intervention." (p. 313)
Relevant Quotes:
1) "This study adopted a longitudinal quasi-experimental design to investigate the role of the ELSA Speak application in enhancing EFL students' pronunciation over a 12-week intervention." (p. 313)
2) "Students' pronunciation performance of both groups was assessed at the beginning (week 1), mid (week 6), and end (week 12)." (p. 314)
3) "Average compliance across the 12 weeks was 94% (no attrition); participants completed a mean of 30.2 minutes of daily practice (SD = 2.1) and a total of 2174 cumulative practice minutes (SD = 148) over the intervention period." (p. 314)
Detailed Analysis:
Criterion T requires outcomes to be measured at least one full academic term (approximately 3-4 months, a semester or equivalent) after the intervention begins. The intervention began at week 1 and the final outcome measurement took place at week 12, i.e. roughly three months after intervention start. Twelve weeks corresponds to the length of a typical academic term/trimester and sits at the lower bound of the "approximately 3-4 months" definition in the standard. The measurement occurred at the end of the intervention with no further delayed follow-up, but the standard only requires the start-to-measurement interval to span at least one term, which a 12-week span just satisfies.
Criterion T is met because the interval from intervention start (week 1) to the final measurement (week 12) covers approximately three months, equivalent to one academic term.
-
D
Documented Control Group
- The control group's size, composition, baseline scores, and exact activities are documented in detail, so criterion D is met.
- "They were further compared on baseline English proficiency, assessed via pre-test CAFIS scores, revealing no statistically significant differences (p > 0.05) between groups, confirming initial equivalence." (p. 313)
Relevant Quotes:
1) "60 second-year English program students from a public university of Saudi Arabia were selected through a convenience sampling and assigned to two groups (experimental: n=30; control: n=30)... All participants shared a relatively homogeneous linguistic background and had no prior experience of using the ELSA Speak App for English learning. They were further compared on baseline English proficiency, assessed via pre-test CAFIS scores, revealing no statistically significant differences (p > 0.05) between groups, confirming initial equivalence." (p. 313)
2) "In contrast, the control group received traditional pronunciation instruction through teacher-led activities without the support of mobile-assisted pronunciation technology. It included repetition drills... oral reading tasks of EFL textbook passages (Headway Intermediate, 5th Ed.), and written pronunciation worksheets (error correction and pattern recognition). Practice compliance was tracked via weekly worksheet submission and signed practice logs (average compliance: 92%)." (p. 314)
3) "At the pre-test stage, the experimental and control groups demonstrated comparable levels of pronunciation performance across all three dimensions, indicating baseline equivalence." (p. 315, with Table 1 reporting control group means and SDs at all three time points)
Detailed Analysis:
Criterion D requires the control group to be well documented: size, baseline performance, characteristics, and the treatment it received. The paper specifies the control group size (n=30), its composition (second-year English programme students at the same public Saudi university, homogeneous linguistic background, no prior ELSA experience), its baseline CAFIS scores (Table 1, with statistical confirmation of baseline equivalence), and a detailed account of exactly what the control condition received (identical 90-minute weekly in-class instruction plus specified traditional out-of-class practice with tracked 92% compliance). This is sufficient documentation to judge comparability, although finer demographic breakdowns (age, gender per group) are not reported.
Criterion D is met because the control group's size, baseline performance, background, and treatment conditions are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- All participants came from a single university and were allocated individually, so there was no school-level randomisation and criterion S is not met.
- "Following a non-equivalent group design, 60 second-year English program students from a public university of Saudi Arabia were selected through a convenience sampling and assigned to two groups (experimental: n=30; control: n=30) with random distribution prior to the course onset." (p. 313)
Relevant Quotes:
1) "Following a non-equivalent group design, 60 second-year English program students from a public university of Saudi Arabia were selected through a convenience sampling and assigned to two groups (experimental: n=30; control: n=30) with random distribution prior to the course onset." (p. 313)
2) "Both groups received identical 90-minute weekly in-class pronunciation instruction... led by the same trained EFL instructor." (p. 314)
Detailed Analysis:
Criterion S requires randomisation among schools or equivalent implementing institutions. This study took place within a single public university in Saudi Arabia, with individual students from one programme allocated to two groups taught by the same instructor. No multiple schools, campuses, or sites were involved, and no school-level assignment of any kind is described.
Criterion S is not met because the study involved only one institution with student-level allocation, so no school-level randomisation occurred.
-
I
Independent Conduct
- The evaluated intervention (the commercial ELSA Speak app) was not designed by the authors, who declare no competing interests, and scoring was done by raters blind to group assignment.
- "Raters scored all samples independently (blind to group assignment and measurement time point) using the CAFIS scoring rubric." (p. 314)
Relevant Quotes:
1) "One such AI-powered application gaining popularity in EFL contexts is ELSA Speak. It is prominent for its focus on improving learners' spoken English through automated speech recognition and real-time feedback (Arbain et al., 2023)." (p. 312)
2) "Three trained EFL instructors (with ≥5 years of pronunciation teaching experience) served as raters for the audio-recorded speech samples." (p. 314)
3) "Raters scored all samples independently (blind to group assignment and measurement time point) using the CAFIS scoring rubric." (p. 314)
4) "Competing interests: The authors declare that they have no competing interests." (p. 318)
5) "Dr. AHA and Dr. HA were responsible for study design, data collection, and revising." (p. 318)
Detailed Analysis:
Criterion I requires that the study be conducted independently from the designers of the intervention. The intervention under test is the ELSA Speak application, a commercial third-party product; the authors are university researchers who did not develop the app, and they explicitly declare no competing interests, indicating no financial or contractual relationship with the app's developer. This parallels the ERCT exception examples in which evaluation independent of the product provider satisfies the criterion. Additionally, outcome scoring was performed by three experienced EFL instructor-raters who were blind to group assignment and time point, with inter-rater reliability reported (all ICCs >= 0.82), which further insulates the outcome measurement from implementer bias. The authors did design the study procedures and collect the data themselves, but relative to the intervention designer (the ELSA Speak provider) the conduct, data collection, analysis, and conclusions were independent, and the provider is nowhere reported as involved.
Criterion I is met because the evaluation was conducted by researchers independent of the app's developer, with no declared competing interests and blinded outcome rating.
-
Y
Year Duration
- The study tracked outcomes for only 12 weeks, far less than 75% of an academic year, so criterion Y is not met.
- "Students' pronunciation performance of both groups was assessed at the beginning (week 1), mid (week 6), and end (week 12)." (p. 314)
Relevant Quotes:
1) "This study adopted a longitudinal quasi-experimental design to investigate the role of the ELSA Speak application in enhancing EFL students' pronunciation over a 12-week intervention." (p. 313)
2) "Students' pronunciation performance of both groups was assessed at the beginning (week 1), mid (week 6), and end (week 12)." (p. 314)
3) "Further research should include a wider diversity of students, long-term outcomes, and approaches to implementation of the tools in classrooms for better validation and generalization." (p. 318)
Detailed Analysis:
Criterion Y requires outcomes to be measured at least 75% of an academic year (roughly 6.75-7.5 months of a 9-10 month year) after intervention start. Here the entire tracked period was 12 weeks (about 3 months) from start to final measurement, with no later follow-up; the authors themselves list long-term outcomes as future work. Three months falls far short of 75% of an academic year.
Criterion Y is not met because the tracking interval was only 12 weeks, well below 75% of an academic year.
-
B
Balanced Control Group
- Both groups received identical class time and matched 30-minute daily practice, differing only in practice mode (the treatment itself), so criterion B is met.
- "Both groups received identical 90-minute weekly in-class pronunciation instruction... led by the same trained EFL instructor. The only variable difference was the mode of 30-minute daily out-of-class practice." (p. 314)
Relevant Quotes:
1) "Both groups received identical 90-minute weekly in-class pronunciation instruction, focusing English word/sentence stress, rhythmic pacing, and intonation contours, led by the same trained EFL instructor. The only variable difference was the mode of 30-minute daily out-of-class practice." (p. 314)
2) "Subsequently, they were made engaged with ELSA Speak app with 30 minutes per day throughout the entire intervention... participants completed a mean of 30.2 minutes of daily practice (SD = 2.1)." (p. 314)
3) "In contrast, the control group received traditional pronunciation instruction through teacher-led activities without the support of mobile-assisted pronunciation technology. It included repetition drills (for segmental/suprasegmental features) oral reading tasks of EFL textbook passages (Headway Intermediate, 5th Ed.), and written pronunciation worksheets (error correction and pattern recognition). Practice compliance was tracked via weekly worksheet submission and signed practice logs (average compliance: 92%), with the same instructor providing written feedback on completed practice materials." (p. 314)
Detailed Analysis:
Criterion B requires the control condition to receive comparable time and resources, unless extra resources are the explicit treatment variable. Comparing the two conditions: both groups received identical 90-minute weekly in-class instruction from the same instructor, and both completed the same dosage of 30-minute daily out-of-class pronunciation practice, with compliance tracked in each group (94% vs 92%). No extra time or budget was given to the experimental group relative to the control group; the only difference was the mode of practice: AI app with real-time feedback versus traditional drills, reading, and worksheets with written instructor feedback. Applying the criterion's decision tree, since the intervention adds no extra time or budget relative to the control condition (both receive matched 90-minute class time and 30-minute daily practice), the balance requirement is trivially satisfied regardless of whether the AI feedback mode is treated as integral to the intervention.
Criterion B is met because the control group received time-matched active practice with instructor feedback, balancing educational inputs so that only the mode of practice (the AI app, integral to the intervention) differed.
-
Level 3 Criteria
-
R
Reproduced
- This recently published study has not been independently replicated, and earlier ELSA Speak studies are not replications of it, so criterion R is not met.
- "Despite the growing adoption of educational technologies in higher education, little studies have examined the effectiveness of AI-driven pronunciation tools, specifically, ELSA Speak." (p. 312)
Relevant Quotes:
1) "Despite the growing adoption of educational technologies in higher education, little studies have examined the effectiveness of AI-driven pronunciation tools, specifically, ELSA Speak." (p. 312)
2) "In a similar vein, a study conducted by Al-Shallakh (2024) at a Jordanian university reported significant enhancement in students' pronunciation after a seven-week intervention with this app." (p. 312)
3) "The study contributes empirical evidence from a Saudi higher education context and offers pedagogical insights." (Abstract, p. 311)
Detailed Analysis:
Criterion R requires that this specific study be independently replicated by a different research team, in a different context, in a peer-reviewed journal. The paper was published online in April 2026 and explicitly positions itself as filling a gap ("little studies have examined... ELSA Speak" in the Saudi context), which signals novelty rather than replication. The literature it cites (e.g., Al-Shallakh 2024, Arbain et al. 2023, Karim et al. 2023) consists of earlier, methodologically different studies of the same app by other teams; these predate the present study and are not replications of this specific 12-week CAFIS-based design. Internet searches (Google Scholar-style queries, publisher sites, PubMed/PMC, ResearchGate, and general web search) for independent replications of this specific 12-week, CAFIS-based, Qassim University trial found only earlier, unrelated ELSA Speak studies (e.g., Pham & Pham, 2025 on learner satisfaction; other Indonesian/Saudi ELSA Speak papers) and no study replicating this trial's design, population, and outcomes by an independent team. Given the very recent publication date (April 2026), no subsequent independent replication of this particular study could be identified.
Criterion R is not met because no independent, peer-reviewed replication of this specific study exists; prior ELSA Speak studies are separate earlier investigations, not replications of this trial.
-
A
All-subject Exams
- Only pronunciation was measured with a non-standardised rubric (E fails), so the all-subject exams criterion is not met.
- "Learners' pronunciation performance was assessed at three time points using the CAFIS analytic framework." (Abstract, p. 311)
Relevant Quotes:
1) "Learners' pronunciation performance was assessed at three time points using the CAFIS analytic framework." (Abstract, p. 311)
2) "From 1 (very poor) to 9 (outstanding or native-like), this scale offers a comprehensive range of rankings, focusing on three core dimensions, accuracy, intonation, and fluency." (p. 314)
Detailed Analysis:
Criterion A requires standardised exam-based assessment across all main subjects, and per the instructions it cannot be met when criterion E fails. Criterion E is not met, since the only outcome measure is the CAFIS pronunciation rubric. Moreover, only English pronunciation was assessed; no other subjects of the students' university programme were measured, and no specialised-intervention justification tied to standardised related-subject exams is offered.
Criterion A is not met because criterion E fails and only a single narrow outcome (pronunciation) was assessed, with no standardised exams across subjects.
-
G
Graduation Tracking
- Measurement stopped at the week-12 post-test with no tracking of the second-year students to graduation, so criterion G is not met.
- "Further research should include a wider diversity of students, long-term outcomes, and approaches to implementation of the tools in classrooms for better validation and generalization." (p. 318)
Relevant Quotes:
1) "Students' pronunciation performance of both groups was assessed at the beginning (week 1), mid (week 6), and end (week 12)." (p. 314)
2) "Further research should include a wider diversity of students, long-term outcomes, and approaches to implementation of the tools in classrooms for better validation and generalization of artificial intelligence in language instruction." (p. 318)
Detailed Analysis:
Criterion G requires participants to be tracked until graduation from their educational stage, and per the instructions it cannot be met when criterion Y fails. Criterion Y is not met. Measurement ended at week 12 with the final post-test; the participants were second-year university students and no follow-up to the end of their degree programme is reported or planned in the paper — indeed the authors identify long-term outcomes as future research. Internet searches for subsequent publications by this author team (Abdelrady, Akram, and co-authors) tracking this same 60-student Qassim University cohort found no follow-up publication reporting graduation outcomes; the paper was only published online in April 2026, making any graduation follow-up implausible at this time regardless.
Criterion G is not met because tracking ended at week 12, well before participants' graduation, criterion Y also fails, and no follow-up publication tracking this cohort was found.
-
P
Pre-Registered
- No pre-registration, registry ID, or registration date is mentioned anywhere in the paper or found via internet search, so criterion P is not met.
Relevant Quotes:
1) "Furthermore, the study followed all ethical requirements set out by the World Medical Association's (2013) declaration of Helsinki... Their participation was voluntary, and informed consent was obtained prior to data collection." (p. 313)
2) "The data were analyzed with SPSS version 26.0." (p. 314)
Detailed Analysis:
Criterion P requires the full study protocol to be pre-registered on a public registry before data collection began, with quoted evidence of the registry and timing. The paper describes ethics compliance and consent procedures but nowhere mentions any registry platform (e.g., ClinicalTrials.gov, OSF, AsPredicted), registration ID, or registration date. No pre-registration statement of any kind appears in the methods, acknowledgments, or declarations sections. Internet searches of common trial and study registries (OSF Registries, ClinicalTrials.gov, AsPredicted, ISRCTN) and general web search for this title, authors, and DOI found no pre-registration record for this study.
Criterion P is not met because the paper contains no mention of any pre-registered protocol or registry entry, and no registry record was located through internet search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.