Exploring the impacts of an AI-driven instructional intervention on Iranian EFL learners' pronunciation skill development

Ismail Xodabande, Sepideh Shiri, Mohammad Zohrabi

Published:
ERCT Check Date:
DOI: 10.1007/s44217-025-00782-2
  • L2 languages
  • adult education
  • Asia
  • homework
  • EdTech app
  • mobile learning
0
  • C

    Randomization was at the individual student level, but the intervention is individualized, self-directed one-to-one pronunciation practice with an AI tutor, so the personal teaching/tutoring exception applies and student-level randomization is acceptable.

    "Participants were randomly assigned to one of two groups: an experimental group (n=30), which used ChatGPT for pronunciation practice, and a waiting control group (n=30), which used traditional tools (i.e. electronic dictionary)." (p. 3)

  • E

    Outcomes were measured with a researcher-designed 30-word read-aloud task scored by three teachers, not a widely recognized standardized exam.

    "The study focused on 30 English words that are common in daily conversation but present pronunciation difficulties for intermediate EFL learners" (pp. 3-4)

  • T

    The intervention lasted three weeks with a delayed post-test about five to six weeks after the start, far shorter than one full academic term.

    "The intervention lasted three weeks, during which both groups practiced 10 target words each week in two self-directed sessions at home" (p. 5)

  • D

    The control group's size, composition, baseline scores, and exact conditions (electronic dictionary practice with the same schedule and logging) are clearly documented.

    "Participants were randomly assigned to one of two groups: an experimental group (n=30), which used ChatGPT for pronunciation practice, and a waiting control group (n=30), which used traditional tools (i.e. electronic dictionary)." (p. 3)

  • S

    Randomization was at the individual student level within a single private language institute, not across schools or institutions.

    "The participants of this study were 60 Iranian EFL learners ... enrolled in a private language institute in Tehran, Iran. ... Participants were randomly assigned to one of two groups" (p. 3)

  • I

    The authors designed the intervention and also collected and analyzed the data themselves, with no independent third-party evaluation team, so independence is not established despite blinded raters.

    "IX : Data analysis, Writing- editing, Supervision: SS: Data collection, Methodology, Writing- original draft; MZ: Writing, Review and editing." (p. 10)

  • Y

    The full study window from intervention start to the delayed post-test was about six weeks, far below 75% of an academic year; also criterion T is not met, which precludes Y.

    "At the end of the three-week intervention, a post-test was administered in the fourth week. ... A delayed post-test was conducted two weeks after the post-test" (p. 5)

  • B

    The control was an active condition with the same word lists, practice schedule, and logging, and the AI tool contrast (plus its brief orientation) is the integral treatment variable being tested against dictionary-based practice.

    "The intervention lasted three weeks, during which both groups practiced 10 target words each week in two self-directed sessions at home, following their weekly on-site classes." (p. 5)

  • R

    No independent replication of this 2025 study by a different research team was found in the paper or in an internet literature search conducted during verification.

  • A

    Only pronunciation accuracy in English was measured, with a custom instrument, so neither the standardized-exam prerequisite nor all-subject coverage is satisfied.

    "Pronunciation performance was assessed over three phases: pre-test, post-test, and delayed post-test." (p. 1)

  • G

    Measurement ended at a delayed post-test two weeks after the intervention, with no tracking of participants to the end of their course of study, and criterion Y is not met, which precludes G.

    "A delayed post-test was conducted two weeks after the post-test to measure pronunciation retention." (p. 5)

  • P

    The paper mentions ethics approval but contains no pre-registration statement, registry name, ID, or registration date, and none was found via registry search.

Abstract

The integration of artificial intelligence (AI) into language education is rapidly transforming instructional practices and learner engagement. Within the domain of second language acquisition, pronunciation plays a crucial role in achieving communicative competence and intelligibility. Recent advancements in AI technologies offer promising opportunities to support pronunciation instruction by providing immediate, individualized, and low-anxiety feedback. This study investigated the effectiveness of AI-driven tools, specifically ChatGPT, in improving the pronunciation accuracy of Iranian EFL learners through a randomized controlled trial. Sixty intermediate learners were randomly assigned to either an experimental group, which practiced pronunciation using ChatGPT, or a control group, which relied on electronic dictionaries. Pronunciation performance was assessed over three phases: pre-test, post-test, and delayed post-test. A repeated measures mixed ANOVA was employed to evaluate group differences and changes over time. Results indicated that the ChatGPT group demonstrated significantly greater improvements in pronunciation accuracy, with gains sustained over time. These findings highlight the potential of interactive AI tools to support both immediate learning and retention in pronunciation instruction and offer pedagogical insights into how AI tools can be meaningfully integrated into EFL pronunciation instruction to promote learner autonomy and retention.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomization was at the individual student level, but the intervention is individualized, self-directed one-to-one pronunciation practice with an AI tutor, so the personal teaching/tutoring exception applies and student-level randomization is acceptable.
      • "Participants were randomly assigned to one of two groups: an experimental group (n=30), which used ChatGPT for pronunciation practice, and a waiting control group (n=30), which used traditional tools (i.e. electronic dictionary)." (p. 3)
      • Relevant Quotes: 1) "Participants were randomly assigned to one of two groups: an experimental group (n=30), which used ChatGPT for pronunciation practice, and a waiting control group (n=30), which used traditional tools (i.e. electronic dictionary)." (p. 3) 2) "Sixty intermediate learners were randomly assigned to either an experimental group, which practiced pronunciation using ChatGPT, or a control group, which relied on electronic dictionaries." (p. 1, Abstract) 3) "The intervention lasted three weeks, during which both groups practiced 10 target words each week in two self-directed sessions at home, following their weekly on-site classes." (p. 5) 4) "Hello, ChatGPT. I'd like to practice my English pronunciation. ... listen carefully to my pronunciation of the word '[target word]' and give me feedback on how to improve it." (pp. 4-5) Detailed Analysis: The study randomized 60 individual learners from one private language institute to two conditions, so the unit of randomization is the student, not the class or school. Ordinarily this would fail criterion C. However, the ERCT standard provides an exception: if the intervention is designed for personal teaching like tutoring, then student-level randomization is acceptable. The intervention here is inherently individual: each learner practiced alone at home in self-directed sessions, interacting one-to-one with ChatGPT, which listened to the learner's pronunciation and gave immediate, individualized corrective feedback, functioning as a personal pronunciation tutor. The control condition was likewise individual practice with an electronic dictionary. The intervention was not delivered to classes at all; there was no classroom teaching component being randomized. This matches the standard's tutoring/personal teaching exception, under which normal student-level RCT is considered acceptable. Verified against the source PDF: all quotes above are verbatim and correctly cited. Criterion C is met because, although randomization was at the student level, the intervention is individualized one-to-one AI-tutored practice, which qualifies for the personal teaching/tutoring exception.
    • E

      Exam-based Assessment

      • Outcomes were measured with a researcher-designed 30-word read-aloud task scored by three teachers, not a widely recognized standardized exam.
      • "The study focused on 30 English words that are common in daily conversation but present pronunciation difficulties for intermediate EFL learners" (pp. 3-4)
      • Relevant Quotes: 1) "The study focused on 30 English words that are common in daily conversation but present pronunciation difficulties for intermediate EFL learners, particularly with segmental and suprasegmental features (e.g., consonant clusters, vowel length). These words were chosen based on phonetic complexity and relevance to learner needs and are listed in Table 1." (pp. 3-4) 2) "Their difficulty was further validated by three experienced EFL instructors, who rated each word based on phonological complexity and observed learner challenges in prior instruction." (p. 4) 3) "For each testing phase (pre-test, post-test, and delayed post-test), participants were provided with a set of 30 sentences, each containing one of the target words. Each sentence was constructed to ensure a natural use of the target word within a typical conversational context" (p. 4) 4) "Three experienced language teachers independently rated the recorded pronunciations of the target words across all testing sessions ... Each target word was scored as either correct or incorrect, with 1 point awarded for each correctly pronounced word." (p. 5) Detailed Analysis: Criterion E requires a standard, widely recognized, standardized exam-based assessment, not an instrument designed specifically for the study. Here the outcome measure was a custom read-aloud task: the researchers themselves selected 30 difficult words, constructed original sentences for each testing phase, and had three teachers rate recordings as correct/incorrect. Although inter-rater reliability was high (Cronbach's alpha = 0.91) and raters were blind, the instrument is entirely researcher-made and is not a recognized standardized examination such as IELTS, TOEFL, or a national curriculum exam. No standardized test is named anywhere in the paper. Verified against the source PDF: all quotes above are verbatim and correctly cited. Criterion E is not met because the assessment was a custom-designed word-reading task created for this study rather than a recognized standardized exam.
    • T

      Term Duration

      • The intervention lasted three weeks with a delayed post-test about five to six weeks after the start, far shorter than one full academic term.
      • "The intervention lasted three weeks, during which both groups practiced 10 target words each week in two self-directed sessions at home" (p. 5)
      • Relevant Quotes: 1) "The intervention lasted three weeks, during which both groups practiced 10 target words each week in two self-directed sessions at home, following their weekly on-site classes." (p. 5) 2) "At the end of the three-week intervention, a post-test was administered in the fourth week." (p. 5) 3) "A delayed post-test was conducted two weeks after the post-test to measure pronunciation retention." (p. 5) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. In this study the intervention ran for three weeks, the post-test was in week four, and the final delayed post-test occurred two weeks later, i.e. roughly five to six weeks after the intervention started. The standard allows short interventions if follow-up tracking extends at least a term from the start, but here the entire measurement window ended after about six weeks, well short of a 3-4 month term. No longer follow-up is reported. Verified against the source PDF: all quotes above are verbatim and correctly cited. Criterion T is not met because the interval from intervention start to the final measurement was only about six weeks, which is shorter than one academic term.
    • D

      Documented Control Group

      • The control group's size, composition, baseline scores, and exact conditions (electronic dictionary practice with the same schedule and logging) are clearly documented.
      • "Participants were randomly assigned to one of two groups: an experimental group (n=30), which used ChatGPT for pronunciation practice, and a waiting control group (n=30), which used traditional tools (i.e. electronic dictionary)." (p. 3)
      • Relevant Quotes: 1) "The participants of this study were 60 Iranian EFL learners (21 males and 39 females) enrolled in a private language institute in Tehran, Iran. They were aged between 19 and 25 years and were at an intermediate proficiency level in English." (p. 3) 2) "Participants were randomly assigned to one of two groups: an experimental group (n=30), which used ChatGPT for pronunciation practice, and a waiting control group (n=30), which used traditional tools (i.e. electronic dictionary)." (p. 3) 3) "In the control group, participants noted the dictionary entries used and logged their repetitions. Although practice occurred at home, learners were instructed to avoid supplementary resources. The control group was explicitly instructed not to use AI tools, and informal post-study interviews confirmed general adherence to these conditions." (p. 5) 4) "Table 2 provides the descriptive statistics for pronunciation scores at each testing phase (pre-test, post-test, and delayed post-test) for both the experimental (ChatGPT) and control (electronic dictionary) groups." (p. 6) Detailed Analysis: Criterion D requires detailed documentation of the control group: who they are, their baseline performance, and what treatment they received. The paper documents the control group's size (n=30), that they were drawn from the same pool of 60 intermediate-level learners aged 19-25 at one Tehran institute, and precisely what they did: two self-directed home sessions per week practicing the same 10 words per week using electronic dictionaries, with logged repetitions and an explicit instruction not to use AI tools. Baseline performance is reported in Table 2 (control pre-test mean 8.20, SD reported, with normality checks), enabling direct baseline comparison with the experimental group. The demographics are given for the full sample rather than per group, but the size, conditions, and baseline scores of the control group are clearly documented. Verified against the source PDF: all quotes above are verbatim and correctly cited. Criterion D is met because the control group's size, source population, baseline scores, and exact conditions are clearly described.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomization was at the individual student level within a single private language institute, not across schools or institutions.
      • "The participants of this study were 60 Iranian EFL learners ... enrolled in a private language institute in Tehran, Iran. ... Participants were randomly assigned to one of two groups" (p. 3)
      • Relevant Quotes: 1) "The participants of this study were 60 Iranian EFL learners (21 males and 39 females) enrolled in a private language institute in Tehran, Iran." (p. 3) 2) "Participants were randomly assigned to one of two groups: an experimental group (n=30), which used ChatGPT for pronunciation practice, and a waiting control group (n=30)" (p. 3) Detailed Analysis: Criterion S requires randomization at the level of schools or equivalent institutional units. In this study all 60 participants came from a single private language institute and were individually randomized to conditions. Only one institution was involved, and no cluster or site-level assignment occurred. While the tutoring exception can excuse the class-level criterion C, criterion S explicitly requires school-level assignment, which is absent here. Verified against the source PDF: all quotes above are verbatim and correctly cited. Criterion S is not met because randomization occurred at the individual student level within one institution, not at the school level.
    • I

      Independent Conduct

      • The authors designed the intervention and also collected and analyzed the data themselves, with no independent third-party evaluation team, so independence is not established despite blinded raters.
      • "IX : Data analysis, Writing- editing, Supervision: SS: Data collection, Methodology, Writing- original draft; MZ: Writing, Review and editing." (p. 10)
      • Relevant Quotes: 1) "IX : Data analysis, Writing- editing, Supervision: SS: Data collection, Methodology, Writing- original draft; MZ: Writing, Review and editing." (p. 10, Author contributions) 2) "Before the implementation of the interventions, participants in the experimental (ChatGPT) group received a one-hour training session on using ChatGPT 4.0 for pronunciation practice." (p. 4) 3) "Three experienced language teachers independently rated the recorded pronunciations ... The raters were blind to group assignments and were instructed not to inquire about or infer treatment conditions from the recordings." (p. 5) Detailed Analysis: Criterion I requires that the study be conducted independently of those who designed the intervention, e.g. by an external evaluation team. Here the author contribution statement shows the same research team designed the instructional procedure (methodology), collected the data, and analyzed the results; there is no external evaluator, agency, or third-party oversight mentioned anywhere in the paper, and funding is "Not applicable." The use of three blind raters for scoring reduces measurement bias but the raters were engaged by the research team within their own study; this does not constitute independent conduct of the trial in the sense required (external design-independent evaluation). Although the underlying tool (ChatGPT) was built by a third party, the instructional intervention tested (the structured prompt, schedule, and practice protocol) was designed and evaluated by the same authors. Verified against the source PDF: quote 1 has been corrected to include a space before the colon ("IX :") to match the exact typesetting of the author contributions statement; quotes 2 and 3 are verbatim and correctly cited. Criterion I is not met because the same team designed, implemented, and analyzed the intervention with no independent third-party evaluation.
    • Y

      Year Duration

      • The full study window from intervention start to the delayed post-test was about six weeks, far below 75% of an academic year; also criterion T is not met, which precludes Y.
      • "At the end of the three-week intervention, a post-test was administered in the fourth week. ... A delayed post-test was conducted two weeks after the post-test" (p. 5)
      • Relevant Quotes: 1) "The intervention lasted three weeks, during which both groups practiced 10 target words each week in two self-directed sessions at home" (p. 5) 2) "At the end of the three-week intervention, a post-test was administered in the fourth week. Participants read aloud a new set of 30 sentences containing the target words" (p. 5) 3) "A delayed post-test was conducted two weeks after the post-test to measure pronunciation retention." (p. 5) Detailed Analysis: Criterion Y requires that outcomes be measured at least 75% of a full academic year (roughly 9-10 months) after the intervention begins. The entire study, from the start of the three-week intervention to the delayed post-test, spanned approximately six weeks. This is an order of magnitude shorter than the required duration. Additionally, per the instructions, since criterion T (Term Duration) is not met, criterion Y cannot be met. Verified against the source PDF: all quotes above are verbatim and correctly cited. Criterion Y is not met because tracking lasted only about six weeks rather than at least 75% of an academic year.
    • B

      Balanced Control Group

      • The control was an active condition with the same word lists, practice schedule, and logging, and the AI tool contrast (plus its brief orientation) is the integral treatment variable being tested against dictionary-based practice.
      • "The intervention lasted three weeks, during which both groups practiced 10 target words each week in two self-directed sessions at home, following their weekly on-site classes." (p. 5)
      • Relevant Quotes: 1) "The intervention lasted three weeks, during which both groups practiced 10 target words each week in two self-directed sessions at home, following their weekly on-site classes." (p. 5) 2) "To maintain consistency, all participants were provided with a printed schedule and log sheet to record their practice sessions." (p. 5) 3) "Before the implementation of the interventions, participants in the experimental (ChatGPT) group received a one-hour training session on using ChatGPT 4.0 for pronunciation practice. ... The control group did not receive any training, as they used familiar tools (electronic dictionaries), ensuring equitable initial conditions." (p. 4) 4) "This study aims to examine the effectiveness of ChatGPT in supporting pronunciation practice for EFL learners and to compare it with the use of electronic dictionaries as a traditional, non-interactive pronunciation tool." (p. 3) 5) "In the experimental group, participants submitted screenshots of their ChatGPT sessions weekly as evidence of engagement. In the control group, participants noted the dictionary entries used and logged their repetitions." (p. 5) 6) "While full control over home environments was not feasible, adherence and exposure were monitored through the submitted logs and follow-up reminders, which minimized discrepancies in practice intensity." (p. 5) Detailed Analysis: Applying the criterion B decision tree: extra resources are present (the experimental group received access to ChatGPT 4.0 and a one-hour orientation session that the control group did not receive). However, these resources are the explicit treatment variable: the study's stated aim is "to examine the effectiveness of ChatGPT in supporting pronunciation practice ... and to compare it with the use of electronic dictionaries," i.e., the contrast between the interactive AI tool and the static dictionary tool is precisely what is being tested, not a separable add-on to a shared core intervention. Both groups otherwise received matched "business as usual" style inputs: the same 10 target words per week, the same two self-directed home sessions per week, the same three-week schedule, and the same logging/monitoring apparatus, so time-on-task was designed to be equivalent. The one-hour ChatGPT orientation is a small onboarding component needed to use the treatment tool itself, and the authors explicitly justify omitting training for the control group because dictionaries were already familiar, "ensuring equitable initial conditions"; this difference is also negligible relative to six weeks of practice. Verified against the source PDF: all quotes above are verbatim and correctly cited. Criterion B is met because the additional resource (ChatGPT access, plus its brief orientation) is the integral treatment variable being tested against an active, resource-matched dictionary-based control, per the criterion B decision tree.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this 2025 study by a different research team was found in the paper or in an internet literature search conducted during verification.
      • Relevant Quotes: 1) "Despite the growing body of research supporting the role of AI tools in language learning [20], the application of AI-driven interventions to enhance EFL pronunciation skills remains underexplored, particularly in controlled experimental settings." (p. 2) 2) "Importantly, this study contributes to the growing literature on AI in language learning by offering controlled experimental evidence of its effectiveness for pronunciation instruction." (p. 8) 3) "While this study contrasted ChatGPT with electronic dictionaries, future research could compare ChatGPT with other AI-powered pronunciation tools" (p. 8) Detailed Analysis: Criterion R requires independent replication of the study by a different team, in a different context, published in a peer-reviewed journal. The paper itself frames the topic as "underexplored" and positions this trial as novel controlled evidence, citing related but methodologically different work (e.g., Fathi et al. on speaking skills, Huang on ChatGPT voice feedback) rather than replications of this design. A dedicated internet search for citing and related work (via Google Scholar and general web search) was carried out during this verification. It identified two papers that cite this study: Y. Shuzhen (2025), "A Study on the Transmission of Cultural-Loaded Words in Cross-Cultural Communication from the Perspective of Translation" (Journal of Sociology and Education), which mentions this paper only in passing and is about translation, not pronunciation RCTs; and A. Kosimov and D. Murotova (2025), "Examining the Effects of AI-Mediated Informal Language Learning on EFL Learners' Academic Writing Performance" (Applied Linguistics Compass), a quasi-experimental study on writing performance, not a replication of this ChatGPT-versus-electronic-dictionary pronunciation trial. Neither paper is an independent reproduction of this specific study's design, sample, or outcome measure. The paper was first published online on 23 August 2025, so an independent peer-reviewed replication is unlikely to exist yet. Criterion R is not met because no independent peer-reviewed replication of this specific study was found in the paper or via internet search.
    • A

      All-subject Exams

      • Only pronunciation accuracy in English was measured, with a custom instrument, so neither the standardized-exam prerequisite nor all-subject coverage is satisfied.
      • "Pronunciation performance was assessed over three phases: pre-test, post-test, and delayed post-test." (p. 1)
      • Relevant Quotes: 1) "Pronunciation performance was assessed over three phases: pre-test, post-test, and delayed post-test." (p. 1, Abstract) 2) "The average of the three raters' scores was calculated for each participant, yielding a pronunciation accuracy score out of 30 for each testing session." (p. 5) Detailed Analysis: Criterion A requires standardized exam-based assessment across all main subjects taught at the educational level, and it presupposes criterion E. Criterion E is not met (a custom word-reading task was used), so criterion A automatically fails. Moreover, the study measured a single narrow outcome, pronunciation accuracy of 30 English words; no other language skills, let alone other subjects, were assessed. The setting is a private adult language institute, where a specialized focus might arguably justify an exception, but even then the prerequisite of a standardized exam is unmet. Verified against the source PDF: all quotes above are verbatim and correctly cited. Criterion A is not met because criterion E fails and only a single custom pronunciation measure was used.
    • G

      Graduation Tracking

      • Measurement ended at a delayed post-test two weeks after the intervention, with no tracking of participants to the end of their course of study, and criterion Y is not met, which precludes G.
      • "A delayed post-test was conducted two weeks after the post-test to measure pronunciation retention." (p. 5)
      • Relevant Quotes: 1) "A delayed post-test was conducted two weeks after the post-test to measure pronunciation retention." (p. 5) 2) "Expanding this line of inquiry to include learners at different proficiency levels and collecting qualitative feedback on learner experiences with AI tools like ChatGPT could offer a richer understanding" (p. 9) Detailed Analysis: Criterion G requires following participants until graduation from their educational stage. The study's final data point is a delayed post-test two weeks after the post-test, roughly six weeks after the intervention began. There is no mention of tracking learners to completion of their language program or any graduation milestone, and the future directions section proposes new studies rather than continued follow-up of this cohort. A search for follow-up publications by the same authors (Xodabande, Shiri, Zohrabi) tracking this cohort further was conducted via Google Scholar and general web search during this verification; no such follow-up papers were found. Additionally, per the instructions, criterion Y is not met, which precludes criterion G. Verified against the source PDF: all quotes above are verbatim and correctly cited. Criterion G is not met because tracking stopped at a two-week delayed post-test with no graduation follow-up, no follow-up publications were found, and criterion Y is not met.
    • P

      Pre-Registered

      • The paper mentions ethics approval but contains no pre-registration statement, registry name, ID, or registration date, and none was found via registry search.
      • Relevant Quotes: 1) "All data collection procedures were conducted in agreement with the ethical standards of the institutional research committee and the 1964 Helsinki Declaration and were approved by the Ethics Committee of Islamic Azad University (No. 2024-116)" (p. 10) Detailed Analysis: Criterion P requires that the full study protocol be pre-registered on a public registry before data collection began, with a verifiable link or ID and date. The paper reports ethics committee approval, but ethics approval is not pre-registration. No registry platform (e.g., ClinicalTrials.gov, OSF, AsPredicted, IRCT), registration number, or registration date is mentioned anywhere in the manuscript. During this verification, a targeted internet search was carried out for a pre-registration record under the authors' names and the study's topic (e.g., OSF Registries, AsPredicted); no registered protocol matching this trial was found. Without quoted evidence of pre-registration prior to data collection, the criterion fails. Verified against the source PDF: the quote above is verbatim and correctly cited. Criterion P is not met because no pre-registration of the study protocol is reported in the paper or discoverable via registry search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.