The impact of automatic speech recognition technology on second language pronunciation and speaking skills of EFL learners: a mixed methods investigation

Weina Sun

Published:
ERCT Check Date:
DOI: 10.3389/fpsyg.2023.1210187
  • L2 languages
  • adult education
  • China
  • EdTech website
  • formative assessment
0
  • C

    Two intact classes were randomly assigned to the experimental and control conditions, satisfying class-level randomisation.

    "The participants were enrolled in a 14-week English pronunciation course, with two classes randomly assigned to two different research conditions." (p. 4)

  • E

    Pronunciation outcomes were measured with a researcher-built read-aloud task and in-house ratings, and the IELTS-format speaking assessment was administered unofficially by the research team, so no recognised standardised exam was used as an official outcome measure.

    "The reading activity involved seven sentences that aimed to assess seven English phonemes that are known to pose challenges for Chinese speakers (Zhang and Yin, 2009)." (p. 6)

  • T

    The intervention and pre/post measurement spanned a 14-week course, which is approximately one academic term.

    "The participants were enrolled in a 14-week English pronunciation course, with two classes randomly assigned to two different research conditions." (p. 4)

  • D

    The control group's size, demographics, baseline scores, and the teacher-led instruction it received are well documented.

    "One class consisting of 32 students served as the control group and received traditional teacher-led feedback and instruction" (p. 4)

  • S

    Randomisation was of two classes within one language training center, not of schools or institutions.

    "The study was conducted at a language training center in Shenzhen, Guangdong province, China, with the approval of the center's authorities and informed consent of the participants." (p. 4)

  • I

    The same researcher designed, oversaw, rated, and analysed the study with no independent third-party evaluation.

    "The researcher worked closely with the instructor to design the intervention protocol, develop the assessment measures, and ensure consistency across groups." (p. 5)

  • Y

    The study spanned only a 14-week course, far less than 75% of an academic year.

    "Secondly, the study was constrained by a relatively brief intervention period, which may have influenced the extent of the observed improvements." (p. 11)

  • B

    Both groups received matched lesson time and structure with an active teacher-feedback control, the ASR tool being the integral treatment variable.

    "To ensure consistency and comparability across all groups, team members in both the EG and CG engaged in pronunciation-based discussions during the practice stage." (p. 5)

  • R

    No independent replication of this specific study is reported; the paper frames itself as filling a research gap, and a targeted internet search found no such replication.

  • A

    Only L2 pronunciation and speaking were assessed; no other subjects or skills were measured.

    "Data collection involved read-aloud tasks, spontaneous conversations, and IELTS speaking tests to evaluate L2 pronunciation and speaking skills." (p. 1)

  • G

    Measurement ended at the close of the 14-week course with no follow-up tracking until graduation, and no such follow-up publication was found via internet search.

    "To gain a more comprehensive understanding of the sustainability and long-term effects of incorporating ASR technology with peer correction, longitudinal investigations should be pursued." (p. 11)

  • P

    The paper mentions only ethics approval and contains no reference to a pre-registered protocol or registry ID; no registration record was found via internet search.

Abstract

This study employed an explanatory sequential design to examine the impact of utilizing automatic speech recognition technology (ASR) with peer correction on the improvement of second language (L2) pronunciation and speaking skills among English as a Foreign Language (EFL) learners. A total of 61 intermediate-level Chinese EFL learners were randomly assigned to either a control group (CG) or an experimental group (EG). The CG received conventional teacher-led feedback and instruction, while the EG used ASR technology with peer correction. Data collection involved read-aloud tasks, spontaneous conversations, and IELTS speaking tests to evaluate L2 pronunciation and speaking skills. Additionally, semi-structured interviews were conducted with a subset of the participants. The quantitative analysis demonstrated that the EG outperformed the CG in all measures of L2 pronunciation, including accentedness and comprehensibility, and exhibited significant improvements in global speaking skill compared to the CG.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Two intact classes were randomly assigned to the experimental and control conditions, satisfying class-level randomisation.
      • "The participants were enrolled in a 14-week English pronunciation course, with two classes randomly assigned to two different research conditions." (p. 4)
      • Relevant Quotes: 1) "A total of 61 intermediate-level Chinese EFL learners were randomly assigned to either a control group (CG) or an experimental group (EG)." (p. 1, Abstract) 2) "The participants were enrolled in a 14-week English pronunciation course, with two classes randomly assigned to two different research conditions. One class consisting of 32 students served as the control group and received traditional teacher-led feedback and instruction, while the other class consisting of 29 students was the experimental group and used automatic speech recognition technology (ASR) with peer correction." (p. 4) Detailed Analysis: The abstract loosely describes learners as "randomly assigned," but the more detailed Methods section clarifies that the unit of assignment was the class: "two classes randomly assigned to two different research conditions." Thus randomisation was carried out at the class level, with one intact class (n=32) allocated to the control condition and the other intact class (n=29) to the experimental condition. This satisfies the ERCT requirement that entire classes, not individual students within a single class, be randomised, which prevents within-class contamination between treatment and control students. The design is weak in that only two classes were randomised (one per condition), so class and condition are confounded, but the criterion as written requires class-level (or stronger) random assignment, which the paper explicitly states was done. Baseline equivalence was additionally checked (t(59)=-0.85, p=0.399 for English proficiency). Criterion C is met because the paper explicitly states that two intact classes were randomly assigned to the experimental and control conditions, i.e., randomisation at the class level.
    • E

      Exam-based Assessment

      • Pronunciation outcomes were measured with a researcher-built read-aloud task and in-house ratings, and the IELTS-format speaking assessment was administered unofficially by the research team, so no recognised standardised exam was used as an official outcome measure.
      • "The reading activity involved seven sentences that aimed to assess seven English phonemes that are known to pose challenges for Chinese speakers (Zhang and Yin, 2009)." (p. 6)
      • Relevant Quotes: 1) "The initial tool used for evaluating the students' pronunciation was a read-aloud task, which has been employed in previous studies (e.g., Thomson, 2011). The reading activity involved seven sentences that aimed to assess seven English phonemes that are known to pose challenges for Chinese speakers (Zhang and Yin, 2009), as well as the accuracy of stress, juncture, and intonation within sentences." (p. 6) 2) "Throughout the task, two instructors evaluated the students' performance using a 9-point scale for accentedness (ranging from 1, indicating heavily accented, to 9, indicating native-like) and comprehensibility (ranging from 1, indicating very difficult to understand, to 9, indicating no effort required to understand)." (p. 6) 3) "To assess the participants' pronunciation in spontaneous conversation, a short conversation including three to four questions was used as the second instrument. The assessment was based on the IELTS speaking skill rubric, which is a widely recognized and standardized assessment tool." (p. 6) 4) "The IELTS Speaking Skill Test is a comprehensive assessment tool that evaluates learners' speaking ability in four areas: fluency and coherence, lexical resources, grammatical range and accuracy, and pronunciation." (p. 6) 5) "To guarantee impartiality and diminish partiality, every participant was evaluated and graded by two skilled and debriefed evaluators, one of whom is a researcher..." (p. 6) Detailed Analysis: Criterion E requires that outcomes be measured with standard, widely recognised, exam-based assessments rather than instruments specially assembled for the study. The primary pronunciation outcomes (accentedness and comprehensibility) were measured with a read-aloud task consisting of seven sentences constructed by the researchers specifically to target phonemes difficult for Chinese speakers, rated on 9-point scales by two instructors; this is a study-specific research instrument, not a standardised exam. The spontaneous conversation task was likewise a short researcher-designed conversation of three to four questions, merely scored using the IELTS speaking rubric. The "IELTS Speaking Skill Test" component borrows the format and band descriptors of a genuinely standardised, widely recognised exam, but it was administered and scored in-house by the study's own evaluators (one of whom was the researcher), not delivered as the official standardised examination by an accredited test centre. The bulk of the outcome battery is therefore custom built, with only an unofficial, researcher-administered approximation of IELTS providing any link to a standardised test. This does not satisfy the requirement to use a standard, widely recognised standardised exam as the outcome measure. Criterion E is not met because the outcome measures were researcher-constructed tasks and in-house ratings, with the IELTS-style test administered unofficially by the research team rather than as a recognised standardised examination.
    • T

      Term Duration

      • The intervention and pre/post measurement spanned a 14-week course, which is approximately one academic term.
      • "The participants were enrolled in a 14-week English pronunciation course, with two classes randomly assigned to two different research conditions." (p. 4)
      • Relevant Quotes: 1) "The participants were enrolled in a 14-week English pronunciation course, with two classes randomly assigned to two different research conditions." (p. 4) 2) "The timing of the interviews was strategically planned in reference to the intervention phase. Specifically, the interviews were conducted after the completion of the intervention to allow participants sufficient exposure and engagement with the ASR technology and peer correction activities." (p. 6) 3) "an analysis of covariance (ANCOVA) was performed with pretest scores as the covariate and posttest scores as the dependent variable" (p. 6) 4) "Secondly, the study was constrained by a relatively brief intervention period, which may have influenced the extent of the observed improvements." (p. 11) Detailed Analysis: The intervention was embedded in a 14-week English pronunciation course, with pretests administered before and posttests after the intervention. Fourteen weeks is approximately 3.2-3.5 months, which corresponds to a semester or one full academic term under the ERCT definition of a term as "approximately 3-4 months." The paper does not give exact calendar dates, but the pre/post design tied to the 14-week course indicates that the interval from intervention start to outcome measurement spanned the length of the course, i.e., at least one term. The author's own limitation note about a "relatively brief intervention period" refers to the absence of longer-term (year-long/longitudinal) tracking rather than contradicting the 14-week duration. Criterion T is met because outcomes were measured at the end of a 14-week (roughly one semester/term) intervention period, meeting the minimum one-term interval from intervention start to measurement.
    • D

      Documented Control Group

      • The control group's size, demographics, baseline scores, and the teacher-led instruction it received are well documented.
      • "One class consisting of 32 students served as the control group and received traditional teacher-led feedback and instruction" (p. 4)
      • Relevant Quotes: 1) "One class consisting of 32 students served as the control group and received traditional teacher-led feedback and instruction" (p. 4) 2) "All participants were native Chinese speakers and had no prior experience studying abroad. Their English proficiency level was intermediate (B1 level in the Common European Framework of Reference for Languages), as assessed by their English language scores on the college entrance examinations" (p. 4-5) 3) "An independent-samples t-test showed that no significant difference was found in English proficiency level between the control and experimental groups, t (59)=-0.85, p =0.399." (p. 5) 4) "Conversely, the CG participants received personalized pronunciation feedback from the teacher, who carefully analyzed their performances and provided constructive comments and guidance for improvement." (p. 5) 5) "The CG, on the other hand, individually focused on implementing the teacher's feedback, engaging in targeted pronunciation practice without direct peer support." (p. 5) 6) "Table 1 presents the descriptive statistics for the variables of the study. The table displays the means and standard deviations for each dependent variable for both the experimental and control groups." (p. 7) Detailed Analysis: The control group is documented in reasonable detail: its size (n=32), demographic profile (shared with the sample: age 20-31, native Chinese speakers, no study-abroad experience, B1 proficiency), baseline comparability verified via t-test on college entrance examination English scores, and baseline pretest means and standard deviations for all outcome variables in Table 1. Crucially, the paper also describes exactly what the control condition received: the same lesson structure (reading stage, feedback stage, practice stage, group discussions, and 40 minutes of general English tasks) but with teacher-led pronunciation feedback instead of ASR with peer correction. This allows a proper comparison between conditions. Criterion D is met because the control group's size, baseline characteristics, baseline scores, and the exact instruction it received are all clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was of two classes within one language training center, not of schools or institutions.
      • "The study was conducted at a language training center in Shenzhen, Guangdong province, China, with the approval of the center's authorities and informed consent of the participants." (p. 4)
      • Relevant Quotes: 1) "The study was conducted at a language training center in Shenzhen, Guangdong province, China, with the approval of the center's authorities and informed consent of the participants." (p. 4) 2) "The participants were enrolled in a 14-week English pronunciation course, with two classes randomly assigned to two different research conditions." (p. 4) Detailed Analysis: The entire study took place within a single language training center, and randomisation was applied to two classes inside that one institution. No schools, centers, or sites were randomised; the unit of randomisation was the class within one center. The ERCT S criterion requires random assignment of whole educational institutions/units (schools, centers, sites), which did not occur here. Criterion S is not met because randomisation occurred at the class level within a single language training center, not at the school/institution level.
    • I

      Independent Conduct

      • The same researcher designed, oversaw, rated, and analysed the study with no independent third-party evaluation.
      • "The researcher worked closely with the instructor to design the intervention protocol, develop the assessment measures, and ensure consistency across groups." (p. 5)
      • Relevant Quotes: 1) "An experienced professional male IELTS teacher, who collaborated closely with the researchers, served as the instructor responsible for delivering the intervention and providing feedback to the participants of both groups. The researcher's primary role was to monitor and oversee the intervention process, ensuring its adherence to the research design and objectives. The researcher worked closely with the instructor to design the intervention protocol, develop the assessment measures, and ensure consistency across groups." (p. 5) 2) "every participant was evaluated and graded by two skilled and debriefed evaluators, one of whom is a researcher" (p. 6) 3) "The author confirms being the sole contributor of this work and has approved it for publication." (p. 11, Author contributions) Detailed Analysis: The same researcher who designed the intervention protocol and developed the assessment measures also monitored and oversaw the intervention, and even served as one of the two evaluators who graded participants' speaking performance. There is no external or third-party evaluation team, no independent data collection agency, and no statement of independent oversight. The sole author designed, supervised, rated, and analysed the study. Although the commercial Speechnotes ASR software was developed by a third party, the ERCT I criterion concerns independence of study conduct (data collection, analysis, conclusions) from the intervention designers, which is absent here. Criterion I is not met because the researcher who designed the intervention also oversaw its delivery, served as one of the outcome evaluators, and conducted the analysis, with no independent third-party involvement.
    • Y

      Year Duration

      • The study spanned only a 14-week course, far less than 75% of an academic year.
      • "Secondly, the study was constrained by a relatively brief intervention period, which may have influenced the extent of the observed improvements." (p. 11)
      • Relevant Quotes: 1) "The participants were enrolled in a 14-week English pronunciation course, with two classes randomly assigned to two different research conditions." (p. 4) 2) "Secondly, the study was constrained by a relatively brief intervention period, which may have influenced the extent of the observed improvements. To gain a more comprehensive understanding of the sustainability and long-term effects of incorporating ASR technology with peer correction, longitudinal investigations should be pursued." (p. 11) Detailed Analysis: The Y criterion requires that outcomes be measured at least 75% of a full academic year (roughly 7-9+ months) after the intervention begins. Here the entire study, from intervention start to posttest, was contained within a 14-week (approximately 3.2-3.5 month) course, and the author explicitly identifies the brief intervention period and the need for longitudinal follow-up as limitations. Fourteen weeks is well below 75% of an academic year. Criterion Y is not met because the tracking interval was only about 14 weeks, far short of 75% of an academic year.
    • B

      Balanced Control Group

      • Both groups received matched lesson time and structure with an active teacher-feedback control, the ASR tool being the integral treatment variable.
      • "To ensure consistency and comparability across all groups, team members in both the EG and CG engaged in pronunciation-based discussions during the practice stage." (p. 5)
      • Relevant Quotes: 1) "The CG received conventional teacher-led feedback and instruction, while the EG used ASR technology with peer correction." (p. 1, Abstract) 2) "To initiate the intervention, the participants were divided into small groups consisting of three or four individuals in both the EG and control group (CG). Each group received a carefully selected text" (p. 5) 3) "Meanwhile, in the CG, the teacher directly provided pronunciation feedback to the PS, employing strategies such as modeling correct pronunciation, offering explicit explanations, and suggesting specific improvement techniques." (p. 5) 4) "To ensure consistency and comparability across all groups, team members in both the EG and CG engaged in pronunciation-based discussions during the practice stage." (p. 5) 5) "It is worth noting that the final 40min of each lesson were dedicated to general English tasks, unrelated to pronunciation, which were carefully designed to maintain consistency and control between the EG and CG." (p. 5) 6) "The intervention utilized the 'Speechnotes - Speech to Text' dictation ASR software, which was accessed by the experimental group (EG) through their individual laptops and a dedicated website." (p. 5) Detailed Analysis: Applying the current criterion B decision tree: the EG's extra resource relative to the CG is access to laptops and the Speechnotes ASR software during the reading/feedback stages. This resource is not a supplementary add-on; it is precisely the treatment variable under investigation, since the study's explicit research question compares "automatic speech recognition technology (ASR) with peer correction" against "traditional teacher-led feedback and instruction." Per the decision tree, when extra resources are integral to the design being tested (RESOURCES_ARE_TREATMENT = true), the criterion is met regardless of whether the control group receives the same technology, provided the control still receives a comparable "business-as-usual" substitute. Here the CG did receive an active substitute: personalized teacher-led pronunciation feedback (modeling, explicit explanation, improvement techniques), matched small-group structure, identical lesson length, the same texts, and identical 40-minute general English tasks each lesson. Time on task and non-treatment educational inputs were deliberately held consistent between groups, and only the treatment-defining resource (ASR access) differed. Criterion B is met because the ASR software/laptop access is the integral treatment variable explicitly being tested against teacher-led feedback, and the control group received a matched, active "business-as-usual" substitute with equivalent instructional time and structure.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study is reported; the paper frames itself as filling a research gap, and a targeted internet search found no such replication.
      • Relevant Quotes: 1) "This study fills a gap by exploring the viewpoints of intermediate-level Chinese EFL learners regarding their experience, contributing insights to the existing research." (p. 4) 2) "limited research has been conducted to evaluate the influence of ASR technology on FL learners' speaking abilities in educational contexts (Jiang et al., 2021; Inceoglu et al., 2023), especially in the Chinese EFL context" (p. 2) 3) "Investigators, for instance, may integrate ASR technologies with native speaker teaching or employ various forms of ASR technology... It also requires longer-term trials with a diverse range of participants." (p. 11) Detailed Analysis: The R criterion requires that this specific study (its central experimental claim, in a comparable design) be independently replicated by a different research team and published in a peer-reviewed journal. The paper positions itself as filling a gap, i.e., as novel rather than replicated. Related prior work on ASR with peer feedback exists (e.g., Evers and Chen, 2022; Dai and Wu, 2021; McCrocklin, 2019a), but those studies predate this 2023 paper and are conceptually related investigations, not replications of this specific ASR-with- peer-correction trial with Chinese intermediate EFL adults. No reference to any independent replication of this particular study is provided, and the paper itself calls for future comparable studies. Internet Search (verification): A targeted internet search for papers citing "Sun (2023)" Frontiers in Psychology 14:1210187, or describing an independent replication of this specific ASR-with-peer- correction, class-level RCT with intermediate Chinese adult EFL learners, was conducted. Search results returned only the original article itself (Frontiers, PMC, ResearchGate, PubMed), a later meta-analysis and a systematic literature review of ASR-in-EFL-pronunciation research (which include this study only as one of many primary studies reviewed, not as a replication), and unrelated ASR/education papers. No paper by a different research team was found that independently reproduces this specific study's design, sample population, and comparison in a different context. No such replication could be confirmed. Criterion R is not met because no independent, peer-reviewed replication of this specific study is reported or identified, including after a dedicated internet search.
    • A

      All-subject Exams

      • Only L2 pronunciation and speaking were assessed; no other subjects or skills were measured.
      • "Data collection involved read-aloud tasks, spontaneous conversations, and IELTS speaking tests to evaluate L2 pronunciation and speaking skills." (p. 1)
      • Relevant Quotes: 1) "Data collection involved read-aloud tasks, spontaneous conversations, and IELTS speaking tests to evaluate L2 pronunciation and speaking skills." (p. 1, Abstract) 2) "The IELTS Speaking Skill Test is a comprehensive assessment tool that evaluates learners' speaking ability in four areas: fluency and coherence, lexical resources, grammatical range and accuracy, and pronunciation." (p. 6) 3) "It is worth noting that the final 40min of each lesson were dedicated to general English tasks, unrelated to pronunciation, which were carefully designed to maintain consistency and control between the EG and CG. These tasks encompassed various language skills, such as reading comprehension, vocabulary acquisition, and grammar exercises." (p. 5) Detailed Analysis: All outcome measures assess a single domain: English pronunciation and speaking. No other subjects, and not even other English skills that were part of the course (reading comprehension, vocabulary, grammar), were measured with standardised exams. The A criterion requires assessment of all main subjects (or a clearly justified specialised exception). While this is an adult EFL pronunciation course at a private language center, where English is the focus, the paper offers no explicit rationale invoking a specialised-intervention exception, and even within English instruction only the speaking/pronunciation dimension was tested, leaving possible spillover effects on other skills unmeasured. Criterion A is not met because only pronunciation and speaking outcomes were assessed, without coverage of other main subjects or an explicitly justified exception.
    • G

      Graduation Tracking

      • Measurement ended at the close of the 14-week course with no follow-up tracking until graduation, and no such follow-up publication was found via internet search.
      • "To gain a more comprehensive understanding of the sustainability and long-term effects of incorporating ASR technology with peer correction, longitudinal investigations should be pursued." (p. 11)
      • Relevant Quotes: 1) "The timing of the interviews was strategically planned in reference to the intervention phase. Specifically, the interviews were conducted after the completion of the intervention" (p. 6) 2) "To gain a more comprehensive understanding of the sustainability and long-term effects of incorporating ASR technology with peer correction, longitudinal investigations should be pursued." (p. 11) Detailed Analysis: Measurement ended with the posttests and interviews at the completion of the 14-week intervention. There is no follow-up tracking of participants after the course, let alone until graduation from any educational stage; the author explicitly recommends future longitudinal investigations, confirming none were conducted here. Per the ranking instructions, criterion G cannot be met when criterion Y is not met, and Y is not met for this study, which is independently sufficient to fail G. Internet Search (verification): A search was conducted for subsequent publications by Weina Sun (School of Foreign Languages, Changchun Institute of Technology) that might track the same 61-participant cohort of Chinese EFL learners through to course/programme completion or graduation. No such follow-up publication tracking this cohort was found; search results returned only the original 2023 article and unrelated third-party ASR/EFL research. No follow-up papers reporting graduation tracking for this cohort could be identified or confirmed. Criterion G is not met because tracking stopped at the end of the 14-week intervention with no graduation follow-up, no follow-up publication was found after a dedicated search, and criterion Y (a prerequisite for G) is not met.
    • P

      Pre-Registered

      • The paper mentions only ethics approval and contains no reference to a pre-registered protocol or registry ID; no registration record was found via internet search.
      • Relevant Quotes: 1) "The studies involving human participants were reviewed and approved by School of Foreign Languages, Changchun Institute of Technology, Changchun. The patients/participants provided their written informed consent to participate in this study." (p. 11, Ethics statement) 2) "This study obtained ethical clearance from the relevant institutional review board before data collection commenced." (p. 5) Detailed Analysis: The paper reports institutional ethics approval and informed consent, but ethics approval is not pre-registration. There is no mention of any trial registry (e.g., ClinicalTrials.gov, OSF, AsPredicted, ISRCTN), no registration ID, no pre-registration date, and no pre-specified analysis plan referenced anywhere in the article. Without a documented pre-registered protocol dated before data collection, the P criterion cannot be satisfied. Internet Search (verification): The published Frontiers article page and full text were checked directly for any pre-registration or trial-registration statement; none was found beyond the institutional ethics clearance already quoted above. A further search for a registered protocol by Weina Sun matching this study's design (Chinese EFL learners, ASR with peer correction, pronunciation outcomes) on common registries did not return any matching pre-registration record. Criterion P is not met because no pre-registration of the study protocol on any registry is mentioned in the paper or found via internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.