Utilizing Automatic Speech Recognition for English Pronunciation Practice and Analyzing its Impact

Katsuyuki Umezawa, Makoto Nakazawa, Michiko Nakano, Shigeichi Hirasawa

Published:
ERCT Check Date:
DOI: 10.1109/ICEED59801.2023.10264049
  • L2 languages
  • adult education
  • Asia
  • EdTech website
  • digital assessment
  • formative assessment
  • mobile learning
0
  • C

    Randomisation was performed at the individual participant level, not by class or school, and no tutoring-style exception applies to this self-study intervention.

    "The experiment consisted of 28 participants, randomly divided into groups of seven individuals each." (Section IV.B)

  • E

    Outcomes were judged by a single native-speaker evaluator using a custom 12-item word/sentence checklist devised by the authors, not a standardised exam.

    "Figure 1 shows the list of words and sentences used in the experiment. This list was created with reference to previous research [8] [1], focusing on minimal pairs and those with 'l' and 'r' that are difficult for Japanese to pronounce." (Section IV.A)

  • T

    The entire intervention and outcome measurement spanned only seven days, far short of one academic term.

    "This process is performed for seven days." (Section IV.B)

  • D

    The paper reports only group size and raw pronunciation scores for the no-intervention group, without demographic or baseline proficiency documentation.

    "In addition, the seventh participants (D1 to D7) practice pronunciation without using the speech recognition function." (Section IV.B)

  • S

    There is no school- or institution-level randomisation; participants were individually and randomly assigned to groups.

    "The experiment consisted of 28 participants, randomly divided into groups of seven individuals each." (Section IV.B)

  • I

    The same author team appears to have designed, run, and evaluated the study, with only a single unspecified native-speaker evaluator and no stated independence or blinding.

    "Participants recorded their pronunciation on the first and last days of the experiment and asked a native English speaker (hereafter referred to as the evaluator) to judge whether the pronunciation was correct." (Section IV.B)

  • Y

    Since the term-duration criterion (T) is not met, the stronger year-duration criterion cannot be met either; the study lasted only seven days.

    "This process is performed for seven days." (Section IV.B)

  • B

    The differing use of speech-recognition checking is the explicit treatment variable being compared; all groups otherwise received the same 1-2 hours of daily practice time and shared pronunciation reference tool.

    "If you don't know the pronunciation of a word, use the Google Text-To-Speech function." / "Pronunciation practice (1-2 hours)." (Fig. 2, Section IV.B)

  • R

    No independent replication of this specific study by a different research team was found in the paper or in external searches.

  • A

    Since criterion E (Exam-based Assessment) is not met, criterion A is automatically not met; only English pronunciation was assessed in any case.

  • G

    Since criterion Y (Year Duration) is not met, the graduation-tracking criterion is automatically not met, and no follow-up tracking beyond Day 7 is described or found in later publications.

  • P

    The paper contains no reference to a pre-registered protocol, registry platform, or registration date, and no such pre-registration was found via internet search.

Abstract

The advancement of AI in recent years has been remarkable, along with the widespread use of speech recognition functions. In addition, an increasing number of people are self-studying in their fields of interest. In a world considered a global society, many people in countries where English is not their first language are learning English. Therefore, in this study, the authors focus on self-study English learning. The pronunciation is assumed to be correct if the speech recognition feature correctly identifies the pronounced English words and sentences. It has been recognized that feedback plays a crucial role in English pronunciation practice. If the speech recognition function can be employed to provide feedback, it would alleviate the burden on teachers and enable students to practice pronunciation independently. In this research, the authors propose practicing pronunciation using each browser's built-in speech recognition features, including Google Chrome, iOS, and Microsoft Edge, and evaluating if one's English pronunciation is correctly recognized. Through a 7-day experiment, the authors clarify the speech recognition function suitable for self-studying English pronunciation. Through this experiment, it was observed that the speech recognition function on iOS (version 15.0) outperforms compared to Google Chrome (version 107.0.5304.63) and Microsoft Edge (version 107.0.1418.24) in accurately understanding speech, even when the pronunciation is incorrect. However, it became apparent that such highly accurate speech recognition capabilities may not be suitable for self-study pronunciation practice.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was performed at the individual participant level, not by class or school, and no tutoring-style exception applies to this self-study intervention.
      • "The experiment consisted of 28 participants, randomly divided into groups of seven individuals each." (Section IV.B)
      • Relevant Quotes: 1) "Specifically, 28 participants were randomly assigned to four groups and engaged in a seven-day randomized controlled trial for English pronunciation practice." (Section III.A) 2) "The experiment consisted of 28 participants, randomly divided into groups of seven individuals each." (Section IV.B) 3) "Seven participants (A1 to A7) practiced pronunciation using Google's speech recognition function, seven participants (B1 to B7) practiced pronunciation using iOS's function, and seven participants (C1 to C7) practiced pronunciation using Microsoft Edge's function. In addition, the seventh participants (D1 to D7) practice pronunciation without using the speech recognition function." (Section IV.B) Detailed Analysis: The unit of randomisation described in every quoted passage is the individual participant, not a class or school. The 28 participants were simply "randomly divided into groups of seven," with no mention of intact classes or institutions being assigned as a block. The ERCT exception for personal teaching/tutoring interventions does not clearly apply here: this is a self-study exercise using browser-based speech recognition tools, performed independently by each participant, not a one-to-one tutoring relationship with an instructor. There is no statement in the paper invoking a tutoring rationale for individual-level assignment. Since randomisation occurred at the student (individual) level without a clearly applicable exception, criterion C is not met.
    • E

      Exam-based Assessment

      • Outcomes were judged by a single native-speaker evaluator using a custom 12-item word/sentence checklist devised by the authors, not a standardised exam.
      • "Figure 1 shows the list of words and sentences used in the experiment. This list was created with reference to previous research [8] [1], focusing on minimal pairs and those with 'l' and 'r' that are difficult for Japanese to pronounce." (Section IV.A)
      • Relevant Quotes: 1) "Figure 1 shows the list of words and sentences used in the experiment. This list was created with reference to previous research [8] [1], focusing on minimal pairs and those with 'l' and 'r' that are difficult for Japanese to pronounce." (Section IV.A) 2) "Participants recorded their pronunciation on the first and last days of the experiment and asked a native English speaker (hereafter referred to as the evaluator) to judge whether the pronunciation was correct." (Section IV.B) 3) "In addition, the criteria for this evaluation are similar to the conventional study [9], and judge whether it sounds like native English rather than Japanese English." (Section V.A) Detailed Analysis: The outcome measure is a bespoke 12-word/sentence checklist assembled by the authors from earlier internal research, judged subjectively by a single native-English-speaker evaluator against informal "sounds native" criteria. This is neither a standardised, widely recognised test nor a validated psychometric instrument; it is a researcher-designed measure created specifically for this and prior studies by the same group. Because the assessment tool is custom-built rather than a recognised standardised exam, criterion E is not met.
    • T

      Term Duration

      • The entire intervention and outcome measurement spanned only seven days, far short of one academic term.
      • "This process is performed for seven days." (Section IV.B)
      • Relevant Quotes: 1) "Specifically, 28 participants were randomly assigned to four groups and engaged in a seven-day randomized controlled trial for English pronunciation practice." (Section III.A) 2) "This process is performed for seven days." (Section IV.B) 3) "Participants recorded their pronunciation on the first and last days of the experiment and asked a native English speaker (hereafter referred to as the evaluator) to judge whether the pronunciation was correct." (Section IV.B) Detailed Analysis: The intervention began on Day 1 and outcomes were measured again on Day 7, a total interval of one week. This is dramatically shorter than the approximately 3-4 month minimum required for a term. There is no indication of any longer-term follow-up beyond the seventh day. Since the tracked interval from intervention start to final measurement is only seven days, criterion T is not met.
    • D

      Documented Control Group

      • The paper reports only group size and raw pronunciation scores for the no-intervention group, without demographic or baseline proficiency documentation.
      • "In addition, the seventh participants (D1 to D7) practice pronunciation without using the speech recognition function." (Section IV.B)
      • Relevant Quotes: 1) "The experiment consisted of 28 participants, randomly divided into groups of seven individuals each." (Section IV.B) 2) "In addition, the seventh participants (D1 to D7) practice pronunciation without using the speech recognition function." (Section IV.B) 3) Table I lists first-day and last-day correct answer counts for participants D1-D7 (average 6.3 first day, 7.4 last day), with no accompanying demographic, English-proficiency, or background information for this group. (Section V.A, Table I) Detailed Analysis: Group D (no speech-recognition function) functions as the closest analogue to a control group. Beyond its size (seven participants) and its raw pronunciation scores on Day 1 and Day 7, the paper gives no demographic information (age, gender, English proficiency level, native-language background) nor any description of what, if anything, these participants received besides the shared self-study protocol. Sparse numeric scores alone do not constitute the detailed documentation the criterion requires. Because no demographic or baseline characteristic documentation is provided for the comparison group, criterion D is not met.
  • Level 2 Criteria

    • S

      School-level RCT

      • There is no school- or institution-level randomisation; participants were individually and randomly assigned to groups.
      • "The experiment consisted of 28 participants, randomly divided into groups of seven individuals each." (Section IV.B)
      • Relevant Quotes: 1) "Specifically, 28 participants were randomly assigned to four groups and engaged in a seven-day randomized controlled trial for English pronunciation practice." (Section III.A) 2) "The experiment consisted of 28 participants, randomly divided into groups of seven individuals each." (Section IV.B) Detailed Analysis: No schools, institutions, or other higher-level units are mentioned as the randomisation unit; individual self-study participants were assigned directly to one of four groups. There is no evidence of school-level recruitment or assignment. Since randomisation was at the individual level with no school-level unit involved, criterion S is not met.
    • I

      Independent Conduct

      • The same author team appears to have designed, run, and evaluated the study, with only a single unspecified native-speaker evaluator and no stated independence or blinding.
      • "Participants recorded their pronunciation on the first and last days of the experiment and asked a native English speaker (hereafter referred to as the evaluator) to judge whether the pronunciation was correct." (Section IV.B)
      • Relevant Quotes: 1) "In this research, the authors propose practicing pronunciation using each browser's built-in speech recognition features, including Google Chrome, iOS, and Microsoft Edge, and evaluating if one's English pronunciation is correctly recognized." (Abstract) 2) "Participants recorded their pronunciation on the first and last days of the experiment and asked a native English speaker (hereafter referred to as the evaluator) to judge whether the pronunciation was correct." (Section IV.B) 3) "All experiments were approved by the Research Ethics Committee of Shonan Institute of Technology." (Research Ethics section) Detailed Analysis: The paper gives no indication that the study was designed, conducted, or analysed by anyone other than the four listed authors. A single "native English speaker" evaluator is used to judge pronunciation, but the paper does not describe this evaluator's relationship to the research team, nor does it state that the evaluator was blinded to group assignment or otherwise independent of the intervention designers. Institutional ethics approval concerns participant protection, not independent conduct of the evaluation. Because there is no statement of independent, third-party conduct of data collection or analysis, criterion I is not met.
    • Y

      Year Duration

      • Since the term-duration criterion (T) is not met, the stronger year-duration criterion cannot be met either; the study lasted only seven days.
      • "This process is performed for seven days." (Section IV.B)
      • Relevant Quotes: 1) "This process is performed for seven days." (Section IV.B) Detailed Analysis: Per the ERCT specification, if criterion T (Term Duration) is not met, criterion Y (Year Duration) is automatically not met. The entire study, including both intervention and outcome measurement, spanned only seven days, nowhere near 75% of an academic year. Criterion Y is not met.
    • B

      Balanced Control Group

      • The differing use of speech-recognition checking is the explicit treatment variable being compared; all groups otherwise received the same 1-2 hours of daily practice time and shared pronunciation reference tool.
      • "If you don't know the pronunciation of a word, use the Google Text-To-Speech function." / "Pronunciation practice (1-2 hours)." (Fig. 2, Section IV.B)
      • Relevant Quotes: 1) "At the beginning of each day's study, group participants using the speech recognition function (Groups A, B, and C) must use the function to check all the checklists. Pronunciation practice with the speech recognition function entails practicing pronunciation until the speech recognition function identifies the words and sentences on the checklist correctly." (Section IV.B) 2) "In addition, the seventh participants (D1 to D7) practice pronunciation without using the speech recognition function." (Section IV.B) 3) "It will be also investigated whether this learning method is effective even for self-study without an evaluator and the difference in the effect of learning using and not using the speech recognition function." (Section III.A) 4) Figure 2 (Experiment flow) shows Groups A, B, and C each "Check own pronunciation" with their respective speech recognition tool, while Group D is "Not checking own pronunciation"; all four groups then share the steps "If you don't know the pronunciation of a word, use the Google Text-To-Speech function" and "Pronunciation practice (1-2 hours)," followed by "Record own pronunciation" and "Evaluation by native English speaker." (Section IV.B, Fig. 2) Detailed Analysis: Applying the ERCT criterion B decision procedure: extra resources are present, since Groups A, B, and C receive an additional daily step (checking pronunciation with a speech-recognition function until it is recognised correctly) that Group D does not receive. However, this differing resource is explicitly the primary treatment variable under investigation: the paper states the study aims to determine "the difference in the effect of learning using and not using the speech recognition function," making the presence or absence of the checking step, and which tool is used, the very object of the experiment rather than a confounding add-on. Aside from this checking step, all four groups follow an identical daily protocol: access to the same Google Text-to-Speech pronunciation reference if unsure, the same 1-2 hours of practice time, and the same recording/evaluation procedure on Day 1 and Day 7. Group D thus represents a "business-as-usual" self-study baseline receiving the same core inputs minus the treatment variable itself. Because the differing resource (speech-recognition checking) is clearly framed as the central treatment variable being tested, and the remaining study inputs (reference tool, practice time, evaluation) are held constant across groups, criterion B is met.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study by a different research team was found in the paper or in external searches.
      • Relevant Quotes: No quotes in the paper reference a prior or subsequent independent replication of this specific 7-day, 4-group ASR pronunciation-checking experiment; the references section [1]-[9] consists of the authors' own prior work and general background literature on ASR and pronunciation, not independent replications of this study's design. Detailed Analysis: An internet search (multiple queries covering the paper's title, authors, and DOI 10.1109/ICEED59801.2023.10264049) did not identify any independent replication of this specific study by a different research team in a peer-reviewed outlet. The search surfaced other, unrelated speech-recognition-assisted pronunciation studies (e.g., "I Can Speak: improving English pronunciation through automatic speech recognition-based language learning systems," Innovation in Language Learning and Teaching, 2024, a different author team), but none of these reproduce this study's four-arm, seven-day design comparing Chrome, iOS, and Edge speech recognition against a no-recognition control; they are independent, differently-designed studies on a related topic, not replications of this paper. Since no independent replication of this specific study was found, criterion R is not met.
    • A

      All-subject Exams

      • Since criterion E (Exam-based Assessment) is not met, criterion A is automatically not met; only English pronunciation was assessed in any case.
      • Relevant Quotes: 1) "Figure 1 shows the list of words and sentences used in the experiment." (Section IV.A) Detailed Analysis: Per the ERCT specification, criterion A requires criterion E to be met as a prerequisite. Since criterion E is not met (a custom, non-standardised checklist was used), criterion A cannot be met. Independently, the study only assessed English pronunciation of 12 words/sentences and did not measure any other subject area. Criterion A is not met.
    • G

      Graduation Tracking

      • Since criterion Y (Year Duration) is not met, the graduation-tracking criterion is automatically not met, and no follow-up tracking beyond Day 7 is described or found in later publications.
      • Relevant Quotes: 1) "This process is performed for seven days." (Section IV.B) 2) No statement in the paper describes any tracking of participants beyond the Day 7 recording and evaluation, nor any planned or published follow-up study of the same cohort. Detailed Analysis: Per the ERCT specification, if criterion Y (Year Duration) is not met, criterion G is automatically not met. Consistent with this, the paper describes no follow-up beyond the seven-day experiment and no graduation tracking of any kind; participants are adult self-study volunteers, not a school cohort being followed to graduation. An internet search for later papers by Umezawa, Nakazawa, Nakano, and Hirasawa did not surface any subsequent publication tracking this same cohort or extending this experiment's follow-up period. Criterion G is not met.
    • P

      Pre-Registered

      • The paper contains no reference to a pre-registered protocol, registry platform, or registration date, and no such pre-registration was found via internet search.
      • Relevant Quotes: No quotes in the paper mention a study registry (e.g., a clinical-trials-style registry, OSF, or similar), a registration ID, or a registration date. The only procedural approvals mentioned are: "All experiments were approved by the Research Ethics Committee of Shonan Institute of Technology. The authors also received written informed consent from the participants and parents or guardians." (Research Ethics section) Detailed Analysis: Ethics committee approval and informed consent are distinct from pre-registration of the study protocol (hypotheses, methods, and planned analyses) on a public registry before data collection began. No such pre-registration is mentioned anywhere in the paper, and an internet search for a pre-registration record by these authors for this study did not locate any registry entry (e.g., on OSF or a comparable platform). Since no evidence of a pre-registered protocol is present, criterion P is not met.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.