The effectiveness of direct articulatory-abdominal pronunciation instruction for English learners in Hong Kong

Michael Yeldham and Vincent Choy

Published:
ERCT Check Date:
DOI: 10.1080/07908318.2021.1978476
  • L2 languages
  • higher education
  • China
  • Asia
0
  • C

    Randomisation was conducted at the individual student level rather than by class or school, and no tutoring exception applies, so criterion C is not met.

    "The study employed an experimental design, where 40 L1 Cantonese-speaking students in their first year at a Hong Kong community college were randomly assigned to either an experimental or a comparison group." (p. 187)

  • E

    The study used a custom reading-aloud task devised by the researchers for their own prior studies rather than a widely recognised standardised exam.

    "The reading aloud task, used pre-test and post-test to assess the learners' pronunciation of the relevant sounds, had been previously piloted and used in the research of other groups of L1 Cantonese EFL learners by Yeldham (2018, 2020, 2021)." (p. 191)

  • T

    The intervention lasted a single four-hour session with outcomes measured immediately afterward, far short of the one-term interval required.

    "Each learner group attended an intensive four-hour pronunciation course (which included short breaks after each hour)..." (p. 189)

  • D

    The comparison group's size, baseline proficiency, and exact instructional treatment are clearly documented.

    "The comparison group shared the same content and classroom activities with the experimental group. However, the critical difference between the two courses was that the teacher did not explain to the comparison group how to pronounce the language features when he modeled them for the class." (pp. 190-191)

  • S

    Randomisation was at the individual-student level within a single college, not at the school level.

    "The study employed an experimental design, where 40 L1 Cantonese-speaking students in their first year at a Hong Kong community college were randomly assigned to either an experimental or a comparison group." (p. 187)

  • I

    The study's own authors designed, delivered, and predominantly scored the intervention, with no independent third-party evaluator involved.

    "The second author taught both the experimental and comparison courses..." (p. 189)

  • Y

    Since criterion T is not met and the study tracked outcomes for only a few hours, criterion Y is not met.

    "A second limitation was the lack of a delayed post-test." (p. 196)

  • B

    Both groups received equal total time, content, and teacher contact, with only the direct/indirect instructional style (the treatment variable) differing.

    "Each learner group attended an intensive four-hour pronunciation course (which included short breaks after each hour), which was taught by the same teacher, and where both groups were taught the same course content and undertook the same tasks." (p. 189)

  • R

    No independent replication of this specific study by a different research team was found; prior related studies were conducted by the same first author, and a citation search located no external replications.

    "In conclusion, future research could replicate the study, while addressing the above limitations." (p. 196)

  • A

    Criterion A is not met because criterion E is not met and the assessment was narrowly limited to targeted pronunciation features.

  • G

    Criterion G is not met because criterion Y is not met and there was no follow-up tracking beyond the immediate post-test, confirmed by a search of the author's subsequent publications.

    "A second limitation was the lack of a delayed post-test." (p. 196)

  • P

    No mention of a pre-registered protocol, registry, or registration date appears anywhere in the paper, and none was found in an external search.

Abstract

The main purpose of this study was to examine the effectiveness for L2 English learners of a new direct approach to segmental pronunciation instruction that combined articulatory instruction with abdominal enhancement techniques. The participants were Cantonese speakers in Hong Kong, where the school curriculum relies chiefly on indirect instruction within a task-based language teaching (TBLT) framework. Thus a second purpose of the study was to examine whether the direct approach may be a useful addition to the Hong Kong curriculum. Randomly-assigned experimental and comparison groups of recent school graduates completed pronunciation tasks embedded within a TBLT framework. However, the experimental group had direct attention drawn to the segmental sounds, including advice and feedback on how to produce them, while the comparison group did not. Both groups completed a pretest/posttest reading-aloud task. The segments targeted in this test (and in the instruction) involved selected long vowel/diphthong sounds, voiced fricative consonants, and /t/ and /d/ in syllable-final consonant clusters. Results showed the experimental group significantly outperformed the comparison group overall and in each of these segmental categories, highlighting the importance of the direct articulatory-abdominal instruction. The results also suggested such instruction should be given greater attention in the Hong Kong curriculum.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was conducted at the individual student level rather than by class or school, and no tutoring exception applies, so criterion C is not met.
      • "The study employed an experimental design, where 40 L1 Cantonese-speaking students in their first year at a Hong Kong community college were randomly assigned to either an experimental or a comparison group." (p. 187)
      • Relevant Quotes: 1) "The study employed an experimental design, where 40 L1 Cantonese-speaking students in their first year at a Hong Kong community college were randomly assigned to either an experimental or a comparison group." (p. 187) 2) "Two learners from the experimental group failed to take the post-test, leaving 18 learners in this group and 20 in the comparison group." (p. 187) 3) "The participants had responded to advertisements offering the pronunciation workshop." (p. 188) Detailed Analysis: The paper explicitly states that randomisation was performed at the level of individual students (40 volunteers recruited via advertisement), not at the level of pre-existing classes or schools. There is no indication that intact classes or whole institutions were the unit of assignment. The exception for personal/one-to-one tutoring interventions does not apply here, since the intervention was a group pronunciation workshop delivered to cohorts of learners, not individual tutoring. Consequently, the randomisation unit fails to meet either the class-level requirement or the tutoring exception. Criterion C is not met because randomisation occurred at the individual student level, not at the class or school level, and the tutoring exception does not apply.
    • E

      Exam-based Assessment

      • The study used a custom reading-aloud task devised by the researchers for their own prior studies rather than a widely recognised standardised exam.
      • "The reading aloud task, used pre-test and post-test to assess the learners' pronunciation of the relevant sounds, had been previously piloted and used in the research of other groups of L1 Cantonese EFL learners by Yeldham (2018, 2020, 2021)." (p. 191)
      • Relevant Quotes: 1) "To collect the primary data for the study, the participants completed a reading aloud task, pre-test and post-test (see the Supplementary file, Part A)." (p. 188) 2) "The reading aloud task, used pre-test and post-test to assess the learners' pronunciation of the relevant sounds, had been previously piloted and used in the research of other groups of L1 Cantonese EFL learners by Yeldham (2018, 2020, 2021)." (p. 191) 3) "It consisted of two passages in which many of the words were embedded with the sounds targeted in our research." (p. 191) Detailed Analysis: The outcome measure was a bespoke reading-aloud task designed and piloted by the first author specifically for his own line of research into these particular phonemes, rather than a widely recognised, externally validated standardised exam. Although it had been reused across several of the author's own prior studies, reuse within one research programme does not constitute standardisation by an independent testing body. This is a custom assessment tailored precisely to the intervention's targeted sounds. Criterion E is not met because the assessment was a custom-built reading-aloud task designed by the researchers, not a standardised exam.
    • T

      Term Duration

      • The intervention lasted a single four-hour session with outcomes measured immediately afterward, far short of the one-term interval required.
      • "Each learner group attended an intensive four-hour pronunciation course (which included short breaks after each hour)..." (p. 189)
      • Relevant Quotes: 1) "Each learner group attended an intensive four-hour pronunciation course (which included short breaks after each hour), which was taught by the same teacher..." (p. 189) 2) "The pre-test and post-test were conducted immediately before and after the instruction courses, respectively, in a large classroom with all the students in each group undertaking the task simultaneously." (p. 191) 3) "A second limitation was the lack of a delayed post-test. In the absence of such a test there is little indication the experimental group would have maintained the benefits they gained from the direct instruction." (p. 196) Detailed Analysis: The entire intervention was delivered in a single four-hour block, and outcomes were measured immediately before and immediately after this single session. There was no follow-up assessment weeks or months later; the authors themselves acknowledge the lack of any delayed post-test as a limitation. The interval between intervention start and outcome measurement was therefore a matter of hours, far short of one academic term. Criterion T is not met because outcomes were measured immediately after a single four-hour session, with no term-length follow-up interval.
    • D

      Documented Control Group

      • The comparison group's size, baseline proficiency, and exact instructional treatment are clearly documented.
      • "The comparison group shared the same content and classroom activities with the experimental group. However, the critical difference between the two courses was that the teacher did not explain to the comparison group how to pronounce the language features when he modeled them for the class." (pp. 190-191)
      • Relevant Quotes: 1) "Personal information elicited through a short pre-instruction questionnaire showed they all spoke Cantonese as their L1, and in the Hong Kong Diploma of Secondary Education Examination (HKDSE) English Speaking Examination, a college entrance exam in Hong Kong, all had attained level 3, except for two learners in the comparison group and one in the experimental group, who had attained level 4." (pp. 188-189) 2) "Two learners from the experimental group failed to take the post-test, leaving 18 learners in this group and 20 in the comparison group." (p. 187) 3) "The comparison group shared the same content and classroom activities with the experimental group. However, the critical difference between the two courses was that the teacher did not explain to the comparison group how to pronounce the language features when he modeled them for the class. Instead, this group was given indirect instruction, including the provision of recasts for their problematic pronunciation without further explanation." (pp. 190-191) Detailed Analysis: The paper documents the comparison group's size (20 learners), their L1 (Cantonese), and their baseline English speaking proficiency (HKDSE level, closely matched to the experimental group). It also clearly describes exactly what the comparison group received: the same content and tasks as the experimental group, but with indirect instruction and recasts instead of direct explanations and feedback, plus extra preparation and task time. This level of detail satisfies the requirement for a well-documented control/comparison group. Criterion D is met because the comparison group's size, baseline characteristics, and exact treatment are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was at the individual-student level within a single college, not at the school level.
      • "The study employed an experimental design, where 40 L1 Cantonese-speaking students in their first year at a Hong Kong community college were randomly assigned to either an experimental or a comparison group." (p. 187)
      • Relevant Quotes: 1) "The study employed an experimental design, where 40 L1 Cantonese-speaking students in their first year at a Hong Kong community college were randomly assigned to either an experimental or a comparison group." (p. 187) 2) "The participants had responded to advertisements offering the pronunciation workshop." (p. 188) Detailed Analysis: The study drew volunteer students from a single community college who individually responded to advertisements, and randomisation was performed at the individual level, not across multiple schools or institutions. There is no mention of whole schools being randomised to conditions. Criterion S is not met because the study involved individual-level randomisation within a single institution, not school-level randomisation.
    • I

      Independent Conduct

      • The study's own authors designed, delivered, and predominantly scored the intervention, with no independent third-party evaluator involved.
      • "The second author taught both the experimental and comparison courses..." (p. 189)
      • Relevant Quotes: 1) "The second author was a part-time teacher at the college, affording us access to this cohort along with insider knowledge of the college and its students." (p. 187) 2) "The second author taught both the experimental and comparison courses, receiving permission from the college to teach the two groups of students outside of both his and their normal class hours." (p. 189) 3) "In marking the tests, based on pronunciations from the IPA sound chart, the targeted sounds in the learners' speech samples were rated by the second author as either: (1) matching the IPA pronunciation, or (2) deviating from it." (p. 192) 4) "To examine for reliability of the marking, the rater and a colleague separately marked approximately the first 25% of the collected speech samples (i.e. 20 recordings)..." (p. 192) Detailed Analysis: The second author, who co-designed the study and is a co-author of the paper, both delivered the instruction to both groups and served as the primary rater of the outcome measure. A colleague only cross-checked a subset (25%) of recordings for reliability purposes, but did not independently conduct or oversee the trial. There is no external, third-party evaluator responsible for implementation or outcome assessment; the study was designed, delivered, and predominantly scored by the study's own authors. Criterion I is not met because the same authors who designed the study also delivered the instruction and served as the primary rater of outcomes.
    • Y

      Year Duration

      • Since criterion T is not met and the study tracked outcomes for only a few hours, criterion Y is not met.
      • "A second limitation was the lack of a delayed post-test." (p. 196)
      • Relevant Quotes: 1) "Each learner group attended an intensive four-hour pronunciation course (which included short breaks after each hour)..." (p. 189) 2) "The pre-test and post-test were conducted immediately before and after the instruction courses, respectively..." (p. 191) 3) "A second limitation was the lack of a delayed post-test." (p. 196) Detailed Analysis: Per the ERCT standard, if criterion T (Term Duration) is not met, criterion Y is automatically not met. Here the entire study, from intervention to outcome measurement, spanned a single four-hour session with immediate testing, nowhere close to even one academic term, let alone 75% of an academic year. Criterion Y is not met both because criterion T is not met and because the actual tracking interval was only a few hours.
    • B

      Balanced Control Group

      • Both groups received equal total time, content, and teacher contact, with only the direct/indirect instructional style (the treatment variable) differing.
      • "Each learner group attended an intensive four-hour pronunciation course (which included short breaks after each hour), which was taught by the same teacher, and where both groups were taught the same course content and undertook the same tasks." (p. 189)
      • Relevant Quotes: 1) "Each learner group attended an intensive four-hour pronunciation course (which included short breaks after each hour), which was taught by the same teacher, and where both groups were taught the same course content and undertook the same tasks." (p. 189) 2) "In the absence of the time taken in the experimental course for direct attention to pronunciation, this was filled in this comparison course by giving the learners more time to prepare for tasks and allowing the tasks to run for slightly longer periods." (p. 190) 3) "The comparison group shared the same content and classroom activities with the experimental group." (p. 190) Detailed Analysis (applying the criterion B decision tree): Step 1 - EXTRA_RESOURCES_PRESENT: both groups received the identical total dose of four hours of instruction from the same teacher, using the same content, exercises, and pedagogical tasks. The intervention added no additional overall time or budget beyond what the comparison group received; the only difference was how that shared time was allocated (direct explanation/feedback for the experimental group versus extra task-preparation and slightly longer task time for the comparison group). Since no extra time or budget was given to the intervention group in total, this branch alone would already satisfy the criterion trivially. Step 2 (secondary support) - even if the direct instruction time were treated as an "extra resource," it is integral to the treatment itself (the direct-versus-indirect delivery style is precisely the variable under test), and the paper documents that the control group's freed-up time was explicitly reallocated to comparable task engagement, so CONTROL_MATCHES_RESOURCES is also satisfied. No extra budget, materials, or overall class time were given only to the experimental group; only the instructional delivery style (direct vs indirect), which is the treatment variable, differed. Criterion B is met because both groups received equal total instructional time, content, and teacher contact, and the only substantive difference is the direct vs indirect delivery style that constitutes the treatment variable itself.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study by a different research team was found; prior related studies were conducted by the same first author, and a citation search located no external replications.
      • "In conclusion, future research could replicate the study, while addressing the above limitations." (p. 196)
      • Relevant Quotes: 1) "Recently, the first author examined the effectiveness of these abdominal techniques for adult Chinese speakers of English in three experimental studies (Yeldham, 2018, 2020, 2021)." (p. 186) 2) "In conclusion, future research could replicate the study, while addressing the above limitations." (p. 196) Internet Search Findings: A citation search (via CrossRef and Semantic Scholar) on the DOI 10.1080/07908318.2021.1978476 found only two citing works: "Pengenalan Kosakata Sederhana pada Siswa TK Islam Bhakti 02 Semarang untuk Menumbuhkan Motivasi Belajar Bahasa Inggris Pasca Covid-19" (Rahayu, Widyaningrum, Purwanto, Rustipa, & Soepriatmadji, 2023), which concerns vocabulary instruction for kindergarten children, and "Improving Students' English Pronunciation Competence by Using Shadowing Technique" (Utami & Morganna, 2022), which tests a shadowing technique. Neither paper attempts to reproduce this study's direct articulatory-abdominal versus indirect TBLT comparison for Cantonese-speaking learners; both are unrelated designs citing the paper only tangentially. No independent replication of this specific study, by a different research team, applying a comparable articulatory-abdominal instructional design was found in any available source. Detailed Analysis: The paper references prior related studies, but these (Yeldham, 2018, 2020, 2021) were conducted by the same first author, not by an independent research team, and they examined related but not identical designs (e.g. different learner cohorts, different sound sets, no direct/indirect TBLT comparison as tested here). The authors explicitly call for future research to replicate this specific study, indicating no independent replication has yet occurred. The citation search confirms no external replication of this particular articulatory-abdominal versus indirect TBLT comparison was identified. Criterion R is not met because no independent replication by a different research team of this specific study has been published.
    • A

      All-subject Exams

      • Criterion A is not met because criterion E is not met and the assessment was narrowly limited to targeted pronunciation features.
      • Relevant Quotes: 1) "The reading aloud task, used pre-test and post-test to assess the learners' pronunciation of the relevant sounds..." (p. 191) 2) "It consisted of two passages in which many of the words were embedded with the sounds targeted in our research." (p. 191) Detailed Analysis: Per the ERCT standard, criterion A requires criterion E to be met as a prerequisite; since criterion E (Exam-based Assessment) is not met here, criterion A automatically fails. Separately, the study assessed only pronunciation of specific targeted phonemes via a custom reading-aloud task, not performance across multiple main academic subjects. Criterion A is not met both because criterion E is not met and because the study narrowly assessed only pronunciation of targeted phonemes.
    • G

      Graduation Tracking

      • Criterion G is not met because criterion Y is not met and there was no follow-up tracking beyond the immediate post-test, confirmed by a search of the author's subsequent publications.
      • "A second limitation was the lack of a delayed post-test." (p. 196)
      • Relevant Quotes: 1) "The pre-test and post-test were conducted immediately before and after the instruction courses, respectively..." (p. 191) 2) "A second limitation was the lack of a delayed post-test. In the absence of such a test there is little indication the experimental group would have maintained the benefits they gained from the direct instruction." (p. 196) Internet Search Findings: A search of the first author's (Michael Yeldham's) subsequent publication record (via Semantic Scholar) found later related papers, including "The impact of abdominal enhancement techniques on L1 Spanish, Japanese and Mandarin speakers' English pronunciation" (2023) and "L2 English pronunciation instruction: techniques that increase expiratory drive through enhanced use of the abdominal muscles, and transfer of learning" (2025). Both examine different participant cohorts (Spanish, Japanese, and Mandarin speakers) and neither is a follow-up study of the 40 Hong Kong Cantonese-speaking community college students from this paper, nor do they mention tracking these participants toward graduation. No follow-up publication tracking this specific cohort was found. Detailed Analysis: Per the ERCT standard, since criterion Y is not met, criterion G is automatically not met. Additionally, the paper explicitly acknowledges the absence of any delayed post-test, let alone tracking through graduation; measurement stopped immediately after the single instructional session, and the internet search found no planned or published follow-up study tracking these participants further. Criterion G is not met because criterion Y is not met and no follow-up tracking beyond the immediate post-test was conducted or reported.
    • P

      Pre-Registered

      • No mention of a pre-registered protocol, registry, or registration date appears anywhere in the paper, and none was found in an external search.
      • Relevant Quotes: No quotes were found anywhere in the paper (including the Method, Disclosure statement, or Funding sections) referencing a study registry, a pre-registration identifier, or a pre-registration date. Internet Search Findings: A search for a registered protocol associated with this study (e.g. OSF, AsPredicted, or a trial registry) found no matching pre-registration record, consistent with the absence of any mention of pre-registration in the paper itself. Detailed Analysis: The paper contains a "Disclosure statement" ("No potential conflict of interest was reported by the author(s).") and a "Funding" statement ("The author(s) reported there is no funding associated with the work featured in this article."), neither of which mentions pre-registration. No section of the paper references a registry platform, protocol publication, or pre-specified analysis plan published before data collection began. Criterion P is not met because no evidence of pre-registration is present anywhere in the paper or in any external registry search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.