The efficacy of the type of instruction on second language pronunciation acquisition

Sharif Alghazo, Marwan Jarrah, Mohd Nour Al Salem

Published:
ERCT Check Date:
DOI: 10.3389/feduc.2023.1182285
  • L2 languages
  • higher education
  • Asia
0
  • C

    Students were randomised individually rather than by class or school, and the classroom-based intervention does not qualify for the tutoring exception.

    "The participants were randomly divided into two groups (32 students in each, based on the initial count)."

  • E

    The study used custom researcher-made pronunciation tasks scored by a single rater on a Likert scale, not a recognised standardised exam.

    "In order to validate the tests, a pilot experiment was conducted, with four students sitting for pre-, post-, and delayed post-tests."

  • T

    Outcomes were tracked from intervention start (Week 1) to a delayed post-test at Week 14, an interval of roughly one academic term (about 3 months).

    "At Week 14, a delayed post-test was also conducted to validate the outcomes and make sure that the gains are in the long-term memory."

  • D

    The comparison group's size, baseline scores, and treatment are documented in tables and text, along with demographic details of the sample.

    "The participants had a low-intermediate to intermediate level of proficiency (based on the results of the pre-test). Their ages ranged between 18 and 21 years. They were all native speakers of JA and had never been to an English-speaking country."

  • S

    The study randomised individual students within a single university, so no school-level randomisation occurred.

    "The study was conducted in an EFL context: a provisional governmental university in Jordan."

  • I

    The same author designed the intervention, taught both groups, and analysed the data, with no independent evaluation team.

    "The study is interventionist. The principal researcher (who is also the teacher) is a specialist in English pronunciation."

  • Y

    The study tracked outcomes for only about 13 weeks, far short of 75% of an academic year.

    "Progress in L2 pronunciation was assessed at three time points (i.e., week 1, week 6, and week 14)."

  • B

    Both groups received identical time, instructor, and content coverage, with only the instructional method differing as the tested variable.

    "Each group received PI for approximately 2.5h a week from the principal researcher which totaled 15h of instruction."

  • R

    No independent replication of this specific study by a different research team has been identified, confirmed via a citation-database search conducted during verification.

  • A

    Only custom pronunciation tasks were used, so with criterion E unmet and no other subjects assessed, this criterion fails.

    "The controlled tasks included 10 items in each. The first task tested the segmental aspect, the second tested the syllabic aspect, the third the prosodic aspect, the fourth the global aspect, and the fifth the temporal aspect."

  • G

    Measurement ended at Week 14 with no tracking of participants to graduation, and prerequisite criterion Y is unmet; no follow-up publication was found via search.

    "Progress in L2 pronunciation was assessed at three time points (i.e., week 1, week 6, and week 14)."

  • P

    The paper contains no mention of a pre-registered protocol, registry, or registration date, and none was found via internet search.

Abstract

This study investigates the efficacy of the type of instruction (i.e., perception-based vs. production-based) on second language (L2) pronunciation acquisition in an English as a foreign language (EFL) context. To achieve this objective, 60 tertiary-level Jordanian learners of English were recruited and put into two groups (30 learners in each group). Group A received 6 weeks of perception-based instruction on both segmental and suprasegmental aspects of English pronunciation, and Group B received production-based instruction over the same period and on the same aspects of pronunciation. Progress in L2 pronunciation was assessed at three time points (i.e., week 1, week 6, and week 14). Pre-, post- and delayed post-tests were run to achieve the study's objective. A statistical analysis was conducted to analyse the data. The results show that both groups demonstrated a significant improvement in L2 pronunciation accuracy; in particular, Group A which received perception-based instruction demonstrated higher gains in segmental, syllabic, and prosodic aspects while Group B which received production-based instruction demonstrated more improvement in both global (i.e., comprehensibility) and temporal (i.e., fluency) aspects of pronunciation. However, both groups demonstrated similar gains on the delayed post-test. The findings provide implications for L2 pronunciation learners and teachers on the impact of the type of instruction on the addressed aspects of pronunciation.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Students were randomised individually rather than by class or school, and the classroom-based intervention does not qualify for the tutoring exception.
      • "The participants were randomly divided into two groups (32 students in each, based on the initial count)."
      • Relevant Quotes: 1) "The participants were randomly divided into two groups (32 students in each, based on the initial count)." (p. 3) 2) "The participants were 64 university students who were enrolled in an English language course." (p. 3) 3) "The principal researcher (who is also the teacher) is a specialist in English pronunciation. He explained to the students in the two sections the aims of the study and requested their written consent to participate in it." (p. 4) 4) "Each group received PI for approximately 2.5h a week from the principal researcher which totaled 15h of instruction." (p. 4) Detailed Analysis: Criterion C requires randomisation at the class level (or stronger, school level) to prevent contamination between treatment and control conditions. The paper states that individual participants "were randomly divided into two groups," i.e., the unit of randomisation was the student, not intact classes. Although the paper mentions "the two sections," it does not state that pre-existing sections were randomly assigned as intact units; rather, the recruited students were divided into two groups at the individual level. The intervention was delivered as group classroom instruction (2.5 hours per week per group), not personal one-to-one tutoring, so the tutoring exception does not apply. Because both groups were taught by the same instructor within the same course context, student-level randomisation leaves contamination risks unaddressed. Criterion C is not met because randomisation was performed at the individual student level for a classroom-based intervention, and no class-level or school-level assignment is described.
    • E

      Exam-based Assessment

      • The study used custom researcher-made pronunciation tasks scored by a single rater on a Likert scale, not a recognised standardised exam.
      • "In order to validate the tests, a pilot experiment was conducted, with four students sitting for pre-, post-, and delayed post-tests."
      • Relevant Quotes: 1) "At Week 1 of the intervention, the researcher conducted a pre-test which included five sections to cover the five aspects of L2 pronunciation targeted in the study." (p. 4) 2) "Two types of tasks were included as outcome measures: controlled and free production tasks." (p. 4) 3) "The controlled tasks included 10 items in each. The first task tested the segmental aspect, the second tested the syllabic aspect, the third the prosodic aspect, the fourth the global aspect, and the fifth the temporal aspect." (p. 4) 4) "In order to assess the recordings of the participants' performance on the five tests, an expert rater with a specialty in English pronunciation was recruited... The rater assessed each utterance using a 10-point Likert scale where a score of 1 meant that the utterance is inaccurate, and a score of 10 is perfectly accurate." (p. 4) 5) "In order to validate the tests, a pilot experiment was conducted, with four students sitting for pre-, post-, and delayed post-tests... As a result of the pilot experiment, some words were replaced by others, and more hints were provided in the picture-description task." (p. 4) Detailed Analysis: Criterion E requires that outcomes be measured with standard, widely recognised standardised exams rather than instruments created specifically for the study. Here, the outcome measures were researcher-designed tasks (five controlled tasks of 10 items each, plus a picture-description free production task) constructed and piloted by the authors for this study, with items revised after piloting. Scoring was done by a single recruited rater using a 10-point Likert scale, not through any recognised standardised examination (e.g., IELTS, TOEFL, or a national curriculum exam). No name of any standardised test appears anywhere in the paper. Criterion E is not met because outcomes were measured with a custom, study-specific instrument scored by a single rater rather than a recognised standardised exam.
    • T

      Term Duration

      • Outcomes were tracked from intervention start (Week 1) to a delayed post-test at Week 14, an interval of roughly one academic term (about 3 months).
      • "At Week 14, a delayed post-test was also conducted to validate the outcomes and make sure that the gains are in the long-term memory."
      • Relevant Quotes: 1) "Group A received 6 weeks of perception-based instruction on both segmental and suprasegmental aspects of English pronunciation, and Group B received production-based instruction over the same period and on the same aspects of pronunciation. Progress in L2 pronunciation was assessed at three time points (i.e., week 1, week 6, and week 14)." (p. 1) 2) "At Week 1 of the intervention, the researcher conducted a pre-test... At Week 6, the same test was run by the researcher to investigate the gains (if any) after instruction had occurred for 6 weeks. At Week 14, a delayed post-test was also conducted to validate the outcomes and make sure that the gains are in the long-term memory." (p. 4) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months, i.e., a semester or equivalent) after the intervention begins. The standard explicitly allows short interventions provided there is term-long follow-up tracking from the intervention start. Here the intervention itself lasted 6 weeks, which alone is well below a term. However, the study also administered a delayed post-test at Week 14, i.e., roughly 13 weeks (about 3.2 months) after the intervention began at Week 1. This interval corresponds approximately to one university semester of tracking from intervention start to final outcome measurement, and falls within the standard's definition of a term as "approximately 3-4 months." The delayed post-test used the same instrument and was an explicit part of the study design to verify retention of gains. Criterion T is met because the final outcome measurement (delayed post-test at Week 14) occurred approximately one academic term (about 13 weeks / 3 months) after the intervention began, satisfying the term-long tracking requirement.
    • D

      Documented Control Group

      • The comparison group's size, baseline scores, and treatment are documented in tables and text, along with demographic details of the sample.
      • "The participants had a low-intermediate to intermediate level of proficiency (based on the results of the pre-test). Their ages ranged between 18 and 21 years. They were all native speakers of JA and had never been to an English-speaking country."
      • Relevant Quotes: 1) "The participants were 64 university students who were enrolled in an English language course. However, four students missed the delayed post-test, and consequently their data were excluded. The participants were majoring in Applied English or English language and literature. The participants had a low-intermediate to intermediate level of proficiency (based on the results of the pre-test). Their ages ranged between 18 and 21 years. They were all native speakers of JA and had never been to an English-speaking country." (p. 3) 2) "Group B received production-based instruction over the same period and on the same aspects of pronunciation." (p. 1) 3) "Perception 30 3.80 0.57 6.10 0.68 / Production 30 3.65 0.59 6.19 0.43" (Table 1, p. 5) 4) "Each group received PI for approximately 2.5h a week from the principal researcher which totaled 15h of instruction." (p. 4) Detailed Analysis: Criterion D requires clear documentation of the comparison (control) group: its composition, size, baseline performance, and treatment received. This study has no untreated control group; it is a two-arm comparative design in which each instruction condition serves as the comparison for the other. The comparison arm (Group B, production-based) is documented: its size (n = 30 analysed), its baseline pre-test means and standard deviations overall (Table 1) and for each of the five sub-tests (Table 3), and the exact treatment it received (production-based PI, 2.5 hours per week for 6 weeks, 15 hours total, same content areas and same instructor as Group A). Shared demographic characteristics (age 18-21, native Jordanian Arabic speakers, majors, proficiency level, no residence abroad) are reported for the full sample. While demographics are not broken down per group, the baseline performance data, group sizes, and condition descriptions provide adequate documentation for comparison. Criterion D is met because the comparison group's size, baseline test performance, and received instruction are documented in detail, together with sample-level demographics.
  • Level 2 Criteria

    • S

      School-level RCT

      • The study randomised individual students within a single university, so no school-level randomisation occurred.
      • "The study was conducted in an EFL context: a provisional governmental university in Jordan."
      • Relevant Quotes: 1) "The participants were randomly divided into two groups (32 students in each, based on the initial count)." (p. 3) 2) "The study was conducted in an EFL context: a provisional governmental university in Jordan." (p. 3) Detailed Analysis: Criterion S requires randomisation at the level of schools or equivalent institutional units. This study took place at a single university, and randomisation occurred at the individual student level within one course. No multiple institutions were involved and no institutional-level assignment is described anywhere in the paper. Criterion S is not met because randomisation was at the student level within a single university, not across schools or institutional units.
    • I

      Independent Conduct

      • The same author designed the intervention, taught both groups, and analysed the data, with no independent evaluation team.
      • "The study is interventionist. The principal researcher (who is also the teacher) is a specialist in English pronunciation."
      • Relevant Quotes: 1) "The study is interventionist. The principal researcher (who is also the teacher) is a specialist in English pronunciation." (p. 4) 2) "Each group received PI for approximately 2.5h a week from the principal researcher which totaled 15h of instruction." (p. 4) 3) "SA conducted the study, analysed the data and wrote the discussion. MJ wrote the introduction and literature review. MA wrote the conclusion and checked references." (p. 9) 4) "The results of all tests were given to a statistician to run the appropriate statistical analyses for the experiment." (p. 4) 5) "In order to assess the recordings of the participants' performance on the five tests, an expert rater with a specialty in English pronunciation was recruited." (p. 4) Detailed Analysis: Criterion I requires that the study be conducted independently of the intervention designers. Here the principal researcher designed the study, delivered all instruction to both groups himself, and, per the author contributions, also "conducted the study, analysed the data and wrote the discussion." The use of an external rater for scoring recordings and a statistician for running analyses provides only partial outsourcing of specific tasks; it does not constitute independent third-party conduct or oversight of the evaluation, since the authors controlled design, delivery, data handling, and interpretation. There is no statement of an external evaluation team or independent oversight body. Criterion I is not met because the intervention designer was also the teacher, data collector, and analyst, with no independent third-party evaluation.
    • Y

      Year Duration

      • The study tracked outcomes for only about 13 weeks, far short of 75% of an academic year.
      • "Progress in L2 pronunciation was assessed at three time points (i.e., week 1, week 6, and week 14)."
      • Relevant Quotes: 1) "Group A received 6 weeks of perception-based instruction... Progress in L2 pronunciation was assessed at three time points (i.e., week 1, week 6, and week 14)." (p. 1) 2) "At Week 14, a delayed post-test was also conducted to validate the outcomes and make sure that the gains are in the long-term memory." (p. 4) Detailed Analysis: Criterion Y requires that outcomes be measured at least 75% of a full academic year (roughly 7+ months of a 9-10 month year) after the intervention begins. In this study, the total tracking window from intervention start (Week 1) to the final delayed post-test (Week 14) is about 13 weeks, i.e., roughly 3 months. This is far short of 75% of an academic year. No longer follow-up is reported or planned. Criterion Y is not met because tracking lasted only about 13 weeks from intervention start, well below 75% of an academic year.
    • B

      Balanced Control Group

      • Both groups received identical time, instructor, and content coverage, with only the instructional method differing as the tested variable.
      • "Each group received PI for approximately 2.5h a week from the principal researcher which totaled 15h of instruction."
      • Relevant Quotes: 1) "Group A received 6 weeks of perception-based instruction on both segmental and suprasegmental aspects of English pronunciation, and Group B received production-based instruction over the same period and on the same aspects of pronunciation." (p. 1) 2) "Each group received PI for approximately 2.5h a week from the principal researcher which totaled 15h of instruction. Both groups received PI on segmental (using minimal pairs) and suprasegmental (using phrases and dialogues) phonology focusing on English and Arabic differences and highlighting potential problems because of interference." (p. 4) 3) "Potential teaching techniques were utilized by the instructor researcher to ensure optimal understanding and performance. Those included body gestures, hand clapping, and communicative activities. If performance was erroneous, the instructor would give corrective feedback in the form of recast and repetition." (p. 4) Detailed Analysis: Criterion B requires that time and resources be balanced between conditions unless the extra resource is itself the treatment variable. Applying the current decision tree: this is a two-arm active comparison, not a treatment-vs-untreated design, so there is no business-as-usual arm to check for matched resources. Both groups received the same dosage (approximately 2.5 hours per week for 6 weeks, 15 hours total), from the same instructor, over the same period, covering the same pronunciation content (segmental and suprasegmental aspects), with the same range of teaching techniques available (gestures, clapping, communicative activities, corrective feedback). The only difference between the arms is the instructional modality (perception-based versus production-based), which is exactly the treatment contrast being tested (EXTRA_RESOURCES_PRESENT is false between the two compared conditions, so the decision tree resolves to "met" at its first branch). Criterion B is met because both groups received identical instructional time (15 hours), the same instructor, and the same content coverage, with only the instructional method differing as the treatment variable.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study by a different research team has been identified, confirmed via a citation-database search conducted during verification.
      • Relevant Quotes: 1) "Of particular importance is that by Lee et al. (2020) who examined the effect of the type of PI (perception-based vs. production-based) on pronunciation acquisition among 115 Japanese university students of English." (p. 3) 2) "This study fills this gap in the literature and ventures to explore the relative effect of the type of instruction on English pronunciation gains among Jordanian Arabic-speaking learners of English at the university level." (p. 2) 3) "This last result is surprising and warrants more studies in different contexts." (p. 3) Detailed Analysis: Criterion R requires that this specific study be independently replicated by a different research team in a different context and published in a peer-reviewed journal. This paper (published May 2023) is itself positioned as filling a gap and extending prior work (Lee et al., 2020) to a new population; prior related studies by other teams (e.g. Lee et al., 2020, on Japanese students) are antecedents, not replications of this study. Internet verification (2026-07-27): a Semantic Scholar citation search for this paper's DOI (10.3389/feduc.2023.1182285) returned 10 citing works as of this check. None of the citing papers is a replication of this specific perception- vs. production-based instruction design among Jordanian EFL university learners by a different research team; the citing works address unrelated topics (e.g., digital scaffolding for collaborative writing, virtual discussion platforms for speaking courses). No independent peer-reviewed replication of this specific study was located. Criterion R is not met because no independent peer-reviewed replication of this specific study by a different research team was found, either referenced in the paper or via citation-database search.
    • A

      All-subject Exams

      • Only custom pronunciation tasks were used, so with criterion E unmet and no other subjects assessed, this criterion fails.
      • "The controlled tasks included 10 items in each. The first task tested the segmental aspect, the second tested the syllabic aspect, the third the prosodic aspect, the fourth the global aspect, and the fifth the temporal aspect."
      • Relevant Quotes: 1) "Therefore, the present study aims to provide further testing of the effect of the type of PI on five aspects of L2 pronunciation: segmental, syllabic (epenthesis), prosodic (stress placement), global (comprehensibility), and temporal (fluency) aspects." (p. 3) 2) "The controlled tasks included 10 items in each. The first task tested the segmental aspect, the second tested the syllabic aspect, the third the prosodic aspect, the fourth the global aspect, and the fifth the temporal aspect." (p. 4) Detailed Analysis: Criterion A requires standardised exam-based assessment of all main subjects taught at the educational level, and explicitly fails if criterion E fails. Criterion E is not met here because all measures were custom, study-specific pronunciation tasks. Furthermore, the study measured only English pronunciation sub-skills; no other academic subjects (or even other English language skills) were assessed, and no rationale invoking the specialised-intervention exception with standardised related-subject exams is provided. Criterion A is not met because criterion E fails and only custom pronunciation measures in a single narrow domain were used.
    • G

      Graduation Tracking

      • Measurement ended at Week 14 with no tracking of participants to graduation, and prerequisite criterion Y is unmet; no follow-up publication was found via search.
      • "Progress in L2 pronunciation was assessed at three time points (i.e., week 1, week 6, and week 14)."
      • Relevant Quotes: 1) "Progress in L2 pronunciation was assessed at three time points (i.e., week 1, week 6, and week 14)." (p. 1) 2) "At Week 14, a delayed post-test was also conducted to validate the outcomes and make sure that the gains are in the long-term memory." (p. 4) Detailed Analysis: Criterion G requires tracking participants until graduation from their educational stage, and per the instructions it cannot be met if criterion Y is not met. Criterion Y fails here. Measurement ended at Week 14 with the delayed post-test; there is no mention of following the university students through to graduation, and no follow-up study on this cohort is referenced or planned in the paper. Internet verification (2026-07-27): a Semantic Scholar citation search for this paper's DOI found 10 citing works. Two involve the same lead author, Sharif Alghazo, in later papers ("Exploring the impact of digital scaffolding on collaborative writing practices," 2025; "Virtual Versus Reality: A Look into the Effects of Discussion Platforms on Speaking Course Achievements in Gather.town," 2024), but both concern different interventions (writing scaffolding; virtual discussion platforms) and different student cohorts, not a continuation of this pronunciation study's original 60 participants through to graduation. No follow-up publication tracking this specific cohort was found. Criterion G is not met because tracking stopped at Week 14, criterion Y is not met, and no graduation-tracking follow-up publication on this cohort was located via search.
    • P

      Pre-Registered

      • The paper contains no mention of a pre-registered protocol, registry, or registration date, and none was found via internet search.
      • Relevant Quotes: 1) "The study is interventionist. The principal researcher (who is also the teacher) is a specialist in English pronunciation. He explained to the students in the two sections the aims of the study and requested their written consent to participate in it." (p. 4) 2) "Data availability statement: The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author/s." (p. 9) Detailed Analysis: Criterion P requires pre-registration of the full study protocol (hypotheses, methods, planned analyses) on a public registry before data collection began. The paper contains no mention of any registry (e.g., ClinicalTrials.gov, OSF, AsPredicted), no registration ID, and no registration date. The methods describe consent and procedures but never reference a pre-registered protocol or analysis plan. Internet verification (2026-07-27): no pre-registration record for this study or its authors was located during verification searches; the published article and its metadata contain no registry link or identifier. Criterion P is not met because no pre-registration statement or registry reference appears anywhere in the paper, and no registry entry was found during internet verification.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.