Reexamining Effects of Form-Focused Instruction on L2 Pronunciation Development: The Role of Explicit Phonetic Information

Kazuya Saito

Published:
ERCT Check Date:
DOI: 10.1017/S0272263112000666
  • L2 languages
  • adult education
  • Asia
0
  • C

    Entire classes of six students were randomly assigned as intact units to one of the three conditions, so no class contained a mix of treatment and control students, avoiding within-class contamination.

    "The 49 participants were randomly assigned to nine classes of 6 students." (p. 10)

  • E

    Outcomes were measured with a custom acoustic pronunciation test devised by the author for this study, not a widely recognized standardized exam.

    "To measure the effects of the two types of FFI (i.e., FFI+EI vs. FFI-only) as compared to the control group, all learners were asked to complete two types of production tests at both pre- and posttest sessions: (a) the controlled production (CP) test (i.e., reading a list of words) and (b) the spontaneous production (SP) test (i.e., describing a set of pictures)." (p. 14)

  • T

    The intervention lasted only two weeks and outcomes were measured two weeks after the intervention ended, far short of a full academic term.

    "Instructional treatments consisted of four 1-hr lessons distributed over 2 weeks (1-hr lesson x 2 times per week x 2 weeks = 4 hr)." (p. 8)

  • D

    The control group's size, gender composition, lesson content, and baseline equivalence to the treatment groups are all explicitly documented.

    "The 14 participants in the control group also received comparable meaning-oriented lessons on English argumentative skills but with neither FFI nor EI." (p. 13)

  • S

    The whole study was run at a single private language institute, with randomization at the class level, not across multiple schools.

    "The project was conducted at a private language institute in Osaka, Japan." (p. 8)

  • I

    The sole author designed the intervention, trained and closely monitored both instructors, observed every class, and performed the acoustic analysis himself, with no independent evaluation team.

    "All classes were videotaped and observed by the researcher, who always sat at the back of the classroom to ensure the consistency of FFI treatment for the entire project." (p. 8)

  • Y

    Because criterion T (Term Duration) is not met, this stronger year-duration criterion cannot be met either; substantively the tracked interval was only about four weeks.

    "Two weeks after the end of the lessons, all students took posttests and were interviewed." (p. 8)

  • B

    All three groups received the identical total amount of instructional time and comparable lesson content, differing only in the presence or absence of pronunciation-focused instruction, so no imbalance in time or resources requires further scrutiny.

    "For the control group, students received meaning- oriented lessons that were comparable in terms of duration and content but without any focus on form (i.e., English /r/)." (p. 8)

  • R

    No independent replication of this specific FFI+EI study by a different research team was found in the paper or via internet search; the only closely related prior study shares the same lead author.

    "Saito and Lyster (2012) took a first step toward testing how a range of FFI techniques can promote the acquisition of the English sound /r/ by adult Japanese learners." (p. 2)

  • A

    Because criterion E (Exam-based Assessment) is not met, this stronger all-subject criterion cannot be met either; the study also measures only pronunciation of a single phone, not other subjects.

    "Acoustic analyses were conducted on the primary acoustic property of /r/—that is, F3 values—in all 2,700 tokens." (p. 16)

  • G

    Because criterion Y (Year Duration) is not met, this stronger graduation-tracking criterion cannot be met either; no follow-up beyond the two-week posttest is reported, and no later paper tracking this cohort was found.

    "Two weeks after the end of the lessons, all students took posttests and were interviewed." (p. 8)

  • P

    No statement of pre-registration, registry platform, or registration date is present anywhere in the paper, and no registry entry for this study was found online.

Abstract

The present study examines whether and to what degree providing explicit phonetic information (EI) at the beginning of form-focused instruction (FFI) on second language pronunciation can enhance the generalizability and magnitude of FFI effectiveness by increasing learners' ability to notice a new phone. Participants were 49 Japanese learners of English in English as a foreign language setting. Whereas the control group (n = 14) received meaning-oriented lessons without any focus on form, the experimental groups received 4 hr of FFI treatment designed to encourage them to practice the target feature of an English /r/ in meaningful discourse. Instructors provided EI (i.e., multiple exposure to an exaggerated model pronunciation of /r/ and rule presentation on the relevant articulatory configurations) to the FFI+EI group (n = 17) but not to the FFI-only group (n = 18). Their pre- and posttest performance was acoustically analyzed according to various lexical, task, and following vowel conditions. The results of the ANOVAs showed that (a) the FFI-only group demonstrated moderate improvement with medium effects (e.g., change from hybrid exemplars to poor exemplars), particularly in familiar lexical contexts, and (b) the FFI+EI group not only demonstrated considerable improvement with large effects (e.g., change from hybrid exemplars to good exemplars) but also generalized the instructional gains to unfamiliar lexical contexts beyond the instructional materials.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Entire classes of six students were randomly assigned as intact units to one of the three conditions, so no class contained a mix of treatment and control students, avoiding within-class contamination.
      • "The 49 participants were randomly assigned to nine classes of 6 students." (p. 10)
      • Relevant Quotes: 1) "The 49 participants were randomly assigned to nine classes of 6 students." (p. 10) 2) "Two treatment groups and one control group, each of which comprised three classes, were formed: (a) FFI+EI group (three classes, n = 17: males, n = 3; females, n = 14), (b) FFI-only group (three classes, n = 18: males, n = 3; females, n = 15), and (c) control group (three classes, n = 14: male, n = 1; females, n = 13)." (p. 10) 3) "All classes were videotaped and observed by the researcher, who always sat at the back of the classroom to ensure the consistency of FFI treatment for the entire project." (p. 8) Detailed Analysis: The ERCT 'C' criterion is concerned with contamination: whether treatment and control participants share the same classroom, allowing intervention behaviors to leak into the control condition. Here, the 49 participants were organized into nine classes of six students each, and every one of those classes was assigned wholesale to a single condition (three classes for FFI+EI, three for FFI-only, three for control). No class mixed students from different conditions. Although these classes were purpose-formed for the study rather than pre-existing intact school classrooms, the randomization procedure still operates at the level of whole class-groups, which satisfies the underlying purpose of the criterion: preventing treatment/control students from being taught together in the same room. The researcher's continuous observation of each class further supports that treatment fidelity, and separation between conditions, was maintained throughout. Criterion C is met because random assignment was performed on whole classes of students, keeping each class uniformly in one condition and preventing within-class contamination.
    • E

      Exam-based Assessment

      • Outcomes were measured with a custom acoustic pronunciation test devised by the author for this study, not a widely recognized standardized exam.
      • "To measure the effects of the two types of FFI (i.e., FFI+EI vs. FFI-only) as compared to the control group, all learners were asked to complete two types of production tests at both pre- and posttest sessions: (a) the controlled production (CP) test (i.e., reading a list of words) and (b) the spontaneous production (SP) test (i.e., describing a set of pictures)." (p. 14)
      • Relevant Quotes: 1) "To measure the effects of the two types of FFI (i.e., FFI+EI vs. FFI-only) as compared to the control group, all learners were asked to complete two types of production tests at both pre- and posttest sessions: (a) the controlled production (CP) test (i.e., reading a list of words) and (b) the spontaneous production (SP) test (i.e., describing a set of pictures)." (p. 14) 2) "Materials. All target words were consonant-vowel-consonant (CVC) singletons except one word, Ryan, which was a CVVC (see the 14 words with asterisks in Table 1)." (p. 14) 3) "Acoustic analyses were conducted on the primary acoustic property of /r/—that is, F3 values—in all 2,700 tokens (i.e., 2,450 words from 49 learners + 250 words from 10 NSs) to assess in depth to what degree Japanese learners exhibited gains resulting from FFI with and without EI in comparison with the NS baseline." (p. 16) Detailed Analysis: The outcome measure is not a standardized, widely recognized exam but a bespoke phonetic production instrument designed specifically by the author for this project: a word-reading list (CP test) and a picture-description task (SP test), scored via acoustic (F3 formant) analysis rather than any external standardized test battery. This instrument was purpose built to elicit the single target sound /r/ in particular phonetic and lexical contexts, and there is no indication it is a recognized, externally validated assessment used beyond this line of research. Criterion E is not met because the assessment is a custom acoustic production test created by the author, not a standardized exam.
    • T

      Term Duration

      • The intervention lasted only two weeks and outcomes were measured two weeks after the intervention ended, far short of a full academic term.
      • "Instructional treatments consisted of four 1-hr lessons distributed over 2 weeks (1-hr lesson x 2 times per week x 2 weeks = 4 hr)." (p. 8)
      • Relevant Quotes: 1) "Instructional treatments consisted of four 1-hr lessons distributed over 2 weeks (1-hr lesson x 2 times per week x 2 weeks = 4 hr)." (p. 8) 2) "Two weeks after the end of the lessons, all students took posttests and were interviewed." (p. 8) 3) "To measure the effectiveness and durability of FFI, the posttest sessions in the current study took place two weeks after instruction." (Note 1, p. 26) Detailed Analysis: The intervention itself spanned only two weeks (four 1-hour lessons), and the posttest measuring the primary outcome was administered two weeks after the lessons ended. The total interval from the start of the intervention to the outcome measurement is therefore approximately four weeks, which is far shorter than the one full academic term (roughly 3-4 months) required by this criterion. The author's own footnote frames the two-week delay only as "short delayed" rather than a term-long follow-up. Criterion T is not met because the total tracked interval from intervention start to outcome measurement is about four weeks, well short of a full academic term.
    • D

      Documented Control Group

      • The control group's size, gender composition, lesson content, and baseline equivalence to the treatment groups are all explicitly documented.
      • "The 14 participants in the control group also received comparable meaning-oriented lessons on English argumentative skills but with neither FFI nor EI." (p. 13)
      • Relevant Quotes: 1) "(c) control group (three classes, n = 14: male, n = 1; females, n = 13)." (p. 10) 2) "The 14 participants in the control group also received comparable meaning-oriented lessons on English argumentative skills but with neither FFI nor EI; the students received feedback not on any pronunciation errors but rather on ungrammatical or inappropriate lexical choices ... as well as the content of the lessons." (p. 13) 3) "As for warm-up games, the participants in the control group were given different communicative games without any focus on pronunciation or listening practice, which the instructor usually used in her regular English conversation classes." (p. 13) 4) "Pretest Data ... Neither main effects of group nor lexis were found significant in any contexts, p = .300-.800. This indicates that any changes in the experimental groups were attributable to neither group nor lexical difference at the onset of the study." (p. 19) Detailed Analysis: The paper documents the control group's exact size and gender breakdown (n = 14: 1 male, 13 females across three classes), describes precisely what the control group did instead of FFI/EI (meaning-oriented lessons with feedback only on grammar/lexis and content, plus different warm-up games), and statistically confirms that the control group did not differ from the treatment groups at pretest. This level of detail satisfies the documentation requirement. Criterion D is met because the control group's composition, treatment, and baseline comparability are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • The whole study was run at a single private language institute, with randomization at the class level, not across multiple schools.
      • "The project was conducted at a private language institute in Osaka, Japan." (p. 8)
      • Relevant Quotes: 1) "The project was conducted at a private language institute in Osaka, Japan." (p. 8) 2) "The 49 participants were randomly assigned to nine classes of 6 students." (p. 10) Detailed Analysis: The entire study, all nine classes across the three conditions, took place within one single private language institute. Randomization occurred at the class level within that one institute rather than across multiple schools or institutions being randomly assigned to conditions. There is no evidence of a school-level (institution-level) random assignment. Criterion S is not met because the study was conducted at a single institute with class-level, not school-level, randomization.
    • I

      Independent Conduct

      • The sole author designed the intervention, trained and closely monitored both instructors, observed every class, and performed the acoustic analysis himself, with no independent evaluation team.
      • "All classes were videotaped and observed by the researcher, who always sat at the back of the classroom to ensure the consistency of FFI treatment for the entire project." (p. 8)
      • Relevant Quotes: 1) "The results reported here are based on my doctoral dissertation submitted to McGill University in 2011." (p. 1, acknowledgments) 2) "All classes were videotaped and observed by the researcher, who always sat at the back of the classroom to ensure the consistency of FFI treatment for the entire project." (p. 8) 3) "Two instructors participated in a total of 4 hr of teacher training led by the researcher over a 2-day period." (p. 13) 4) "The author and one experienced phonetician conducted acoustic analyses separately," used only for a 10% intercoder-reliability check, after which "the author analyzed the remaining dataset." (p. 17-18) Detailed Analysis: This is a single-authored paper in which the researcher personally designed the FFI and EI materials, personally trained the two instructors on how and when to deliver corrective feedback and EI, personally observed and videotaped every single class session to enforce treatment fidelity, and personally conducted essentially all of the acoustic data analysis (an independent phonetician was involved only in a small reliability check on 10% of tokens). There is no external or third-party evaluation team independent of the intervention's designer overseeing data collection or analysis. Criterion I is not met because the same researcher who designed the intervention also trained the instructors, supervised implementation, and analyzed the data, with no independent third-party conduct.
    • Y

      Year Duration

      • Because criterion T (Term Duration) is not met, this stronger year-duration criterion cannot be met either; substantively the tracked interval was only about four weeks.
      • "Two weeks after the end of the lessons, all students took posttests and were interviewed." (p. 8)
      • Relevant Quotes: 1) "Instructional treatments consisted of four 1-hr lessons distributed over 2 weeks." (p. 8) 2) "Two weeks after the end of the lessons, all students took posttests and were interviewed." (p. 8) Detailed Analysis: Per the ERCT rule, Y cannot be met when the weaker T criterion is not met. Substantively, the total interval from intervention start to outcome measurement was only about four weeks, nowhere near 75% of an academic year (roughly 9-10 months). Criterion Y is not met because criterion T is not met, and the actual tracked duration is far shorter than a year.
    • B

      Balanced Control Group

      • All three groups received the identical total amount of instructional time and comparable lesson content, differing only in the presence or absence of pronunciation-focused instruction, so no imbalance in time or resources requires further scrutiny.
      • "For the control group, students received meaning- oriented lessons that were comparable in terms of duration and content but without any focus on form (i.e., English /r/)." (p. 8)
      • Relevant Quotes: 1) "For the control group, students received meaning- oriented lessons that were comparable in terms of duration and content but without any focus on form (i.e., English /r/)." (p. 8) 2) "Instructional treatments consisted of four 1-hr lessons distributed over 2 weeks (1-hr lesson x 2 times per week x 2 weeks = 4 hr)." (p. 8) [applied uniformly to all three groups per Figure 1] 3) "To ensure that all groups received the same amount of instruction (i.e., 4 hr), the instructors were asked to spend more time on warm-up games (for the FFI-only group) and small talk (for the control group) instead of providing EI." (p. 13) Detailed Analysis (re-applying the current criterion B decision tree): Step 1, are extra time/budget resources present in the intervention conditions relative to control? No. All three groups (FFI+EI, FFI-only, control) received the identical total instructional time (4 hr over 2 weeks) and comparable lesson materials and topics (English argumentative skills); the only systematic difference is the pedagogical focus (form vs. meaning) and the brief EI segments, which the authors explicitly substituted with equivalent warm-up games or small talk time so that total instructional time remained matched across groups. Because no group received additional time, budget, or materials beyond what the others received, this criterion is met at the first branch of the decision tree (no extra resources present), without needing to invoke the "integral to treatment" exception. Criterion B is met because the intervention and control conditions received matched instructional time and comparable lesson content, with only the pronunciation-focused component varied and no net difference in resources provided to any group.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific FFI+EI study by a different research team was found in the paper or via internet search; the only closely related prior study shares the same lead author.
      • "Saito and Lyster (2012) took a first step toward testing how a range of FFI techniques can promote the acquisition of the English sound /r/ by adult Japanese learners." (p. 2)
      • Relevant Quotes: 1) "Saito and Lyster (2012) took a first step toward testing how a range of FFI techniques can promote the acquisition of the English sound /r/ by adult Japanese learners." (p. 2) 2) "The results of the current study, with beginner- intermediate Japanese EFL learners in Japan, echoed those of the original study, with intermediate Japanese ESL learners in Canada." (p. 23) Detailed Analysis: The current paper positions itself as a reexamination and extension of Saito and Lyster (2012), but that earlier study shares the same lead author (Saito), so it does not constitute independent replication by a different research team. Internet verification (2026-07-30): a citation-database search (Semantic Scholar citation graph for DOI 10.1017/S0272263112000666, cross-checked against general web search) was conducted for papers citing this study or replicating its FFI+EI design with adult Japanese learners of English /r/. The citing literature identified (e.g., Saito & Plonsky, 2019, "Effects of Second Language Pronunciation Teaching Revisited: A Meta-Analysis"; Kang, Sok, & Han, 2019, "Thirty-Five Years of ISLA on Form-Focused Instruction: A Meta-Analysis"; Crowther & Loewen, "Instructed Second Language Acquisition and Second Language Pronunciation"; and various 2020-2025 L2-speech-perception papers) consists of meta-analyses and conceptually related studies rather than independent replications of this specific study's design, sample, and measures by a separate, unaffiliated team. No paper reporting a direct replication of this Osaka private language institute FFI+EI /r/ study by researchers unaffiliated with Saito was found. Criterion R is not met because no independent replication of this specific study by a different research team was identified either in the paper or through internet searching.
    • A

      All-subject Exams

      • Because criterion E (Exam-based Assessment) is not met, this stronger all-subject criterion cannot be met either; the study also measures only pronunciation of a single phone, not other subjects.
      • "Acoustic analyses were conducted on the primary acoustic property of /r/—that is, F3 values—in all 2,700 tokens." (p. 16)
      • Relevant Quotes: 1) "To measure the effects of the two types of FFI ... all learners were asked to complete two types of production tests at both pre- and posttest sessions." (p. 14) 2) "Acoustic analyses were conducted on the primary acoustic property of /r/—that is, F3 values—in all 2,700 tokens." (p. 16) Detailed Analysis: Per the ERCT rule, criterion A cannot be met when criterion E is not met, and E was not satisfied here because the outcome instrument is a custom acoustic pronunciation test rather than a standardized exam. Substantively, the study also measures only production of a single target phone (/r/) and does not assess any other subject areas, so there is no all-subject coverage regardless of the E prerequisite. Criterion A is not met because the prerequisite criterion E is not met and the assessment covers only a single narrow pronunciation target.
    • G

      Graduation Tracking

      • Because criterion Y (Year Duration) is not met, this stronger graduation-tracking criterion cannot be met either; no follow-up beyond the two-week posttest is reported, and no later paper tracking this cohort was found.
      • "Two weeks after the end of the lessons, all students took posttests and were interviewed." (p. 8)
      • Relevant Quotes: 1) "Two weeks after the end of the lessons, all students took posttests and were interviewed." (p. 8) Detailed Analysis: Per the ERCT rule, criterion G cannot be met when criterion Y is not met. Substantively, the paper reports no follow-up data collection beyond the single posttest session two weeks after the intervention ended. Internet verification (2026-07-30): a search was conducted for subsequent papers by Kazuya Saito (or co-authors) tracking the same cohort of 49 adult private- language-institute learners toward course completion or any later milestone. No such follow-up publication was found; Saito's later citing/related works (e.g., Saito & Plonsky, 2019 meta-analysis) do not report longitudinal tracking of this specific 2011/2013 sample. In addition, because these participants were adult learners at a private language institute rather than students in a K-12 or degree program, "graduation tracking" in the sense intended by the criterion does not clearly apply to this population, and no relevant tracking data exist in any available source. Criterion G is not met because criterion Y is not met, no tracking beyond the immediate posttest is reported in the paper, and no follow-up publication tracking this cohort was found via internet search.
    • P

      Pre-Registered

      • No statement of pre-registration, registry platform, or registration date is present anywhere in the paper, and no registry entry for this study was found online.
      • Relevant Quotes: No quotes were found anywhere in the manuscript (including the Method, Notes, and References sections) referencing a study registry, a pre-registration platform, or a registration date for hypotheses or analysis plans. Detailed Analysis: The paper describes a dissertation-based experimental study (data collected circa 2010-2011, published 2013), predating the widespread adoption of pre-registration norms in second language acquisition research. Absent any quoted evidence of a registry ID or registration timing in the manuscript, and with no entry for this study found in a search of registry-related web sources (e.g., OSF) during the 2026-07-30 verification, this criterion cannot be considered satisfied. Criterion P is not met because no pre-registration is referenced anywhere in the paper and none was located through internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.