Effects of Form-Focused Instruction and Corrective Feedback on L2 Pronunciation Development of /ɹ/ by Japanese Learners of English

Kazuya Saito, Roy Lyster

Published:
ERCT Check Date:
DOI: 10.1111/j.1467-9922.2011.00639.x
  • L2 languages
  • adult education
  • Canada
0
  • C

    Individual students, not intact classes or schools, were the true unit of random assignment to conditions, with delivery groups formed only afterward.

    "In the instructional phase, 65 learners were randomly divided into three groups (i.e., FFI-only group, FFI + CF group, and control group)..." (p. 604)

  • E

    The study used custom word/sentence-reading and picture-description tasks with acoustic and listener ratings devised solely for this study, not a standardized exam.

    "Three tests targeting 14 of the 38 words that had appeared in the instructional treatment were used to evaluate the learners' improvement in their pronunciation of English /ɹ/ between pretests and posttests: a word-reading task, a sentence-reading task, and a timed picture-description task." (p. 610)

  • T

    Outcomes were measured only about two weeks after a 4-hour, two-week-long intervention, far short of a full academic term.

    "Two weeks after the 4 hr of instruction, the students individually completed posttests as well as final interviews." (p. 608)

  • D

    Detailed demographic data (Table 1) and a clear description of the control group's comparable instructional conditions are provided.

    "Although those in the control group received comparable instruction (English argumentative skills), the target pronunciation feature of their FFI was different (English vowel sounds)." (p. 607)

  • S

    Participants were individually recruited adult volunteers randomized as individuals, not schools or intact institutional units.

    "The ad was posted on several community Web sites specifically for Japanese people studying abroad, with hardcopy versions also distributed at many language institutes in Montreal." (p. 604-605)

  • I

    The authors designed the intervention and personally supervised instruction, conducted all testing, and performed all data analysis with no independent evaluator.

    "The testing sessions were completed in a quiet room in one-on-one meetings with the first author, a NS of Japanese." (p. 610)

  • Y

    The tracked duration (about one month) is far shorter than the required academic-year threshold, and the term-duration criterion it depends on is also not met.

    "Two weeks after the 4 hr of instruction, the students individually completed posttests as well as final interviews." (p. 608)

  • B

    All three groups received the identical 4-hour instructional allotment; only the pronunciation-focused content/feedback embedded within that time differed, so no extra resources were provided to the intervention groups.

    "the participants in the control group received 4 hr of comparable instruction also on the topic of 'developing English argumentative skills' but without form-focused instruction on English /r/ and with no exposure to the 38 target words." (p. 608)

  • R

    The authors describe the study as unprecedented and call for future replication; a later adaptation of the design was conducted by the same original authors rather than an independent research team.

    "Given that our findings were based on the specific case of L2 speech acquisition of English /ɹ/ by Japanese learners, we call for future research to replicate and extend the current research framework..." (p. 627)

  • A

    Since criterion E is not met and only a single pronunciation feature was assessed, the all-subject exam requirement is not satisfied.

  • G

    No follow-up beyond a two-week-delayed posttest was conducted, its prerequisite criterion Y is not met, and no later publication tracks this cohort further.

    "it would be important for future research to investigate the sustainability of FFI effectiveness over a longer period of time" (p. 627)

  • P

    The paper contains no mention of a pre-registered protocol, registry, or date.

Abstract

Sixty-five Japanese learners of English participated in the current study, which investigated the acquisitional value of form-focused instruction (FFI) with and without corrective feedback (CF) on learners' pronunciation development. All students received a 4-hr FFI treatment designed to encourage them to notice and practice the target feature of English /ɹ/ in meaningful discourse, except those in the control group (n = 11), who received comparable instruction but without FFI on English /ɹ/. During FFI, the instructors provided CF only to students in the FFI + CF group (n = 29) by recasting their mispronunciation or unclear pronunciation of /ɹ/, whereas no CF was provided to those in the FFI-only group (n = 25). Acoustic analyses were conducted on frequency values of the third formant (F3) of English /ɹ/ tokens elicited via pretest and posttest measures targeting familiar items and a generalizability test targeting unfamiliar items. The results showed that: (a) F3 values of the FFI + CF group significantly declined after the intervention, not only at a controlled-speech level but also a spontaneous-speech level, regardless of following vowel contexts; (b) change in F3 values of the FFI-only group and the control group was not statistically significant; and (c) the generalizability of FFI to novel tokens remained unclear.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Individual students, not intact classes or schools, were the true unit of random assignment to conditions, with delivery groups formed only afterward.
      • "In the instructional phase, 65 learners were randomly divided into three groups (i.e., FFI-only group, FFI + CF group, and control group)..." (p. 604)
      • Relevant Quotes: 1) "In the instructional phase, 65 learners were randomly divided into three groups (i.e., FFI-only group, FFI + CF group, and control group) and received 4 hr of meaning-based instruction about argumentative skills taught by the two ESL teachers." (p. 604) 2) "After the interviews, the 65 students were first randomly divided into 12 classes (6 students per class; cf. e.g., Sheen, 2006), and then these classes were designated as one of three groups: (a) FFI-only group (five classes, n = 29), (b) FFI + CF group (five classes, n = 25), and (c) control group (n = 11)." (p. 605) 3) "The ad was posted on several community Web sites specifically for Japanese people studying abroad, with hardcopy versions also distributed at many language institutes in Montreal. Interested participants contacted the first author through email or by phone..." (p. 604-605) Detailed Analysis: The randomization procedure described shows that individual students were first randomly allocated to condition ("65 learners were randomly divided into three groups"), and only afterward organized into artificial groups of six for logistical delivery of instruction (the paper's "classes"). These were not pre-existing, naturally occurring classrooms drawn from a school; they were newly formed ad hoc groups assembled purely for the study, from individually recruited volunteers responding to advertisements. Because the true unit of random allocation was the individual learner rather than an intact class or school, and because the "classes" were an artifact of administrative convenience rather than the object of randomization itself, this design does not meet the ERCT requirement for class-level (or stronger) randomization. The tutoring/personal-teaching exception does not apply either, since instruction was delivered to groups of six, not one-to-one. Criterion C is not met because randomization occurred at the individual-student level, with class groupings created afterward purely for instructional delivery.
    • E

      Exam-based Assessment

      • The study used custom word/sentence-reading and picture-description tasks with acoustic and listener ratings devised solely for this study, not a standardized exam.
      • "Three tests targeting 14 of the 38 words that had appeared in the instructional treatment were used to evaluate the learners' improvement in their pronunciation of English /ɹ/ between pretests and posttests: a word-reading task, a sentence-reading task, and a timed picture-description task." (p. 610)
      • Relevant Quotes: 1) "Three tests targeting 14 of the 38 words that had appeared in the instructional treatment were used to evaluate the learners' improvement in their pronunciation of English /ɹ/ between pretests and posttests: a word-reading task, a sentence-reading task, and a timed picture-description task." (p. 610) 2) "Learners read 25 individual words, which included 10 target words and 15 distracters. All target words were consonant-vowel-consonant (CVC) singleton tokens beginning with /ɹ/..." (p. 610) 3) "One hundred speech tokens were randomly chosen from the original data pool of the learners' performance in the pretest sessions and then presented to the five NS listeners to rate." (p. 612) 4) "the NS listener was presented 100 speech tokens in a randomized order and asked to rate them on a 9-point scale, with 1 as 'very good English /ɹ/' and 9 as 'very poor English /ɹ/,'" (p. 613) Detailed Analysis: The outcome measures in this study consist entirely of researcher-designed word-reading, sentence-reading, and picture-description tasks targeting a narrow set of minimally paired words containing /ɹ/, evaluated through acoustic (formant) analysis and native-speaker listener ratings devised specifically for this project. None of these instruments are standardized, widely recognized exams; they were custom-built by the authors to elicit and measure a single phonological feature. Criterion E is not met because the assessments are bespoke acoustic/listener-rating instruments created for this study rather than a standardized, widely recognized exam.
    • T

      Term Duration

      • Outcomes were measured only about two weeks after a 4-hour, two-week-long intervention, far short of a full academic term.
      • "Two weeks after the 4 hr of instruction, the students individually completed posttests as well as final interviews." (p. 608)
      • Relevant Quotes: 1) "Each class consisted of four 1-hr lessons and took place twice a week, finishing within 2 weeks (1 hr×2 lessons per week×2 weeks=4 hr)." (p. 608) 2) "Two weeks after the 4 hr of instruction, the students individually completed posttests as well as final interviews." (p. 608) Detailed Analysis: The instructional treatment itself lasted only 4 hours spread across two weeks, and the posttest measuring outcomes was administered just two weeks after instruction ended. The total interval from intervention start to outcome measurement is therefore approximately one month, far short of the one academic term (roughly 3-4 months) required by criterion T. No longer-term follow-up beyond this short-delayed posttest is reported anywhere in the paper. Criterion T is not met because the interval between intervention start and outcome measurement was only about a month, well below the term-length threshold.
    • D

      Documented Control Group

      • Detailed demographic data (Table 1) and a clear description of the control group's comparable instructional conditions are provided.
      • "Although those in the control group received comparable instruction (English argumentative skills), the target pronunciation feature of their FFI was different (English vowel sounds)." (p. 607)
      • Relevant Quotes: 1) "Table 1 provides the details of the 65 participants' information according to the three groups." (p. 605), with Table 1 listing the Control group (n = 11): Age M = 30.9 (SD 9.1), LOR M = 18.0 months (SD 40.8), Age of arrival M = 29.3 (SD 6.4), TOEIC M = 425.0 (SD 176.7), and "2 males and 9 females in the Control group." (Table 1, p. 606) 2) "Although those in the control group received comparable instruction (English argumentative skills), the target pronunciation feature of their FFI was different (English vowel sounds)." (p. 607) 3) "the participants in the control group received 4 hr of comparable instruction also on the topic of 'developing English argumentative skills' but without form-focused instruction on English /r/ and with no exposure to the 38 target words." (p. 608) Detailed Analysis: The control group (n = 11) is clearly documented in terms of demographics (age, gender, length of residence, age of arrival, TOEIC scores) via Table 1, and the paper explicitly describes the instructional content the control group received (the same argumentative-skills lessons and the same 4 hours of instruction, but with a different pronunciation focus and no exposure to the 38 target /ɹ/ words). This gives readers sufficient detail to assess comparability of the groups. Criterion D is met because the control group's demographics, size, and instructional conditions are thoroughly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Participants were individually recruited adult volunteers randomized as individuals, not schools or intact institutional units.
      • "The ad was posted on several community Web sites specifically for Japanese people studying abroad, with hardcopy versions also distributed at many language institutes in Montreal." (p. 604-605)
      • Relevant Quotes: 1) "For the purpose of student recruitment, the first author created ads that advertised an opportunity to participate in a 4-hr free course on argumentative skills in English..." (p. 604) 2) "The ad was posted on several community Web sites specifically for Japanese people studying abroad, with hardcopy versions also distributed at many language institutes in Montreal." (p. 604-605) 3) "After the interviews, the 65 students were first randomly divided into 12 classes (6 students per class...), and then these classes were designated as one of three groups..." (p. 605) Detailed Analysis: There is no school-level randomization in this study; participants were individually recruited volunteers from the community (via advertisements and language institutes), not students drawn from and randomized within or across existing schools. As established under criterion C, the randomization unit was the individual learner, organized post hoc into small ad hoc instructional groups. Since even the weaker class-level criterion is not met, the stronger school-level criterion cannot be met either. Criterion S is not met because the study recruited and randomized individual adult volunteers outside any school structure, rather than randomizing schools.
    • I

      Independent Conduct

      • The authors designed the intervention and personally supervised instruction, conducted all testing, and performed all data analysis with no independent evaluator.
      • "The testing sessions were completed in a quiet room in one-on-one meetings with the first author, a NS of Japanese." (p. 610)
      • Relevant Quotes: 1) "the first author sat at the back of the classroom to ensure that the consistency of the instruction was maintained within groups by the two instructors." (p. 608) 2) "The testing sessions were completed in a quiet room in one-on-one meetings with the first author, a NS of Japanese." (p. 610) 3) "the first author measured F1, F2, and F3 values in hertz (Hz) and transitional duration of F3 in milliseconds (ms) for all speech tokens." (p. 613) 4) "In preparation for the proposed research, the first author administered an 'expert judgment' questionnaire..." (p. 602) Detailed Analysis: The same research team that designed the FFI/CF intervention (the two authors) also recruited participants, trained and monitored the instructors, personally observed every instructional session, conducted all pretest/posttest data collection one-on-one with participants, and personally performed all acoustic measurements and statistical analyses. There is no external, third-party evaluator or independent agency involved in delivery, data collection, or analysis at any stage. Criterion I is not met because the intervention's designers also implemented, monitored, tested, and analyzed the entire study themselves.
    • Y

      Year Duration

      • The tracked duration (about one month) is far shorter than the required academic-year threshold, and the term-duration criterion it depends on is also not met.
      • "Two weeks after the 4 hr of instruction, the students individually completed posttests as well as final interviews." (p. 608)
      • Relevant Quotes: 1) "Each class consisted of four 1-hr lessons and took place twice a week, finishing within 2 weeks..." (p. 608) 2) "Two weeks after the 4 hr of instruction, the students individually completed posttests as well as final interviews." (p. 608) Detailed Analysis: As established under criterion T, the interval from intervention start to outcome measurement was only about a month (4 hours of instruction across two weeks, followed by posttesting two weeks later), nowhere near the 75% of an academic year (about 9-10 months) required by criterion Y. Since criterion T is not met, criterion Y cannot be met either per the standard's rule that Y requires T to be satisfied. Criterion Y is not met because the tracked duration was only about a month, and its prerequisite, criterion T, is also not met.
    • B

      Balanced Control Group

      • All three groups received the identical 4-hour instructional allotment; only the pronunciation-focused content/feedback embedded within that time differed, so no extra resources were provided to the intervention groups.
      • "the participants in the control group received 4 hr of comparable instruction also on the topic of 'developing English argumentative skills' but without form-focused instruction on English /r/ and with no exposure to the 38 target words." (p. 608)
      • Relevant Quotes: 1) "the participants in the control group received 4 hr of comparable instruction also on the topic of 'developing English argumentative skills' but without form-focused instruction on English /r/ and with no exposure to the 38 target words." (p. 608) 2) "Although those in the control group received comparable instruction (English argumentative skills), the target pronunciation feature of their FFI was different (English vowel sounds)." (p. 607) 3) "Each class consisted of four 1-hr lessons and took place twice a week, finishing within 2 weeks (1 hr×2 lessons per week×2 weeks=4 hr)." (p. 608), applying equally to FFI-only, FFI+CF, and control classes. Detailed Analysis: Applying the criterion B decision procedure: does the intervention add extra time or budget relative to control? No - all three groups (FFI-only, FFI+CF, control) received exactly the same amount of instructional time (4 hours across four 1-hr lessons over two weeks) with the same two trained instructors and the same "English argumentative skills" curriculum framework. The only difference between conditions is the content/focus embedded within that identical time allotment: the FFI groups had target /ɹ/ words highlighted and practiced (and, for FFI+CF, received corrective recasts), whereas the control group focused on English vowel sounds instead and did not encounter the 38 target words. Because no additional class time, materials, or budget were provided to the intervention groups beyond what the control group received, this satisfies the "no extra resources present" branch of the criterion B decision tree. Criterion B is met because all groups received an equal amount of instructional time and comparable overall lesson content, with only the specific pronunciation-focused content/feedback differing within that shared time.
  • Level 3 Criteria

    • R

      Reproduced

      • The authors describe the study as unprecedented and call for future replication; a later adaptation of the design was conducted by the same original authors rather than an independent research team.
      • "Given that our findings were based on the specific case of L2 speech acquisition of English /ɹ/ by Japanese learners, we call for future research to replicate and extend the current research framework..." (p. 627)
      • Relevant Quotes: 1) "To the best of our knowledge, there exist no experimental studies with a pretest and posttest design that investigate the acquisitional value of FFI with and without CF for L2 phonological development, especially during a set of meaningful FFI tasks." (p. 597) 2) "Given that our findings were based on the specific case of L2 speech acquisition of English /ɹ/ by Japanese learners, we call for future research to replicate and extend the current research framework..." (p. 627) Detailed Analysis: The authors themselves note that this is a novel study design with no prior precedent, and in the conclusion they explicitly call for future research to "replicate and extend the current research framework," confirming that no independent replication existed at the time of publication. Internet search for independent replications identified Gooch, R., Saito, K., and Lyster, R. (2016), "Effects of recasts and prompts on L2 pronunciation development: Teaching English /ɹ/ to Korean adult EFL learners" (System, 60, 68-78), which adapts this study's FFI-plus- corrective-feedback design to 22 Korean adult EFL learners. However, this follow-up is co-authored by Saito and Lyster, the same two authors who designed and conducted the original 2012 study, so it does not constitute an independent reproduction by a different research team. A further related study, Saito, K. (2013), "The Acquisitional Value of Recasts in Instructed Second Language Speech Learning: Teaching the Perception and Production of English /ɹ/ to Adult Japanese Learners" (Language Learning, 63, 499-529), is likewise single-authored by the original first author, using a new sample of 45 Japanese learners. No replication of this specific study by researchers unaffiliated with Saito or Lyster was found in any available source. Criterion R is not met because, although related follow-up studies exist, they were conducted by the same original authors rather than an independent research team, so no independent reproduction of this study was found.
    • A

      All-subject Exams

      • Since criterion E is not met and only a single pronunciation feature was assessed, the all-subject exam requirement is not satisfied.
      • Relevant Quotes: 1) "The results showed that: (a) F3 values of the FFI + CF group significantly declined after the intervention... (b) change in F3 values of the FFI-only group and the control group was not statistically significant; and (c) the generalizability of FFI to novel tokens remained unclear." (p. 595-596) Detailed Analysis: This study's outcome measures were narrowly restricted to acoustic and perceptual assessment of a single English phoneme, /ɹ/, in word-reading, sentence-reading, and picture-description tasks. It did not assess achievement in any other subject, let alone use standardized exams across the main curriculum. Because criterion E (Exam-based Assessment) is not met - the assessments were custom acoustic/rating instruments rather than a standardized exam - criterion A is automatically not met per the standard's rule that A requires E to be satisfied. Independent of that rule, the study also measured only a single narrow pronunciation feature, not all-subject outcomes. Criterion A is not met because the assessments are not standardized exams (failing the E prerequisite) and covered only a single phonetic feature rather than all core subjects.
    • G

      Graduation Tracking

      • No follow-up beyond a two-week-delayed posttest was conducted, its prerequisite criterion Y is not met, and no later publication tracks this cohort further.
      • "it would be important for future research to investigate the sustainability of FFI effectiveness over a longer period of time" (p. 627)
      • Relevant Quotes: 1) "Two weeks after the 4 hr of instruction, the students individually completed posttests as well as final interviews." (p. 608) 2) "it would be important for future research to investigate the sustainability of FFI effectiveness over a longer period of time (see Derwing & Munro, 2005)." (p. 627) Detailed Analysis: The study's only outcome measurement occurred approximately two weeks after the end of the brief instructional treatment; there is no indication of any tracking of participants beyond this short-delayed posttest, let alone through to program or educational-stage graduation. The authors themselves acknowledge this as a limitation, calling for future research on long-term sustainability. In addition, since criterion Y (Year Duration) is not met, criterion G is automatically not met per the standard's rule. An internet search for later publications by Saito and/or Lyster tracking this same cohort of 65 Montreal- based Japanese learners through to course completion or any later stage found none. The related follow-up studies by the original authors (Saito, 2013; Gooch, Saito, & Lyster, 2016) each use different, newly recruited samples (45 Japanese learners and 22 Korean learners, respectively) rather than extending observation of the original 65 participants, so they do not constitute graduation-tracking data for this study. No evidence of longer-term or graduation-stage follow-up on this specific cohort was found in any available source. Criterion G is not met because there is no tracking of participants beyond a short-delayed posttest, its prerequisite criterion Y is also not met, and no follow-up publication tracking this cohort was found.
    • P

      Pre-Registered

      • The paper contains no mention of a pre-registered protocol, registry, or date.
      • Relevant Quotes: No quotes were found anywhere in the methods, procedure, funding, or acknowledgments sections referencing a pre-registration registry, protocol ID, or pre-registration date. Detailed Analysis: A thorough review of the paper reveals no statement about pre-registration of the study's hypotheses, design, or analysis plan on any public registry prior to data collection. Given the publication timeframe (revised version accepted October 2010, published June 2012, with data collection and conference presentations occurring in 2009) and the nature of the study (a classroom-based instructed-SLA experiment), pre- registration would be unusual for this research area and era and, indeed, none is mentioned. An internet search of study-registry conventions in instructed SLA research from this period confirms pre-registration was not yet a standard practice in this field, consistent with the absence of any registry reference in the paper. Criterion P is not met because no pre-registration statement, registry link, or date is provided anywhere in the paper.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.