Abstract
The current study examines in depth how two types of form-focused instruction (FFI), which are FFI with and without corrective feedback (CF), can facilitate second language speech perception and production of /r/ by 49 Japanese learners in English as a Foreign Language settings. FFI effectiveness was assessed via three outcome measures (perception, controlled production, and spontaneous production) and also according to two lexical contexts (trained and untrained items). Two experimental groups received 4 hr of FFI treatment to notice and practice the target feature of /r/ (but without any explicit instruction) in meaningful discourse. A control group (n = 14) received comparable instruction in the absence of FFI. During FFI, the instructors provided CF only to students in the FFI + CF group (n = 18) by recasting their mispronunciations of /r/, while no CF was provided to those in the FFI-only group (n = 17). Analyses of pre- and posttests showed that FFI itself can sufficiently promote the development of speech perception and production of /r/ and the acquisitional value of CF in second language speech learning remains unclear. The results suggest that the beginner learners without much phonetic knowledge on how to repair their mispronunciation of /r/ should be encouraged to learn the target sound only through FFI in a receptive mode without much pressure for modified output.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Whole intact classes, not individual students within one class, were randomly assigned to the three study conditions.
- "Students were randomly assigned to nine classes of six students each, and then each of the classes was assigned to one of three groups, with each group consisting of three classes." (p. 384)
Relevant Quotes:
1) "Students were randomly assigned to nine classes of six students each, and then each of the classes was assigned to one of three groups, with each group consisting of three classes." (p. 384)
2) "The three groups were the FFI + CF group (three classes, n = 17), the FFI-only group (three classes, n = 18), and the control group (three classes, n = 14)." (p. 384)
Detailed Analysis:
The paper explicitly states that the unit assigned to each of the three experimental conditions was the intact class (nine classes total, three per condition), not individual students within a shared classroom. Because every student in a given class received the same condition, there is no risk of within-class contamination between treatment and control students, which is exactly the concern the ERCT "C" criterion is designed to rule out. This is not a one-to-one tutoring intervention, so the tutoring exception does not need to be invoked; the class-level assignment itself is sufficient. Note that quote 2 numerically swaps the n values relative to Table 1 on the same page (which lists FFI + CF as n = 18 and FFI-only as n = 17); this appears to be a labeling inconsistency in the original paper itself, verified verbatim against the PDF, and does not affect the class-level randomisation finding.
Final: Criterion C is met because randomisation was conducted by assigning whole classes, not individual students within a class, to conditions.
-
E
Exam-based Assessment
- All outcome measures (a forced-choice perception test, two production tests, and a teacher rating scale) were custom-built by the author for this study, not standardised exams.
- "The two-alternative forced-choice identification task was used to measure the learners' perception performance of /r/." (p. 387)
Relevant Quotes:
1) "The two-alternative forced-choice identification task was used to measure the learners' perception performance of /r/." (p. 387)
2) "The two production tests (word reading and timed picture description) were designed to measure the learners' pronunciation performance of /r/ at two different levels of processing (controlled- vs. spontaneous-speech levels)." (p. 388)
3) "A descriptor of a 9-point scale was adapted and modified from Flege, Takagi, et al.'s (1995) 6-point scale, and the rating criteria were explained as follows: 1 (very good /r/) ... 9 (very good /l/)." (p. 391)
Detailed Analysis:
None of the outcome instruments used in this study are widely recognised, standardised exams. The perception test is a bespoke two-alternative forced-choice identification task built around 39 target words selected for this study; the controlled- and spontaneous-production tests are researcher- designed word-reading and picture-description tasks; and the teacher-evaluation measure is a 9-point rating scale adapted from an earlier laboratory study's 6-point scale. All are purpose-built psycholinguistic instruments aligned tightly to the specific intervention (/r/ pronunciation), not general, standardised, widely recognised achievement tests.
Final: Criterion E is not met because the study relies exclusively on custom-made perception, production, and rating instruments rather than a standardised exam.
-
T
Term Duration
- Outcomes were measured about four weeks after the intervention began, far short of one academic term.
- "After completing pretests, all students in the experimental groups received four 1-hr English communication lessons over 2 weeks ... Two weeks after the end of the lessons, all students completed posttests." (p. 383)
Relevant Quotes:
1) "After completing pretests, all students in the experimental groups received four 1-hr English communication lessons over 2 weeks in which a range of FFI activities were embedded to encourage students to notice and practice the target pronunciation features of /r/ ... during meaningful discourse." (p. 383)
2) "Two weeks after the end of the lessons, all students completed posttests and participated in a final interview." (p. 383)
Detailed Analysis:
The intervention itself consisted of four 1-hr lessons delivered over a 2-week span, and the posttest/final outcome measurement occurred two further weeks after the lessons ended. The full interval from intervention start to final measurement is therefore roughly four weeks, which is far below the one full academic term (approximately 3-4 months) required by the standard. No later follow-up measurement is reported.
Final: Criterion T is not met because the interval from intervention start to outcome measurement was about one month, not a full academic term.
-
D
Documented Control Group
- The control group's size, demographics, and the instruction it received (comparable duration and content, minus the FFI focus) are described in detail.
- "The 14 learners in the control group received the same instruction on English argumentative skills as the other groups but without any FFI component on /r/. All target words in the instructional materials were replaced with comparable words ('rain' -> 'typhoon,' 'run' -> 'jog')." (p. 387)
Relevant Quotes:
1) "Students in the control group received comparable communicative language instruction in terms of duration and content but with no focus on phonetic form." (p. 383)
2) "The 14 learners in the control group received the same instruction on English argumentative skills as the other groups but without any FFI component on /r/. All target words in the instructional materials were replaced with comparable words ('rain' -> 'typhoon,' 'run' -> 'jog')." (p. 387)
3) "For the warm-up games, they did other communicative activities with no emphasis on pronunciation or listening practice." (p. 387)
4) Table 1 (p. 384) reports control-group (n = 14) age (M = 28.7, SD = 8.0), gender (1 male, 13 females), and length of residence (LOR) abroad (M = 5.6, SD = 9.9 months) alongside the same statistics for the two experimental groups.
Detailed Analysis:
The paper documents the control group's size (n = 14, three classes), demographic profile (age, gender, length of residence, shown alongside the treatment groups in Table 1), and precisely what instructional content it received (the same argumentative-skills lessons and warm-up game structure, with target words swapped for non-minimal-pair equivalents and no pronunciation/listening emphasis). This level of detail allows a reader to judge comparability of the control group to the treatment groups at baseline and during the study.
Final: Criterion D is met because the control group's composition, baseline characteristics, and exact instructional content are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred among classes within a single language institute, not among separate schools or sites.
- "This study took place at a private language institute in Osaka, Japan, which granted access to its students, classrooms, and teachers for research purposes." (p. 383)
Relevant Quotes:
1) "This study took place at a private language institute in Osaka, Japan, which granted access to its students, classrooms, and teachers for research purposes." (p. 383)
2) "Students were randomly assigned to nine classes of six students each, and then each of the classes was assigned to one of three groups." (p. 384)
Detailed Analysis:
All nine classes and all 49 participating learners were drawn from a single institution (one private language institute in Osaka). The unit of randomisation was the class, not the school/institute, and there was no attempt to randomise across multiple schools, sites, or institutes to capture school-level or real-world implementation variance.
Final: Criterion S is not met because the study was conducted, and randomised, entirely within one language institute rather than across multiple schools.
-
I
Independent Conduct
- The sole author designed the materials, trained the teachers, personally monitored lesson delivery, and coded the interaction data himself, with no independent evaluator involved.
- "The author observed the instruction from the back of the classroom to ensure consistency of treatments." (p. 383)
Relevant Quotes:
1) "The author observed the instruction from the back of the classroom to ensure consistency of treatments." (p. 383)
2) "The researcher provided the instructors with 4 hr of training sessions over 2 days before the intervention commenced. They were given a package of instructional materials with a list of the target words and then they were told the content and purpose of each activity as well as the way they were to provide CF following students' pronunciation errors." (p. 387)
3) "To examine the relationship between recasts and self-modified output, the author carefully watched 12 hr of videotaped FFI + CF lessons ... and checked the number of times the teachers provided recasts and to what degree they elicited students' repetition ... the author (a NS of Japanese) made a form of dichotomous coding on whether students made clear efforts to approximate English /r/." (p. 386)
Detailed Analysis:
This is a single-author study. The same person (Kazuya Saito) designed the FFI activities and target-word materials, personally trained the two instructors on how to deliver treatment and administer CF, sat in on lessons to monitor treatment fidelity, and later personally coded the recast-repair video data used in the analysis. There is no mention anywhere in the paper of an external, third-party evaluator or independent data-collection/analysis team; the author was involved at every stage from design through data coding.
Final: Criterion I is not met because the intervention was designed, delivered (via author-trained teachers), monitored, and analysed by the same single author with no independent oversight.
-
Y
Year Duration
- Because the term-duration criterion T is not met (total tracking was only about four weeks), the year-duration criterion is automatically not met.
- "Two weeks after the end of the lessons, all students completed posttests and participated in a final interview." (p. 383)
Relevant Quotes:
1) "After completing pretests, all students in the experimental groups received four 1-hr English communication lessons over 2 weeks." (p. 383)
2) "Two weeks after the end of the lessons, all students completed posttests and participated in a final interview." (p. 383)
Detailed Analysis:
Per the ERCT specification, if the weaker "T - Term Duration" criterion is not met, the stronger "Y - Year Duration" criterion cannot be met either. Independently, the actual tracked interval here (roughly one month from intervention start to final posttest) is far short of even one full academic term, let alone 75% of an academic year (~9-10 months).
Final: Criterion Y is not met, both because criterion T is not met and because the study's total duration is only about one month.
-
B
Balanced Control Group
- The control group received the same duration and general type of communicative instruction as the treatment groups, differing only in the absence of the FFI focus on /r/, which is itself the treatment variable being tested.
- "Students in the control group received comparable communicative language instruction in terms of duration and content but with no focus on phonetic form." (p. 383)
Relevant Quotes:
1) "Students in the control group received comparable communicative language instruction in terms of duration and content but with no focus on phonetic form." (p. 383)
2) "The 14 learners in the control group received the same instruction on English argumentative skills as the other groups but without any FFI component on /r/. All target words in the instructional materials were replaced with comparable words ('rain' -> 'typhoon,' 'run' -> 'jog')." (p. 387)
3) "For the warm-up games, they did other communicative activities with no emphasis on pronunciation or listening practice." (p. 387)
Detailed Analysis:
Applying the decision procedure: no extra time or budget was given to the intervention groups relative to the control group -- all three arms (FFI + CF, FFI-only, and control) received the identical four 1-hr lessons of communicative argumentative-skills instruction over the same two weeks, taught by the same pool of instructors. The only difference is the content focus: experimental groups practiced target /r/ words and /r/-focused warm-up games (plus, for one group, corrective feedback), while the control group used content-matched non-target words and unrelated communicative warm-up activities. Since EXTRA_RESOURCES_PRESENT is false (no additional time, budget, or materials were allocated to the treatment groups), the decision procedure resolves to "met" at the first branch, independent of whether the FFI focus itself would otherwise be considered the treatment variable.
Final: Criterion B is met because no extra time or resources were given to the treatment groups relative to the control group; all groups received the same instructional dosage.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study by another research team was reported in the paper or found in external follow-up literature.
Relevant Quotes:
1) "To isolate and reexamine the pure effects of FFI and CF on L2 speech learning, the study needs to be replicated, especially in conjunction with learners with homogeneous proficiency levels (e.g., LOR < 1 year) living in an English as a foreign language (EFL) environment." (p. 382, referring to the author's own precursor study, Saito & Lyster, 2012)
2) "Given that the current study took an exploratory approach to examining the effects of FFI on L2 speech learning, certain shortcomings must be acknowledged with an eye toward future replication." (p. 404)
Detailed Analysis:
The paper itself is framed as an extension of the author's own earlier precursor study (Saito & Lyster, 2012) rather than as a replication of it conducted by an independent team -- it is the same author (Saito) revisiting and refining his own prior design. The paper explicitly calls for future replication of its own findings but does not report that such replication has occurred. Checking the works citing this paper via OpenAlex (id W2098665834, 43 citing works as of this check) surfaced further pronunciation-instruction studies by Saito himself (e.g., "Re-examining effects of form-focused instruction on L2 pronunciation development," SSLA, 2013; "The acquisitional value of recasts in instructed second language speech learning," Language Learning, 2013, DOI 10.1111/lang.12015, cited in this paper as "Saito, in press") and pronunciation studies by other author teams (e.g., Lee et al.; Wisniewska & Mora; Martin & Sippel) that are topically related to L2 /r/-/l/ or pronunciation training but do not replicate this specific FFI/CF classroom design in a different context by an independent team. No independent replication of this specific study was found.
Final: Criterion R is not met because no independent replication of this specific study by a different research team was found.
-
A
All-subject Exams
- Because the exam-based-assessment criterion E is not met, this criterion automatically fails as well; in any case only /r/ pronunciation was assessed, not other subjects.
- "The current study examines in depth how two types of form-focused instruction (FFI) ... can facilitate second language speech perception and production of /r/ by 49 Japanese learners." (p. 377)
Relevant Quotes:
1) "The current study examines in depth how two types of form-focused instruction (FFI) ... can facilitate second language speech perception and production of /r/ by 49 Japanese learners in English as a Foreign Language settings." (p. 377)
Detailed Analysis:
Per the ERCT specification, criterion A cannot be met unless criterion E (standardised exam-based assessment) is first met, and E was not met here since all measures were custom-built. Independently, the study's entire outcome battery (perception task, controlled/spontaneous production tests, teacher ratings) targets only the pronunciation of a single phoneme, /r/; no other subject areas or broader language skills were assessed.
Final: Criterion A is not met, both because criterion E is not satisfied and because only a single narrow phonetic outcome was measured.
-
G
Graduation Tracking
- Because the year-duration criterion Y is not met, this criterion automatically fails; tracking stopped about two weeks after the lessons ended with no long-term follow-up.
- "Two weeks after the end of the lessons, all students completed posttests and participated in a final interview." (p. 383)
Relevant Quotes:
1) "Two weeks after the end of the lessons, all students completed posttests and participated in a final interview." (p. 383)
2) "Given that the current study took an exploratory approach to examining the effects of FFI on L2 speech learning, certain shortcomings must be acknowledged with an eye toward future replication." (p. 404, no mention of any longer-term or graduation follow-up)
Detailed Analysis:
Per the specification, because criterion Y (Year Duration) is not met, criterion G cannot be met either. Independently, data collection in this study concluded with a single posttest session two weeks after the four-lesson intervention ended; there is no mention anywhere in the paper of any further follow-up. The participants are adult learners at a private language institute rather than students in a K-12 or degree program, so "graduation tracking" in the ERCT sense does not naturally apply here in any case. A check of subsequent Saito publications citing this study (via OpenAlex, id W2098665834) found further pronunciation-instruction studies by the same author but no paper tracking this specific cohort of 49 learners toward any later milestone.
Final: Criterion G is not met because tracking ended shortly after the intervention with no longer-term or graduation follow-up was found in this paper or in subsequent literature, and because criterion Y is not met.
-
P
Pre-Registered
- No statement of pre-registration, registry identifier, or pre-registration date appears anywhere in the paper.
Relevant Quotes:
1) No quotes were found anywhere in the methods, acknowledgments, or notes sections referencing a trial registry (e.g., ClinicalTrials.gov, AEA RCT Registry, OSF) or a pre-registered protocol, hypotheses, or analysis plan.
2) "This study was funded by the Government of Canada Post-Doctoral Research Fellowship." (p. 406, Acknowledgments section, containing only funding and thanks, no registry reference)
Detailed Analysis:
A full read of the methods, acknowledgments, and notes sections reveals no reference to any pre-registration platform, registration ID, or date of registration. There is no evidence the study's hypotheses or analysis plan were published before data collection began. The study was accepted for publication in January 2013, predating the widespread adoption of pre-registration norms in applied linguistics/ education research; no registry record for this study was located via DOI/Crossref/OpenAlex metadata checks either.
Final: Criterion P is not met because no pre-registration reference of any kind is present in the paper, and none was located externally.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.