Abstract
To what extent do second language (L2) learners benefit from instruction that includes corrective feedback (CF) on L2 speech perception? This article addresses this question by reporting the results of a classroom-based experimental study conducted with 32 young adult Korean learners of English. An instruction-only group and an instruction + CF group were exposed to five 1-hr form-focused lessons that drew learners' attention to the nonnative phonemic contrast /i/-/ɪ/, but only the instruction + CF group was given relevant feedback. Forced-choice identification tasks were completed by participants in a pretest, an immediate posttest, and a delayed posttest. The two groups showed similar accuracy on the pretest; however, the instruction + CF group outperformed the instruction-only group on the immediate and delayed posttests as well as on unfamiliar words. The significant predictors for these differences turned out to be perceptual accuracy vis-a-vis /ɪ/-natural and /ɪ/-synthesized sounds. These findings are discussed in terms of the pivotal role played by CF in developing accuracy in L2 speech perception.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- The study is explicitly quasi-experimental with matched (non-random) assignment of learners, not a randomised trial.
- "The current study was based on a quasi-experimental design implemented in simulated classrooms rather than in intact classes."
Relevant Quotes:
1) "The current study was based on a quasi-experimental design implemented in simulated classrooms rather than in intact classes." (p. 43)
2) "Taking into account the L2 learners' age, gender, length of residence (LOR), and percent correct identification scores yielded by the pretest, we assigned the 32 nonnative speakers to either the instruction-only group or the instruction + CF group (see Table 2). The groups were matched so that there were no significant differences on any of these variables." (p. 42-43)
3) "A total of four classes participated (i.e., two classes per group), each with eight learners." (p. 43)
Detailed Analysis:
The paper explicitly self-labels its design as quasi-experimental, and the described procedure for placing participants into conditions is a matching procedure based on age, gender, LOR, and pretest score, not a randomisation procedure. There is no quote anywhere in the Methodology describing random assignment, a random number generator, or any randomisation method; instead, deliberate matching on demographic and baseline variables is described. Since the ERCT standard requires evidence of random assignment (not matching) to conditions, and the authors themselves label the design quasi-experimental, this study does not qualify as a Randomised Controlled Trial at any level. Even setting aside the randomisation question, the four "classes" are ad hoc simulated groupings of eight learners recruited for the study, not intact pre-existing classes, and it is the matched, non-random assignment of individual learners to condition that determines group membership. The intervention is classroom-based pronunciation instruction for groups of learners, not personal one-to-one tutoring, so the tutoring exception does not apply.
Final: Criterion C is not met because the paper explicitly describes a quasi-experimental design using matched, non-random assignment of learners rather than randomisation at the class level.
-
E
Exam-based Assessment
- Outcomes were measured with a custom-built, study-specific forced-choice identification test rather than a standardised exam.
- "A computerized test was designed specifically for the purpose of this study to assess the treatment effects."
Relevant Quotes:
1) "A computerized test was designed specifically for the purpose of this study to assess the treatment effects." (p. 47)
2) "Using the PsyScope software (Cohen, MacWhinney, Flatt, & Provost, 1993), forced-choice identification tasks were designed in which a minimal pair of stimuli and a relevant sound were given." (p. 47)
3) "Twelve sets of /i/-/ɪ/ pairs were collected from the Longman Communication 3000 corpus (Longman Dictionary of Contemporary English, 2009, pp. 2044-2059), which is a list of the 3,000 most frequent words in both spoken and written English." (p. 47)
Detailed Analysis:
The outcome measure is a purpose-built, laboratory-style forced-choice identification task created specifically for this experiment using custom software (PsyScope) and duration-manipulated synthesized stimuli produced with a Klatt synthesizer. Although the target words were sourced from a word-frequency corpus, the assessment instrument itself is not a standardised, widely recognised exam of the kind required by criterion E; it is an experimenter-designed psycholinguistic instrument tailored precisely to this study's phonemic targets.
Final: Criterion E is not met because outcomes were measured using a custom-built forced-choice identification task designed specifically for this study, not a standardised exam.
-
T
Term Duration
- The delayed posttest occurred only two weeks after a five-day intervention, far short of a full academic term.
- "the participants also took part in a delayed posttest 2 weeks later"
Relevant Quotes:
1) "The instructional sessions consisted of five 1-hr pronunciation lessons, which were conducted for five consecutive days (i.e., 1 hr per day)." (p. 43)
2) "Finally, after the five 1-hr lessons, the learners took an immediate posttest. To determine whether the effects of the instruction and CF were retained after the instruction, the participants also took part in a delayed posttest 2 weeks later." (p. 43)
Detailed Analysis:
The intervention lasted only five consecutive days (one hour per day), and the final (delayed) outcome measurement occurred just two weeks after the intervention ended. The total elapsed time from intervention start to final measurement is therefore roughly three weeks, far short of the "at least one full academic term (approximately 3-4 months)" required by criterion T. No quote in the paper indicates any longer follow-up period.
Final: Criterion T is not met because the total tracking period from intervention start to final measurement was only about three weeks, well short of one academic term.
-
D
Documented Control Group
- The instruction-only comparison group's demographics, baseline scores, and exact procedures are explicitly documented in Table 2 and the text.
- "Table 2. Participant information by group ... Age (years) 28.5 (SD = 5.9) [Instruction-only] ... Total pretest scores (%) 58.2 (SD = 11.6)"
Relevant Quotes:
1) "Table 2. Participant information by group ... Age (years) 28.5 (SD = 5.9) [Instruction-only] 27.9 (SD = 5.3) [Instruction + CF]; Male:female ratio (%) 25:75 for both groups; LOR (months) 12.6 (SD = 13.7) vs. 14.0 (SD = 14.0); Total pretest scores (%) 58.2 (SD = 11.6) vs. 58.9 (SD = 11.2)." (Table 2, p. 43)
2) "Taking into account the L2 learners' age, gender, length of residence (LOR), and percent correct identification scores yielded by the pretest, we assigned the 32 nonnative speakers to either the instruction-only group or the instruction + CF group ... The groups were matched so that there were no significant differences on any of these variables." (p. 42-43)
3) "The instruction-only group participated in the same tasks, but the instructors provided no CF and no clues that might have lead students to believe that they were right or wrong." (p. 44)
Detailed Analysis:
Although this study lacks a true no-treatment control group (both groups receive instruction; only the presence of corrective feedback differs), the "instruction-only" group functions as the reference condition for isolating the added effect of CF. Its demographic and baseline characteristics (age, gender ratio, length of residence, pretest scores) are explicitly and quantitatively documented in Table 2, and the paper explicitly describes what this group received: the same five lessons and awareness tasks, but with no corrective feedback and no confirmation of correctness. This level of detail satisfies the documentation requirement.
Final: Criterion D is met because the comparison group's demographics, baseline scores, and exact treatment conditions are explicitly documented in Table 2 and the Methodology.
-
Level 2 Criteria
-
S
School-level RCT
- There is no randomisation at any level, let alone school-level; assignment was via matching, not randomisation.
- "The current study was based on a quasi-experimental design implemented in simulated classrooms rather than in intact classes."
Relevant Quotes:
1) "The current study was based on a quasi-experimental design implemented in simulated classrooms rather than in intact classes." (p. 43)
2) "A total of four classes participated (i.e., two classes per group), each with eight learners." (p. 43)
Detailed Analysis:
There is no school-level randomisation described anywhere in the paper; indeed there is no randomisation at all, and the "classes" used are simulated groupings created for the purposes of the study rather than intact schools or even intact classes. Since the weaker class-level criterion (C) is not met, the stronger school-level requirement cannot be met either.
Final: Criterion S is not met because the study used non-random, matched assignment of learners in simulated classroom groupings, with no school-level randomisation.
-
I
Independent Conduct
- The study's designers (the authors) also conducted the trial themselves, with no independent third-party evaluator described.
- "Neither of the authors participated in this study as a L2 student, L1 English speaker, L1 English listener, or instructor."
Relevant Quotes:
1) "Neither of the authors participated in this study as a L2 student, L1 English speaker, L1 English listener, or instructor." (p. 41)
2) "The three ESL instructors all met the criteria described for the L1 English listeners group and had ESL teaching experience." (p. 41)
3) "All lessons were video recorded to confirm the consistency of the implementation of the two treatments." (p. 43)
Detailed Analysis:
The two authors designed the study, the instructional materials, the feedback protocols, and the testing instruments, and there is no statement anywhere indicating that an independent, third-party organisation conducted data collection or analysis separate from the researchers who designed the intervention. The statement that the authors did not personally serve as instructors or participants addresses role contamination within the study, not independence of the evaluation team from the intervention's designers. Because the designers of the CF intervention (the authors) also ran the study and presumably analysed the data and interpreted results, criterion I is not met.
Final: Criterion I is not met because the same researchers who designed the corrective feedback intervention also conducted the study, with no independent third-party evaluator described.
-
Y
Year Duration
- Criterion Y is not met because criterion T is not met and the total study duration was only about three weeks.
- "the participants also took part in a delayed posttest 2 weeks later"
Relevant Quotes:
1) "The instructional sessions consisted of five 1-hr pronunciation lessons, which were conducted for five consecutive days (i.e., 1 hr per day)." (p. 43)
2) "the participants also took part in a delayed posttest 2 weeks later." (p. 43)
Detailed Analysis:
Per the ERCT standard's specific instruction, if criterion T (Term Duration) is not met, criterion Y (Year Duration) is automatically not met. As established for criterion T, the entire study spanned roughly three weeks from intervention start to final measurement, far short of even one academic term, let alone 75% of an academic year.
Final: Criterion Y is not met because criterion T is not met and the total study duration (~3 weeks) is far shorter than the required academic-year timeframe.
-
B
Balanced Control Group
- Both groups received identical lesson time and tasks; the only difference (corrective feedback) was the explicit treatment variable being tested, not an added resource.
- "During every class, the respective instructional treatment was provided to both groups, whereas CF was given only to the instruction + CF group during the instructional tasks."
Relevant Quotes:
1) "During every class, the respective instructional treatment was provided to both groups, whereas CF was given only to the instruction + CF group during the instructional tasks." (p. 43)
2) "The instruction-only group participated in the same tasks, but the instructors provided no CF and no clues that might have lead students to believe that they were right or wrong." (p. 44)
3) "The instructional sessions consisted of five 1-hr pronunciation lessons, which were conducted for five consecutive days (i.e., 1 hr per day)." (p. 43)
Detailed Analysis:
Applying the criterion B decision procedure: both the instruction-only and instruction + CF groups received exactly the same five 1-hr lessons, the same pronunciation-focused explicit instruction, the same input enhancement, and the same three awareness tasks (pick-a-card, bingo, fill-in-the- blank), for the same amount of class time and with the same three instructors. No extra instructional time, budget, or materials were allocated to either group; the intervention does not add time or budget beyond what the control condition receives, so the "no extra resources present" branch of the decision procedure applies and the criterion is trivially satisfied. In addition, the sole difference between conditions - whether the instructor provided corrective feedback during the shared awareness tasks - is the explicit, integral treatment variable the study was designed to test (Research Question 2: "To what extent do the training effects differ between a group receiving instruction only and a group receiving instruction plus CF?"), which independently satisfies the criterion even if CF were considered a resource.
Final: Criterion B is met because both groups received identical instructional time, tasks, and materials, and the only difference (corrective feedback) was the explicit treatment variable being tested, not an additional resource.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this study's specific classroom-based CF and speech-perception design was found in any available source.
- "Given that there are no previous studies investigating the effects of feedback on L2 speech perception, this study is expected to expand horizons in regard to the roles attributed to CF."
Relevant Quotes:
1) "Given that there are no previous studies investigating the effects of feedback on L2 speech perception, this study is expected to expand horizons in regard to the roles attributed to CF." (p. 61)
2) "To the best of our knowledge, there is no specific classroom-based research that has investigated whether L2 learners benefit from classroom-based perception training including L2 instruction and oral feedback provided during instructor-student interaction." (p. 40)
Detailed Analysis:
The authors themselves state that no prior research exists investigating the effects of feedback on L2 speech perception, framing their own study as the first of its kind. An internet search of citation databases (Semantic Scholar) for papers citing this study found related but independent research on corrective feedback and L2 speech (e.g., work by Saito on CF and L2 pronunciation development, and Felker et al. on corrective feedback and perceptual learning of a novel L2 accent), but none of these are a replication of this specific classroom-based, Korean-learner /i/-/ɪ/ perception design by a different research team in a different context; they investigate different populations, targets, or feedback mechanisms. No independent replication of this particular study's design and findings was identified in any available source.
Final: Criterion R is not met because no independent replication of this study's specific design and findings by a different research team was found.
-
A
All-subject Exams
- Criterion A is not met because criterion E is not met and only a single phonemic contrast, not multiple subjects, was assessed.
Relevant Quotes:
1) "Forced-choice identification tasks were completed by participants in a pretest, an immediate posttest, and a delayed posttest." (Abstract, p. 35)
2) "A computerized test was designed specifically for the purpose of this study to assess the treatment effects." (p. 47)
Detailed Analysis:
Per the ERCT standard's specific instruction, if criterion E is not met, criterion A cannot be met either. Additionally, on substantive grounds, the study measured only perception of a single narrow phonemic contrast (/i/ versus /ɪ/) rather than performance across the main subjects of a curriculum, so it would not satisfy an "all-subject" requirement independent of the E-prerequisite failure.
Final: Criterion A is not met both because criterion E is not met and because only a single narrow phonemic contrast, not multiple subjects, was assessed.
-
G
Graduation Tracking
- Criterion G is not met because criterion Y is not met, and no follow-up publication tracking this cohort was found.
Relevant Quotes:
1) "the participants also took part in a delayed posttest 2 weeks later." (p. 43)
2) "First, a sample size larger than 32 would, of course, strengthen the statistical analyses and findings." (p. 60, Limitations and Future Directions)
Detailed Analysis:
Per the ERCT standard's specific instruction, if criterion Y is not met, criterion G cannot be met either. Substantively, data collection ends with a delayed posttest only two weeks after the five-day intervention; the Limitations section discusses sample size, generalizability-trial timing, and simulated-versus-real classrooms as limitations but makes no mention of any further follow-up, let alone tracking students until graduation. A search for subsequent publications by Lee or Lyster tracking the same cohort of 32 Korean adult learners found no such follow-up study; this was a one-off classroom experiment with adult learners outside a graduation-bound school system, so no graduation-tracking follow-up is expected to exist.
Final: Criterion G is not met because criterion Y is not met, and no follow-up publication tracking this cohort until any form of graduation was found.
-
P
Pre-Registered
- No mention of pre-registration on a public registry is found anywhere in the paper or elsewhere.
Detailed Analysis:
There is no mention anywhere in the paper - including the Methodology, Procedure, or funding acknowledgment sections - of any pre-registration of the study's hypotheses, methods, or analysis plan on a public registry prior to data collection. The paper only references funding sources (an SSHRC grant and a fellowship), not a registered protocol. No pre-registration record for this study was found in available registries.
Final: Criterion P is not met because no pre-registration statement, registry link, or registration date is present anywhere in the paper, and none could be located elsewhere.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.