Effects of Different Types of Corrective Feedback on Receptive Skills in a Second Language: A Speech Perception Training Study

Andrew H. Lee, Roy Lyster

Published:
ERCT Check Date:
DOI: 10.1111/lang.12167
  • L2 languages
  • adult education
  • Canada
  • EdTech platform
  • digital assessment
0
  • C

    Randomisation occurred at the individual student level in a laboratory setting, not at the class or school level, and no valid tutoring/personal-teaching exception was invoked.

    "The L2 learners were then randomly assigned to one of four treatment groups or the control group (n = 20 per group) before participating in eight computer-assisted perception training sessions distributed over a 2-week period." (p. 815)

  • E

    Outcomes were measured using a custom-built forced-choice identification test created by the researchers for this study, not a standardised, widely recognised exam.

    "Forced-choice identification tests designed using Praat (Boersma & Weenink, 2013) were used to measure learners' perceptual accuracy..." (p. 818)

  • T

    The entire study, from pretest to final (delayed) posttest, spanned only about four weeks, far short of one academic term.

    "...eight computer-assisted perception training sessions distributed over a 2-week period... After the last training session, the learners completed an immediate posttest and, 2 weeks later, a delayed posttest." (p. 815)

  • D

    The control group's size, procedure (same training minus CF), and baseline/outcome scores are clearly documented in the text and in Tables 1-3.

    "Although the control group engaged in the same perception training, there was no CF regardless of the response alternative selected." (p. 817)

  • S

    Randomisation was conducted at the individual participant level across multiple different institutes, not at the school/institute level.

    "The L2 learners were then randomly assigned to one of four treatment groups or the control group (n = 20 per group)..." (p. 815)

  • I

    The same authors who designed the intervention and materials also conducted data collection and analysis themselves, with no independent evaluator involved.

    "Andrew H. Lee and Roy Lyster, McGill University" (title page)

  • Y

    Since criterion T (Term Duration) is not met, and the whole study spanned only about four weeks, criterion Y is not met.

    "...eight computer-assisted perception training sessions distributed over a 2-week period... 2 weeks later, a delayed posttest." (p. 815)

  • B

    The presence/type of corrective feedback is the explicit treatment variable being tested, so the no-CF control condition is the appropriate "business-as-usual" baseline rather than an unbalanced resource allocation.

    "To what extent do the training effects differ according to the type of auditory or visual CF...?" (p. 814)

  • R

    No evidence was found of independent replication of this specific study by a different research team; the authors themselves call for future replication.

    "...it would be important to replicate this study, controlling the amount of repeated experience with training materials during L2 speech perception training." (p. 826)

  • A

    Criterion E (standardised exam) is not met, so criterion A is automatically not met; additionally, only two specific vowel contrasts were assessed rather than a broad set of subjects/skills.

    "Two English vowel contrasts (/i/-/I/ and /E/-/ae/) were selected as targets." (p. 816)

  • G

    Tracking ended entirely within about a month of the first training session, with no continued follow-up and no graduation tracking; criterion Y is also not met, which further precludes G.

    "...the delayed posttest 2 weeks after the immediate posttest." (p. 818)

  • P

    The paper documents open sharing of materials but contains no statement of pre-registration of the study's hypotheses, design, or analysis plan before data collection.

Abstract

This study investigated the effects of different types of corrective feedback (CF) provided during second language (L2) speech perception training. One hundred Korean learners of L2 English, randomly assigned to five groups (n = 20 per group), participated in eight computer-assisted perception training sessions targeting two minimal pairs of English vowels. Four treatment groups each received a different type of CF; three groups received one of three types of auditory CF and a fourth group received a visual type of CF; the control group did not receive CF. Results of pretests, immediate posttests, and delayed posttests showed that, in comparison to the control group, the groups that received auditory CF improved significantly in trained over untrained words, whereas the group that received visual CF fared less well. These results are discussed in terms of the benefits of auditory CF types, especially CF combining target and nontarget forms.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation occurred at the individual student level in a laboratory setting, not at the class or school level, and no valid tutoring/personal-teaching exception was invoked.
      • "The L2 learners were then randomly assigned to one of four treatment groups or the control group (n = 20 per group) before participating in eight computer-assisted perception training sessions distributed over a 2-week period." (p. 815)
      • Relevant Quotes: 1) "One hundred Korean learners of L2 English, randomly assigned to five groups (n = 20 per group), participated in eight computer-assisted perception training sessions targeting two minimal pairs of English vowels." (p. 809, Abstract) 2) "Before the training sessions, 100 L2 learners and 20 L1 listeners completed a pretest at a research laboratory. The L2 learners were then randomly assigned to one of four treatment groups or the control group (n = 20 per group)..." (p. 815) 3) "Most participants were students learning English in private or university language institutes located in the Montreal area." (p. 815) Detailed Analysis: The unit of randomisation was the individual participant, not an intact class or school. The quotes make clear that 100 individually recruited adult learners (drawn from various private and university language institutes, not a single class) were randomly allocated one-by-one to one of five conditions and tested individually at a research laboratory using a computer program. There is no description of classes or schools being randomised as units. The paper also does not explicitly invoke the "personal teaching/tutoring" exception; the intervention is automated computer-delivered CF rather than one-to-one human tutoring, so this exception cannot be clearly applied either. Although the lab-based, individually administered design means classic within-classroom contamination is unlikely, the criterion as written requires either class/school- level randomisation or an explicit tutoring exception, neither of which is documented here. Criterion C is not met because randomisation was conducted at the individual participant level in a laboratory setting without a documented class/school unit or a clearly applicable tutoring exception.
    • E

      Exam-based Assessment

      • Outcomes were measured using a custom-built forced-choice identification test created by the researchers for this study, not a standardised, widely recognised exam.
      • "Forced-choice identification tests designed using Praat (Boersma & Weenink, 2013) were used to measure learners' perceptual accuracy..." (p. 818)
      • Relevant Quotes: 1) "Forced-choice identification tests designed using Praat (Boersma & Weenink, 2013) were used to measure learners' perceptual accuracy, given the importance of using assessment tasks that are compatible with training tasks (Hardison, 2012)." (p. 818) 2) "In the pretest, immediate posttest, and delayed posttest, the learners completed a total of 384 trials, which included trained stimuli spoken in familiar voices (48 words x 4 speakers), trained stimuli spoken in unfamiliar voices (48 words x 2 speakers), and untrained stimuli spoken in familiar voices (24 words x 4 speakers)." (p. 818) Detailed Analysis: The assessment instrument was purpose-built by the researchers using recordings from their own study speakers and Praat software, explicitly designed to mirror the training task's format and stimuli. There is no mention of any nationally or internationally recognised standardised test (e.g., a standardised language proficiency exam) being used to measure the outcome. The test is a custom perceptual-identification instrument tailored specifically to this experiment's target contrasts and word lists. Criterion E is not met because the outcome measure is a custom, researcher-designed test rather than a standardised, widely recognised exam.
    • T

      Term Duration

      • The entire study, from pretest to final (delayed) posttest, spanned only about four weeks, far short of one academic term.
      • "...eight computer-assisted perception training sessions distributed over a 2-week period... After the last training session, the learners completed an immediate posttest and, 2 weeks later, a delayed posttest." (p. 815)
      • Relevant Quotes: 1) "The L2 learners were then randomly assigned to one of four treatment groups or the control group (n = 20 per group) before participating in eight computer-assisted perception training sessions distributed over a 2-week period." (p. 815) 2) "After the last training session, the learners completed an immediate posttest and, 2 weeks later, a delayed posttest." (p. 815) 3) "The L2 participants took the pretest 1 day before the first training session, the immediate posttest on the same day of the last training session, and the delayed posttest 2 weeks after the immediate posttest." (p. 818) Detailed Analysis: From the day before the first training session (pretest) to the delayed posttest is roughly four weeks (2 weeks of training plus a 2-week delay before the final measurement). This is far shorter than the 3-4 month minimum required for a full academic term. There is no indication of any longer follow-up beyond the delayed posttest. Criterion T is not met because the interval from intervention start to final outcome measurement was approximately four weeks, well short of one academic term.
    • D

      Documented Control Group

      • The control group's size, procedure (same training minus CF), and baseline/outcome scores are clearly documented in the text and in Tables 1-3.
      • "Although the control group engaged in the same perception training, there was no CF regardless of the response alternative selected." (p. 817)
      • Relevant Quotes: 1) "Although the control group engaged in the same perception training, there was no CF regardless of the response alternative selected. As soon as a response was made, the next trial began." (p. 817) 2) "In contrast, the next trial was automatically played without any CF in the control group." (p. 817) 3) Table 1 reports, for the "Control" group (n = 20): pretest, immediate posttest, and delayed posttest mean accuracy scores and SDs for both target contrasts (e.g., "Control 66.2 (16.7) 59.7 (15.9) 67.8 (13.8) 60.7 (10.2) 64.7 (12.5) 59.7 (12.8)"). (p. 819) 4) "Participants in this study included 100 Korean learners of L2 English (73 females, 27 males), with a mean age of 30.3 years (SD = 9.69)... their average length of residence in English-speaking countries, including Canada, was 19.2 months (SD = 19.1)." (p. 815, sample-wide characteristics applying across the five equally-sized n = 20 groups, including control) Detailed Analysis: The control group's composition (n = 20), procedure (identical training and testing trials but with no CF at all after incorrect responses), and baseline/ outcome performance (explicit pretest, immediate, and delayed posttest means and SDs in Tables 1-3) are all clearly reported, allowing readers to assess baseline comparability and exactly what the control condition did and did not receive. Criterion D is met because the control group's size, procedure, and baseline/outcome data are documented in sufficient detail.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was conducted at the individual participant level across multiple different institutes, not at the school/institute level.
      • "The L2 learners were then randomly assigned to one of four treatment groups or the control group (n = 20 per group)..." (p. 815)
      • Relevant Quotes: 1) "The L2 learners were then randomly assigned to one of four treatment groups or the control group (n = 20 per group) before participating in eight computer-assisted perception training sessions distributed over a 2-week period." (p. 815) 2) "Most participants were students learning English in private or university language institutes located in the Montreal area." (p. 815) Detailed Analysis: Participants were recruited individually from various private and university language institutes and randomly assigned one-by-one to conditions; there is no description of whole institutes/schools being randomised as a unit. Criterion S is not met because randomisation occurred at the individual level, not the school/institute level.
    • I

      Independent Conduct

      • The same authors who designed the intervention and materials also conducted data collection and analysis themselves, with no independent evaluator involved.
      • "Andrew H. Lee and Roy Lyster, McGill University" (title page)
      • Relevant Quotes: 1) "Andrew H. Lee and Roy Lyster, McGill University" (title page) 2) "This research was supported by a Social Sciences and Humanities Research Council of Canada grant (410-2011-0671) awarded to Roy Lyster." (p. 809) 3) "We are also grateful to the following research assistants, who contributed at various phases of this study: Benjamin Gormley, James Mutch, and Sunyoung Park." (p. 809) Detailed Analysis: The intervention design (CF types), stimuli construction, training software, testing protocol, and data analysis were all carried out by the two authors, assisted only by research assistants under their own supervision. There is no mention anywhere in the paper of an independent or third-party organisation conducting data collection or analysis separate from the study designers. Criterion I is not met because the same research team that designed the intervention also conducted the entire study without independent oversight.
    • Y

      Year Duration

      • Since criterion T (Term Duration) is not met, and the whole study spanned only about four weeks, criterion Y is not met.
      • "...eight computer-assisted perception training sessions distributed over a 2-week period... 2 weeks later, a delayed posttest." (p. 815)
      • Relevant Quotes: 1) "The L2 learners were then randomly assigned to one of four treatment groups or the control group (n = 20 per group) before participating in eight computer-assisted perception training sessions distributed over a 2-week period." (p. 815) 2) "After the last training session, the learners completed an immediate posttest and, 2 weeks later, a delayed posttest." (p. 815) Detailed Analysis: Per the ERCT specification, if criterion T is not met, criterion Y is automatically not met. Independently, the total tracked period (roughly four weeks from pretest to delayed posttest) is nowhere close to 75% of an academic year (~9-10 months). Criterion Y is not met both because criterion T is not met and because the actual duration is far short of a full academic year.
    • B

      Balanced Control Group

      • The presence/type of corrective feedback is the explicit treatment variable being tested, so the no-CF control condition is the appropriate "business-as-usual" baseline rather than an unbalanced resource allocation.
      • "To what extent do the training effects differ according to the type of auditory or visual CF...?" (p. 814)
      • Relevant Quotes: 1) "Participants in the four treatment groups received a specific type of CF when they made perceptual errors during the training sessions..." (p. 815) 2) "Once they selected their answer, a CF intervention followed, after which the next trial began. In contrast, the next trial was automatically played without any CF in the control group." (p. 817) 3) "To what extent do the training effects differ according to the type of auditory or visual CF (i.e., rejection plus target form, rejection plus nontarget form, rejection plus target and nontarget forms, or wrong shown on the computer screen)?" (p. 814) 4) "Overall, the participants in the four CF groups received an average of 100.5 CF instances (SD = 42.6) per training session (i.e., 384 trials)." (p. 819) 5) "Although the control group engaged in the same perception training, there was no CF regardless of the response alternative selected." (p. 817) Detailed Analysis: All five groups completed the identical number of trials (384 per session), the same stimuli, and the same overall training/testing structure; the only difference is whether, and what type of, corrective feedback message followed an incorrect response. The study's second research question is explicitly framed around comparing different CF types against a no-CF control, meaning the presence/absence and type of CF is itself the treatment variable under investigation, not a supplementary add-on resource. The brief CF messages (a few seconds per instance) do not constitute a separate block of extra instructional time, budget, or materials beyond the manipulation itself, so under the ERCT decision tree there are no extra time/budget resources present in the first place. Even under the alternative reading that the CF itself is an "additional resource," the study's own framing shows that CF type is explicitly the treatment variable being tested (RESOURCES_ARE_TREATMENT), which independently satisfies criterion B. Criterion B is met because the differing CF (or its absence) is explicitly the treatment variable under study, and the control condition mirrors all other aspects of the procedure.
  • Level 3 Criteria

    • R

      Reproduced

      • No evidence was found of independent replication of this specific study by a different research team; the authors themselves call for future replication.
      • "...it would be important to replicate this study, controlling the amount of repeated experience with training materials during L2 speech perception training." (p. 826)
      • Relevant Quotes: 1) "Lee and Lyster (2015) demonstrated the effects of CF in classroom-based L2 speech perception training." (p. 812) [an earlier, different study by the same authors, not an independent replication of this one] 2) "...it would be important to replicate this study, controlling the amount of repeated experience with training materials during L2 speech perception training." (p. 826) Detailed Analysis: The paper itself, in its "Conclusion and Future Directions" section, explicitly calls for future replication of this specific study, which indicates that no independent replication existed at the time of writing. An internet search (general web search plus citation-tracking on ResearchGate/Wiley) for later studies citing Lee and Lyster (2016) and attempting to reproduce its specific design (four types of auditory/ visual CF versus a no-CF control on Korean learners' perception of /i/-/I/ and /E/-/ae/) did not locate any independent replication by a different research team in a different context published in a peer-reviewed outlet. Later related work in the CF/speech-perception literature builds on or cites this study but does not replicate it. I could not find evidence of independent reproduction; I am not aware of any such study and did not fabricate one. Criterion R is not met because no independent replication of this specific study was found; the authors explicitly flag replication as future work.
    • A

      All-subject Exams

      • Criterion E (standardised exam) is not met, so criterion A is automatically not met; additionally, only two specific vowel contrasts were assessed rather than a broad set of subjects/skills.
      • "Two English vowel contrasts (/i/-/I/ and /E/-/ae/) were selected as targets." (p. 816)
      • Relevant Quotes: 1) "Two English vowel contrasts (/i/-/I/ and /E/-/ae/) were selected as targets." (p. 816) 2) "Forced-choice identification tests designed using Praat (Boersma & Weenink, 2013) were used to measure learners' perceptual accuracy..." (p. 818) Detailed Analysis: Per the ERCT specification, criterion A requires criterion E to be met as a prerequisite; since E is not met (the assessment is a custom, non-standardised instrument), A cannot be met either. Additionally, the outcome measure targets only two specific phonemic vowel contrasts rather than a broad set of subjects or skills, further reinforcing that the "all-subject" breadth requirement is not satisfied. Criterion A is not met because criterion E is not met and the assessed outcomes are narrowly limited to two vowel contrasts.
    • G

      Graduation Tracking

      • Tracking ended entirely within about a month of the first training session, with no continued follow-up and no graduation tracking; criterion Y is also not met, which further precludes G.
      • "...the delayed posttest 2 weeks after the immediate posttest." (p. 818)
      • Relevant Quotes: 1) "The L2 participants took the pretest 1 day before the first training session, the immediate posttest on the same day of the last training session, and the delayed posttest 2 weeks after the immediate posttest." (p. 818) 2) No further follow-up beyond the delayed posttest is described anywhere in the paper, including the Discussion and Conclusion sections. Detailed Analysis: Per the ERCT specification, if criterion Y is not met, criterion G is automatically not met. Independently, the paper reports no follow-up whatsoever beyond the delayed posttest roughly four weeks after the study began, and there is no cohort of school-age students being tracked toward any graduation milestone in this adult-learner laboratory study; participants are adult language-institute learners (mean age 30.3 years), not students progressing through a K-12 or degree program whose graduation could be tracked. An internet search for subsequent follow-up publications by Lee and Lyster tracking this same cohort of 100 Korean adult learners found no such papers; I could not find evidence of any follow-up graduation-tracking study and did not fabricate one. Criterion G is not met because criterion Y is not met and there is no follow-up beyond a few weeks after the intervention, nor any applicable graduation milestone for this adult-learner sample.
    • P

      Pre-Registered

      • The paper documents open sharing of materials but contains no statement of pre-registration of the study's hypotheses, design, or analysis plan before data collection.
      • Relevant Quotes: 1) "This article has been awarded an Open Materials badge. All materials are publicly accessible in the IRIS digital repository at http://www.iris-database. org." (p. 809) 2) No mention anywhere in the Method, Data Analysis, or other sections of a pre-registration platform, registry ID, or pre-registration date. Detailed Analysis: The paper discloses an Open Materials badge, indicating that study materials were later made publicly available, but this is distinct from pre-registering the study's hypotheses and analysis plan before data collection began. No registry reference (e.g., OSF Registries, ClinicalTrials.gov, AsPredicted) or pre-registration date is provided anywhere in the text. An internet search of standard pre-registration registries (OSF Registries, AsPredicted, ClinicalTrials.gov) for a protocol matching this study's title, authors, or design did not locate any pre-registration entry; I could not find evidence of pre-registration and did not fabricate one. Criterion P is not met because no evidence of pre-registration prior to data collection was found in the paper or in available registries.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.