The Effects of Explicit and Implicit Oral Corrective Feedback on L2 Learning: The Case of That-Trace Filter

Sam Salmi, Mohammad Taghi Farvardin

Published:
ERCT Check Date:
DOI: 10.32038/ltrq.2025.47.03
  • L2 languages
  • adult education
  • Asia
0
  • C

    The Method section explicitly states participants came from "non-randomized, intact classes," directly contradicting the abstract's claim of random assignment, so no credible randomisation is documented.

    "They were from five non-randomized, intact classes at an English institute, with each class consisting of 23 to 25 students."

  • E

    Both outcome measures (the GJT and WPT) were custom instruments created by the researcher specifically for this study, not a widely recognised standardised exam.

    "For the written production test (WPT) test, 32 declarative sentences were developed by the researcher (20 target and 12 distractor items, each containing an underlined NP)."

  • T

    The entire study, from pretest through treatment to posttest, spanned only about two weeks, far short of the one-term minimum required.

    "Data were collected across four sessions, held on three separate days over a two-week period, in a classroom setting."

  • D

    The control group's size, treatment condition (no CF), and baseline scores are clearly reported and statistically confirmed comparable to the other groups.

    "The results of one-way ANOVAs on pre-tests illustrate no statistically significant between-group differences with regard to GJT (F4, 115 = .062, p = .993) and also WPT (F4, 115 = .053, p = .995). Therefore, it can be concluded that the groups began the experiment with almost equal performance."

  • S

    The study was conducted within a single English institute using intact classes, with no school/institute-level randomisation.

    "They were from five non-randomized, intact classes at an English institute, with each class consisting of 23 to 25 students."

  • I

    The same researcher who designed the intervention also personally delivered it as the teacher, with no independent evaluator involved.

    "Second, the researcher worked as the teacher in this study, which may have influenced the effectiveness of the CF."

  • Y

    The study spanned only about two weeks in total, far short of the 75%-of-a-year requirement, and criterion T was also not met.

    "Data were collected across four sessions, held on three separate days over a two-week period, in a classroom setting."

  • B

    No group received extra time, budget, or materials; the only differing element (corrective feedback) is explicitly the treatment variable, with identical activities and duration otherwise across all groups.

    "During the treatment tasks, participants received one type of CF based on their group membership, following any incorrect utterances ... The control group, however, received no CF."

  • R

    No independent replication of this specific study was found in the paper or via internet search; it is a recently published original study with no cited or located reproduction.

  • A

    Criterion E is not met, and the study assessed only a single narrow grammatical structure with non-standardised instruments, so this stronger criterion is not satisfied either.

  • G

    Criterion Y is not met, the study explicitly lacked any delayed or long-term follow-up, and no subsequent publication tracking this cohort toward graduation was found via internet search.

    "First, the time allotted for the study was relatively short, which limited data collection. Given more time, we could have administered a delayed post-test one month after the treatment, potentially yielding more valid results."

  • P

    No statement referencing pre-registration on any registry platform was found in the paper, and none was located via internet search.

Abstract

The purpose of this study was to examine the effectiveness of explicit corrective feedback (CF) strategies (i.e., metalinguistic feedback and explicit correction) versus implicit CF methods (i.e., recasts and explanation questions) in helping English language learners acquire the that-trace filter. To this end, one hundred twenty intermediate English learners were recruited. Each participant was randomly assigned to one of five groups: control, recast, metalinguistic feedback, explicit correction, and clarification request. The participants were given a written production test (WPT) and a grammaticality judgment test (GJT) as pre- and posttests, respectively. The experimental groups underwent two sessions of treatment that included interactive activities. These activities were designed as information-gap tasks that utilized the that-trace filter. Participants received CF that focused on their wrong answers based on the groups they were assigned to. The posttest results indicated that the experimental groups significantly outperformed the control group. When comparing the experimental groups' results on receptive and productive tests, however, no statistically significant differences were found. In addition, a semi-structured interview was conducted with 12 individuals. The findings indicate that teachers' CF has a significant role in enhancing students' acquisition of L2 complex grammar structures.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • The Method section explicitly states participants came from "non-randomized, intact classes," directly contradicting the abstract's claim of random assignment, so no credible randomisation is documented.
      • "They were from five non-randomized, intact classes at an English institute, with each class consisting of 23 to 25 students."
      • Relevant Quotes: 1) "Each participant was randomly assigned to one of five groups: control, recast, metalinguistic feedback, explicit correction, and clarification request." (Abstract, p. 40) 2) "They were from five non-randomized, intact classes at an English institute, with each class consisting of 23 to 25 students." (p. 42) 3) "The participants were divided into five groups: one served as a control group (n = 25), while the other four were given different types of feedback: metalinguistic (n = 23), explicit correction (n = 23), recast (n = 24), and clarification request (n = 25)." (pp. 42-43) Detailed Analysis: Criterion C requires that random assignment to condition be clearly described, with the unit of randomisation being at least the class (not individual students within a single class, unless the tutoring exception applies). The abstract asserts that "each participant was randomly assigned," which on its face reads as student-level randomisation. However, the Method section directly contradicts this, explicitly stating that participants "were from five non-randomized, intact classes." Critically, the five group sizes reported (25, 23, 23, 24, 25) match almost exactly the stated per-class enrolment range of "23 to 25 students," strongly indicating that each of the five pre-existing, intact classes was used wholesale as one of the five study conditions rather than students being individually and randomly allocated across classes. No randomisation procedure (e.g., a random number generator, coin flip, or independent allocation scheme) is described anywhere in the paper for either students or classes. The paper's own terminology, "non-randomized, intact classes," is a direct admission that random assignment was not actually implemented, despite the abstract's claim. This internal contradiction, combined with the absence of any described randomisation mechanism, means the study cannot be verified as a genuine randomised controlled trial at the class level (or any level). This intervention is not a personal one-to-one tutoring intervention, so the tutoring exception in the standard does not apply. Because the paper does not credibly document a randomisation process, and its own text indicates a quasi-experimental, intact-groups design, criterion C is not met.
    • E

      Exam-based Assessment

      • Both outcome measures (the GJT and WPT) were custom instruments created by the researcher specifically for this study, not a widely recognised standardised exam.
      • "For the written production test (WPT) test, 32 declarative sentences were developed by the researcher (20 target and 12 distractor items, each containing an underlined NP)."
      • Relevant Quotes: 1) "The researcher generated 48 pairings of sentences, which were subsequently divided into two sets of 24 pairs (Form A and Form B): one set was designated for the pretest, while the other was designated for the posttest." (Grammaticality judgment test, p. 44) 2) "For the written production test (WPT) test, 32 declarative sentences were developed by the researcher (20 target and 12 distractor items, each containing an underlined NP)." (p. 45) 3) "Their results on the institute-administered Oxford Placement Test indicated that they were proficient at the intermediate level." (p. 42) Detailed Analysis: Criterion E requires that outcomes be measured with a widely recognised, standardised exam rather than an instrument custom-built for the study. Both primary outcome measures here, the Grammaticality Judgment Test (GJT) and the Written Production Test (WPT), were explicitly created by the researcher solely to target the that-trace filter structure under investigation ("generated," "developed by the researcher"). Neither is a recognised, widely used standardised assessment; both were purpose-built, narrow instruments tailored to this specific study's grammatical target. The Oxford Placement Test is mentioned, but only as a screening tool used by the institute to confirm participants' general intermediate proficiency level prior to the study; it was not used to measure the study's outcomes (pre/post gains from the intervention). Therefore, the actual dependent measures used to assess intervention effects were both non-standardised, custom-made tests. Since the outcome instruments were researcher-created rather than standardised, criterion E is not met.
    • T

      Term Duration

      • The entire study, from pretest through treatment to posttest, spanned only about two weeks, far short of the one-term minimum required.
      • "Data were collected across four sessions, held on three separate days over a two-week period, in a classroom setting."
      • Relevant Quotes: 1) "Data were collected across four sessions, held on three separate days over a two-week period, in a classroom setting." (p. 46) 2) "They then attended two treatment sessions on separate days (sessions two and three)." (p. 46) 3) "We administered the WPT and GJT posttests immediately after the second treatment in the third session, each taking nearly 40 minutes." (p. 46) 4) "First, the time allotted for the study was relatively short, which limited data collection. Given more time, we could have administered a delayed post-test one month after the treatment, potentially yielding more valid results." (p. 53-54, Conclusion) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Here, the entire study -- pretest, two treatment sessions, and posttest -- was compressed into "three separate days over a two-week period," with the posttest administered "immediately after the second treatment" in the same week as the last treatment session. There is no delayed measurement; the authors themselves acknowledge in the limitations that a delayed post-test (even just one month later) was not conducted and would have strengthened the study. This interval (well under two weeks from intervention start to outcome measurement) falls far short of the one-term minimum required by the standard, and the authors' own admission confirms no term-long (or even month-long) follow-up occurred. Because the interval between intervention start and outcome measurement is only about two weeks, far short of a full academic term, criterion T is not met.
    • D

      Documented Control Group

      • The control group's size, treatment condition (no CF), and baseline scores are clearly reported and statistically confirmed comparable to the other groups.
      • "The results of one-way ANOVAs on pre-tests illustrate no statistically significant between-group differences with regard to GJT (F4, 115 = .062, p = .993) and also WPT (F4, 115 = .053, p = .995). Therefore, it can be concluded that the groups began the experiment with almost equal performance."
      • Relevant Quotes: 1) "one served as a control group (n = 25), while the other four were given different types of feedback: metalinguistic (n = 23), explicit correction (n = 23), recast (n = 24), and clarification request (n = 25)." (pp. 42-43) 2) "The control group, however, received no CF." (p. 46) 3) "GJ Pretest ... Control ... M = 6.24, SD = 2.204, N = 25" and "WP Pretest ... Control ... M = .72, SD = .980, N = 25" and "GJ Posttest ... Control ... M = 7.16, SD = 2.135, N = 25" and "WP Posttest ... Control ... M = 1.49, SD = 1.734, N = 25" (Table 2, p. 47) 4) "The results of one-way ANOVAs on pre-tests illustrate no statistically significant between-group differences with regard to GJT (F4, 115 = .062, p = .993) and also WPT (F4, 115 = .053, p = .995). Therefore, it can be concluded that the groups began the experiment with almost equal performance." (p. 48) Detailed Analysis: Criterion D requires clear documentation of the control group's size, baseline characteristics, and conditions. The paper specifies the control group's exact size (n = 25) and its treatment condition (it "received no CF" but otherwise underwent the same pretest/posttest procedure as the other groups). Full descriptive statistics (mean, SD, N) for the control group are reported separately for both outcome measures at both pretest and posttest (Table 2), and a one-way ANOVA explicitly confirms the control group's baseline scores did not differ significantly from the other four groups, establishing comparability at baseline. While the paper does not provide separate demographic breakdowns (age, gender) by group, overall sample demographics are given (all male, ages 16-25, M = 19.50 years) and the control group's baseline performance, size, and condition (no feedback) are clearly and explicitly documented, which meets the core intent of this criterion. Because the control group's size, condition, and baseline performance are clearly documented and statistically confirmed comparable to the treatment groups, criterion D is met.
  • Level 2 Criteria

    • S

      School-level RCT

      • The study was conducted within a single English institute using intact classes, with no school/institute-level randomisation.
      • "They were from five non-randomized, intact classes at an English institute, with each class consisting of 23 to 25 students."
      • Relevant Quotes: 1) "They were from five non-randomized, intact classes at an English institute, with each class consisting of 23 to 25 students." (p. 42) 2) "The participants were studying English in an EFL context." (p. 42) Detailed Analysis: Criterion S requires randomisation at the level of the whole school or implementing institution (e.g., multiple schools/institutes randomly assigned to condition), which is stronger than class-level assignment. Here, all 120 participants were drawn from a single English institute, with five intact classes at that one institute each apparently mapped onto one of the five study conditions. There is no randomisation across multiple schools or institutes; the entire study takes place within one institution, and, as established under criterion C, even the class-level allocation was explicitly described as "non-randomized." Since only a single institute was involved and no school/institute-level randomisation is described or possible with this single-site design, criterion S is not met.
    • I

      Independent Conduct

      • The same researcher who designed the intervention also personally delivered it as the teacher, with no independent evaluator involved.
      • "Second, the researcher worked as the teacher in this study, which may have influenced the effectiveness of the CF."
      • Relevant Quotes: 1) "Second, the researcher worked as the teacher in this study, which may have influenced the effectiveness of the CF. This raises a possible limitation regarding the study's pedagogical implications." (p. 53-54) 2) "Nevertheless, such a researcher-as-teacher arrangement was necessary since the researcher designed the tasks, knew how to provide the required feedback for the experimental groups, and was more familiar with the instructional materials than anyone else." (p. 54) 3) "In each treatment session, the learners participated in interactional activities with the researcher for about 30 minutes." (p. 46) Detailed Analysis: Criterion I requires that the study be conducted independently from the people who designed the intervention, typically via a third-party evaluator/implementer, to reduce bias in delivery and analysis. Here, the paper explicitly states that "the researcher worked as the teacher," meaning the same person who designed the tasks, the CF strategies, and the study also personally delivered all corrective feedback to participants and (as author) analysed the results. There is no external evaluator, independent data collector, or blinded administrator described anywhere in the study. The authors themselves acknowledge this as a limitation, confirming there was no independent conduct: the same individual designed, delivered, and evaluated the intervention. Because the researcher personally designed and delivered the intervention as the teacher, with no independent third-party involvement, criterion I is not met.
    • Y

      Year Duration

      • The study spanned only about two weeks in total, far short of the 75%-of-a-year requirement, and criterion T was also not met.
      • "Data were collected across four sessions, held on three separate days over a two-week period, in a classroom setting."
      • Relevant Quotes: 1) "Data were collected across four sessions, held on three separate days over a two-week period, in a classroom setting." (p. 46) 2) "First, the time allotted for the study was relatively short, which limited data collection. Given more time, we could have administered a delayed post-test one month after the treatment, potentially yielding more valid results." (p. 53-54) Detailed Analysis: Per the criterion-specific instructions, if criterion T (Term Duration) is not met, criterion Y is automatically not met. As established above, the entire study spanned only about two weeks from pretest to posttest, nowhere near the 75% of an academic year (approximately 9-10 months) required by this stronger criterion. The authors' own limitations section confirms that even a one-month delayed posttest was not feasible within the study's timeframe, let alone year-long tracking. Because criterion T is not met and the study duration (about two weeks) is drastically shorter than the required academic-year benchmark, criterion Y is not met.
    • B

      Balanced Control Group

      • No group received extra time, budget, or materials; the only differing element (corrective feedback) is explicitly the treatment variable, with identical activities and duration otherwise across all groups.
      • "During the treatment tasks, participants received one type of CF based on their group membership, following any incorrect utterances ... The control group, however, received no CF."
      • Relevant Quotes: 1) "They then attended two treatment sessions on separate days (sessions two and three). In each treatment session, the learners participated in interactional activities with the researcher for about 30 minutes." (p. 46) 2) "During the treatment tasks, participants received one type of CF based on their group membership, following any incorrect utterances. For each experimental group, only one type of CF was provided. ... The control group, however, received no CF." (p. 46) 3) "To this end, the following research questions were developed: RQ1: Does implicit corrective feedback ... significantly affect EFL learners' that-trace filter learning? RQ2: Does explicit corrective feedback ... significantly affect EFL learners' that-trace filter learning?" (p. 43) Detailed Analysis: Applying the criterion B decision procedure: no group received any extra instructional time, materials, or budget relative to any other group. All five groups (control included) attended the identical pretest and the same two 30-minute treatment sessions built around the same information-gap activities; the only documented difference between groups is the presence or type of corrective feedback delivered on incorrect utterances during those otherwise-identical sessions. Since EXTRA_RESOURCES_PRESENT (additional time, budget, or materials) is not the case here, the decision tree's first branch ("no extra time/budget at all") is satisfied on its own. Independently, the sole differing element (CF itself) is also explicitly the treatment variable under investigation, as stated directly in the research questions ("Does implicit/explicit corrective feedback ... significantly affect EFL learners' that-trace filter learning?"), which is also sufficient for the criterion under the "resources are the treatment" branch. The control group's "business as usual" condition (identical activity, identical duration, no feedback) is the appropriate baseline for isolating CF's effect, and instructional time/task exposure is matched across all five groups. Because no group received additional time, budget, or materials relative to any other group, and the sole differing element (CF) is explicitly the treatment variable being tested, criterion B is met.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study was found in the paper or via internet search; it is a recently published original study with no cited or located reproduction.
      • Relevant Quotes: 1) "Salmi, S., & Farvardin, M. T. (2025). The effects of explicit and implicit oral corrective feedback on L2 learning: The case of that-trace filter. Language Teaching Research Quarterly, 47, 40-56." (Citation, p. 40) No quotes describing an independent replication of this specific study were found in the paper. Internet Search for Independent Replication: A search was conducted for independent replications of this specific study (the explicit-vs-implicit oral CF comparison on the that-trace filter with intermediate EFL learners at a single Iranian institute). One tangentially related work was located: a paper describing itself as a "conceptual replication of Goo (2012)" on corrective feedback and phonological short-term memory in acquiring the that-trace filter. However, this work replicates a different, earlier (2012) original study and was itself published in 2014, over a decade before Salmi and Farvardin (2025), so it cannot constitute a replication of this paper. A Google Scholar search for citations of this specific 2025 paper returned no results. No paper replicating this specific study by a different research team was found. Detailed Analysis: Criterion R requires that this specific study (its central experimental comparison of explicit versus implicit CF for the that-trace filter) be independently replicated by a different research team in a peer-reviewed outlet. This paper was published in 2025 and, being newly published, does not reference any prior replication of itself (as would be expected, since it is the original study). No evidence of an independent reproduction of this specific experiment/design was found within the document itself or via internet search. Given the very recent publication date and the absence of any citation trail indicating a separate team has since replicated this exact study (same target structure, same CF comparisons, same population), there is no basis to conclude this criterion is satisfied. Because no independent replication of this specific study was found in the paper or via internet search, criterion R is not met.
    • A

      All-subject Exams

      • Criterion E is not met, and the study assessed only a single narrow grammatical structure with non-standardised instruments, so this stronger criterion is not satisfied either.
      • Relevant Quotes: 1) "The purpose of this study was to examine the effectiveness of explicit corrective feedback (CF) strategies ... in helping English language learners acquire the that-trace filter." (Abstract, p. 40) 2) "The participants were given a written production test (WPT) and a grammaticality judgment test (GJT) as pre- and posttests, respectively." (Abstract, p. 40) Detailed Analysis: Per the criterion-specific instructions, criterion A is automatically not met if criterion E (Exam-based Assessment) is not met, since only standardised exam-based assessments count toward this stronger criterion. As established above, both outcome instruments (GJT, WPT) are researcher-made, non-standardised tests, so criterion E already fails. Independently of that dependency, the study also measures only a single, narrow grammatical structure (the that-trace filter) within one subject domain (L2 English grammar), rather than assessing outcomes across multiple core subjects. Because criterion E is not met and only one narrow construct within a single domain was assessed with non-standardised instruments, criterion A is not met.
    • G

      Graduation Tracking

      • Criterion Y is not met, the study explicitly lacked any delayed or long-term follow-up, and no subsequent publication tracking this cohort toward graduation was found via internet search.
      • "First, the time allotted for the study was relatively short, which limited data collection. Given more time, we could have administered a delayed post-test one month after the treatment, potentially yielding more valid results."
      • Relevant Quotes: 1) "We administered the WPT and GJT posttests immediately after the second treatment in the third session ... In the fourth session, the semi-structured interview was carried out on 12 participants from the experimental groups." (p. 46) 2) "First, the time allotted for the study was relatively short, which limited data collection. Given more time, we could have administered a delayed post-test one month after the treatment, potentially yielding more valid results." (p. 53-54) Internet Search for Follow-up Publications: A search was conducted for subsequent papers by Salmi and/or Farvardin tracking this same cohort of 120 EFL learners toward course completion or graduation. No such follow-up publication was found; the study is newly published (2025) and no later studies by these authors extending this cohort's tracking were located. Detailed Analysis: Per the criterion-specific instructions, criterion G is automatically not met if criterion Y (Year Duration) is not met, which is the case here. Beyond that dependency, the study's own data collection timeline confirms there was no long-term follow-up of any kind: measurement ended with the posttest and interview in the same two-week window as the intervention, and the authors explicitly flag the absence of even a one-month delayed posttest as a limitation. No follow-up publication tracking this cohort toward graduation or course completion was found via internet search, and there is no mention of any planned follow-up in the paper itself. Because criterion Y is not met and no long-term or graduation-tracking follow-up was conducted, planned, or found in subsequent publications, criterion G is not met.
    • P

      Pre-Registered

      • No statement referencing pre-registration on any registry platform was found in the paper, and none was located via internet search.
      • Relevant Quotes: 1) "Competing Interests: No, there are no conflicting interests." (p. 54) 2) "Funding: Not applicable." (p. 54) 3) "Acknowledgements: Not applicable." (p. 54) No statement referencing a study registry, protocol pre-registration, or registration date was found anywhere in the paper. Internet Search for Pre-Registration: A search was conducted for a pre-registration of this study's protocol (e.g., on OSF or a comparable public registry) under the authors' names and the study title/topic. No registry entry was found. Detailed Analysis: Criterion P requires evidence that the full study protocol (hypotheses, methods, planned analyses) was registered on a public registry before data collection began. The paper includes standard end-matter sections (ORCID, Acknowledgements, Funding, Ethics Declarations, Competing Interests, Rights and Permissions) but none of these, nor any part of the Method section, mentions a pre-registration platform, registry ID, or registration date. No such statement could be located in the text, and no registry entry was found via internet search. Because no reference to pre-registration of the study protocol was found anywhere in the paper or online, criterion P is not met.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.