The Effect of Focused Written Corrective Feedback and Language Aptitude on ESL Learners' Acquisition of Articles

Younghee Sheen

Published:
ERCT Check Date:
DOI: 10.1002/j.1545-7249.2007.tb00059.x
  • L2 languages
  • adult education
  • US
0
  • C

    The paper explicitly describes a quasi-experimental design using intact classes assigned to conditions, not a randomised controlled trial.

    "This study used a quasi-experimental research design with a pretest-treatment-posttest-delayed posttest structure, using intact ESL classrooms." (p. 260-261)

  • E

    All outcome measures (dictation, writing, error correction, language analysis tests) were custom instruments built or adapted by the researcher, not standardised, widely recognised exams.

    "This test was adapted from one of Muranoi's (2000) test instruments for English articles." (p. 266)

  • T

    The full study, from intervention start to the delayed posttest, spanned only about two months, short of a full academic term.

    "During the 2-month period of this study, the teacher did not explicitly teach or correct articles outside the treatment." (p. 275)

  • D

    The control group's composition, size, baseline scores, and lack of any feedback treatment are clearly documented, and pretest scores show no significant baseline differences between groups.

    "The control group completed the tests only. It did not perform the narrative tasks and did not receive any feedback but instead followed normal classes." (p. 262)

  • S

    Randomisation was not conducted at all (the design is explicitly quasi-experimental with intact classes), so the stronger school-level requirement is also not satisfied.

    "This study used a quasi-experimental research design with a pretest-treatment-posttest-delayed posttest structure, using intact ESL classrooms." (p. 260-261)

  • I

    The same researcher who designed the study also served as the sole corrector of student errors and analysed the results, with no independent third- party evaluator involved.

    "Given the amount of time and labour involved, it was decided that the researcher would serve as the corrector." (p. 264)

  • Y

    Since criterion T (Term Duration) is not met, this stronger criterion is automatically not met; the study duration was about two months, far short of a full academic year.

    "During the 2-month period of this study, the teacher did not explicitly teach or correct articles outside the treatment." (p. 275)

  • B

    The narrative-writing and feedback sessions are integral to delivering the corrective-feedback treatment being tested, so the control group's "business as usual" classes (with no narrative tasks or feedback) represent an acceptable baseline rather than an unbalanced comparison.

    "The control group completed the tests only. It did not perform the narrative tasks and did not receive any feedback but instead followed normal classes." (p. 262)

  • R

    No independent replication of this specific study (same design comparing direct-only vs. direct metalinguistic feedback on article acquisition with language aptitude) was found in the paper or in a subsequent literature search.

  • A

    Since criterion E (Exam-based Assessment) is not met, this stronger criterion is automatically not met; only article usage was assessed, not all main subjects.

  • G

    Since criterion Y (Year Duration) is not met, this stronger criterion is automatically not met; there is no tracking of students beyond the delayed posttest a few weeks after treatment.

  • P

    The paper contains no statement of pre-registration of the study protocol, hypotheses, or analysis plan prior to data collection.

Abstract

This study examines the differential effect of two types of written corrective feedback (CF) and the extent to which language analytic ability mediates the effects of CF on the acquisition of articles by adult intermediate ESL learners of various L1 backgrounds (N = 91). Three groups were formed: a direct-only correction group, a direct metalinguistic correction group, and a control group. The study found that both treatment groups performed much better than the control group on the immediate posttests, but the direct metalinguistic group performed better than the direct-only correction group in the delayed posttests. It also found a significantly positive association between students' gains and their aptitude for language analysis. Moreover, language analytic ability was more strongly related to acquisition in the direct metalinguistic group than in the direct-only group. The results showed that written CF targeting a single linguistic feature improved learners' accuracy, especially when metalinguistic feedback was provided and the learners had high language analytic ability.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • The paper explicitly describes a quasi-experimental design using intact classes assigned to conditions, not a randomised controlled trial.
      • "This study used a quasi-experimental research design with a pretest-treatment-posttest-delayed posttest structure, using intact ESL classrooms." (p. 260-261)
      • Relevant Quotes: 1) "This study used a quasi-experimental research design with a pretest-treatment-posttest-delayed posttest structure, using intact ESL classrooms." (p. 260-261) 2) "The current study was conducted in six intact classrooms in the American Language Program (ALP) of a community college in the United States." (p. 261) 3) "Out of six intact classrooms, three groups were formed: the direct-only correction group (n = 31), direct metalinguistic group (n = 32), and the control group (n = 28)." (p. 261) Detailed Analysis: The ERCT Standard requires that the study be a genuine Randomised Controlled Trial, with a clearly described randomisation process, at minimum at the class level. The author explicitly labels the design "quasi-experimental" and states that it used "intact ESL classrooms," meaning existing class groups were assigned to conditions rather than classes being randomly allocated to treatment or control. No quote anywhere in the Method section describes a random-number-based or lottery-based assignment procedure for the six classrooms; the three conditions were simply formed "out of six intact classrooms." There is also no exception here for one-to-one tutoring, since this is a classroom- based intervention. Because the paper itself disclaims random assignment and instead reports a quasi-experimental design based on pre-existing intact classes, criterion C is not met.
    • E

      Exam-based Assessment

      • All outcome measures (dictation, writing, error correction, language analysis tests) were custom instruments built or adapted by the researcher, not standardised, widely recognised exams.
      • "This test was adapted from one of Muranoi's (2000) test instruments for English articles." (p. 266)
      • Relevant Quotes: 1) "During each testing session, three subtests were administered: a speeded dictation test, a writing test, and an error correction test." (p. 261) 2) "This test consisted of 14 items, each of which contained one or two sentences involving the use of indefinite and definite articles." (p. 265, Speeded Dictation Test) 3) "This test was adapted from one of Muranoi's (2000) test instruments for English articles. It consisted of four sequential pictures; the students were asked to write one coherent story based on them." (p. 266, Writing Test) 4) "This test consisted of 17 items ... The items were adapted from test instruments used in Liu and Gleason (2002) and Muranoi (2000)." (p. 267, Error Correction Test) 5) "The instrument was based on a language analysis test developed by Ottó and used previously by Schmitt, Dörnyei, Adolphs, and Durow (2003)." (p. 267, Language Analytic Ability Test) Detailed Analysis: Criterion E requires a widely recognised, standardised exam-based assessment rather than an instrument built specifically for the study. Every outcome measure used here (speeded dictation test, narrative writing test, error correction test) was constructed or adapted by the researcher specifically to elicit article usage for this study, drawing on prior research instruments (e.g. Muranoi, 2000; Liu & Gleason, 2002) rather than on any nationally or internationally recognised standardised exam. There is no mention of a standardised, externally validated language proficiency or achievement test being used as the dependent measure. The language analytic ability test is likewise a researcher-administered aptitude instrument, not an exam of educational achievement. Because none of the assessments used to measure the outcome (article acquisition) are standardised, widely recognised exams, criterion E is not met.
    • T

      Term Duration

      • The full study, from intervention start to the delayed posttest, spanned only about two months, short of a full academic term.
      • "During the 2-month period of this study, the teacher did not explicitly teach or correct articles outside the treatment." (p. 275)
      • Relevant Quotes: 1) "Two weeks prior to the start of the corrective feedback treatments, the participating students took a language analytic ability test. In the following week, they completed the pretests. The immediate posttests were completed following the two CF sessions and the delayed posttests 3 to 4 weeks later." (p. 261) 2) "There was a one-week interval between the first and second narrative tasks." (p. 263) 3) "During the 2-month period of this study, the teacher did not explicitly teach or correct articles outside the treatment." (p. 275) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Here, the treatment itself consisted of only two narrative-writing sessions one week apart, followed by an immediate posttest and a delayed posttest administered "3 to 4 weeks later." The author explicitly characterises the entire study window as a "2-month period," which is shorter than the minimum one-term threshold the ERCT standard requires. There is no indication that outcome measurement extended into a second academic term. Because the interval from intervention start to final (delayed) measurement was only about two months, shorter than one academic term, criterion T is not met.
    • D

      Documented Control Group

      • The control group's composition, size, baseline scores, and lack of any feedback treatment are clearly documented, and pretest scores show no significant baseline differences between groups.
      • "The control group completed the tests only. It did not perform the narrative tasks and did not receive any feedback but instead followed normal classes." (p. 262)
      • Relevant Quotes: 1) "The control group completed the tests only. It did not perform the narrative tasks and did not receive any feedback but instead followed normal classes." (p. 262) 2) "Out of six intact classrooms, three groups were formed: the direct-only correction group (n = 31), direct metalinguistic group (n = 32), and the control group (n = 28)." (p. 261) 3) "A one-way ANOVA showed no statistically significant group differences in the pretest total scores among the three groups, F(2, 88) = 1.23, ns." (p. 269) 4) Table 1 reports pretest M and SD for the control group (M = 48.3, SD = 14.2) alongside the two treatment groups. (Table 1, p. 269) Detailed Analysis: Criterion D requires clear documentation of the control group's size, characteristics, baseline performance, and treatment conditions. The paper specifies the control group's exact size (n = 28), describes precisely what the control group did and did not receive ("completed the tests only," "followed normal classes"), and reports its descriptive baseline statistics in Table 1 alongside confirmation via ANOVA that the three groups did not differ significantly at pretest. This level of detail allows readers to assess baseline comparability and the nature of the control condition. Because the control group's size, baseline scores, and conditions are all clearly documented, criterion D is met.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was not conducted at all (the design is explicitly quasi-experimental with intact classes), so the stronger school-level requirement is also not satisfied.
      • "This study used a quasi-experimental research design with a pretest-treatment-posttest-delayed posttest structure, using intact ESL classrooms." (p. 260-261)
      • Relevant Quotes: 1) "This study used a quasi-experimental research design ... using intact ESL classrooms." (p. 260-261) 2) "The current study was conducted in six intact classrooms in the American Language Program (ALP) of a community college in the United States." (p. 261) Detailed Analysis: Criterion S requires random assignment at the school (or equivalent institutional) level, which is a stronger requirement than criterion C. Since the study used a single institution (one community college's language program) with intact classes assigned to conditions in a quasi-experimental (non- randomised) design, there is no school-level randomisation, nor any randomisation at all. Because there is no evidence of randomisation at any level, let alone the school level, criterion S is not met.
    • I

      Independent Conduct

      • The same researcher who designed the study also served as the sole corrector of student errors and analysed the results, with no independent third- party evaluator involved.
      • "Given the amount of time and labour involved, it was decided that the researcher would serve as the corrector." (p. 264)
      • Relevant Quotes: 1) "Given the amount of time and labour involved, it was decided that the researcher would serve as the corrector. The researcher corrected all the article errors in the learners' narratives." (p. 264) 2) "The teacher collected the students' written narratives which were then handed to the researcher. The researcher corrected the narratives focusing mainly on article errors based on the correction guidelines." (p. 264) 3) "Prior to the current study, the researcher visited the site many times and observed and piloted a number of instruments in several classes at different levels." (p. 261) 4) Acknowledgments thank faculty and a research assistant for "data coding" but do not describe any independent agency conducting or analysing the trial. (p. 278, Acknowledgments) Detailed Analysis: Criterion I requires that the study be conducted independently from those who designed the intervention, typically via an external evaluator or agency handling implementation, data collection, or analysis. In this study, the sole author designed the corrective feedback treatments, designed the testing instruments, personally corrected all student errors according to her own guidelines, and analysed the results. There is no mention of any independent evaluator, agency, or blinded administrator being responsible for delivering treatment, scoring, or analysis; a second researcher only cross-checked a 25% reliability sample, which does not amount to independent conduct of the study as a whole. Because the same researcher designed, implemented (as corrector), and analysed the study without independent third-party conduct, criterion I is not met.
    • Y

      Year Duration

      • Since criterion T (Term Duration) is not met, this stronger criterion is automatically not met; the study duration was about two months, far short of a full academic year.
      • "During the 2-month period of this study, the teacher did not explicitly teach or correct articles outside the treatment." (p. 275)
      • Relevant Quotes: 1) "During the 2-month period of this study, the teacher did not explicitly teach or correct articles outside the treatment." (p. 275) 2) "The immediate posttests were completed following the two CF sessions and the delayed posttests 3 to 4 weeks later." (p. 261) Detailed Analysis: Criterion Y requires that outcomes be tracked for at least 75% of a full academic year (roughly 9-10 months). Per the criterion-specific instruction, if criterion T (Term Duration) is not met, then Y is automatically not met. As established above under criterion T, the entire study spanned only about two months from pretest to delayed posttest, which is far shorter than even one academic term, let alone three-quarters of an academic year. Because criterion T is not met and the observed study duration (about two months) is far below the year-long threshold, criterion Y is not met.
    • B

      Balanced Control Group

      • The narrative-writing and feedback sessions are integral to delivering the corrective-feedback treatment being tested, so the control group's "business as usual" classes (with no narrative tasks or feedback) represent an acceptable baseline rather than an unbalanced comparison.
      • "The control group completed the tests only. It did not perform the narrative tasks and did not receive any feedback but instead followed normal classes." (p. 262)
      • Relevant Quotes: 1) "The two experimental groups completed the treatments and tests. The control group completed the tests only. It did not perform the narrative tasks and did not receive any feedback but instead followed normal classes." (p. 262) 2) "The students in all three groups were of the same level of proficiency and received the same amount and type of instruction involving identical writing and reading materials." (p. 275) 3) "This study constitutes an attempt to address some of the perceived problems in written CF research ... The study considers the following research questions: 1. Does focused written corrective feedback have an effect on intermediate ESL learners' acquisition of English articles?" (p. 260) Detailed Analysis: Under criterion B, additional time or resources given only to the intervention group do not violate the balance requirement when those resources are themselves the explicit treatment variable under study, with the control group representing a standard "business as usual" baseline. Here, the two narrative-writing/feedback sessions are the mechanism by which the treatment variable (written corrective feedback) is delivered; the study's central research question is explicitly whether this CF has an effect, so the additional sessions are integral to, not separable from, the intervention being tested. The control group's lack of narrative tasks and feedback is therefore the intended business-as-usual comparison, not an incidental confound of an unrelated resource. Importantly, the paper also documents that regular classroom instruction, materials, and proficiency level were identical across all three groups outside of the treatment sessions, which supports that no other unrelated resource advantage was given to the treatment groups. The authors do acknowledge a general "test practice effect" present across all groups from repeated testing, but attribute the differential gains specifically to the CF treatment over and above that shared practice effect. Applying the criterion-B decision tree: extra time and attention are present (EXTRA_RESOURCES_PRESENT), and this is not a negligible difference, but these resources are explicitly the treatment variable being tested (RESOURCES_ARE_TREATMENT), so the criterion is met regardless of whether the control group received a matching resource, provided (as here) the study frames this clearly as its intent. Because the additional treatment sessions are integral to the CF treatment variable being tested, and non-treatment instruction was reported as identical across groups, criterion B is met.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study (same design comparing direct-only vs. direct metalinguistic feedback on article acquisition with language aptitude) was found in the paper or in a subsequent literature search.
      • Relevant Quotes: 1) No quote in the paper itself references a prior or planned replication of this specific study design, as this is the original 2007 publication. Detailed Analysis: Criterion R requires independent replication of this specific study by a different research team in a different context, published in a peer-reviewed outlet. An internet search for replications of Sheen (2007) confirms that the paper is very widely cited in the written corrective feedback literature, with later studies (e.g., on unfocused written corrective feedback for countable/uncountable nouns, and on feedback explicitness in L2 morpheme acquisition) building on its methodology and citing it as background. However, none of these studies replicate this specific three-group design (direct-only vs. direct metalinguistic correction vs. control) on English article acquisition with the same language analytic ability mediation focus; they instead extend the paradigm to new linguistic targets or feedback modalities. No study was identified, including in a search repeated as part of this verification, that reproduces this specific study's design and findings under a different research team. Because no independent replication of this specific study was identified, criterion R is not met.
    • A

      All-subject Exams

      • Since criterion E (Exam-based Assessment) is not met, this stronger criterion is automatically not met; only article usage was assessed, not all main subjects.
      • Relevant Quotes: 1) "During each testing session, three subtests were administered: a speeded dictation test, a writing test, and an error correction test." (p. 261) Detailed Analysis: Per the criterion-specific instruction, criterion A requires criterion E to be met as a prerequisite; since E was not met (all instruments were researcher-built, non-standardised measures), A cannot be met either. Separately, the study measured only English article usage, a narrow grammatical target within L2 English, and did not assess performance across other main academic subjects. Because criterion E is not met and only a single narrow linguistic target was assessed, criterion A is not met.
    • G

      Graduation Tracking

      • Since criterion Y (Year Duration) is not met, this stronger criterion is automatically not met; there is no tracking of students beyond the delayed posttest a few weeks after treatment.
      • Relevant Quotes: 1) "The immediate posttests were completed following the two CF sessions and the delayed posttests 3 to 4 weeks later." (p. 261) 2) No mention anywhere in the paper of any further follow-up beyond the delayed posttest, nor of any subsequent follow-up publication tracking the same cohort. Detailed Analysis: Per the criterion-specific instruction, if criterion Y (Year Duration) is not met, criterion G is automatically not met. As established above, the study's total duration was only about two months, and data collection ended at the delayed posttest 3 to 4 weeks after the immediate posttest, with no further tracking of students' progress, let alone through to graduation from their ESL program. A search for subsequent publications by Younghee Sheen tracking this same cohort of intermediate ESL learners toward graduation did not identify any such follow-up study. Because criterion Y is not met and no graduation- level tracking occurred, criterion G is not met.
    • P

      Pre-Registered

      • The paper contains no statement of pre-registration of the study protocol, hypotheses, or analysis plan prior to data collection.
      • Relevant Quotes: 1) No quote anywhere in the Method, Analysis, or Acknowledgments sections mentions a registry platform, a pre-registration identifier, or a date of pre-registration. Detailed Analysis: Criterion P requires quoted evidence of a published protocol, including hypotheses and planned analyses, registered before data collection began. This 2007 paper predates the widespread adoption of pre- registration practice in applied linguistics/SLA research, and no reference to any registry or protocol pre-registration appears anywhere in the text. No pre-registration record for this study was located in a search of trial/protocol registries. Because there is no evidence of pre-registration, criterion P is not met.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.