Abstract
Truscott [Truscott, J., 1996. The case against grammar correction in L2 writing classes. Language Learning 46, 327-369; Truscott, J., 1999. The case for ''the case for grammar correction in L2 writing classes'': a response to Ferris. Journal of Second Language Writing 8, 111-122] laid down the challenge to teacher educators and teachers to justify their faith in written corrective feedback (CF) with hard evidence from studies that have investigated its effects on subsequent writing. The study reported in this article set out to provide evidence that CF is effective in an EFL context. Using a pre-test- immediate post-test-delayed post-test design, it compared the effects of focused and unfocused written CF on the accuracy with which Japanese university students used the English indefinite and definite articles to denote first and anaphoric reference in written narratives. The focused group received correction of just article errors on three written narratives while the unfocused group received correction of article errors alongside corrections of other errors. Both groups gained from pre-test to post-tests on both an error correction test and on a test involving a new piece of narrative writing and also outperformed a control group, which received no correction, on the second post-test. The CF was equally effective for the focused and unfocused groups. This study, together with a few other recent studies, indicates that written CF is effective, at least where English articles are concerned, and thus strengthens the case for teachers providing written CF.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- The study explicitly used a quasi-experimental design with pre-existing intact classes rather than random assignment to conditions.
- "The study used a quasi-experimental design involving intact classes serving as two experimental groups - focused CF (N = 18), unfocused CF (N = 18) - and a control group (N = 13)." (p. 357)
Relevant Quotes:
1) "The study used a quasi-experimental design involving intact classes serving as two experimental groups - focused CF (N = 18), unfocused CF (N = 18) - and a control group (N = 13)." (Section 4.1, p. 357)
2) "The participants were 49 students enrolled in general English classes in a national university in Japan." (Section 4.2, p. 357)
3) "The students in the focussed group were all male and were studying aviation or technology as their major. The students in the unfocussed group were predominantly male (only two females) and were studying industrial design as their major. The students in the control group were between 19 and 20, were in their second year of study at the university and were studying agriculture as their major." (pp. 357-358)
4) "Due to timetabling constraints, separate English classes were held for students from different departments. The experimental groups were taking a reading class... The control group was an oral communication class." (Section 4.3, p. 358)
Detailed Analysis:
The ERCT 'C' criterion requires a clearly described randomisation process at the class level or stronger. Here the authors explicitly label the design "quasi-experimental" and state that it involved "intact classes" rather than randomly constituted groups. The three groups differ systematically in ways that confirm no randomisation occurred: they are drawn from different academic majors (aviation/technology, industrial design, agriculture), different course types (reading vs. oral communication), and even different year-groups (first-year experimental groups vs. second-year control group). No quote anywhere in the paper describes a random allocation procedure of classes, schools, or students to conditions - allocation was determined entirely by which pre-existing class students happened to be enrolled in. The tutoring/ personal-teaching exception does not apply, as this is classroom-based written feedback, not one-to-one tutoring.
Criterion C is not met because the study used non-randomised, pre-existing intact classes rather than a randomised allocation to conditions.
-
E
Exam-based Assessment
- Outcomes were measured with researcher-designed narrative writing tasks and a custom error correction test, not standardised exams.
- "This was a version of the testing instrument used in Sheen (2007). It consisted of 16 items, each containing two related statements, one of which was underlined." (p. 360)
Relevant Quotes:
1) "Three different picture compositions were used, all taken from Byrne, 1967... For each test, students were given one of the stories and a sheet of paper and asked to write a title for the story and a detailed story." (Section 4.6.1, p. 359)
2) "Writing test scores were calculated by means of obligatory occasion analysis (Ellis and Barkhuizen, 2005)... An accuracy score was then calculated for each learner by dividing the total number of correctly supplied articles by the total number of obligatory occasions." (Section 4.10, p. 361)
3) "This was a version of the testing instrument used in Sheen (2007). It consisted of 16 items, each containing two related statements, one of which was underlined. The underlined sentence contained an error." (Section 4.7, p. 360)
Detailed Analysis:
Neither outcome measure is a standardised, widely-recognised exam. The narrative writing tests are researcher-devised picture-composition tasks scored via a bespoke "obligatory occasion analysis" created specifically to count article accuracy. The error correction test is an adapted version of an instrument used in a prior study by one of the same research groups (Sheen, 2007), not a standard, externally validated assessment. Both instruments were purpose-built to measure the very structure targeted by the intervention, which is precisely the kind of custom, intervention-aligned assessment the E criterion is designed to exclude.
Criterion E is not met because both outcome measures are custom-built research instruments rather than standardised, widely recognised exams.
-
T
Term Duration
- From the start of the intervention to the final measurement was about eight to nine weeks, short of a full academic term.
- "The entire study was spread over a period of 10 weeks." (p. 360)
Relevant Quotes:
1) "Week 1: Writing pre-test: error correction pre-test... Week 2: Task 1... Week 6: Feedback on task 3: exit questionnaire: writing post-test 1: error correction post-test... Week 10: Writing post-test 2." (Table 1, p. 360)
2) "The entire study was spread over a period of 10 weeks. There was a gap of 3 weeks between the Writing Post-test 1 and the Writing Post-test 2..." (Section 4.9, p. 360)
3) "The delayed post-test was administered approximately 4 weeks later." (Section 4.6.1, p. 360)
Detailed Analysis:
Treatment began in Week 2 (Task 1) and the final outcome measure (delayed writing post-test 2) was administered in Week 10, an interval of roughly eight weeks (about two months) from intervention start to final measurement. This is well short of the one full academic term (approximately 3-4 months) required by the standard. The overall study window of 10 weeks likewise falls short of a term.
Criterion T is not met because the interval from intervention start to final outcome measurement was only about eight weeks, shorter than one academic term.
-
D
Documented Control Group
- The control group's demographics, course context, baseline scores, and exact procedure are described in detail.
- "The students in the control group were between 19 and 20, were in their second year of study at the university and were studying agriculture as their major. Six were female and 11 male." (pp. 357-358)
Relevant Quotes:
1) "The students in the control group were between 19 and 20, were in their second year of study at the university and were studying agriculture as their major. Six were female and 11 male. They had completed two previous English classes at the university." (Section 4.2, pp. 357-358)
2) "The control group was an oral communication class. The aim of this class was to foster communicative ability by means of task-based instruction. Care was taken to ensure that, during the period of study, no explicit attention was paid to articles." (Section 4.3, p. 358)
3) "The procedure for the control group was the same except that the students received no corrections of their linguistic errors. Instead they received a simple general comment or question (e.g. 'Good!', 'Are they happy then?', 'What happened then?')." (Section 4.4, p. 359)
4) "Control 11 0.588 0.238 0.740 0.218 0.663 0.187" and "Control 13 4.385 3.453 13 5.385 3.990" (Tables 3 and 4, pp. 362-363)
Detailed Analysis:
The paper provides a clear demographic profile of the control group (age, gender split, year of study, major, prior English classes), an explicit description of what course they were otherwise enrolled in, a precise description of exactly what they received instead of correction, and baseline (pre-test) performance figures in Tables 3 and 4. This level of detail satisfies the documentation requirement.
Criterion D is met because the control group's composition, baseline scores, and treatment (no correction, general comments) are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- There was no randomisation at any level - the study compared pre-existing intact classes, not randomised schools.
- "The study used a quasi-experimental design involving intact classes serving as two experimental groups... and a control group." (p. 357)
Relevant Quotes:
1) "The study used a quasi-experimental design involving intact classes serving as two experimental groups - focused CF (N = 18), unfocused CF (N = 18) - and a control group (N = 13)." (Section 4.1, p. 357)
2) "The participants were 49 students enrolled in general English classes in a national university in Japan." (Section 4.2, p. 357)
Detailed Analysis:
Criterion S requires randomisation at the school (or higher) level. This study involved a single university with three pre-existing, non-randomised intact classes assigned to conditions based on their major and enrolled course, not schools randomly allocated to intervention or control. Since even the weaker class-level randomisation (criterion C) was not met, the stronger school-level requirement is necessarily unmet too.
Criterion S is not met because there was no school-level (or any level) randomisation; a single institution's intact classes were compared.
-
I
Independent Conduct
- One of the study's authors was the teacher who delivered the treatment, and the researchers themselves scored the outcome measures.
- "The teacher was one of the researchers. She was an experienced non-native speaking teacher of English as a foreign language..." (p. 358)
Relevant Quotes:
1) "The teacher was one of the researchers. She was an experienced non-native speaking teacher of English as a foreign language and possessed a masters degree in English language teaching. She was the class's normal teacher." (Section 4.2, p. 358)
2) "The teacher corrected the narratives in the two experimental groups in accordance with the correction guidelines (see below)." (Section 4.4, p. 359)
3) "...scores for each of the narrative writing tests... and also for the two administrations of the error correction test were obtained by one of the researchers." (Section 4.10, p. 361)
Detailed Analysis:
The intervention (written CF) was designed and delivered by one of the paper's authors acting as the students' regular class teacher, and outcome scoring was performed by one of the researchers themselves rather than an independent, blinded evaluator. There is no mention anywhere in the paper of an external or third-party team conducting data collection, delivery, or analysis independently of the research/design team.
Criterion I is not met because the same research team designed, delivered, and scored the study with no independent third-party conduct.
-
Y
Year Duration
- The study spanned only about 10 weeks in total, nowhere near a full academic year, and criterion T was also not met.
- "The entire study was spread over a period of 10 weeks." (p. 360)
Relevant Quotes:
1) "The entire study was spread over a period of 10 weeks." (Section 4.9, p. 360)
2) "The delayed post-test was administered approximately 4 weeks later." (Section 4.6.1, p. 360)
Detailed Analysis:
Per the ERCT rules, since criterion T (Term Duration) was not met, criterion Y is automatically not met. Independently, the roughly eight-to-ten week total span of the study (from intervention start to final delayed post-test) is far short of the 75% of an academic year (~9-10 months) required.
Criterion Y is not met because the study duration was only about 10 weeks, far short of a year, and because criterion T was not met.
-
B
Balanced Control Group
- All three groups followed an identical schedule, teacher, and writing procedure, so no extra time or budget was directed to the intervention groups; the only difference was the content of feedback (correction vs. a matched general comment), which is itself the treatment variable being tested.
- "The procedure for the control group was the same except that the students received no corrections of their linguistic errors. Instead they received a simple general comment or question." (p. 359)
Relevant Quotes:
1) "The students in all three groups wrote the same three narratives in separate lessons and received feedback from the same teacher on each piece of writing." (Section 4.4, p. 358)
2) "The procedure for the control group was the same except that the students received no corrections of their linguistic errors. Instead they received a simple general comment or question (e.g. 'Good!', 'Are they happy then?', 'What happened then?')." (Section 4.4, p. 359)
3) "The students were asked to look over their errors and the corrections carefully for at least five minutes." (Section 4.4, p. 359, describing the experimental groups only)
Detailed Analysis:
Applying the Criterion B decision procedure: the first question is whether the intervention added extra time or budget relative to the control. Here it did not - all three groups wrote the same three narratives, in the same lessons, with the same teacher, on the same schedule (one 90-minute class per week over the same period), and every group's work was returned with feedback. The only variable is the content of that feedback: explicit article/ error correction for the two experimental groups versus a "simple general comment or question" for the control group. Since no additional instructional time, sessions, or resources were given to the intervention groups, EXTRA_RESOURCES_PRESENT is false and the criterion is met on that basis alone. As a secondary, reinforcing point, even if the correction itself were viewed as an "additional resource," it is precisely the treatment variable under investigation (Research Question 1: "Does written CF help Japanese learners... to become more accurate"), and the control's matched, minimal- difference "business as usual" comment is an appropriate comparison condition rather than an uncontrolled confound.
Criterion B is met because no extra time or budget was given to the intervention groups relative to the control - the only difference across the identical procedure was the content of the feedback, which is itself the treatment variable under investigation.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study is reported in the text, and none was found via internet search of the subsequent literature.
Relevant Quotes:
(No quotes describing an independent replication of this specific study were found in the paper.)
Detailed Analysis:
The paper situates itself among related but distinct studies (Sheen, 2007; Bitchener et al., 2005; Bitchener, forthcoming) that investigated CF on articles, but these are cited as prior, complementary work rather than as independent replications of this specific focused-vs-unfocused CF comparison in an EFL context. No subsequent replication by an independent team of this precise study design is referenced or evident from the paper's own text. An internet search for later studies citing or reproducing this specific design found only papers that cite Ellis et al. (2008) as background/motivation (e.g., a 2009 commentary, "Overgeneralization from a narrow focus: A response to Ellis et al. (2008) and Bitchener (2008)," in the Journal of Second Language Writing) or that test related but different interventions/populations (e.g., studies of focused vs. unfocused CF with Iranian EFL learners). None of these constitute an independent reproduction of this specific Japanese EFL university study by a different research team in a different context.
Criterion R is not met because no independent replication of this specific study was found in the paper or through internet search of the subsequent literature.
-
A
All-subject Exams
- Only article accuracy in English writing was assessed; no other subjects were measured, and criterion E was not met.
- "It compared the effects of focused and unfocused written CF on the accuracy with which Japanese university students used the English indefinite and definite articles..." (abstract)
Relevant Quotes:
1) "It compared the effects of focused and unfocused written CF on the accuracy with which Japanese university students used the English indefinite and definite articles to denote first and anaphoric reference in written narratives." (Abstract)
2) Per criterion A's prerequisite rule, criterion E must be met for A to be met.
Detailed Analysis:
The study measured a single, narrowly defined linguistic outcome (article accuracy in English writing) using non-standardised instruments; no other subjects were assessed. In addition, because criterion E (Exam-based Assessment) was not met, criterion A is automatically not met per the specification's explicit rule.
Criterion A is not met because only one narrow outcome (article use) was measured with non-standardised tools, and criterion E was not met.
-
G
Graduation Tracking
- Follow-up ended at the four-week delayed post-test with no tracking toward graduation, criterion Y was not met, and no follow-up publications tracking this cohort were found via internet search.
- "The delayed post-test was administered approximately 4 weeks later." (p. 360)
Relevant Quotes:
1) "The delayed post-test was administered approximately 4 weeks later." (Section 4.6.1, p. 360)
2) No further data collection beyond the delayed post-test (Week 10) is reported anywhere in the paper.
Detailed Analysis:
Tracking ended at the delayed post-test, roughly four weeks after the immediate post-test, with no mention of any longer-term or graduation-level follow-up. An internet search for subsequent papers by Ellis, Sheen, Murakami, or Takashima tracking this same cohort of Japanese university students toward graduation did not locate any such follow-up study; no paper id, quotes, or analysis of a graduation-tracking follow-up can therefore be provided. Per the specification, since criterion Y was not met, criterion G is automatically not met as well.
Criterion G is not met because tracking stopped at a four-week delayed post-test with no follow-up toward graduation found in this paper or in any subsequent publication located via internet search, and criterion Y was not met.
-
P
Pre-Registered
- The paper contains no mention of a pre-registered protocol, registry, or registration date, and no evidence of pre-registration was found via internet search.
Relevant Quotes:
(No quotes referencing pre-registration, a trial registry, or a published protocol were found anywhere in the paper.)
Detailed Analysis:
There is no statement in the Method, Results, or any other section referring to a pre-registration platform, a registered hypothesis/analysis plan, or a registration date. An internet search for a pre-registration record of this study (e.g., on trial or study registries) did not locate any such record. This is consistent with the paper's 2008 publication date, prior to widespread adoption of pre-registration practice in applied linguistics/ education research, but the absence of any such statement or record means the criterion cannot be considered met.
Criterion P is not met because no pre-registration of the study protocol is mentioned anywhere in the paper, and none was found through internet search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.