Level 1 Criteria
-
C Class-level RCT
- Although allocation was at the class level, the three intact classes were "arbitrarily assigned" with no described randomisation procedure, so a properly implemented class-level RCT is not demonstrated.
- "Classes were arbitrarily assigned to one of the two treatment options (group 1 = 12 students, group 2 = 12 students) or to the control group option (group 3 = 10 students)." (p. 351)
- Relevant Quotes: 1) "Classes were arbitrarily assigned to one of the two treatment options (group 1 = 12 students, group 2 = 12 students) or to the control group option (group 3 = 10 students)." (p. 351) 2) "The study was conducted in a private language school in New Zealand. Three classes of students (n = 34) were involved." (p. 350) 3) "Also, we were forced to use intact groups, with the result that the groups were not equivalent at the commencement of the study, thus obligating the use of ANCOVA." (p. 366) 4) "In an experimental design (two experimental groups and a control group), low-intermediate learners of second language English completed two communicative tasks..." (p. 339, abstract) Detailed Analysis: Criterion C requires a randomised controlled trial with randomisation clearly described and properly implemented at the class level or higher. In this study the unit of assignment was indeed the intact class (three whole classes were allocated to the three conditions), so the unit of allocation would satisfy the class-level requirement. However, the paper explicitly states that classes were "arbitrarily assigned," not randomly assigned, and no randomisation procedure (random number generation, coin toss, stratification, etc.) is described anywhere in the paper. The authors themselves acknowledge in the limitations that they "were forced to use intact groups" and that the groups "were not equivalent at the commencement of the study" - significant pretest differences were found on the oral imitation and grammaticality judgment tests. With only three clusters allocated arbitrarily, the study is better characterised as a quasi-experiment at the class level than a properly implemented class-level RCT. The intervention is whole-class corrective feedback, not one-to-one tutoring, so the tutoring exception does not apply. This lack of true randomisation is also why the top-level "rct" field has been corrected to false during this verification (see 2.9 change summary): nowhere in the paper is a genuine random-assignment procedure described, only an "arbitrary" one. Criterion C is not met because assignment of the three intact classes is described as "arbitrary" rather than random, and no randomisation procedure is described or properly implemented.
-
E Exam-based Assessment
- All three outcome measures were custom instruments designed by the researchers for this study, not widely recognised standardised exams.
- "Acquisition was measured by means of an oral imitation test (designed to measure implicit knowledge) and both an untimed grammaticality judgment test and a metalinguistic knowledge test (both designed to measure explicit knowledge)." (p. 339)
- Relevant Quotes: 1) "Acquisition was measured by means of an oral imitation test (designed to measure implicit knowledge) and both an untimed grammaticality judgment test and a metalinguistic knowledge test (both designed to measure explicit knowledge)." (p. 339, abstract) 2) "Oral Imitation Test. This test consisted of a set of 36 belief statements. Statements were grammatically correct (n = 18) or incorrect (n = 18)." (p. 354) 3) "Grammaticality Judgment Test. This was a pen-and-paper test consisting of 45 sentences. Fifteen sentences targeted past tense -ed, and the remainder targeted 30 other structures." (p. 355) 4) "Metalinguistic Knowledge Test. Learners were presented with five sentences and were told that they were ungrammatical. Two of the sentences contained errors in past tense -ed." (p. 356) 5) "For more information about the theoretical rationale for this test and its design, see Erlam (in press)." (p. 355) Detailed Analysis: Criterion E requires that outcomes be measured with widely recognised standardised exams rather than instruments created for the study. All three outcome measures here (the oral elicited imitation test, the untimed grammaticality judgment test, and the metalinguistic knowledge test) were purpose-built research instruments designed by the researchers to separate implicit from explicit knowledge of a single grammatical structure (past tense -ed), drawing on the authors' own prior instrument-development work (R. Ellis 2005; Erlam, in press). No national curriculum exam, state-wide achievement test, or other recognised standardised examination was used. These are exactly the kind of custom, intervention-aligned measures the criterion is designed to guard against. Criterion E is not met because all outcome measures were custom researcher-designed tests rather than recognised standardised exams.
-
T Term Duration
- Outcomes were last measured about two weeks after the 1-hour, 2-day intervention began, far short of the required full academic term of follow-up.
- "The tests were administered prior to the instruction, 1 day after the instruction, and again 2 weeks later." (p. 339)
- Relevant Quotes: 1) "For the purposes of the study, each experimental group received the same amount of instruction (i.e., a total of 1 hr over 2 consecutive days during which they completed two different half-hour communicative tasks)." (pp. 351-352) 2) "Five days prior to the start of the instructional treatments, the learners involved in the study signed the consent forms ... and completed all of the pretests. The immediate posttesting was completed the day after the second (and last) day of instruction, and the delayed posttesting was completed 12 days later." (p. 353) 3) "The tests were administered prior to the instruction, 1 day after the instruction, and again 2 weeks later." (p. 339, abstract) 4) "Third, the length of the treatments was very short (approximately 1 hr)." (p. 366) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Here the intervention consisted of just 1 hour of instruction spread over 2 consecutive days, the immediate posttest was administered 1 day after instruction ended, and the final delayed posttest was administered only about 2 weeks after the intervention began. The total interval from intervention start to final outcome measurement is approximately two weeks, which falls far short of one academic term. The authors themselves flag the very short treatment as a limitation. Criterion T is not met because the final outcome measurement occurred about 2 weeks after the intervention began, far less than one academic term.
-
D Documented Control Group
- The control group's size (n = 10), condition (testing-only with normal instruction and no feedback), and baseline test scores are clearly documented in the text and tables.
- "The control group continued with their normal instruction. They did not complete the tasks and did not receive any feedback on past tense -ed errors." (p. 352)
- Relevant Quotes: 1) "Group 1 received implicit feedback (recast group), group 2 received explicit feedback (metalinguistic group), and group 3 (a testing control) had no opportunity to practice the target structure and, thus, received no feedback." (p. 350) 2) "Classes were arbitrarily assigned to one of the two treatment options (group 1 = 12 students, group 2 = 12 students) or to the control group option (group 3 = 10 students)." (p. 351) 3) "The control group continued with their normal instruction. They did not complete the tasks and did not receive any feedback on past tense -ed errors." (p. 352) 4) "Information obtained from a background questionnaire showed that the majority of learners (77%) were of East Asian origin. Most of them had spent less than a year in New Zealand (the mean length of stay was just over 6 months). The mean age of all participants was 25 years." (p. 350) 5) "Test-retest reliability (Pearson's r) was calculated for the control group (n = 10) only." (p. 355) 6) Tables 3, 4, 5, 6, and 7 (pp. 357-360) report the control group's pretest, immediate posttest, and delayed posttest means and standard deviations on all three outcome measures. Detailed Analysis: Criterion D requires clear documentation of the control group: who they are, their size, baseline performance, and what treatment they received. The paper specifies the control group's size (one class of 10 students), its condition (a testing control that continued normal instruction, completed all three tests at all three time points, did not perform the treatment tasks, and received no feedback on the target structure), and its baseline performance (pretest means and SDs reported separately for the control group in Tables 3-5). Sample-level demographics (origin, age, length of English study, proficiency level) are documented, although not broken down by group, and the authors transparently note baseline non-equivalence and handle it with ANCOVA. Taken together, this constitutes adequate documentation of the control group for comparison purposes. Criterion D is met because the control group's size, condition (testing-only, normal instruction, no feedback), and baseline performance are clearly documented.