Level 1 Criteria
-
C Class-level RCT
- Students were randomly assigned individually to the PI, TI and MOI groups within each school rather than whole classes or schools being randomized, and no tutoring exception applies.
- "They were split into three groups using a random procedure from the beginning of the experiment: PI group (n = 15); TI group (n = 15); MOI group (n = 17)." (p. 74-75)
- Relevant Quotes: 1) "A parallel study was carried out among first-semester school students in two different secondary schools in China and Greece." (p. 74) 2) "They were split into three groups using a random procedure from the beginning of the experiment: PI group (n = 15); TI group (n = 15); MOI group (n = 17)." (p. 74-75) 3) "In the second school, the participants were 30 in total. The subjects were studying English in a secondary Greek school... They were also split into three groups using the same random procedure: PI group (n = 10); TI group (n = 10): MOI group (n = 10)." (p. 75) 4) "In both classroom studies, the regular classroom teachers also served as the instructors for the present study." (p. 75) Detailed Analysis: The paper describes a parallel design carried out in two secondary schools (one in China, one in Greece), where within each school the pool of consenting students (47 in China, 30 in Greece) was split into three instructional groups (PI, TI, MOI) "using a random procedure." There is no statement that whole classes were randomized as intact units; instead, groups of 10-17 individual students were formed from the eligible pool at each school and taught by the regular classroom teachers acting as instructors. This matches the ERCT Standard's described problem of assigning treatment at the individual level within the same classroom/school population, which risks contamination between groups (e.g., students discussing material, or the same teacher inadvertently blending approaches). The intervention is not a personal tutoring intervention (each condition involves group classroom instruction of 10-17 students), so the tutoring exception in the ERCT Standard does not apply. No stronger school-level randomization is present either, since both schools implement all three conditions. Criterion C is not met because randomization occurred at the individual-student level within each school rather than at the class or school level, and no tutoring exception applies.
-
E Exam-based Assessment
- Outcomes were measured with researcher-designed interpretation and production tasks piloted by the author, not with a widely recognised standardised exam.
- "The interpretation task (see Appendix D) consisted of 20 sentences (10 distractor items in the present regular form) of which 10 were in the English past simple tense (only regular forms)." (p. 79)
- Relevant Quotes: 1) "Two versions of the tests were designed, one interpretation task and one written production task, in order to form a split block design." (p. 79) 2) "The interpretation task (see Appendix D) consisted of 20 sentences (10 distractor items in the present regular form) of which 10 were in the English past simple tense (only regular forms). The participants had to listen to the sentences and indicate (interpret) whether the sentence they heard was related to a past or present action." (p. 79) 3) "The written production task (see Appendix E) was developed and used to measure learner's ability to produce correct sentences using the English past tense. The students were required to look at 10 pictures and use the verbs provided to produce a sentence for each of the picture." (p. 79) 4) "Tests were balanced in terms of difficulty and vocabulary (high frequency items) in a previous pilot experiment." (p. 79) Detailed Analysis: Both outcome measures (a listening interpretation task and a written production task) were purpose-built by the author for this experiment and piloted internally, rather than being widely recognised, externally validated standardised exams (e.g., a national English proficiency test). The paper explicitly describes constructing the items (20 sentences; 10 pictures), balancing them for difficulty/vocabulary in "a previous pilot experiment," and creating two parallel versions (A/B) for a split-block design. There is no reference to an established, standardised assessment instrument being used or adapted. Criterion E is not met because the assessments used are custom, researcher-designed interpretation and production tasks rather than standardised exams.
-
T Term Duration
- Outcomes were measured immediately after a three-day, six-hour instructional period, with no delayed post-test, far short of one academic term.
- "In both studies, the three groups were taught for three consecutive days for a total of six hours of instruction (two hours per day) on the target feature." (p. 75)
- Relevant Quotes: 1) "In both studies, the three groups were taught for three consecutive days for a total of six hours of instruction (two hours per day) on the target feature." (p. 75) 2) "IMMEDIATE POST-TESTS (Immediately after the end of the instructional period) Interpretation and production tasks" (Figure 1, p. 77) 3) "Despite the outcomes of the present study, long-term effects of the variables under investigation should be re-examined as delayed post-tests were not available (the semester system in the two countries where data were collected did not allow for this)." (p. 85) Detailed Analysis: The intervention itself lasted only three consecutive days (six hours total), and the outcome measures were administered immediately after the instructional period ended, as shown in Figure 1's overview of the procedure. The authors explicitly acknowledge as a limitation that no delayed post-test was administered, meaning there is no tracking of outcomes even weeks later, let alone across a full academic term (approximately 3-4 months). This falls far short of the ERCT requirement of tracking outcomes at least one term after intervention start. Criterion T is not met because outcomes were measured immediately after a three-day intervention, with no term-long (or any delayed) follow-up.
-
D Documented Control Group
- The study compares three active instructional treatments with no documented untreated or business-as-usual control group.
- "The first group received processing instruction; the second group was exposed to traditional instruction; the third group received meaning-based output instruction." (p. 67, abstract)
- Relevant Quotes: 1) "The first group received processing instruction; the second group was exposed to traditional instruction; the third group received meaning-based output instruction." (p. 67, abstract) 2) "The same instructional treatments were used in both experiments... The first group, received the PI treatment; the second group the TI treatment; and the third group the MOI treatment." (p. 76) 3) "All the three treatments were exposed to the same amount of explicit information... regarding the target feature so that the only difference between the three treatments was limited to the nature of the practice." (p. 76) Detailed Analysis: This study compares three active instructional treatments (PI, TI, MOI) against one another; there is no untreated, business-as-usual, or no-instruction control group in this experiment. All three groups received formal instruction targeting the identical grammatical feature over the same three-day period. While participant demographics and baseline scores for each of the three groups are documented in Tables 1-4, none of the three arms functions as a genuine control group receiving no intervention or standard instruction unrelated to the target feature; TI is itself an active treatment being compared, not a documented control condition. Criterion D is not met because the study contains no documented control group; all three groups received an active instructional intervention on the same target feature.