Level 1 Criteria
-
C Class-level RCT
- Intact classes (not individual students within a class) were randomly assigned to the two task-repetition conditions, so the unit of randomisation was the class.
- "Intact classes of EFL learners at a private English language school in Central London were randomly assigned to take part in either a same-task repetition (STR) poster carousel training session (n = 24) or a parallel-task repetition (PTR) carousel training session (n = 22) during their normal scheduled English class." (pp. 6-8)
- Relevant Quotes: 1) "Intact classes of EFL learners at a private English language school in Central London were randomly assigned to take part in either a same-task repetition (STR) poster carousel training session (n = 24) or a parallel-task repetition (PTR) carousel training session (n = 22) during their normal scheduled English class." (pp. 6-8) 2) "A mixed between- and within-subjects design was used to look for fluency changes within-subjects across three task performances and between-subjects to compare changes between performances for the two different groups." (p. 8) 3) "A preliminary power analysis with G-Power recommended a total sample size of 86. This was not a possibility in the current study for reasons of practicality (intact classes) and availability of time for manual analysis." (p. 8, footnote 2) Detailed Analysis: Criterion C requires randomisation at the class level (or stronger) rather than randomising individual students within a single classroom. The paper explicitly states that "intact classes" were randomly assigned to the STR or PTR condition, which means the unit of randomisation was the whole class, preventing within-class contamination between the two conditions. The description of the randomisation process is brief (the exact number of classes and the randomisation method are not reported), but the quote clearly establishes that entire classes, not individual students, were the assigned unit, which is what the criterion requires. The exception for tutoring does not apply here because this is a whole-class activity. Criterion C is met because intact classes were randomly assigned to conditions, satisfying the class-level randomisation requirement.
-
E Exam-based Assessment
- Outcomes were custom research measures of utterance fluency derived from manual PRAAT analysis of recorded speech, not any standardised, widely recognised exam.
- "their performances were audio-recorded, transcribed, and analyzed using PRAAT to detail their L2 utterance fluency." (Abstract, p. 1)
- Relevant Quotes: 1) "their performances were audio-recorded, transcribed, and analyzed using PRAAT to detail their L2 utterance fluency." (Abstract) 2) "This study adopts three dependent variables which are traditionally associated with speed, breakdown and repair aspects of utterance fluency, respectively." (p. 7) 3) "articulation rate was calculated as the total number of raw syllables produced in the 1-minute sample from each performance that was analyzed, divided by total speaking time (excluding all pauses >250 ms), and multiplied by 60." (p. 7) 4) "each extract was annotated manually by the author using the standard textgrid feature." (p. 14) 5) "A unique PRAAT script was developed which would generate frequencies and durations for the intervals that had been manually created. This output was then used to calculate the specific fluency measures as outlined above." (p. 14) Detailed Analysis: Criterion E requires that outcomes be measured with a standard, widely recognised, standardised exam rather than instruments created for the study. Here the outcomes were three researcher-defined utterance-fluency metrics (articulation rate, frequency of mid-clause pauses, frequency of repetitions) computed from one-minute speech samples that were manually annotated by the author in PRAAT using a bespoke script. These are research measures specific to this study's analytical framework, not a named standardised examination such as IELTS, TOEFL, or a national curriculum exam. The only standardised reference is the CEFR level used for course placement via an in-house school test, which is itself not a standardised exam and was not an outcome measure. Criterion E is not met because outcomes were measured with custom PRAAT-based fluency metrics rather than a recognised standardised exam.
-
T Term Duration
- The entire intervention and all outcome measurement took place within a single scheduled class session, far short of one academic term of follow-up.
- "In both conditions, participants gave three oral performances (a narrative task) ... during their normal scheduled English class." (p. 8)
- Relevant Quotes: 1) "Intact classes of EFL learners ... were randomly assigned to take part in either a same-task repetition (STR) poster carousel training session ... or a parallel-task repetition (PTR) carousel training session ... during their normal scheduled English class." (pp. 6-8) 2) "In both conditions, participants gave three oral performances (a narrative task) - in the STR group, the three performances related to the same task and in the PTR group they were three different tasks of the same type." (p. 8) 3) "Participants were given up to 2 minutes to tell their stories, but they did not always use all the allotted time so a 1-minute extract was taken from the beginning of each performance." (p. 14) 4) "Steps 5 and 6 are repeated until each A has told the same story 3 times to 3 different visitors." (p. 9) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. In this study the intervention was a single poster carousel training session conducted "during their normal scheduled English class," and the outcome data were the three task performances recorded within that same session. There was no delayed post-test and no follow-up measurement of any kind; the interval from intervention start to final measurement is on the order of one lesson, not one term. The paper itself notes that longer-term or transfer benefits were not captured: fluency changes were observed only across immediate repetitions. Criterion T is not met because all outcomes were measured within the same single class session in which the intervention took place, far short of a term.
-
D Documented Control Group
- There is no business-as-usual control group, and the comparison (PTR) group's demographics and baseline characteristics are not documented separately from the pooled sample.
- "46 ESL students took part in the study and the age range of participants was 18-42 (M = 20.1 years)." (p. 8)
- Relevant Quotes: 1) "46 ESL students took part in the study and the age range of participants was 18-42 (M = 20.1 years). First languages spoken were: Arabic (3), Kasakh (1), Portuguese (2), Slovenian (1), Serbian (1), Thai (1), Ukrainian (1), French (8), German (7), Spanish (4), Chinese (2), Japanese (4), Korean (4), Swedish (1), Dutch (1), and Italian (5)." (p. 8) 2) "The proficiency level of participants was Intermediate (B1, CEFR), and the school allocated new students to a proficiency level based on their performance on an in-house grammar test, a short oral interview, and a writing sample." (p. 8) 3) "Intact classes of EFL learners ... were randomly assigned to take part in either a same-task repetition (STR) poster carousel training session (n = 24) or a parallel-task repetition (PTR) carousel training session (n = 22)." (pp. 6-8) Detailed Analysis: Criterion D requires detailed documentation of the control group, including demographic information, baseline performance, and the conditions it received. This study has no untreated or business-as-usual control group at all: both arms received a version of the poster carousel intervention (same-task vs parallel-task repetition), so the PTR group is at best an active comparison condition. Even treating the PTR group as the control, the paper reports demographics (age, first languages, proficiency) only for the pooled sample of 46 students, with no breakdown by group, no per-group demographic table, and no baseline equivalence analysis. Group sizes (n = 24 and n = 22) and the PTR procedure are described, and Table 1 reports per-group performance-1 fluency descriptives, but there is no documentation showing that the two intact classes were comparable in composition at baseline, which is the core purpose of this criterion. Criterion D is not met because the study lacks a documented control group: demographics and baseline characteristics are reported only for the whole sample, not for the comparison group specifically, and no business-as-usual control exists.