Level 1 Criteria
-
C Class-level RCT
- Randomisation was at the individual student level rather than at the class or school level, and the computerized practice intervention does not qualify for the tutoring exception.
- "Then, the 63 learners were randomly assigned to two groups to practice the two target pragmatic features."
- Relevant Quotes: 1) "After the instruction, they were randomly assigned to an interleaved-practice group (n = 31) or a blocked-practice group (n = 32)." (p. 812, Abstract) 2) "Then, the 63 learners were randomly assigned to two groups to practice the two target pragmatic features. The interleaved-practice (IP) group consisted of 31 learners, whereas the blocked-practice (BP) group consisted of 32 learners." (p. 818) 3) "Sixty-three sophomores (35 females, 28 males) from EFL classes at a public university in China participated in this study." (p. 818) 4) "After the pretest, the learners were randomly assigned to an IP group or a BP group." (p. 819) Detailed Analysis: Criterion C requires randomisation at the class level (or stronger school level), with an exception for one-to-one tutoring or personal-teaching interventions. The paper clearly states that the 63 individual learners, drawn from EFL classes at a single public university, were randomly assigned to the two practice conditions. The unit of randomisation was the individual student, not intact classes or schools. The intervention (computerized corpus-based discourse completion practice completed individually at computers) is not a personal tutoring or one-to-one teaching intervention, so the tutoring exception does not apply. Individual randomisation of students drawn from the same classes risks the contamination that class-level randomisation is designed to prevent. Criterion C is not met because randomisation was conducted at the individual student level within a single university, and no valid tutoring exception applies.
-
E Exam-based Assessment
- Outcomes were assessed with multimedia discourse completion tasks and a rubric custom-designed for this study, not with a widely recognised standardised exam.
- "To address the drawbacks of traditional DCTs, MMDCTs ... were designed for the current study."
- Relevant Quotes: 1) "To address the drawbacks of traditional DCTs, MMDCTs (i.e., encompassing text-formatted prompts as well as dialogues presented aurally and visually via pictures) were designed for the current study." (p. 822) 2) "Moreover, three equivalent versions of the tests were designed for the pretest, immediate posttest, and delayed posttest to avoid practice effects." (p. 823) 3) "With respect to the design of the target test items, first, each of them was designed based on the authentic data related to making requests or suggestions from MICASE (Simpson et al., 2002)." (p. 823) 4) "a four-level rubric (see Table 3) was adopted to measure the two groups' speech acts of making requests and suggestions" (p. 824) 5) "They were intermediate English learners based on their scores (M = 84.39, SD = 3.95) on the TOEFL (Test of English as a Foreign Language), which they took three months prior to this study." (p. 818) Detailed Analysis: Criterion E requires that outcomes be measured with widely recognised standardised exams, not instruments specially designed for the study. The outcome measures here were multimedia discourse completion tasks (MMDCTs) that were "designed for the current study" by the researcher, based on the MICASE corpus, and scored with a custom four-level pragmatic accuracy rubric plus a manually calculated speech rate measure. Although the standardised TOEFL is mentioned, it was used only to describe participants' baseline proficiency, not as an outcome measure. The researcher-built MMDCTs, however carefully piloted and reliability-checked, are custom study-specific instruments closely aligned with the intervention content, which is exactly the situation the criterion is designed to guard against. Criterion E is not met because outcomes were measured with custom-designed MMDCTs and a researcher-developed rubric rather than a recognised standardised exam.
-
T Term Duration
- The interval from intervention start to the final delayed posttest was only about five to seven weeks, which is shorter than one full academic term.
- "both groups completed a delayed posttest five weeks later."
- Relevant Quotes: 1) "During Week 1, all participants completed a background questionnaire. During Week 2, all students received two hours of pragmatics instruction on making requests and suggestions delivered by the researcher." (p. 819) 2) "The entire practice session (i.e., 20 practice items) lasted one hour" (p. 820) 3) "After the practice session, both groups completed an immediate posttest. Moreover, to explore the retention of the benefits of the interleaved or blocked practice, both groups completed a delayed posttest five weeks later." (p. 820) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. In this study, the intervention consisted of two hours of instruction in Week 2 followed by a one-hour practice session, with an immediate posttest right after practice and a delayed posttest five weeks later. Even counting from the background questionnaire in Week 1 to the final delayed posttest, the total span of the study is roughly seven weeks, well short of a 3-4 month academic term. The standard allows short interventions but insists on term-long follow-up tracking from intervention start, which this five-week delayed posttest does not provide. Criterion T is not met because the final outcome measurement occurred only about five weeks after the practice intervention, far less than one full academic term.
-
D Documented Control Group
- The comparison group's size, demographic characteristics, baseline pretest performance, and exact treatment conditions are documented in detail.
- "no noticeable difference in pragmatic accuracy of the target speech acts on the pretest between the IP group (M = 39.98, SD = 3.71) and the BP group (M = 39.01, SD = 4.03)"
- Relevant Quotes: 1) "Sixty-three sophomores (35 females, 28 males) from EFL classes at a public university in China participated in this study. Their ages ranged from 19.5 to 22.6 years (Mage = 20.7). They were intermediate English learners based on their scores (M = 84.39, SD = 3.95) on the TOEFL" (p. 818) 2) "The interleaved-practice (IP) group consisted of 31 learners, whereas the blocked-practice (BP) group consisted of 32 learners." (p. 818) 3) "In contrast, the BP group first engaged in 10 practice tasks pertinent to making requests and then participated in 10 practice tasks involving making suggestions." (pp. 819-820) 4) "The results from an independent samples t-test demonstrated no noticeable difference in pragmatic accuracy of the target speech acts on the pretest between the IP group (M = 39.98, SD = 3.71) and the BP group (M = 39.01, SD = 4.03), t(61) = .991, p = .326, d = .250." (pp. 826-827) 5) "Moreover, no significant difference in fluency of the target pragmatic features was found on the pretest between the IP group (M = 2.69, SD = .32) and the BP group (M = 2.63, SD = .35), t(61) = .786, p = .435, d = .198." (p. 827) Detailed Analysis: Criterion D requires detailed documentation of the comparison group: who they are, their baseline performance, and the treatment they received. This study uses a comparison-group design in which the blocked-practice (BP) group serves as the comparison condition for the interleaved-practice (IP) group. The paper documents the BP group's size (n = 32), the demographic profile of the sample (age, gender, TOEFL-based proficiency including speaking and listening subscores, no residence abroad), the exact treatment the BP group received (the same instruction plus 20 practice tasks in blocked order), and its baseline performance on both outcome measures (pretest means, SDs, and confidence intervals for pragmatic accuracy and fluency), together with statistical checks confirming baseline equivalence, normality, and homogeneity of variances. This constitutes clear, quotable documentation of the comparison group's characteristics, size, and conditions, sufficient for proper comparison. Criterion D is met because the comparison (blocked-practice) group's size, demographics, baseline performance, and exact treatment are clearly documented.