Level 1 Criteria
-
C Class-level RCT
- Randomisation was conducted at the individual student level among students pooled from the same three intact classes, not at the class or school level, and no explicit tutoring exception was invoked by the authors.
- "The participants from the three intact classes were randomly divided into three groups to assess the impact of different interaction configurations on English-speaking skills." (p. 4)
- Relevant Quotes: 1) "A total of 85 sixth-grade Taiwanese students participated in a three-week summer program, where students engaged in 45-min interactions with CoolE Bot for five days a week. The participants from the three intact classes were randomly divided into three groups to assess the impact of different interaction configurations on English-speaking skills..." (p. 4) 2) "The participants included 85 Taiwanese EFL 6-graders recruited from three classes and taught by the same English teacher." (p. 4) Detailed Analysis: The paper states participants were drawn from three intact classes but were then "randomly divided into three groups," indicating that randomisation occurred at the individual student level rather than by preserving whole classes as the unit of assignment. All 85 students, originally distributed across three classes attending the same summer program, were pooled and reshuffled into the I-Bot, P-Bot, and No-Bot conditions. This creates a genuine contamination risk, since students assigned to different conditions remain part of the same overall program and could interact informally outside their assigned sessions. The ERCT standard allows an exception when an intervention is designed for personal teaching such as tutoring, under which student-level randomisation is acceptable. Two of the three arms (I-Bot: individual chatbot interaction; P-Bot: paired chatbot interaction) do resemble individualized or small-group practice sessions. However, the authors never explicitly frame the study as a tutoring intervention, and the design also includes a third arm (No-Bot) drawing on the same intact classes for a whole-class, teacher-led comparison condition, which is a standard classroom-level activity rather than one-to-one tutoring. Because the paper's own methodological description emphasizes individual-level random assignment without invoking or describing any tutoring-style rationale for bypassing class-level randomisation, and because genuine cross-condition contamination risk exists within a single shared summer program, this criterion is judged not met. Final sentence: Criterion C is not met because randomisation was carried out at the individual student level among students drawn from a shared pool of classes, without an explicit tutoring exception being invoked.
-
E Exam-based Assessment
- The speaking test was adapted from GEPT Kids, a nationally recognised standardised assessment developed by Taiwan's Language Training and Testing Center, using its official elementary speaking rating scale.
- "The speaking tests used in this study were adapted from the General English Proficiency Test (GEPT) Kids. GEPT Kids, a test specifically tailored to elementary school students, was developed by the Language Training and Testing Center (LTTC) in Taiwan." (p. 5)
- Relevant Quotes: 1) "The participants' English-speaking skills were evaluated before and after the experiment using a pretest and post-test, respectively." (p. 5) 2) "The speaking tests used in this study were adapted from the General English Proficiency Test (GEPT) Kids. GEPT Kids, a test specifically tailored to elementary school students, was developed by the Language Training and Testing Center (LTTC) in Taiwan." (p. 5) 3) "The quality of the assessment was further verified through evaluation by one assessment specialist." (p. 5) 4) "The GEPT's elementary speaking rating scale was used by two English teachers who coded the results. The inter-rater reliability was 0.85." (p. 5) Detailed Analysis: The instrument used is adapted from GEPT Kids, a nationally recognised, standardised test developed by LTTC, a well-established Taiwanese testing organisation comparable to a national exam board. Although the specific test content (topics/prompts) was adapted for the study, both the underlying test framework and, critically, the official GEPT elementary speaking rating scale were used for scoring, with formal inter-rater reliability (0.85) and specialist review reported. This differs materially from a wholly researcher- invented custom test, since the assessment is anchored to a widely recognised standardised instrument and its official scoring rubric. Final sentence: Criterion E is met because the speaking assessment is grounded in GEPT Kids, a nationally recognised standardised test, using its official rating scale.
-
T Term Duration
- Outcomes were measured immediately after a three-week summer intervention, far short of the one-academic-term minimum required by the standard.
- "A total of 85 sixth-grade Taiwanese students participated in a three-week summer program, where students engaged in 45-min interactions with CoolE Bot for five days a week." (p. 4)
- Relevant Quotes: 1) "A total of 85 sixth-grade Taiwanese students participated in a three-week summer program, where students engaged in 45-min interactions with CoolE Bot for five days a week." (p. 4) 2) "3 weeks prior study: Meetings with school stakeholders & Parental consent forms collected... 3-week Intervention... Post study survey: English speaking posttest... Interviews." (Fig. 4, p. 7) 3) "Following the intervention, all the participants underwent a speaking post-test, and those in the I-Bot and P-Bot groups were interviewed." (p. 7) 4) "The brief duration of the intervention raises concerns about potential novelty effects by introducing a variable that may affect the reliability of the results." (p. 12) Detailed Analysis: The intervention was a three-week summer program, and the speaking post-test (the primary outcome measure) was administered immediately following the end of this three-week period, with no subsequent follow-up interval reported. Three weeks is far below the "at least one full academic term (approximately 3-4 months)" required by criterion T. The authors themselves flag this as a limitation, explicitly noting concerns about novelty effects due to the brief intervention duration, confirming there was no extended tracking period. Final sentence: Criterion T is not met because outcomes were measured immediately after only a three-week intervention.
-
D Documented Control Group
- The control (No-Bot) group's size, demographics, and the parallel topic-matched activities they completed are clearly documented in the paper, including in Table 1.
- "The No-Bot group, which performed similar speaking activities, received worksheets with the same designated topics for interactions with their peers. The teacher facilitated and guided learners through these interactive activities to ensure a consistent experience across all participant groups." (p. 7)
- Relevant Quotes: 1) "Table 1 Demographic information of the participants... No-Bot 29 12.01 [age] 14 [M] 15 [F] 3.78 [years of English learning]." (Table 1, p. 5) 2) "Participants in the No-Bot group completed similar speaking tasks on the same topics as the two Bot groups, except with the teacher and their peers." (p. 4) 3) "The No-Bot group, which performed similar speaking activities, received worksheets with the same designated topics for interactions with their peers. The teacher facilitated and guided learners through these interactive activities to ensure a consistent experience across all participant groups." (p. 7) 4) "However, it is important to note that the No-Bot group consisted of students who were the youngest and had the fewest years of learning English compared to the Bot groups." (p. 11) Detailed Analysis: The paper documents the control group's size (N = 29), age, gender split, and years of English learning experience in Table 1, and clearly describes the parallel activities the control group undertook (identical topics, worksheets, guided by the teacher and peers instead of the chatbot). The authors further transparently acknowledge a baseline demographic imbalance (No-Bot being youngest with least English learning experience), which is itself evidence of thorough documentation of the control group's characteristics rather than an absence of description. Final sentence: Criterion D is met because the control group's size, demographics, and parallel activities are clearly documented, including quantitative baseline data.