Improving elementary EFL speaking skills with generative AI chatbots: Exploring individual and paired interactions

Tzu-Yu Tai, Howard Hao-Jan Chen

Published:
ERCT Check Date:
DOI: 10.1016/j.compedu.2024.105112
  • L2 languages
  • K12
  • Asia
  • EdTech website
0
  • C

    Randomisation was conducted at the individual student level among students pooled from the same three intact classes, not at the class or school level, and no explicit tutoring exception was invoked by the authors.

    "The participants from the three intact classes were randomly divided into three groups to assess the impact of different interaction configurations on English-speaking skills." (p. 4)

  • E

    The speaking test was adapted from GEPT Kids, a nationally recognised standardised assessment developed by Taiwan's Language Training and Testing Center, using its official elementary speaking rating scale.

    "The speaking tests used in this study were adapted from the General English Proficiency Test (GEPT) Kids. GEPT Kids, a test specifically tailored to elementary school students, was developed by the Language Training and Testing Center (LTTC) in Taiwan." (p. 5)

  • T

    Outcomes were measured immediately after a three-week summer intervention, far short of the one-academic-term minimum required by the standard.

    "A total of 85 sixth-grade Taiwanese students participated in a three-week summer program, where students engaged in 45-min interactions with CoolE Bot for five days a week." (p. 4)

  • D

    The control (No-Bot) group's size, demographics, and the parallel topic-matched activities they completed are clearly documented in the paper, including in Table 1.

    "The No-Bot group, which performed similar speaking activities, received worksheets with the same designated topics for interactions with their peers. The teacher facilitated and guided learners through these interactive activities to ensure a consistent experience across all participant groups." (p. 7)

  • S

    The study involved students from a single summer program taught by one teacher, with randomisation at the individual student level rather than across schools.

    "The participants included 85 Taiwanese EFL 6-graders recruited from three classes and taught by the same English teacher." (p. 4)

  • I

    The same research team that designed and built CoolE Bot also implemented, supervised, and analysed the study, with no independent third-party evaluator.

    "One of the two researchers was present in the class to assist students with technology-related issues." (p. 7)

  • Y

    The study duration (three weeks) is far below the year-long requirement, and since criterion T is not met, criterion Y is automatically not met.

    "A total of 85 sixth-grade Taiwanese students participated in a three-week summer program..." (p. 4)

  • B

    All three groups received equal instructional time and topic-matched worksheet materials, differing only in the conversational partner (chatbot vs. teacher/peers), so resource allocation was balanced.

    "Participants in the No-Bot group completed similar speaking tasks on the same topics as the two Bot groups, except with the teacher and their peers." (p. 4)

  • R

    No independent replication of this study or its custom-built CoolE Bot intervention was found; the paper reports a novel, recently published intervention with no prior replication.

  • A

    Only English-speaking outcomes were assessed; no other core subjects were measured, and no exception rationale is provided.

    "The participants' English-speaking skills were evaluated before and after the experiment using a pretest and post-test, respectively." (p. 5)

  • G

    Tracking ended at the immediate post-test with no long-term follow-up, and since criterion Y is not met, criterion G is automatically not met.

  • P

    No statement of pre-registration on any public registry, nor any registration date, is present anywhere in the paper, and no matching registry record was found online.

Abstract

Generative artificial intelligence (GAI) and automatic speech recognition (ASR) have ushered in promising tools for foreign language learning, notably GAI chatbots. This study investigated the impact of GAI chatbots on elementary school English as a foreign language (EFL) learners' speaking skills, focusing on two interaction configurations - individual and paired. Eighty-five elementary school EFL learners participated in a three-week summer program, engaging in daily 45-min interactions with CoolE Bot. The participants were randomly assigned to three groups: (1) individual interaction with CoolE Bot (I-Bot group), (2) paired interaction with CoolE Bot (P-Bot group), and (3) interaction with teachers and peers in a conventional English classroom (No-Bot group). Quantitative (English-speaking tests) and qualitative data (semi-structured interviews) were collected and analyzed. Results revealed that the I-Bot and P-Bot groups' post-test speaking skills were significantly higher than those of the No-Bot group. Both individual and paired interactions with CoolE Bot demonstrated positive effects, with no significant differences between groups.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was conducted at the individual student level among students pooled from the same three intact classes, not at the class or school level, and no explicit tutoring exception was invoked by the authors.
      • "The participants from the three intact classes were randomly divided into three groups to assess the impact of different interaction configurations on English-speaking skills." (p. 4)
      • Relevant Quotes: 1) "A total of 85 sixth-grade Taiwanese students participated in a three-week summer program, where students engaged in 45-min interactions with CoolE Bot for five days a week. The participants from the three intact classes were randomly divided into three groups to assess the impact of different interaction configurations on English-speaking skills..." (p. 4) 2) "The participants included 85 Taiwanese EFL 6-graders recruited from three classes and taught by the same English teacher." (p. 4) Detailed Analysis: The paper states participants were drawn from three intact classes but were then "randomly divided into three groups," indicating that randomisation occurred at the individual student level rather than by preserving whole classes as the unit of assignment. All 85 students, originally distributed across three classes attending the same summer program, were pooled and reshuffled into the I-Bot, P-Bot, and No-Bot conditions. This creates a genuine contamination risk, since students assigned to different conditions remain part of the same overall program and could interact informally outside their assigned sessions. The ERCT standard allows an exception when an intervention is designed for personal teaching such as tutoring, under which student-level randomisation is acceptable. Two of the three arms (I-Bot: individual chatbot interaction; P-Bot: paired chatbot interaction) do resemble individualized or small-group practice sessions. However, the authors never explicitly frame the study as a tutoring intervention, and the design also includes a third arm (No-Bot) drawing on the same intact classes for a whole-class, teacher-led comparison condition, which is a standard classroom-level activity rather than one-to-one tutoring. Because the paper's own methodological description emphasizes individual-level random assignment without invoking or describing any tutoring-style rationale for bypassing class-level randomisation, and because genuine cross-condition contamination risk exists within a single shared summer program, this criterion is judged not met. Final sentence: Criterion C is not met because randomisation was carried out at the individual student level among students drawn from a shared pool of classes, without an explicit tutoring exception being invoked.
    • E

      Exam-based Assessment

      • The speaking test was adapted from GEPT Kids, a nationally recognised standardised assessment developed by Taiwan's Language Training and Testing Center, using its official elementary speaking rating scale.
      • "The speaking tests used in this study were adapted from the General English Proficiency Test (GEPT) Kids. GEPT Kids, a test specifically tailored to elementary school students, was developed by the Language Training and Testing Center (LTTC) in Taiwan." (p. 5)
      • Relevant Quotes: 1) "The participants' English-speaking skills were evaluated before and after the experiment using a pretest and post-test, respectively." (p. 5) 2) "The speaking tests used in this study were adapted from the General English Proficiency Test (GEPT) Kids. GEPT Kids, a test specifically tailored to elementary school students, was developed by the Language Training and Testing Center (LTTC) in Taiwan." (p. 5) 3) "The quality of the assessment was further verified through evaluation by one assessment specialist." (p. 5) 4) "The GEPT's elementary speaking rating scale was used by two English teachers who coded the results. The inter-rater reliability was 0.85." (p. 5) Detailed Analysis: The instrument used is adapted from GEPT Kids, a nationally recognised, standardised test developed by LTTC, a well-established Taiwanese testing organisation comparable to a national exam board. Although the specific test content (topics/prompts) was adapted for the study, both the underlying test framework and, critically, the official GEPT elementary speaking rating scale were used for scoring, with formal inter-rater reliability (0.85) and specialist review reported. This differs materially from a wholly researcher- invented custom test, since the assessment is anchored to a widely recognised standardised instrument and its official scoring rubric. Final sentence: Criterion E is met because the speaking assessment is grounded in GEPT Kids, a nationally recognised standardised test, using its official rating scale.
    • T

      Term Duration

      • Outcomes were measured immediately after a three-week summer intervention, far short of the one-academic-term minimum required by the standard.
      • "A total of 85 sixth-grade Taiwanese students participated in a three-week summer program, where students engaged in 45-min interactions with CoolE Bot for five days a week." (p. 4)
      • Relevant Quotes: 1) "A total of 85 sixth-grade Taiwanese students participated in a three-week summer program, where students engaged in 45-min interactions with CoolE Bot for five days a week." (p. 4) 2) "3 weeks prior study: Meetings with school stakeholders & Parental consent forms collected... 3-week Intervention... Post study survey: English speaking posttest... Interviews." (Fig. 4, p. 7) 3) "Following the intervention, all the participants underwent a speaking post-test, and those in the I-Bot and P-Bot groups were interviewed." (p. 7) 4) "The brief duration of the intervention raises concerns about potential novelty effects by introducing a variable that may affect the reliability of the results." (p. 12) Detailed Analysis: The intervention was a three-week summer program, and the speaking post-test (the primary outcome measure) was administered immediately following the end of this three-week period, with no subsequent follow-up interval reported. Three weeks is far below the "at least one full academic term (approximately 3-4 months)" required by criterion T. The authors themselves flag this as a limitation, explicitly noting concerns about novelty effects due to the brief intervention duration, confirming there was no extended tracking period. Final sentence: Criterion T is not met because outcomes were measured immediately after only a three-week intervention.
    • D

      Documented Control Group

      • The control (No-Bot) group's size, demographics, and the parallel topic-matched activities they completed are clearly documented in the paper, including in Table 1.
      • "The No-Bot group, which performed similar speaking activities, received worksheets with the same designated topics for interactions with their peers. The teacher facilitated and guided learners through these interactive activities to ensure a consistent experience across all participant groups." (p. 7)
      • Relevant Quotes: 1) "Table 1 Demographic information of the participants... No-Bot 29 12.01 [age] 14 [M] 15 [F] 3.78 [years of English learning]." (Table 1, p. 5) 2) "Participants in the No-Bot group completed similar speaking tasks on the same topics as the two Bot groups, except with the teacher and their peers." (p. 4) 3) "The No-Bot group, which performed similar speaking activities, received worksheets with the same designated topics for interactions with their peers. The teacher facilitated and guided learners through these interactive activities to ensure a consistent experience across all participant groups." (p. 7) 4) "However, it is important to note that the No-Bot group consisted of students who were the youngest and had the fewest years of learning English compared to the Bot groups." (p. 11) Detailed Analysis: The paper documents the control group's size (N = 29), age, gender split, and years of English learning experience in Table 1, and clearly describes the parallel activities the control group undertook (identical topics, worksheets, guided by the teacher and peers instead of the chatbot). The authors further transparently acknowledge a baseline demographic imbalance (No-Bot being youngest with least English learning experience), which is itself evidence of thorough documentation of the control group's characteristics rather than an absence of description. Final sentence: Criterion D is met because the control group's size, demographics, and parallel activities are clearly documented, including quantitative baseline data.
  • Level 2 Criteria

    • S

      School-level RCT

      • The study involved students from a single summer program taught by one teacher, with randomisation at the individual student level rather than across schools.
      • "The participants included 85 Taiwanese EFL 6-graders recruited from three classes and taught by the same English teacher." (p. 4)
      • Relevant Quotes: 1) "The participants included 85 Taiwanese EFL 6-graders recruited from three classes and taught by the same English teacher." (p. 4) 2) "The participants from the three intact classes were randomly divided into three groups to assess the impact of different interaction configurations on English-speaking skills." (p. 4) Detailed Analysis: The study was conducted with a single cohort of students (from three classes, all taught by the same teacher) inside one summer program, with no indication of multiple schools being involved or randomised. The unit of randomisation was the individual student, far below the school level required by this stronger criterion. Final sentence: Criterion S is not met because there was no school-level randomisation; all participants came from a single program taught by one teacher.
    • I

      Independent Conduct

      • The same research team that designed and built CoolE Bot also implemented, supervised, and analysed the study, with no independent third-party evaluator.
      • "One of the two researchers was present in the class to assist students with technology-related issues." (p. 7)
      • Relevant Quotes: 1) "This study developed CoolE Bot, a GAI-based chatbot that incorporates OpenAI's advanced language model, Microsoft's text-to-speech (TTS), and ASR techniques." (p. 2) 2) "As illustrated in Fig. 4, the researchers first held a meeting with school stakeholders, including the principal, academic affairs director, curriculum chief, and English teachers, and introduced the objectives of the study." (p. 7) 3) "One of the two researchers was present in the class to assist students with technology-related issues." (p. 7) 4) "Tzu-Yu Tai: Writing – review & editing, Writing – original draft, Validation, Supervision, Project administration, Methodology, Investigation, Formal analysis, Data curation, Conceptualization, Resources, Visualization. Howard Hao-Jan Chen: Conceptualization, Methodology, Formal analysis, Writing – review & editing, Supervision, Project administration, Funding acquisition." (CRediT statement, p. 13) Detailed Analysis: The same two authors who conceived and built CoolE Bot also organised the study logistics, were physically present during intervention sessions to assist students, and (per the CRediT statement) both personally conducted the "Formal analysis." While two English teachers coded the speaking test recordings (providing partial separation for scoring), overall study design, implementation oversight, and statistical analysis were carried out by the same team that designed the intervention, with no external, independent evaluator or third-party organisation responsible for data collection or analysis oversight. Final sentence: Criterion I is not met because the intervention designers also implemented, supervised, and analysed the study without independent third-party oversight.
    • Y

      Year Duration

      • The study duration (three weeks) is far below the year-long requirement, and since criterion T is not met, criterion Y is automatically not met.
      • "A total of 85 sixth-grade Taiwanese students participated in a three-week summer program..." (p. 4)
      • Relevant Quotes: 1) "A total of 85 sixth-grade Taiwanese students participated in a three-week summer program, where students engaged in 45-min interactions with CoolE Bot for five days a week." (p. 4) 2) "The brief duration of the intervention raises concerns about potential novelty effects by introducing a variable that may affect the reliability of the results." (p. 12) Detailed Analysis: The entire intervention and outcome-measurement window spanned only three weeks, nowhere near the required 75% of an academic year (~9-10 months). Per the ERCT specification, if criterion T (Term Duration) is not met, criterion Y is automatically not met as well; here neither the term nor the year duration requirement is satisfied. Final sentence: Criterion Y is not met because the study lasted only three weeks and criterion T is also not met.
    • B

      Balanced Control Group

      • All three groups received equal instructional time and topic-matched worksheet materials, differing only in the conversational partner (chatbot vs. teacher/peers), so resource allocation was balanced.
      • "Participants in the No-Bot group completed similar speaking tasks on the same topics as the two Bot groups, except with the teacher and their peers." (p. 4)
      • Relevant Quotes: 1) "A 45-min English class session was conducted daily." (p. 7) 2) "Bot group participants received worksheets in each class with a designated topic, prompts, and vocabulary to guide their interactions with CoolE Bot... The No-Bot group, which performed similar speaking activities, received worksheets with the same designated topics for interactions with their peers." (p. 7) 3) "Participants in the No-Bot group completed similar speaking tasks on the same topics as the two Bot groups, except with the teacher and their peers." (p. 4) 4) "The teacher facilitated and guided learners through these interactive activities to ensure a consistent experience across all participant groups." (p. 7) Detailed Analysis: Applying the criterion B decision procedure: no extra time or budget was allocated to the intervention arms relative to the control. All three groups (I-Bot, P-Bot, No-Bot) attended the same 45-minute daily class sessions across the same three-week period, using matched topics and comparable worksheet materials; no additional class time, budget, or materials were given only to the Bot groups. The No-Bot group received an equivalent amount of structured speaking practice time on the same topics, simply substituting a teacher/peer conversational partner for the chatbot. Since EXTRA_RESOURCES_PRESENT is false (no extra time or budget was introduced by the intervention), the criterion is met at the first branch of the decision procedure without needing to assess whether the difference is integral to the treatment; the only substantive difference between arms is the conversational-partner medium (chatbot vs. human), with time, topics, and materials held constant. Final sentence: Criterion B is met because instructional time, topics, and materials were equivalent across all three groups, with only the conversational partner varying and no extra resources introduced for any group.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this study or its custom-built CoolE Bot intervention was found; the paper reports a novel, recently published intervention with no prior replication.
      • Relevant Quotes: 1) "This study developed CoolE Bot, a GAI-based chatbot that incorporates OpenAI's advanced language model, Microsoft's text-to-speech (TTS), and ASR techniques." (p. 2) No quotes were found referencing any prior or subsequent independent replication of this specific study or of CoolE Bot by a different research team. Detailed Analysis: The paper introduces CoolE Bot as a chatbot newly developed by this specific research team and reports on its first empirical trial, published in July 2024. There is no reference within the paper to any independent replication effort. An internet search of the citation network of this paper (over 150 citing works as of mid-2026, including systematic reviews and several newer AI-chatbot speaking studies such as a chatbot-versus-human-peer comparison, a Microsoft 365 Copilot voice-chat study, a vocational-college GenAI chatbot study, and an EAP postgraduate chatbot study) found no study that independently replicates this specific design (comparing individual and paired GAI-chatbot interaction against a conventional class) in a different context or by a different research team. The only closely related follow-up located, "Impact of generative AI chatbots and interaction modes on the speaking proficiency of adolescent EFL learners" (Tai & Chen, 2025, Computer Assisted Language Learning, doi: 10.1080/09588221.2025.2572999), was conducted by the SAME two authors with a different cohort of 88 eighth-grade students and a different comparison (speech-only vs. speech-and-text interaction), so it does not constitute independent replication either. Final sentence: Criterion R is not met because no independent replication of this study or its chatbot intervention by a different research team was found.
    • A

      All-subject Exams

      • Only English-speaking outcomes were assessed; no other core subjects were measured, and no exception rationale is provided.
      • "The participants' English-speaking skills were evaluated before and after the experiment using a pretest and post-test, respectively." (p. 5)
      • Relevant Quotes: 1) "This study investigated the impact of GAI chatbots (CoolE Bot) on elementary school EFL learners' speaking skills." (p. 4) 2) "The participants' English-speaking skills were evaluated before and after the experiment using a pretest and post-test, respectively." (p. 5) Detailed Analysis: The study measured only English-speaking skills (fluency, content, pronunciation, grammar, vocabulary within the single domain of EFL speaking); no other core subjects such as mathematics, science, or general language arts were assessed. This is an elementary-level general EFL speaking-practice intervention rather than a highly specialised upper-secondary or vocational program, so the exception allowing a narrower subject focus does not apply. Because only the intervention's own target subject was measured, this criterion is not met. Final sentence: Criterion A is not met because only English-speaking outcomes were assessed, with no other core subjects measured.
    • G

      Graduation Tracking

      • Tracking ended at the immediate post-test with no long-term follow-up, and since criterion Y is not met, criterion G is automatically not met.
      • Relevant Quotes: 1) "Following the intervention, all the participants underwent a speaking post-test, and those in the I-Bot and P-Bot groups were interviewed." (p. 7) No mention of any tracking beyond the immediate post-test and interviews is present anywhere in the paper. Detailed Analysis: Data collection concluded with the post-test and interviews immediately following the three-week intervention; there is no indication of any longer-term follow-up, let alone tracking through graduation. Per the ERCT specification, since criterion Y (Year Duration) is not met, criterion G is automatically not met as well. An internet search for follow-up publications by the same authors (Tzu-Yu Tai, Howard Hao-Jan Chen) tracking this same cohort of 85 sixth-grade participants found none. A related paper, "Impact of generative AI chatbots and interaction modes on the speaking proficiency of adolescent EFL learners" (Tai & Chen, 2025, Computer Assisted Language Learning, doi: 10.1080/09588221.2025.2572999), uses the same CoolE Bot platform but studies a different cohort of 88 eighth-grade students under a different design, so it is not a graduation-tracking follow-up of the participants in this paper. Final sentence: Criterion G is not met because there was no follow-up beyond the immediate post-test, no graduation tracking of this cohort was found in subsequent publications, and criterion Y is also not met.
    • P

      Pre-Registered

      • No statement of pre-registration on any public registry, nor any registration date, is present anywhere in the paper, and no matching registry record was found online.
      • Relevant Quotes: No quotes referencing pre-registration, a trial registry (e.g., ClinicalTrials.gov, OSF, AsPredicted), or a pre-specified analysis plan were found anywhere in the Method, Results, or Discussion sections of the paper. Detailed Analysis: The methods and data-analysis sections describe the study procedures directly, with no reference to any registered protocol, hypotheses, or analysis plan published prior to data collection. An internet search of major trial registries and the Open Science Framework for a pre-registration record matching this study's title, authors, or the CoolE Bot intervention found no matching entry. Absent any such statement or registry record, this criterion cannot be considered met. Final sentence: Criterion P is not met because no pre-registration statement, registry reference, or matching registry record was found for this study.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.