Evaluating the Impact of Integrating Intelligent Educational Chatbots in English Conversation-Based Classes on Reducing Speaking Anxiety Among Adult Language Learners

Akbar Molaei

Published:
ERCT Check Date:
DOI: 10.61838/kman.ijecs.84
  • L2 languages
  • adult education
  • Asia
  • EdTech app
0
  • C

    Randomisation was at the individual learner level rather than by class or school, and no one-to-one tutoring exception applies.

    "A total of 30 adult learners (aged 20-35) who met the inclusion criteria ... were randomly assigned into two groups: an experimental group (n = 15) ... and a control group (n = 15)..."

  • E

    The study used only the FLCAS self-report anxiety questionnaire and no standardised academic exam-based assessment.

    "Speaking anxiety was measured using the Foreign Language Classroom Anxiety Scale (FLCAS) at three time points: pre-test, post-test, and five-month follow-up."

  • T

    Outcome tracking spanned about seven months from intervention start (10 weekly sessions plus a five-month follow-up), exceeding one academic term.

    "The intervention lasted for 10 weekly sessions of 90 minutes each, followed by a five-month follow-up to examine the sustainability of the effects."

  • D

    The control group's size, condition, and baseline FLCAS scores are documented with equivalence checks, despite demographics being pooled across groups.

    "...a control group (n = 15), which participated in regular conversation classes without chatbot integration."

  • S

    Individual learners, not schools or language institutes, were the unit of randomisation.

    "A total of 30 adult learners ... were randomly assigned into two groups: an experimental group (n = 15) ... and a control group (n = 15)..."

  • I

    The single author designed, conducted, and analysed the trial himself with no documented external or independent evaluation.

    "All authors significantly contributed to this study."

  • Y

    Total tracking of roughly seven months without specific dates falls short of a confirmable 75% of a full academic year.

    "The intervention lasted for 10 weekly sessions of 90 minutes each, followed by a five-month follow-up to examine the sustainability of the effects."

  • B

    Both groups had comparable conversation classes, and the chatbot access (with home practice) was the integral treatment variable tested against business-as-usual.

    "This study employed a randomized controlled trial (RCT) design to evaluate the effectiveness of integrating intelligent educational chatbots into English conversation-based classrooms..."

  • R

    The study is framed as novel, citation databases show zero citing works, and no independent published replication of this specific trial exists.

    "Given this background, the current study seeks to fill a critical gap in the literature by investigating the effect of integrating intelligent chatbots into English conversation-based classrooms..."

  • A

    No standardised exams in any subject were used (criterion E fails), so all-subject exam coverage is impossible.

    "Speaking anxiety was measured using the Foreign Language Classroom Anxiety Scale (FLCAS) at three time points: pre-test, post-test, and five-month follow-up."

  • G

    Tracking ended at a five-month follow-up with no graduation endpoint defined, reached, or found in follow-up literature, and criterion Y also fails.

    "...while the five-month follow-up provided insights into long-term impact, even longer-term evaluations would be needed to assess the sustainability of anxiety reduction and language proficiency growth."

  • P

    No pre-registration statement, registry ID, or registration date is provided in the paper, and none was found in the IRCT or ClinicalTrials.gov registries via internet search.

Abstract

Purpose: This study aimed to evaluate the effectiveness of integrating intelligent educational chatbots into English conversation-based classrooms in reducing speaking anxiety among adult language learners. Methods and Materials: A randomized controlled trial was conducted with 30 adult intermediate-level English learners from Tehran, who were randomly assigned to either an experimental group (n = 15) receiving a 10-session chatbot-assisted intervention or a control group (n = 15) participating in traditional conversation classes. Each session lasted 90 minutes and focused on interactive speaking tasks. Speaking anxiety was measured using the Foreign Language Classroom Anxiety Scale (FLCAS) at three time points: pre-test, post-test, and five-month follow-up. Data were analyzed using repeated measures ANOVA and Bonferroni post-hoc tests via SPSS-27. Findings: The results showed a significant time x group interaction effect on speaking anxiety scores (F(2, 56) = 35.41, p < .001), indicating that the reduction in anxiety was significantly greater in the experimental group than in the control group. Conclusion: The integration of AI-powered educational chatbots in conversation-focused English classes appears to be an effective and lasting strategy for reducing speaking anxiety among adult learners.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was at the individual learner level rather than by class or school, and no one-to-one tutoring exception applies.
      • "A total of 30 adult learners (aged 20-35) who met the inclusion criteria ... were randomly assigned into two groups: an experimental group (n = 15) ... and a control group (n = 15)..."
      • Relevant Quotes: 1) "A randomized controlled trial was conducted with 30 adult intermediate-level English learners from Tehran, who were randomly assigned to either an experimental group (n = 15) receiving a 10-session chatbot-assisted intervention or a control group (n = 15) participating in traditional conversation classes." (p. 1, Abstract) 2) "A total of 30 adult learners (aged 20-35) who met the inclusion criteria--being at an intermediate proficiency level and reporting moderate to high speaking anxiety--were randomly assigned into two groups: an experimental group (n = 15), which received the chatbot-enhanced conversation intervention, and a control group (n = 15), which participated in regular conversation classes without chatbot integration." (p. 3) 3) "Participants were recruited from private language institutes in Tehran, Iran, through announcements and purposive screening." (p. 3) Detailed Analysis: The unit of randomisation was the individual learner: 30 recruited adults were "randomly assigned into two groups" of 15 each. There is no mention of intact classes, sites, or institutes being randomised; participants came from "various language institutes in Tehran" and were allocated as individuals. Criterion C requires randomisation of entire classes (or schools) to prevent contamination. The tutoring exception does not apply: the intervention was delivered within conversation-based classes with interactive speaking tasks, and the paper never describes it as one-to-one personal tutoring. Individual chatbot practice was a component embedded in group classes, not a personal-teaching design that would trigger the exception. Criterion C is not met because randomisation was performed at the individual student level, not at the class or school level, and no tutoring exception applies.
    • E

      Exam-based Assessment

      • The study used only the FLCAS self-report anxiety questionnaire and no standardised academic exam-based assessment.
      • "Speaking anxiety was measured using the Foreign Language Classroom Anxiety Scale (FLCAS) at three time points: pre-test, post-test, and five-month follow-up."
      • Relevant Quotes: 1) "Speaking anxiety was measured using the Foreign Language Classroom Anxiety Scale (FLCAS) at three time points: pre-test, post-test, and five-month follow-up." (p. 1, Abstract) 2) "The Foreign Language Classroom Anxiety Scale (FLCAS), developed by Elaine K. Horwitz, Michael B. Horwitz, and Joann Cope in 1986, is one of the most widely used instruments for assessing anxiety in the context of foreign language learning, particularly speaking anxiety." (p. 3) 3) "Responses are rated on a 5-point Likert scale, ranging from 1 (strongly disagree) to 5 (strongly agree). The total score ranges from 33 to 165, with higher scores indicating greater levels of language anxiety." (p. 3) Detailed Analysis: The only outcome instrument in the study is the FLCAS, a 33-item self-report Likert questionnaire measuring anxiety. Although the FLCAS is a well-known validated psychometric scale, it is not a standardised academic exam: it does not assess educational achievement or language proficiency, and no recognised standardised examination (e.g., IELTS, TOEFL, national curriculum exam) was administered at any time point. Criterion E requires that educational outcomes be measured with standard, widely recognised exam-based assessments; a self-report anxiety questionnaire does not satisfy this, regardless of its psychometric quality. No speaking performance or achievement measure of any kind was collected. Criterion E is not met because outcomes were measured solely with a self-report anxiety questionnaire (FLCAS), not a standardised exam-based assessment of educational achievement.
    • T

      Term Duration

      • Outcome tracking spanned about seven months from intervention start (10 weekly sessions plus a five-month follow-up), exceeding one academic term.
      • "The intervention lasted for 10 weekly sessions of 90 minutes each, followed by a five-month follow-up to examine the sustainability of the effects."
      • Relevant Quotes: 1) "The intervention lasted for 10 weekly sessions of 90 minutes each, followed by a five-month follow-up to examine the sustainability of the effects." (p. 3) 2) "To evaluate the effectiveness of the intervention over time, a repeated measures analysis of variance (ANOVA) was conducted with three time points: pre-test, post-test, and five-month follow-up." (p. 4) 3) "Participants in the chatbot-assisted group experienced a significant decline in anxiety from pre-test (M = 126.47) to post-test (M = 97.85), which was sustained at follow-up (M = 99.42)." (p. 1, Abstract) Detailed Analysis: The intervention ran for 10 weekly sessions, i.e., roughly 2.5 months from intervention start to post-test. On its own the post-test falls slightly short of a full 3-4 month term. However, the ERCT standard allows short interventions provided outcome tracking from intervention start reaches at least one term, and the study additionally measured the primary outcome at a five-month follow-up after the intervention, giving a total interval of approximately 7 months from intervention start to the final measurement. This clearly exceeds one academic term (3-4 months). The exact calendar dates are not stated, but the quoted design (10 weeks of sessions plus a five-month follow-up) is sufficient to establish an interval of well over one term. Criterion T is met because outcomes were measured at a five-month follow-up after the 10-week intervention, approximately 7 months after intervention start, exceeding one full academic term.
    • D

      Documented Control Group

      • The control group's size, condition, and baseline FLCAS scores are documented with equivalence checks, despite demographics being pooled across groups.
      • "...a control group (n = 15), which participated in regular conversation classes without chatbot integration."
      • Relevant Quotes: 1) "...a control group (n = 15), which participated in regular conversation classes without chatbot integration." (p. 3) 2) "The final sample consisted of 30 adult language learners, with 56.7% (n = 17) identifying as female and 43.3% (n = 13) as male. Regarding age distribution, 36.7% (n = 11) were between 20-25 years, 46.7% (n = 14) were between 26-30 years, and 16.6% (n = 5) were aged 31-35." (p. 4) 3) "Control | Pre-test | 124.66 | 8.09; Control | Post-test | 122.88 | 7.65; Control | Follow-up | 123.53 | 7.81" (Table 1, p. 4) 4) "Levene's test for equality of error variances indicated no significant differences between groups at any time point (pre-test: F(1,28) = 1.42, p = .24...)." (p. 4) Detailed Analysis: The control group is documented in several respects: its size (n = 15), the condition it received (regular conversation classes without chatbot integration, i.e., no unintended extra treatment), and its baseline performance on the outcome measure (Table 1: pre-test M = 124.66, SD = 8.09), together with statistical checks showing baseline equivalence between groups (Levene's tests, normality checks). Demographic information (gender, age, education) is reported only for the pooled sample of 30, not broken down by group, which is a weakness; however, the combination of documented group size, described control condition, group-specific baseline scores, and equivalence testing provides adequate documentation for comparison purposes in the sense required by criterion D. Criterion D is met because the control group's size, condition, and baseline outcome scores are clearly documented, although demographics are reported only at the whole-sample level.
  • Level 2 Criteria

    • S

      School-level RCT

      • Individual learners, not schools or language institutes, were the unit of randomisation.
      • "A total of 30 adult learners ... were randomly assigned into two groups: an experimental group (n = 15) ... and a control group (n = 15)..."
      • Relevant Quotes: 1) "A total of 30 adult learners ... were randomly assigned into two groups: an experimental group (n = 15) ... and a control group (n = 15)..." (p. 3) 2) "Participants were recruited from private language institutes in Tehran, Iran, through announcements and purposive screening." (p. 3) 3) "All participants were currently enrolled in intermediate-level English courses at various language institutes in Tehran." (p. 4) Detailed Analysis: Criterion S requires randomisation among schools or equivalent implementing institutions. Here, individual learners drawn from various private language institutes were pooled and randomly assigned as individuals to the two conditions. No institutes, sites, or centres were randomised; the institutes served only as recruitment sources. Since only individual-level randomisation is described, the school-level requirement is plainly not satisfied. Criterion S is not met because randomisation occurred at the individual learner level and no schools, institutes, or sites were randomised.
    • I

      Independent Conduct

      • The single author designed, conducted, and analysed the trial himself with no documented external or independent evaluation.
      • "All authors significantly contributed to this study."
      • Relevant Quotes: 1) "Akbar. Molaei1* 1 Department of English Language Teaching, Farhangian University, Tehran, Iran" (p. 1) 2) "All authors significantly contributed to this study." (Authors' Contributions, p. 7) 3) "This study employed a randomized controlled trial (RCT) design to evaluate the effectiveness of integrating intelligent educational chatbots into English conversation-based classrooms..." (p. 3) 4) "According to the authors, this article has no financial support." (Funding, p. 7) Detailed Analysis: Criterion I requires that the study be conducted independently of the intervention's designers, or at minimum that an external evaluation team handled data collection and analysis. This is a single-author study in which the same author designed the chatbot-enhanced intervention, ran the trial, collected the FLCAS data, and analysed the results. The paper contains no statement of any external evaluator, independent data collectors, blinded assessors, or third-party oversight of any part of the study. Nothing in the acknowledgments, contributions, or methods indicates independence between intervention design and evaluation. Criterion I is not met because the sole author designed, implemented, and evaluated the intervention with no documented independent or third-party involvement.
    • Y

      Year Duration

      • Total tracking of roughly seven months without specific dates falls short of a confirmable 75% of a full academic year.
      • "The intervention lasted for 10 weekly sessions of 90 minutes each, followed by a five-month follow-up to examine the sustainability of the effects."
      • Relevant Quotes: 1) "The intervention lasted for 10 weekly sessions of 90 minutes each, followed by a five-month follow-up to examine the sustainability of the effects." (p. 3) 2) "...a repeated measures analysis of variance (ANOVA) was conducted with three time points: pre-test, post-test, and five-month follow-up." (p. 4) 3) "Additionally, while the five-month follow-up provided insights into long-term impact, even longer-term evaluations would be needed to assess the sustainability of anxiety reduction and language proficiency growth." (p. 7) Detailed Analysis: Criterion Y requires outcome measurement at least 75% of a full academic year (roughly 9-10 months, i.e., about 6.75-7.5 months minimum) after intervention start, with clear start and end dates. Here the best estimate of total tracking is about 10 weeks of intervention plus a five-month follow-up, i.e., roughly 7.3 months if the follow-up is counted from the post-test. This sits exactly at the boundary: it would reach 75% only under the shorter (9-month) definition of an academic year, and falls below 75% of a 10-month year. The paper provides no calendar dates, does not state whether the five-month follow-up was timed from the post-test or from baseline, and the adult private-institute setting has no defined academic year to anchor the calculation. The authors themselves characterise the follow-up as insufficient for long-term conclusions. Given the absence of specific dates and an interval that cannot be confirmed to reach 75% of a full academic year, the criterion cannot be judged as satisfied. Criterion Y is not met because the documented tracking (about 7 months, with no specific dates) cannot be confirmed to reach 75% of a full academic year.
    • B

      Balanced Control Group

      • Both groups had comparable conversation classes, and the chatbot access (with home practice) was the integral treatment variable tested against business-as-usual.
      • "This study employed a randomized controlled trial (RCT) design to evaluate the effectiveness of integrating intelligent educational chatbots into English conversation-based classrooms..."
      • Relevant Quotes: 1) "...who were randomly assigned to either an experimental group (n = 15) receiving a 10-session chatbot-assisted intervention or a control group (n = 15) participating in traditional conversation classes. Each session lasted 90 minutes and focused on interactive speaking tasks." (p. 1, Abstract) 2) "...a control group (n = 15), which participated in regular conversation classes without chatbot integration." (p. 3) 3) "This study employed a randomized controlled trial (RCT) design to evaluate the effectiveness of integrating intelligent educational chatbots into English conversation-based classrooms for reducing speaking anxiety in adult language learners." (p. 3) 4) "In the current study, participants appreciated being able to choose conversation topics, repeat exercises, and practice at home, all of which fostered a sense of ownership over their progress..." (p. 6) Detailed Analysis: Applying the ERCT criterion B decision tree: extra resources are present in the experimental arm (chatbot access, plus the ability to "repeat exercises, and practice at home" outside of the 90-minute sessions), so the check does not stop at step 1. However, these resources ARE the treatment variable being tested -- the study's explicit purpose is "to evaluate the effectiveness of integrating intelligent educational chatbots" into conversation classes, tested against business-as-usual "traditional conversation classes." Both arms otherwise attended comparable structured conversation classes, so core instructional time was broadly matched, and the chatbot (with its home-practice affordance) is integral to the intervention package under evaluation rather than a separable, confounding add-on. Under the standard's exception (and the decision-tree branch for RESOURCES_ARE_TREATMENT), the control group may remain at the standard business-as-usual level when the added resource is itself the tested treatment; this is explicitly noted here so the additional home-practice access is not overlooked. Criterion B is met because both arms received comparable conversation classes and the added chatbot access (including home practice) is the integral treatment variable being tested against business-as-usual classes.
  • Level 3 Criteria

    • R

      Reproduced

      • The study is framed as novel, citation databases show zero citing works, and no independent published replication of this specific trial exists.
      • "Given this background, the current study seeks to fill a critical gap in the literature by investigating the effect of integrating intelligent chatbots into English conversation-based classrooms..."
      • Relevant Quotes: 1) "Given this background, the current study seeks to fill a critical gap in the literature by investigating the effect of integrating intelligent chatbots into English conversation-based classrooms on reducing speaking anxiety among adult learners in Tehran." (p. 3) 2) "Ballida and Aydin (2025) found that students who engaged in structured chatbot dialogues demonstrated significantly lower levels of speaking anxiety compared to those relying solely on peer interaction." (p. 2) 3) "Yet despite these advances, research remains limited on the long-term impact of chatbot-assisted speaking practice in adult learner populations, particularly in non-Western educational contexts." (p. 3) Internet Search Findings: OpenAlex (work ID W7117533542) lists a cited_by_count of 0 for this paper, and the Semantic Scholar record for DOI 10.61838/kman.ijecs.84 likewise reports a citationCount of 0 with no listed citing papers, so no external database currently shows any subsequent study engaging with, let alone replicating, this trial. No independent replication of this specific Tehran chatbot-versus-traditional-classes RCT (same intervention, population, and FLCAS outcome) could be located in any source. Detailed Analysis: Criterion R requires that this specific study be independently replicated by a different team, in a different context, in a peer-reviewed journal. The paper positions itself as filling a gap, i.e., as novel rather than replicated. Related studies cited (e.g., Ballidag & Aydin 2025; Duong & Suppasetseree 2024) are prior or contemporary investigations of chatbots and speaking anxiety with different designs, populations, and outcome focuses; they are not replications of this Tehran adult-learner RCT. The article reached final publication in February 2026, citation databases show zero citing works as of this check, and no published independent replication of this particular trial (its intervention, population, and design) could be identified in the paper or via internet search, which is unsurprising given how recently the study was published. Criterion R is not met because no independent replication of this specific study exists; the paper explicitly frames itself as filling a novel gap, and citation databases (OpenAlex, Semantic Scholar) confirm zero citing works to date.
    • A

      All-subject Exams

      • No standardised exams in any subject were used (criterion E fails), so all-subject exam coverage is impossible.
      • "Speaking anxiety was measured using the Foreign Language Classroom Anxiety Scale (FLCAS) at three time points: pre-test, post-test, and five-month follow-up."
      • Relevant Quotes: 1) "Speaking anxiety was measured using the Foreign Language Classroom Anxiety Scale (FLCAS) at three time points: pre-test, post-test, and five-month follow-up." (p. 1, Abstract) 2) "Second, the study relied on self-report measures of anxiety, which, although validated, are susceptible to response biases and may not fully capture the complexity of learners' emotional experiences." (p. 7) Detailed Analysis: Criterion A requires standardised exam-based assessment across all main subjects, and explicitly presupposes criterion E. Criterion E is not met here (the sole outcome is a self-report anxiety questionnaire), so criterion A automatically fails. Moreover, the study measured no academic subject at all -- not even English proficiency -- let alone the range of subjects. As an adult private language-institute programme, a specialised-context exception could in principle justify measuring only English-related outcomes, but no standardised exam of any subject was administered, so the exception cannot rescue the criterion. Criterion A is not met because criterion E fails and no standardised exam in any subject was administered.
    • G

      Graduation Tracking

      • Tracking ended at a five-month follow-up with no graduation endpoint defined, reached, or found in follow-up literature, and criterion Y also fails.
      • "...while the five-month follow-up provided insights into long-term impact, even longer-term evaluations would be needed to assess the sustainability of anxiety reduction and language proficiency growth."
      • Relevant Quotes: 1) "The intervention lasted for 10 weekly sessions of 90 minutes each, followed by a five-month follow-up to examine the sustainability of the effects." (p. 3) 2) "Additionally, while the five-month follow-up provided insights into long-term impact, even longer-term evaluations would be needed to assess the sustainability of anxiety reduction and language proficiency growth." (p. 7) Internet Search Findings: Searches of OpenAlex and Semantic Scholar for the author "Akbar Molaei" and for citing works of DOI 10.61838/kman.ijecs.84 (cited_by_count / citationCount = 0) returned no subsequent or follow-up publications tracking this same cohort of 30 Tehran language learners. No such follow-up paper could be found in any searched source. Detailed Analysis: Criterion G requires tracking participants until graduation from their educational stage. Participants were adults (20-35) enrolled in ongoing intermediate courses at private language institutes; the paper defines no graduation endpoint and tracks learners only to a five-month follow-up. The authors themselves state that longer-term evaluations would still be needed, confirming tracking ended at follow-up. No follow-up publications tracking this cohort further are referenced in the paper or found via internet search. In addition, under the ranking rules, criterion G cannot be met when criterion Y is not met, and Y is not met here. Criterion G is not met because tracking stopped at the five-month follow-up with no graduation endpoint, no follow-up publication was found by internet search, and criterion Y is not met.
    • P

      Pre-Registered

      • No pre-registration statement, registry ID, or registration date is provided in the paper, and none was found in the IRCT or ClinicalTrials.gov registries via internet search.
      • Relevant Quotes: 1) "Informed consent was obtained from all participants, and ethical approval was secured from the institutional review board." (p. 3) 2) "In this study, to observe ethical considerations, participants were informed about the goals and importance of the research before the start of the interview and participated in the research with informed consent." (p. 7) Internet Search Findings: Searches of the Iranian Registry of Clinical Trials (IRCT, irct.ir) for "Molaei chatbot" and related terms returned zero results ("Displaying 0-0 of 0 results"), and a search of ClinicalTrials.gov for chatbot/speaking-anxiety/language learner trials likewise returned no matching entry for this study. No trial registry entry for this study could be located. Detailed Analysis: Criterion P requires pre-registration of the full study protocol (hypotheses, methods, planned analyses) on a registry before data collection began, with a verifiable ID and date. The paper mentions only institutional ethical approval and informed consent, with no trial registry (e.g., ClinicalTrials.gov, IRCT, OSF) reference, registration number, or registration date anywhere in the article. Internet search of the IRCT and ClinicalTrials.gov registries found no matching pre-registration for this trial. Ethics approval is not a substitute for pre-registration. Criterion P is not met because the paper contains no reference to any pre-registered protocol or registry entry, and no such registration was located via internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.