Abstract
With the advancement of artificial intelligence, natural language processing, and speech recognition, conversational agents have emerged as promising tools for second language acquisition. This study designed an English conversational bot to help learners of English as a second language among Chinese university students learning English. The bot was implemented as a web application using Python Flask. A six-day comparative study was conducted with 56 students from the English department, who were randomly assigned to either the conversational bot or a traditional listen-and-repeat interface. To evaluate its feasibility, we compared it with an old listen-and-repeat interface using 56 Chinese college students from the English department over six days. We conducted Shapiro-Wilk normality tests, followed by t-tests to evaluate time spent on the application, engagement, anxiety change, vocabulary, and speaking tests. Findings shows that the conversational bot group reported significantly higher engagement scores (M = 4.37, SD = 0.44) compared to the control group (M = 3.87, SD = 0.51; t (26) = 2.8, p <.05). While both groups showed a reduction in language anxiety, the difference was not statistically significant. Participants using the bot demonstrated greater vocabulary use as responders (t (26) = 3.5, p <.005) and higher speaking test gains (p <.0001). In conclusion our findings suggest that the conversational bot is a more engaging and effective platform for improving spoken English proficiency. The tool shows potential for supporting learners in preparing for international academic environments.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was at the student level, but the intervention is a one-to-one tutoring-style conversational bot, so the personal-teaching exception applies.
- "Xuéxí plays the main role of being a female tutor as well as the practice companion to the student."
Relevant Quotes:
1) "A six-day comparative study was conducted with 56 students from the English department, who were randomly assigned to either the conversational bot or a traditional listen-and-repeat interface." (p. 23393, Abstract)
2) "The users were then randomly selected to and allocated the interface for use, either an old-style or a Conversational bot for the period of evaluation." (p. 23399-23400, Section 3.5)
3) "Xuéxí plays the main role of being a female tutor as well as the practice companion to the student." (p. 23402, Section 3.9)
4) "Therefore, Chatbots help bridge this gap by offering a safe space in which college learners can practice without feeling judged for the errors they make also through the application, the students can make numerous repeats and practice without the bot feeling fatigued or occupied" (p. 23394, Introduction)
Detailed Analysis:
Randomisation was performed at the individual student level: 56 students were "randomly assigned" to one of the two interfaces, with no mention of classes or schools as the unit of allocation. On a strict reading this fails the class-level requirement. However, the ERCT standard provides an exception: "If an intervention is designed for personal teaching like tutoring then this Class-level RCT criterion isn't applicable and even normal student-level RCT is considered OK." The intervention here is a one-to-one conversational tutoring bot: the paper explicitly frames the system as a personal "tutor as well as the practice companion" that substitutes for one-on-one practice with a fluent speaker, and each learner interacts with it individually on their own machine. The control condition is likewise individually delivered software, and access to each interface was controlled by the researchers, limiting cross-group contamination. Because the intervention is personal, tutoring-style one-to-one teaching, the exception applies and student-level randomisation is acceptable.
Criterion C is met via the tutoring/personal-teaching exception because the bot is an individually delivered one-to-one tutoring intervention, making student-level randomisation acceptable under the standard.
-
E
Exam-based Assessment
- Outcomes were measured with custom speaking and vocabulary tests built by the authors (only loosely adapted from IELTS materials), not a recognised standardised exam.
- "Each tutor graded the learners participating using the rubric which we adapted from (IELTS)."
Relevant Quotes:
1) "We sourced the learning materials from IELTS a standardized English examination platform that covers a wide range of scenario scenes that international students could get in English university tests." (p. 23395)
2) "Each tutor graded the learners participating using the rubric which we adapted from (IELTS). The tutors graded each depending on the grammatical coverage, fluency, pronunciation and accurate level using a scale of 0 to 9." (p. 23399, Section 3.4)
3) "In the self, the learner made a conversation with an English fluent-speaking tester randomly which are in the pre-designed pool of 9 questions." (p. 23399, Section 3.4)
4) "To do this, we gave 36 Chinese words and the participants needed to translate them into English." (p. 23399, Section 3.4)
5) "We set two test materials for both the scripted and self-speaking tests." (p. 23399, Section 3.4)
Detailed Analysis:
The outcome measures were all designed by the researchers for this study: a self-speaking test drawn from a "pre-designed pool of 9 questions", a scripted speaking test using pre-written scripts created by the authors, and a 36-item Chinese-to-English vocabulary translation test built from the lesson content. Although the learning materials were sourced from the IELTS website and the grading rubric was "adapted from (IELTS)", the participants did not sit an actual standardised, widely recognised exam such as the official IELTS test administered under standard conditions. Borrowing a rubric and materials from a standardised exam does not make the researcher-constructed tests themselves standardised; the assessments were custom instruments aligned to the taught content, which is exactly the bias the E criterion guards against.
Criterion E is not met because outcomes were measured with custom, researcher-designed speaking and vocabulary tests rather than an actual standardised exam.
-
T
Term Duration
- The intervention and outcome measurement spanned only six days plus a three-week delayed vocabulary test, far less than one academic term.
- "A six-day comparative study was conducted with 56 students from the English department"
Relevant Quotes:
1) "A six-day comparative study was conducted with 56 students from the English department" (p. 23393, Abstract)
2) "Each of them took 6 days consecutively. On the first day, users were given a pre-survey that contained an anxiety measure for a foreign language." (p. 23399, Section 3.5)
3) "On the study's last day, post-survey was conducted using the same tester whereby participants were subjected to the same anxiety measure as on the first day, including the vocabulary and speaking tests and a similar engagement style." (p. 23400, Section 3.5)
4) "And out of which 26 and 19 from group 1 and 2 respectively finished the three-week afterward follow-up study." (p. 23398, Section 3.2)
5) "Figure 6 also shows results from a three-week delayed vocabulary test" (p. 23409, Section 4.4)
Detailed Analysis:
The intervention lasted six consecutive days, with primary outcomes (speaking, vocabulary, anxiety, engagement) measured on day 6, immediately at the end of the intervention. The longest follow-up was a delayed vocabulary test three weeks later (week 4). The total interval from intervention start to the final measurement is therefore roughly one month, far short of the required full academic term (approximately 3-4 months). The authors themselves acknowledge "the short-term nature of the study" (p. 23414).
Criterion T is not met because outcomes were measured six days to three weeks after the intervention began, well short of one full academic term.
-
D
Documented Control Group
- The control group's condition, baseline measures and treatment are documented in the methods and result tables, meeting the documentation requirement.
- "In this study we created a non-AI style which has an interface that listens and repeats based on English learning software (Lanren-English) used for speaking"
Relevant Quotes:
1) "In this study we created a non-AI style which has an interface that listens and repeats based on English learning software (Lanren-English) used for speaking" (p. 23404, Section 3.12)
2) "The system only allows the student to replay as many times as they want but they cannot input anything to the system." (p. 23404, Section 3.12)
3) "Their age was averagely 18.36 years (sigma=3.92) and they studied in 12 different colleges in China and all of the schools offered English classes which equips them to join Western schools outside China. The learners are Chinese students who started learning English at their college level and have never been to English-speaking nations before." (p. 23398, Section 3.2)
4) "Old system before test ... Cognitive anxiety 4.3 ... Somatic anxiety 4.2 ... Behavioral anxiety 4.2" (p. 23408, Table 3)
5) "Both the old listen and repeat interface which we use as our baseline and our designed AI English conversational bot are loaded with the same four conversations which act as our lessons for tutoring." (p. 23401, Section 3.8)
Detailed Analysis:
The control condition is described in detail: control learners used a listen-and-repeat interface modelled on Lanren-English, loaded with the same four IELTS-derived conversations, over the same six days. Baseline data for the control group are reported: per-student time-on-task (Table 1), baseline anxiety sub-scores for the old system group (Table 3), pre-test vocabulary and speaking scores used to compute gains, and engagement ratings (Table 2). Demographics (age, colleges, prior English exposure) are reported for the sample as a whole rather than per group, and the participant flow reporting is somewhat inconsistent (60 recruited, 4 pilots, groups of 28, 26 and 19 completing follow-up). Despite these reporting weaknesses, the control group's size, condition, baseline performance and the fact that it received no special treatment beyond the alternative interface are documented, which satisfies the criterion's core requirement of enabling proper comparison.
Criterion D is met because the control condition, its baseline scores and its treatment (same materials via a listen-and-repeat interface) are clearly documented, though demographics are only reported at the whole-sample level.
-
Level 2 Criteria
-
S
School-level RCT
- Individual students, not schools or institutional units, were randomly assigned to the two interfaces.
- "The users were then randomly selected to and allocated the interface for use, either an old-style or a Conversational bot for the period of evaluation."
Relevant Quotes:
1) "A six-day comparative study was conducted with 56 students from the English department, who were randomly assigned to either the conversational bot or a traditional listen-and-repeat interface." (p. 23393, Abstract)
2) "The users were then randomly selected to and allocated the interface for use, either an old-style or a Conversational bot for the period of evaluation." (p. 23399-23400, Section 3.5)
3) "they studied in 12 different colleges in China" (p. 23398, Section 3.2)
Detailed Analysis:
Randomisation was carried out at the individual student level. Although participants came from 12 different colleges, there is no indication anywhere in the paper that colleges, schools, sites or any institutional units were the unit of random assignment; individuals were allocated to interfaces. The S criterion requires randomisation among schools or equivalent implementing units, which clearly did not occur.
Criterion S is not met because randomisation was performed at the individual student level, not at the school or institutional level.
-
I
Independent Conduct
- The same authors designed the bot and also ran, measured and analysed the evaluation, with no independent evaluators documented.
- "We implemented the conversational bot system using Python Flask as a web application with two independent panes and named it Xuéxí"
Relevant Quotes:
1) "In this study, we developed an artificially intelligent English conversational bot which is to be used in college students who are learning English" (p. 23395)
2) "We implemented the conversational bot system using Python Flask as a web application with two independent panes and named it Xuéxí" (p. 23401, Section 3.9)
3) "For testers, we involved three native tutors who conducted evaluations through Zoom video calls. Each tutor graded the learners participating using the rubric which we adapted from (IELTS)." (p. 23399, Section 3.4)
4) "Competing interests The authors have no competing interest to declare." (p. 23414, Declarations)
Detailed Analysis:
The same two-author team designed the intervention (the conversational bot), built both systems, ran the experiment, administered the measures and analysed the data. There is no statement of an external evaluation team, third-party oversight, or independent data collection and analysis. The three native-speaker tutors who graded the speaking tests were engaged by the authors and are not described as independent of the study team or blinded to condition. A generic "no competing interests" declaration does not establish independent conduct as defined by the criterion.
Criterion I is not met because the intervention designers themselves conducted, measured and analysed the study with no documented independent oversight.
-
Y
Year Duration
- The study covered only six days plus a three-week follow-up, far below the required 75% of an academic year.
- "Furthermore, the short-term nature of the study may not capture long-term engagement, retention, or fluency development associated with the use of conversational bots."
Relevant Quotes:
1) "Each of them took 6 days consecutively." (p. 23399, Section 3.5)
2) "And out of which 26 and 19 from group 1 and 2 respectively finished the three-week afterward follow-up study." (p. 23398, Section 3.2)
3) "Furthermore, the short-term nature of the study may not capture long-term engagement, retention, or fluency development associated with the use of conversational bots." (p. 23414, Section 7)
Detailed Analysis:
The Y criterion requires outcome tracking covering at least 75% of an academic year (roughly 9-10 months) from intervention start. This study spans six days of intervention plus a three-week delayed vocabulary test, about one month in total. The authors themselves flag the short-term nature as a limitation and call for longitudinal studies. Additionally, since criterion T (term duration) is not met, criterion Y cannot be met per the instructions.
Criterion Y is not met because tracking lasted about one month, nowhere near 75% of an academic year.
-
B
Balanced Control Group
- Both groups used software with identical content over the same period; the bot's interactive features are the integral treatment variable, so the groups are balanced.
- "Both the old listen and repeat interface which we use as our baseline and our designed AI English conversational bot are loaded with the same four conversations which act as our lessons for tutoring."
Relevant Quotes:
1) "Both the old listen and repeat interface which we use as our baseline and our designed AI English conversational bot are loaded with the same four conversations which act as our lessons for tutoring." (p. 23401, Section 3.8)
2) "In the first study (Fixed usage), participants were instructed to follow the guidelines of the learning systems precisely." (p. 23400, Section 3.6)
3) "In the second study 2 (Voluntary usage), participants engaged with the systems at their discretion to improve their oral English skills over six days. Each day, learners were granted access to a unit ... Usage time was not restricted" (p. 23400, Section 3.6)
4) "This article makes an experimental study of an Artificial intelligence conversational bot against a listen and repeated old interface for the spoken English language." (p. 23395)
5) "The findings show that in the first study (Fixed), controlled use condition, learners spent average total of 45.07 (sigma=11.14) and 103.74 (sigma=29.63) minutes on the app respectively on old and cconversational bot" (p. 23405, Section 4.1)
Detailed Analysis:
Applying the criterion B decision tree: the control group received an active alternative system, not nothing. Both groups received software access over the same six days, loaded with identical learning materials (the same four IELTS-derived conversations), the same daily unit schedule, the same hardware setting and the same assessments. No extra budget, teacher time or materials were given to the intervention group beyond the bot's interactive features (speech recognition, adaptive feedback), and those features are precisely the treatment variable being tested - the study is explicitly framed as evaluating the AI conversational interface against the listen-and-repeat interface. Bot users did accumulate more time on the application (e.g., 103.74 vs 45.07 minutes in the fixed study), but this arises from the interactive nature of the intervention itself (responding, recording, receiving feedback) and, in the voluntary study, from learners' self-chosen engagement, which is an outcome of interest rather than a separately provided resource. The additional interactive capability is integral to the intervention being tested against a matched active control using identical content.
Criterion B is met because the control group received an active comparison system with identical learning materials and schedule, and the interactive bot features (and the engagement time they generate) are the integral treatment variable being tested.
-
Level 3 Criteria
-
R
Reproduced
- The study is presented as novel, was published in mid-2025, and an internet citation search found no independent replication by another team.
- "This study addresses that gap by designing and evaluating a novel English conversational bot tailored for college students learning English as a second language."
Relevant Quotes:
1) "Received: 16 November 2024 / Accepted: 27 May 2025 / Published online: 10 July 2025" (p. 23393)
2) "This study addresses that gap by designing and evaluating a novel English conversational bot tailored for college students learning English as a second language." (p. 23395)
3) "Thus, future research should focus on long-term studies to assess the effectiveness of conversational bots in language learning." (p. 23413, Section 5)
Detailed Analysis:
The paper presents itself as a novel system and a new evaluation; it does not reference any independent replication of this specific study. Related chatbot studies cited (e.g., Yuan, 2024 on Liulishuo; Fryer et al., 2019) are different interventions by different teams and do not constitute replications of this particular Xuéxí bot trial. An internet citation-index check (OpenAlex work W4412175336, checked 2026-07-27) found nine works citing this paper: Liu, Keane, Sun, Yang & Yang (2025, an AI-enhanced system for maths word problems), Samala & Rawas (2025, an AI/AR/NLP integration review), Yang, Zhang, Liu, Yan, Tu, Yang & Zeng (2025, a reflective feedback dialogue model for EFL speaking), Lu & Zhang (2025), Yang (2025), Wang (2025), Bui & Hoa (2025), Zhang (2025), and Guo, Swaran Singh & Li (2025/2026). None of these authors are Dai or Wu, and each addresses a different system, subject, or a broad review topic rather than reproducing this specific Xuéxí-versus-listen-and-repeat trial. No independent replication of this specific study was found in any source checked.
Criterion R is not met because no independent replication of this specific study by a different team has been published or referenced.
-
A
All-subject Exams
- Only custom English vocabulary and speaking measures were used; no standardised exams across all main subjects, and prerequisite criterion E is not met.
- "In both studies, we evaluated three main metrics: Time spent on the application, engagement and anxiety change, Vocabulary Test and lastly Speaking Test."
Relevant Quotes:
1) "In both studies, we evaluated three main metrics: Time spent on the application, engagement and anxiety change, Vocabulary Test and lastly Speaking Test." (p. 23405, Section 4)
2) "To do this, we gave 36 Chinese words and the participants needed to translate them into English." (p. 23399, Section 3.4)
Detailed Analysis:
The study measured only English speaking proficiency, vocabulary, engagement and anxiety. No other academic subjects taken by these college students were assessed. Moreover, criterion E is a prerequisite for criterion A, and E is not met because the assessments were researcher-designed rather than standardised exams; per the instructions, A therefore cannot be met. While one could argue the intervention is specialised to English speaking, the paper provides no standardised all-subject assessment and no explicit rationale framed against other subjects.
Criterion A is not met because only custom English speaking and vocabulary measures were used, criterion E fails, and no other subjects were assessed.
-
G
Graduation Tracking
- Participants were followed for only three weeks after the six-day intervention, with no tracking to graduation, and no follow-up publication by the authors was found.
- "And out of which 26 and 19 from group 1 and 2 respectively finished the three-week afterward follow-up study."
Relevant Quotes:
1) "And out of which 26 and 19 from group 1 and 2 respectively finished the three-week afterward follow-up study." (p. 23398, Section 3.2)
2) "Furthermore, the short-term nature of the study may not capture long-term engagement, retention, or fluency development associated with the use of conversational bots." (p. 23414, Section 7)
3) "Thus, future research should focus on long-term studies to assess the effectiveness of conversational bots in language learning." (p. 23413, Section 5)
Detailed Analysis:
The longest follow-up was a three-week delayed vocabulary test; there is no tracking of participants to graduation from their degree programmes, and the authors explicitly recommend future long-term studies, confirming none exists for this cohort. No follow-up publications tracking these participants are referenced in the paper. An internet search (OpenAlex citation index for this paper, checked 2026-07-27) for subsequent publications by Lili Dai or Fengming Wu that track this cohort toward graduation returned no results; none of the nine papers citing this study are authored by Dai or Wu. Additionally, criterion Y is not met, so per the instructions criterion G cannot be met.
Criterion G is not met because tracking ended three weeks after the intervention, with no graduation follow-up found in this paper or in any subsequent publication.
-
P
Pre-Registered
- The paper contains no reference to any pre-registered protocol, registry ID, or registration date, and an internet search found no such record either.
Relevant Quotes:
1) "All students were asked for their consent before the start of the study and advised they are kept anonymous in the process of publishing." (p. 23398, Section 3.1)
2) "Data availability Data used are within this manuscript." (p. 23414, Declarations)
3) "Competing interests The authors have no competing interest to declare." (p. 23414, Declarations)
Detailed Analysis:
The paper contains no mention of a pre-registered protocol, no trial registry name or ID (e.g., ClinicalTrials.gov, OSF, AsPredicted, ISRCTN), and no registration date. The declarations section covers only consent, data availability and competing interests. An internet search for a pre-registration record under the authors' names, the institution (Hunan University of Arts and Science), or the study title (checked 2026-07-27) returned no matches on trial registries or protocol repositories. Without any quoted evidence or external record that hypotheses, methods and analysis plans were registered before data collection, the criterion fails by default.
Criterion P is not met because no pre-registration of the study protocol is mentioned anywhere in the paper, and none was found via internet search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.