Abstract
This study investigated the effectiveness of artificial intelligence-based instruction in improving second language (L2) speaking skills and speaking self-regulation in a natural setting. The research was conducted with 93 Chinese English as a foreign language (EFL) students, randomly assigned to either an experimental group receiving AI-based instruction or a control group receiving traditional instruction. The AI-based instruction leveraged the Duolingo application, incorporating natural language processing technology, interactive exercises, personalized feedback, and speech recognition technology. Pre- and post-tests were conducted to assess L2 speaking skills and self-regulation abilities. The results demonstrated that the experimental group exhibited significantly greater improvement in L2 speaking skills compared to the control group, and reported higher levels of self-regulation.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomization was conducted at the individual student level within institutes, not at the class or school level, and the classroom-delivered group intervention does not qualify for the one-to-one tutoring exception.
- "In total, 47 students were randomly assigned to the AI-based instruction (age: M = 20.2, SD = 1.8; 45.5% female), and 46 students were assigned to the traditional speaking course (age: M = 20.1, SD = 1.7; 37.5% female)." (p. 5)
Relevant Quotes:
1) "The research was conducted with 93 Chinese English as a foreign language (EFL) students, randomly assigned to either an experimental group receiving AI-based instruction or a control group receiving traditional instruction." (p. 1)
2) "To ensure a fair and unbiased distribution of students between the control and experimental groups across all participating institutes, a blocked randomization technique, aided by computer-generated random numbers, was applied." (p. 5)
3) "In total, 47 students were randomly assigned to the AI-based instruction (age: M = 20.2, SD = 1.8; 45.5% female), and 46 students were assigned to the traditional speaking course (age: M = 20.1, SD = 1.7; 37.5% female)." (p. 5)
4) "Additionally, in order to maintain experimental fidelity and prevent potential contamination of the control group, experimental students were placed in separate classrooms from the control group." (p. 7)
Detailed Analysis:
The unit of randomization was clearly the individual student: 47 students were assigned to the AI-based course and 46 to the traditional course via blocked randomization with computer-generated random numbers. Entire classes or schools were not randomized; instead, students within the participating institutes were individually allocated and then grouped into separate classrooms after assignment. Placing the groups in separate classrooms mitigates but does not replace class-level randomization. The exception for personal tutoring does not apply: although Duolingo gives individualized feedback, the intervention was delivered as a teacher-led course with group activities and discussions in classrooms, not as one-to-one personal teaching.
Criterion C is not met because students, not classes or schools, were the unit of randomization and the tutoring exception does not apply.
-
E
Exam-based Assessment
- Speaking outcomes were measured with the IELTS speaking examination scored with the official IELTS Speaking Band Descriptors, a widely recognized standardized exam.
- "The speaking abilities of Chinese EFL learners were evaluated using the IELTS speaking skill examination." (p. 5)
Relevant Quotes:
1) "The speaking abilities of Chinese EFL learners were evaluated using the IELTS speaking skill examination. This assessment encompassed four equally weighted components, namely fluency and coherence, vocabulary, grammatical range and accuracy, and pronunciation." (p. 5)
2) "The learners' performance in each area was evaluated based on the topics provided in the IELTS speaking test. The IELTS Speaking Band Descriptors were employed to assign scores ranging from 1 to 9 to each learner in each speaking skill category." (p. 5)
3) "The inter-rater reliability, assessed using the Cohen's kappa coefficient, yielded a satisfactory value of 0.87." (p. 6)
4) "Firstly, the global English proficiency of participants was assessed using the College English Test 3 (CET-3). CET-3 is a well-established standardized English proficiency examination widely recognized in China." (p. 6)
Detailed Analysis:
The primary outcome, L2 speaking skills, was assessed with the IELTS speaking test, an internationally recognized, standardized examination, scored with the official IELTS Speaking Band Descriptors (1-9 bands) rather than a custom-made instrument, with good inter-rater reliability (kappa = 0.87). In addition, the standardized CET-3 was used to measure global English proficiency as a control variable. The secondary self-regulation outcome used a validated questionnaire (SRLLQ), but the education outcome of interest (speaking skill) rests on a standard exam framework, which is what the criterion requires.
Criterion E is met because outcomes were measured with the standardized IELTS speaking examination rather than a researcher-made test.
-
T
Term Duration
- The intervention and outcome tracking spanned 13 weeks from the start of instruction to the posttest, which covers approximately one academic term (about three months).
- "Both groups received instruction across a span of 13 weeks, with each week comprising one session." (p. 5)
Relevant Quotes:
1) "Both groups received instruction across a span of 13 weeks, with each week comprising one session." (p. 5)
2) "Pretest measurements were collected before the intervention started, embedded within the first two course sessions, and posttest measurements were collected in the last course session." (p. 5)
Detailed Analysis:
The intervention began with the first course sessions (which also embedded the pretest) and ran for 13 weekly sessions, with the posttest collected in the last course session. The interval from intervention start to outcome measurement is therefore about 13 weeks, i.e., roughly three months. The ERCT standard defines a term as a semester or equivalent (approximately 3-4 months), so a 13-week span from start to measurement reaches the lower bound of one full academic term.
Criterion T is met because outcomes were measured 13 weeks after the intervention began, which corresponds to approximately one academic term.
-
D
Documented Control Group
- The control group's size, demographics, baseline scores, and business-as-usual traditional speaking course are documented in detail, with baseline equivalence tests reported.
- "In contrast, the control group received instruction through a more traditional speaking course. This course focused on facilitating speaking skills development through group discussions, activities, role plays, and presentations." (p. 7)
Relevant Quotes:
1) "46 students were assigned to the traditional speaking course (age: M = 20.1, SD = 1.7; 37.5% female)." (p. 5)
2) "In contrast, the control group received instruction through a more traditional speaking course. This course focused on facilitating speaking skills development through group discussions, activities, role plays, and presentations. Although this course provided learners with opportunities to practice their speaking skills in a supportive and structured environment, it did not utilize AI technology or offer personalized feedback on learners' speaking performance." (p. 7)
3) "Table 1 shows the means and standard deviations for each group in the pre-and post-tests." (p. 8)
4) "The results of these tests indicated that there were no statistically significant differences between the experimental and control groups on any of the pre-test variables (p > 0.05)." (p. 8)
Detailed Analysis:
The paper documents the control group's size (n = 46), age, and gender composition, describes precisely what instruction the control group received (a traditional speaking course with group discussions, activities, role plays, and presentations, without AI), and reports baseline means and standard deviations for all outcome and control variables in Table 1, together with t-tests confirming baseline equivalence. Overall sample demographics (mean age 21.36, 97% Chinese L1, intermediate proficiency) are also given. This satisfies the requirement for detailed control group documentation.
Criterion D is met because the control group's composition, baseline performance, and conditions are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomization occurred at the individual student level within institutes; entire schools or institutes were not randomly assigned to conditions.
- "To ensure a fair and unbiased distribution of students between the control and experimental groups across all participating institutes, a blocked randomization technique, aided by computer-generated random numbers, was applied." (p. 5)
Relevant Quotes:
1) "These students were enrolled in one of four conversation courses offered by five different institutes that offered the AI-based speaking instruction." (p. 5)
2) "At each institute, one control group participated in a traditional speaking course, while the intervention group received the AI-based speaking instruction." (p. 5)
3) "To ensure a fair and unbiased distribution of students between the control and experimental groups across all participating institutes, a blocked randomization technique, aided by computer-generated random numbers, was applied." (p. 5)
Detailed Analysis:
School-level randomization requires that entire educational institutions or implementation units be randomly assigned to conditions. Here, every participating institute hosted both an intervention group and a control group, and individual students were randomized (in blocks) into one or the other within each institute. No institutes were assigned wholly to treatment or control, so the design is a student-level RCT stratified by institute, not a school-level RCT.
Criterion S is not met because randomization was at the student level within institutes rather than across schools or institutes.
-
I
Independent Conduct
- The researchers themselves designed the study, delivered the instruction with teachers they trained, and one author personally co-scored the speaking tests, with no independent third-party evaluation.
- "Both courses were offered simultaneously, with the AI-based instruction being delivered by a team of researchers and two trained English teachers who collaborated with the present researcher." (p. 5)
Relevant Quotes:
1) "Both courses were offered simultaneously, with the AI-based instruction being delivered by a team of researchers and two trained English teachers who collaborated with the present researcher." (p. 5)
2) "To ensure consistency, the speaking skills of the learners were evaluated by two proficient assessors, consisting of the researcher and another experienced instructor specialized in teaching EFL speaking." (p. 5)
3) "The proficiency of the English teachers in delivering the intervention and control courses was regularly evaluated by the research team." (p. 7)
4) "The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest." (p. 12)
Detailed Analysis:
Although the Duolingo application itself was developed by a third party, the tested intervention (the AI-based speaking course) was designed, delivered, monitored, and analyzed by the authors' own team: instruction was delivered by "a team of researchers and two trained English teachers who collaborated with the present researcher," treatment fidelity was monitored by the research team, and outcome scoring was done partly by the researcher personally. There is no external evaluation agency, no blinded independent data collectors, and no statement of third-party oversight of data collection or analysis. The absence of a financial conflict of interest does not establish independent conduct of the evaluation.
Criterion I is not met because the same team that designed and delivered the intervention also collected, scored, and analyzed the outcome data without independent oversight.
-
Y
Year Duration
- The study lasted only 13 weeks from intervention start to posttest, far short of 75% of an academic year.
- "Both groups received instruction across a span of 13 weeks, with each week comprising one session." (p. 5)
Relevant Quotes:
1) "Both groups received instruction across a span of 13 weeks, with each week comprising one session." (p. 5)
2) "Pretest measurements were collected before the intervention started, embedded within the first two course sessions, and posttest measurements were collected in the last course session." (p. 5)
3) "Another limitation pertains to the duration of the AI-based instruction and the length of exposure to the intervention, as the impact on speaking skills and self-regulation could vary depending on the intervention's duration." (p. 11)
Detailed Analysis:
Criterion Y requires outcome measurement at least 75% of a full academic year (roughly 9-10 months, i.e., about 7+ months minimum) after the intervention begins. Here the whole study, from the first session to the posttest in the last session, spanned 13 weeks (about 3 months). The authors themselves list the short duration as a limitation and call for longer intervention periods. No longer-term follow-up measurement is reported.
Criterion Y is not met because the 13-week tracking period falls well short of 75% of an academic year.
-
B
Balanced Control Group
- Both groups received an equal amount of instructional time in comparable contexts, with the control group taking an active traditional speaking course, so inputs were balanced.
- "To ensure the integrity and validity of our study, both the intervention and control groups received an equal amount of instructional time and were exposed to comparable learning contexts, except for the differing instructional methods described above." (p. 7)
Relevant Quotes:
1) "To ensure the integrity and validity of our study, both the intervention and control groups received an equal amount of instructional time and were exposed to comparable learning contexts, except for the differing instructional methods described above." (p. 7)
2) "In contrast, the control group received instruction through a more traditional speaking course. This course focused on facilitating speaking skills development through group discussions, activities, role plays, and presentations." (p. 7)
3) "Both groups received instruction across a span of 13 weeks, with each week comprising one session." (p. 5)
4) "This checklist covered essential aspects such as the duration of the speaking activities, the specific types of activities employed, and the nature of feedback provided to learners. English teachers responsible for delivering the speaking instruction diligently completed this checklist for each session." (p. 7)
Detailed Analysis:
Following the decision tree: the intervention group's extra input relative to the control is the Duolingo application with its AI chatbot and personalized feedback. Both groups received the same number of weekly sessions over 13 weeks, the same overall instructional time, and comparable learning contexts, with an active control course containing parallel speaking activities (discussions, role plays, presentations). Adherence checklists monitored activity duration in both arms. The only systematic difference - use of the AI application and its feedback - is the treatment variable itself, integral to the intervention being tested (AI-based versus traditional instruction). No additional instructional time or budgeted educational resources were given to the intervention group beyond the AI tool under test.
Criterion B is met because instructional time and learning contexts were explicitly equalized, and the AI application itself is the integral treatment variable being tested against an active traditional course.
-
Level 3 Criteria
-
R
Reproduced
- No independent, peer-reviewed replication of this specific Duolingo-based RCT with Chinese EFL learners is reported or known; related AI-speaking studies are different designs by different teams, not replications of this study.
Relevant Quotes:
1) "However, despite the promising strides made in understanding the impact of AI on language learning, a notable gap persists in comprehending its effects in specific contexts, such as the Chinese EFL environment examined in this study." (p. 4)
2) "The study also recognizes the need for further research to investigate the long-term effects of AI-based instruction and delve into the specific mechanisms underlying the observed improvements." (p. 11)
Detailed Analysis:
The paper positions itself as filling a research gap, not as a replicated study. Prior related work cited by the authors (e.g., Junaidi 2020 with the Lyra app, Maknun 2020 with Orai, Kang 2022, El Shazly 2021, Hsu et al. 2021 with Amazon Alexa) involves different applications, designs, and populations, and predates this study, so it does not constitute a replication of this specific Duolingo RCT in the Chinese EFL context. No subsequent independent, peer-reviewed replication of this particular study by a different team is referenced in the paper or identified.
A targeted search of the papers citing this study (45 citing works identified via Semantic Scholar as of July 2026) found no independent replication of this specific Duolingo-based AI speaking intervention with Chinese EFL learners by a different research team. Related work in the citing literature (e.g., studies on generative AI for speaking proficiency among Chinese vocational students) represents distinct research initiatives with different designs, applications, and populations, not replications of this study.
Criterion R is not met because no independent replication of this specific study has been documented.
-
A
All-subject Exams
- Only English speaking skills (plus self-regulation) were assessed; no other main academic subjects were measured with standardized exams and no specialization rationale covering this is given.
- "In this Randomized Controlled Trial (RCT), we employed two dependent variables: speaking skills, operationalized as a composite of skills encompassing fluency, vocabulary, accuracy, and pronunciation, and self-regulation." (p. 2)
Relevant Quotes:
1) "In this Randomized Controlled Trial (RCT), we employed two dependent variables: speaking skills, operationalized as a composite of skills encompassing fluency, vocabulary, accuracy, and pronunciation, and self-regulation." (p. 2)
2) "To assess the effectiveness of the AI-based speaking instruction on the students' speaking skills, four measures of speaking skill components were used: fluency, vocabulary, accuracy, and pronunciation." (p. 5)
Detailed Analysis:
The study measured outcomes exclusively in one domain: L2 English speaking (four sub-components of one IELTS speaking score) plus self-reported self-regulation. No other core subjects studied by these students at their institutes and universities (e.g., other academic disciplines) were assessed, and the paper offers no explicit rationale of the kind the standard's exception envisions for measuring only related subjects. While the participants were adults in language courses, the paper does not argue or document that English speaking is the sole relevant subject of their education, so the specialized-intervention exception cannot be confidently applied. Criterion E is met, so the prerequisite holds, but coverage of all main subjects does not.
Criterion A is not met because only English speaking outcomes were assessed, without justification covering all main subjects of the participants' education.
-
G
Graduation Tracking
- Measurement stopped at the 13-week posttest with no tracking of participants to graduation, and criterion G also fails automatically because criterion Y is not met.
- "The study also recognizes the need for further research to investigate the long-term effects of AI-based instruction and delve into the specific mechanisms underlying the observed improvements." (p. 11)
Relevant Quotes:
1) "Pretest measurements were collected before the intervention started, embedded within the first two course sessions, and posttest measurements were collected in the last course session." (p. 5)
2) "The study also recognizes the need for further research to investigate the long-term effects of AI-based instruction and delve into the specific mechanisms underlying the observed improvements." (p. 11)
3) "However, further research is needed to explore the long-term effects and specific mechanisms underlying these observed improvements." (p. 1)
Detailed Analysis:
Outcome data collection ended with the posttest in the final (13th) course session. Although the authors mention a planned crossover ("After the study, all students were invited to participate in the respective other course"), no graduation-tracking data or follow-up outcomes are reported, and the authors explicitly flag long-term effects as an open question for future research. No follow-up publications tracking this cohort to graduation are referenced.
A targeted search for subsequent publications by Hongliang Qiao or Aruna Zhao tracking this same cohort of 93 Chinese EFL students found no such follow-up study as of July 2026.
In addition, per the ranking instructions, criterion G cannot be met when criterion Y is not met.
Criterion G is not met because tracking stopped at the immediate posttest, no follow-up publication tracking this cohort was found, and criterion Y is also unmet.
-
P
Pre-Registered
- The paper contains no mention of a pre-registered protocol, registry platform, registration ID, or registration date, and no external registry record for this study was found.
Relevant Quotes:
1) "The studies involving humans were approved by Department of Foreign language, Baotou Teachers' College, Inner Mongolia University of Science and Technology, China." (p. 12)
2) "To evaluate the effectiveness of the AI-based speaking instruction, a randomized controlled trial with repeated measures was conducted (Deaton and Cartwright, 2018)." (p. 5)
Detailed Analysis:
The paper reports institutional ethics approval but nowhere mentions pre-registration of the study protocol on any registry (e.g., ClinicalTrials.gov, OSF, AEA registry), provides no registration ID or link, and gives no registration date. Without quoted evidence of a protocol registered before data collection began, the criterion cannot be satisfied.
A check of the article's official landing page and associated metadata found no reference to a pre-registration record for this study.
Criterion P is not met because no pre-registration statement or registry reference appears anywhere in the paper or in the article's public metadata.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.