Abstract
This randomized controlled trial explored the effects of employing AI-driven methodologies on enhancing listening comprehension, flow experience, and alleviation of listening anxiety among English as a foreign language (EFL) learners. A cohort of 84 Chinese university students, primarily associated with English language programs, participated and were randomly assigned to either an experimental group utilizing AI-driven speech recognition technology or a control group receiving similar instruction without AI integration. Pre- and post-intervention assessments evaluated participants' listening comprehension abilities, flow state, and listening anxiety. To assess the sustainability of the intervention effects, a follow-up assessment was conducted 3 weeks after the post-intervention assessment. The linear mixed model analyses demonstrated sustained benefits associated with the AI-driven approach. The experimental group exhibited notable improvements in listening skill scores, significant enhancements in the sense of flow, and decreased levels of listening anxiety across all three assessment points. In contrast, the control group showed comparatively moderate advancements in listening skills, relatively stable flow experience, and marginal changes in listening anxiety over the same duration. These outcomes highlight the efficacy of AI-driven technology in optimizing language learning outcomes for EFL students, emphasizing its potential to enhance listening comprehension, cultivate a more immersive learning experience, and alleviate listening-related apprehensions among learners.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was performed at the individual student level within one university cohort, not at the class or school level, and the intervention is not one-to-one tutoring, so the class-level RCT requirement is not met.
- "The participants were randomly assigned to either the experimental group (n = 42) or the control group (n = 42) using a random number generator." (p. 5)
Relevant Quotes:
1) "A total of 84 university students of Chinese nationality enrolled in EFL courses at Hainan University in Haikou City, Hainan Province, China, participated in this randomized controlled trial." (p. 4)
2) "The participants were randomly assigned to either the experimental group (n = 42) or the control group (n = 42) using a random number generator." (p. 5)
3) "The study took place over eight weeks, employing a randomized assignment of participants into two groups: an experimental group utilizing AI-driven speech recognition technology and a control group receiving comparable instruction without AI integration." (p. 5)
4) "The intervention was conducted by two experienced EFL instructors--one for each group--who were regular faculty members at the university..." (p. 6)
Detailed Analysis:
The unit of randomisation is clearly the individual student: 84 students were allocated to the two arms with a random number generator. The randomised students were then taught in two separate group classes (one instructor per group), so this is a student-level RCT, not a design where pre-existing classes or schools were the randomised units. The ERCT exception for class-level randomisation applies only to personal tutoring or one-to-one teaching interventions; here the intervention is group classroom instruction plus app-based self-study, which does not qualify for the exception. Because the two arms were formed by individual-level assignment within the same university population, contamination between students in the two groups cannot be excluded by design.
Criterion C is not met because randomisation occurred at the individual student level and the tutoring exception does not apply.
-
E
Exam-based Assessment
- The primary educational outcome was measured with the IELTS Listening test, a widely recognised standardised exam format, rather than a researcher-made instrument.
- "To evaluate the participants' listening comprehension abilities, the original version of the IELTS Listening test (Scovell et al. 2004) was employed as both a pre- and post-intervention assessment tool." (p. 5)
Relevant Quotes:
1) "To evaluate the participants' listening comprehension abilities, the original version of the IELTS Listening test (Scovell et al. 2004) was employed as both a pre- and post-intervention assessment tool." (p. 5)
2) "This test encompassed diversified tasks, including multiple-choice queries, gap-filling exercises, and matching tasks, replicating the authentic format of the IELTS Listening sections." (p. 5)
3) "The post-test replicated the pre-test structure and complexity and was administered after the intervention to gauge the intervention's impact on participants' listening skills. The overall reliability of the test was computed using Cronbach's alpha, yielding values of 0.81 for the pre-test and 0.86 for the post-test..." (p. 5)
4) "These students were predominantly associated with English language courses tailored for proficiency enhancement, including preparation for standardized tests like the IELTS." (p. 4)
Detailed Analysis:
The listening comprehension outcome was assessed with IELTS Listening test material taken from a published IELTS test collection (Scovell et al. 2004), explicitly described as "the original version of the IELTS Listening test" and as "replicating the authentic format of the IELTS Listening sections". IELTS is an internationally recognised standardised English proficiency examination, and the instrument was not specially constructed by the author for this study; reliability figures are also reported. The flow and anxiety measures (PPL-FSQ, FLLAS) are self-report questionnaires, but they are secondary psychological outcomes rather than the educational achievement measure, which rests on the IELTS-based test. A minor caveat is that the test came from a published IELTS practice-test book rather than an official operational IELTS administration, but it is standard, widely recognised exam material rather than a custom study-specific test.
Criterion E is met because the educational outcome was measured with standardised IELTS Listening test material, not a custom-made assessment.
-
T
Term Duration
- Outcomes were measured at the end of an eight-week intervention and at a three-week follow-up, about 11 weeks after the start, which is shorter than a full academic term of roughly 3-4 months.
- "The study took place over eight weeks... To assess the sustainability of the intervention effects, a follow-up assessment was conducted 3 weeks after the post-intervention assessment." (pp. 1, 5)
Relevant Quotes:
1) "The study took place over eight weeks, employing a randomized assignment of participants into two groups..." (p. 5)
2) "Each group attended two in-class sessions per week, each lasting 90 min, totaling 16 sessions per group throughout the study period." (p. 5)
3) "Following the intervention, both groups underwent a post-assessment to gauge changes in listening comprehension, flow experience, and listening anxiety. Additionally, a follow-up session conducted three weeks later assessed the sustainability of any observed changes using similar assessments." (p. 6)
4) "To assess the sustainability of the intervention effects, a follow-up assessment was conducted 3 weeks after the post-intervention assessment." (p. 1)
5) "Firstly, the relatively short duration of the intervention and the specific demographic of participants drawn from a single educational setting may limit the generalizability of the results..." (p. 11)
Detailed Analysis:
The intervention lasted eight weeks, with the post-test (T2) administered immediately at the end of the intervention and the final follow-up (T3) three weeks later. The maximum interval from intervention start to the last outcome measurement is therefore approximately 11 weeks (about 2.5 months). The ERCT standard defines a term as approximately 3-4 months (a semester or equivalent), so an 11-week start-to-measurement window falls short of one full term. The author also explicitly acknowledges the "relatively short duration of the intervention" as a limitation. No longer-term tracking is reported.
Criterion T is not met because the interval from intervention start to the final measurement (about 11 weeks) is shorter than one full academic term.
-
D
Documented Control Group
- The control group's size, demographics, baseline scores, and the instruction it received are documented in detail, including baseline equivalence tests against the experimental group.
- "In the domain of listening skills, the experimental group exhibited a mean score of 4.23 (SD = 0.36), while the control group demonstrated a mean score of 4.11 (SD = 0.40)." (p. 8)
Relevant Quotes:
1) "The participants were randomly assigned to either the experimental group (n = 42) or the control group (n = 42) using a random number generator." (p. 5)
2) "The sample comprised 57% female students (48 participants) and 43% male students (36 participants)... Their ages ranged from 20 to 28 years old, with an average age of 21.06 years (SD = 2.08)..." (p. 4)
3) "The control group underwent a series of instructional sessions aimed at improving their listening comprehension skills, mirroring the experimental group's schedule and content. They also attended two 90-min in-class sessions per week over the eight-week period." (p. 7)
4) "With regard to the baseline comparisons, a set of independent samples t-test (see Table 3) was utilized to compare the initial attributes between the experimental and control groups across three pivotal variables: listening skill, sense of flow, and listening anxiety." (p. 8)
5) "In the domain of listening skills, the experimental group exhibited a mean score of 4.23 (SD = 0.36), while the control group demonstrated a mean score of 4.11 (SD = 0.40). The t-test, however, identified no significant variance in the baseline listening skill scores between the experimental and control groups (t = 0.97, p = 0.341)." (p. 8)
Detailed Analysis:
The paper documents the control group thoroughly: its size (n = 42), the demographic composition of the sample, the exact instructional regime the control group received (identical schedule, tasks and materials without the AI tool, described at length in Table 1 and the "Control group" subsection), and its baseline scores on all three outcome variables with statistical comparisons confirming baseline equivalence (Table 3). Descriptive statistics for the control group at all three time points are given in Table 2. This satisfies the requirement for detailed control-group documentation covering demographics, baseline performance and treatment received.
Criterion D is met because the control group's composition, baseline performance, and conditions are clearly and fully documented.
-
Level 2 Criteria
-
S
School-level RCT
- The trial randomised individual students within a single university, so there was no school-level randomisation.
- "The participants were randomly assigned to either the experimental group (n = 42) or the control group (n = 42) using a random number generator." (p. 5)
Relevant Quotes:
1) "A total of 84 university students of Chinese nationality enrolled in EFL courses at Hainan University in Haikou City, Hainan Province, China, participated in this randomized controlled trial." (p. 4)
2) "The participants were randomly assigned to either the experimental group (n = 42) or the control group (n = 42) using a random number generator." (p. 5)
Detailed Analysis:
The study took place at one institution (Hainan University) and the randomised units were individual students, not schools, campuses, centres or other implementing institutions. There is no description of multiple sites or of any institution-level assignment. Since only student- level randomisation within a single university is described, the school-level RCT requirement cannot be satisfied.
Criterion S is not met because randomisation occurred at the individual student level within a single university rather than across schools or sites.
-
I
Independent Conduct
- The sole author designed the study and conducted the data analysis, with no external or third-party evaluation team, so independent conduct is not demonstrated.
- "The current author is the sole contributor to this manuscript." (p. 14)
Relevant Quotes:
1) "The current author is the sole contributor to this manuscript." (Author contributions, p. 14)
2) "The intervention was conducted by two experienced EFL instructors--one for each group--who were regular faculty members at the university but not involved in the research design to prevent bias." (p. 6)
3) "Both instructors received training on the study protocols to ensure consistency in teaching methods and materials used across groups." (p. 6)
4) "This work was supported by the 2021 Hainan Higher Education and Teaching Reform Research Project..." (Acknowledgements, p. 14)
Detailed Analysis:
Criterion I requires that the evaluation be conducted independently of those who designed the intervention, for example by an external evaluation team. Here a single author designed the study, defined the AI-driven instructional intervention, and carried out the data collection oversight and statistical analysis. While the classroom delivery was done by two faculty instructors "not involved in the research design", these instructors implemented the intervention rather than independently evaluating it; the outcome measurement and analysis remained with the author. The underlying tool (Google TTS) is a commercial product not developed by the author, but the intervention as tested (the AI-integrated instructional programme) was designed by the same person who evaluated it, and no third-party oversight, external evaluator, or independent data-collection agency is mentioned anywhere in the paper.
Criterion I is not met because the intervention designer and the evaluator are the same sole author, with no documented independent oversight.
-
Y
Year Duration
- The total tracking window of about 11 weeks from intervention start is far below 75% of an academic year, and criterion T is already not met.
- "The study took place over eight weeks..." (p. 5)
Relevant Quotes:
1) "The study took place over eight weeks, employing a randomized assignment of participants into two groups..." (p. 5)
2) "Additionally, a follow-up session conducted three weeks later assessed the sustainability of any observed changes using similar assessments." (p. 6)
3) "Firstly, the relatively short duration of the intervention and the specific demographic of participants drawn from a single educational setting may limit the generalizability of the results..." (p. 11)
Detailed Analysis:
The Y criterion requires outcomes to be measured at least 75% of a full academic year (roughly 9-10 months) after the intervention begins. Here the entire study, from intervention start to the final follow-up assessment, spanned approximately 11 weeks (eight-week intervention plus a three-week follow-up), which is only about a quarter of an academic year. In addition, per the ranking rules, criterion Y cannot be met when criterion T (Term Duration) is not met, and T fails here.
Criterion Y is not met because the roughly 11-week tracking period is far shorter than 75% of an academic year.
-
B
Balanced Control Group
- Both groups received identical instructional time, sessions, tasks and self-study loads, with the only difference being the AI tool and its immediate feedback, which are the integral treatment variable being tested.
- "Both groups were subjected to identical teaching methods, tasks, and durations of instructional sessions, with the sole distinction being the presence or absence of AI technology." (pp. 5-6)
Relevant Quotes:
1) "Each group attended two in-class sessions per week, each lasting 90 min, totaling 16 sessions per group throughout the study period. Both groups were subjected to identical teaching methods, tasks, and durations of instructional sessions, with the sole distinction being the presence or absence of AI technology." (pp. 5-6)
2) "To ensure consistent engagement in self-study activities, both groups were assigned identical after-class tasks each week, carefully designed to match in content and difficulty." (p. 6)
3) "We monitored compliance through these logs and addressed any discrepancies promptly to maintain parity between the groups. This approach allowed us to ensure that both groups received an equal amount of instruction and practice time." (p. 6)
4) "To balance the feedback provision, the control group received timely and detailed feedback on their after-class assignments at the beginning of each subsequent class, minimizing delays and keeping them engaged." (p. 6)
5) "The time devoted to these activities was consistent with that of the experimental group, ensuring an equitable distribution of instructional time." (p. 7)
6) "It is acknowledged that the instant feedback provided by the AI tool in the experimental group could have influenced the participants' motivation and engagement levels. While efforts were made to ensure equal instructional time and practice opportunities, the nature of the feedback differed between the groups..." (p. 6)
Detailed Analysis:
Following the criterion B decision tree: the intervention group did receive an additional input (the Google TTS AI application and its immediate feedback), so extra resources are present. Time and dosage, however, were deliberately equalised: both groups had two 90-minute sessions per week for eight weeks, identical lesson plans, identical after-class tasks, roughly 30 minutes of daily self-study, and compliance monitoring via weekly logs. The control group also received active instructor and peer feedback, with deliberate efforts to minimise feedback delay. The remaining difference - the AI tool with real-time feedback - is precisely the treatment variable being tested; the study's stated aim was "to isolate the impact of this technology on the learning outcomes", making the AI tool integral to the intervention rather than a separable, confounding add-on (analogous to the DPL and RAMSR exception examples in the standard). The paper honestly notes that feedback immediacy differed between groups, but that immediacy is an inherent property of the AI feedback being evaluated. It should be clearly noted that the experimental group had access to the AI application (a free, widely available tool) during self-study while the control group used traditional materials; this device/tool difference is the integral core of the tested intervention.
Criterion B is met because instructional time, tasks and practice loads were explicitly equalised across groups and the only added resource (the AI tool with immediate feedback) is the integral treatment variable being tested.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific 2025 trial by a different research team was found via internet search; only unrelated studies using different tools and designs exist.
Relevant Quotes:
1) "This randomized controlled trial explored the effects of employing AI-driven methodologies on enhancing listening comprehension, flow experience, and alleviation of listening anxiety among English as a foreign language (EFL) learners." (p. 1)
2) "Future research encompassing diverse learner populations and extended intervention periods would help validate and enhance the robustness and transferability of these findings." (p. 11)
Detailed Analysis:
Criterion R requires that this specific study (its central experimental claim, design and context) be independently replicated by a different research team and published in a peer-reviewed journal. The paper itself makes no claim of replication; on the contrary, it calls for future research to validate the findings. The literature it cites (e.g., Bashori et al. 2021; Tsai 2023; Mirzaei et al. 2018) consists of related but distinct studies of ASR/TTS in language learning that predate this trial and use different designs, populations and outcome batteries - they are background evidence, not replications of this eight-week Hainan University RCT.
An internet search (Google, Google Scholar) for independent replications published since the article appeared online (25 March 2025) did not locate a peer-reviewed study that replicates this specific trial's design (Google TTS speech-recognition tool, IELTS Listening test, PPL-FSQ flow scale, FLLAS anxiety scale, 8-week intervention plus 3-week follow-up). Two related, but non-replicating, studies citing or resembling this work were identified: (a) a quasi-experimental study on NotebookLM-generated podcasts and EFL listening/flow/anxiety, published in the ACM GBA EDCSIC 2025 proceedings, which cites Xiao (2025) as background but tests a different AI tool (NotebookLM podcasts, not Google TTS ASR) with a quasi-experimental rather than randomised design; and (b) "The effect of AI-powered speech recognition tools on listening comprehension and pronunciation awareness in EFL learners," involving sixty Arabic-L1 female students aged 18-22 in Baghdad assigned to an ASR-mediated practice group versus a time-matched control group, which studies a different ASR tool, population, and outcome battery (no flow or anxiety scales) and does not present itself as a replication of this study. Neither paper reproduces this trial's specific design and context, so they do not constitute an independent reproduction of this particular study, analogous to the Indigenous PAX-GBG exception example in the ERCT standard.
Criterion R is not met because no independent, peer-reviewed replication of this specific study was found.
-
A
All-subject Exams
- Only English listening comprehension was assessed with a standardised exam; no other main academic subjects were measured and no explicit rationale for a specialised- intervention exception is given.
- "To evaluate the participants' listening comprehension abilities, the original version of the IELTS Listening test (Scovell et al. 2004) was employed as both a pre- and post-intervention assessment tool." (p. 5)
Relevant Quotes:
1) "To evaluate the participants' listening comprehension abilities, the original version of the IELTS Listening test (Scovell et al. 2004) was employed as both a pre- and post-intervention assessment tool." (p. 5)
2) "Pre- and post-intervention assessments evaluated participants' listening comprehension abilities, flow state, and listening anxiety." (p. 1)
Detailed Analysis:
Criterion A requires standardised exam-based measurement across all main subjects taught at the educational level in question, to detect possible negative spillovers on non-target subjects. This study measured only one academic outcome - English listening comprehension - plus two psychological self-report constructs (flow and anxiety). No other university subjects, nor even other English skills (reading, writing, speaking) via standardised exams, were assessed. The standard's exception covers highly specialised interventions in upper secondary or vocational education with a clearly explained rationale; this is a university EFL course intervention and the paper offers no explicit rationale for restricting measurement to listening alone. While one could argue that for an EFL listening programme the relevant domain is English, the paper does not provide the justification the exception requires, and even within English only the listening section was tested.
Criterion A is not met because only a single subject (English listening) was assessed and no justified exception for the narrow assessment scope is provided.
-
G
Graduation Tracking
- Tracking ended three weeks after the intervention with no follow-up to graduation, and the prerequisite criterion Y is not met.
- "Additionally, a follow-up session conducted three weeks later assessed the sustainability of any observed changes using similar assessments." (p. 6)
Relevant Quotes:
1) "Additionally, a follow-up session conducted three weeks later assessed the sustainability of any observed changes using similar assessments." (p. 6)
2) "However, our findings regarding learner engagement and autonomy are limited to the intervention period, and we cannot generalize these effects beyond that timeframe without further evidence." (p. 10)
3) "Nevertheless, the long-term benefits and sustainability of these effects require further investigation." (p. 10)
Detailed Analysis:
Criterion G requires participants to be followed until graduation from their educational stage. The last measurement in this study occurred only three weeks after the eight-week intervention ended; the university students were not tracked to the completion of their degree programmes, and the author explicitly states that long-term effects require further investigation.
An internet search for subsequent papers by the same author (Yanling Xiao, College of Foreign Languages, Hainan University) tracking this same cohort toward graduation did not locate any such follow-up publication; no additional papers by this author on this topic or cohort were found via Google Scholar or general web search as of this check. In addition, per the ranking rules, criterion G cannot be met when criterion Y is not met, and Y fails here.
Criterion G is not met because tracking stopped three weeks after the intervention, with no graduation follow-up located in this paper or in any subsequent publication.
-
P
Pre-Registered
- The paper reports IRB ethical approval but no pre-registration of the study protocol on any trial registry before data collection.
Relevant Quotes:
1) "The study protocol, including participant recruitment, data collection, and confidentiality measures, was formally reviewed and approved by the IRB on June 15, 2023, under the approval number HNU-CFL-2023-067." (Ethical approval, p. 14)
2) "Data used and/or analyzed during the current study are available from the corresponding author upon reasonable request." (Data availability, p. 11)
Detailed Analysis:
Criterion P requires pre-registration of hypotheses, methods and planned analyses on a public registry (e.g., ClinicalTrials.gov, OSF, a trial registry) before data collection begins. The paper mentions only institutional IRB ethical approval (June 2023), which is an ethics review, not a public pre-registration of the study protocol and analysis plan. No registry name, registration number, or registration date is provided anywhere in the article, and no pre-registered analysis plan is referenced.
An internet search for a pre-registration record for this study (e.g., on OSF Registries or ClinicalTrials.gov, searching by author name, institution, and study topic) did not locate any pre-registration entry for this trial.
Criterion P is not met because no pre-registration on a public registry is reported or found, only IRB ethical approval.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.