Abstract
This study addresses a critical gap in second language acquisition (SLA) research: the lack of integration between implicit comprehensible input and explicit metacognitive strategy training in AI-driven EFL listening instruction. With a mixed-methods design, it conducted a 16-week randomized controlled trial (RCT) with 120 Chinese EFL undergraduates (Mage = 19.2, CEFR A2-B1), comparing an AI-driven listening system (experimental group, n = 60) with traditional classroom instruction (control group, n = 60). The AI operationalized three core SLA frameworks: (1) Vygotsky's Zone of Proximal Development (ZPD) via dynamically fading strategy prompts; (2) Nation's affective filter hypothesis through real-time anxiety monitoring (heart rate sensors) and culturally familiar content curation; (3) Richards' metacognitive framework via context-specific prompts for prediction, inference, and summarization. Key findings: (1) The experimental group's post-test listening scores were 11.2 points higher (M = 79.5 vs. 68.3, d = 1.02, p < 0.001); (2) Their FLCAS anxiety dropped 5.2 points (M = 24.5 vs. control M = 29.7, d=-1.05, p < 0.001); (3) Their top-down strategy use doubled (3.5 to 7.1 weekly uses), vs. a 10% increase in the control group. These findings suggest AI may modulate auditory cognition by synergizing implicit affective support and explicit scaffolding, challenging the input-strategy dichotomy. It advances SLA theory with empirical evidence for a tech-mediated framework, informing scalable, personalized EFL listening instruction.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was performed at the individual student level within a single university cohort, not at the class or school level, and the paper does not frame the intervention as one-to-one tutoring that would qualify for the exception.
- "A total of 120 participants were randomly assigned to two groups (n=60 each): the experimental group used the AI-driven listening system, while the control group received traditional teacher-led listening instruction" (p. 8)
Relevant Quotes:
1) "Participants were assigned to either an experimental group (n=60) that used the AI-driven listening system or a control group (n=60) that received traditional teacher-led listening instruction." (p. 2)
2) "A total of 120 participants were randomly assigned to two groups (n=60 each): the experimental group used the AI-driven listening system, while the control group received traditional teacher-led listening instruction (e.g., static audio materials+post-task teacher feedback)." (p. 8)
3) "Of the 120 Chinese EFL undergraduates initially recruited for the randomized controlled trial, 82 Mandarin-speaking undergraduates (aged 18-22, CEFR B1-B2) from a large public university in eastern China completed all 16 weeks of intervention and data collection." (p. 8)
4) "Frequency (Weeks 2-15): 2x45-minute sessions/week (self-paced via desktop app) [experimental] ... 2x45-minute sessions/week (teacher-led in classroom) [control]" (Table 3, p. 12)
Detailed Analysis:
The unit of randomisation was the individual student: 120 undergraduates from a single university were randomly assigned to two groups of 60. No classes or schools were randomised.
The exception for personal tutoring was considered. The AI system is used individually and self-paced, and the paper emphasises "personalized listening instruction." However, the intervention is framed as a replacement for regular classroom listening instruction (the control received 2x45-minute teacher-led classroom sessions per week), not as a one-to-one tutoring programme, and the paper contains no statement claiming a personal-teaching/tutoring design. Students from the same cohort attended the same university and could share strategies learned from AI prompts, so contamination is possible in principle.
Criterion C is not met because randomisation was conducted at the individual student level within one institution without a clearly stated tutoring exception.
-
E
Exam-based Assessment
- Outcomes were measured with researcher-administered, CEFR-aligned pre/post listening tests and the FLCAS anxiety scale, with no named, widely recognized standardized exam.
- "Learning outcomes were measured using pre-test and post-test listening proficiency scores, the Foreign Language Classroom Anxiety Scale (FLCAS) for assessing anxiety levels, and weekly strategy logs to track metacognitive strategy use." (p. 2)
Relevant Quotes:
1) "Learning outcomes were measured using pre-test and post-test listening proficiency scores, the Foreign Language Classroom Anxiety Scale (FLCAS) for assessing anxiety levels, and weekly strategy logs to track metacognitive strategy use." (p. 2)
2) "Quantitative data (pre/post listening test scores, Foreign Language Classroom Anxiety Scale [FLCAS] scores, weekly strategy use frequency) and qualitative data (semi-structured interviews, open-ended strategy log comments) were collected in parallel." (p. 8)
3) "The experimental group's post-test listening scores were 11.2 points higher (M = 79.5 vs. 68.3, d = 1.02, p < 0.001)" (Abstract, p. 1)
4) "listening proficiency scores (using the same CEFR-aligned pre/post-tests)" (p. 27, describing proposed future work with the same instruments)
Detailed Analysis:
The ERCT E criterion requires that outcomes be measured with widely recognized standardized exams (e.g., national or state-wide tests such as CET-4, IELTS, or TOEFL for EFL contexts), not instruments assembled for the study. The paper never names any recognized standardized listening examination. The primary outcome is described only as "pre-test and post-test listening proficiency scores" on unspecified "CEFR-aligned" tests administered by the researchers; no test name, source, validation history, or psychometric properties of the listening measure are reported. CEFR alignment is a difficulty-calibration framework, not itself a standardized exam. The FLCAS is a validated anxiety questionnaire, but it measures affect, not educational achievement, and strategy logs are self-report behavioral counts. Under the standard, the assessment appears to be a study-specific instrument rather than a standard, widely recognized exam.
Criterion E is not met because the listening outcome was measured with an unnamed, apparently study-specific test rather than a recognized standardized exam.
-
T
Term Duration
- The intervention ran 16 weeks (about four months, a full semester) with post-tests within one week of its end, so the interval from intervention start to outcome measurement covers at least one full academic term.
- "it conducted a 16-week randomized controlled trial (RCT) with 120 Chinese EFL undergraduates" (Abstract, p. 1)
Relevant Quotes:
1) "it conducted a 16-week randomized controlled trial (RCT) with 120 Chinese EFL undergraduates (Mage = 19.2, CEFR A2-B1)" (Abstract, p. 1)
2) "The 16-week intervention period was divided into 3 phases (baseline, intervention, post-intervention) to track dynamic changes in learning outcomes" (p. 8)
3) "The study only assessed learning outcomes via immediate post-tests (conducted within 1 week of the 16-week intervention)" (p. 24)
4) "Frequency (Weeks 2-15): 2 x 45-minute sessions/week" (Table 3, p. 12)
Detailed Analysis:
The T criterion requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. Here the intervention spanned 16 weeks (Weeks 1-15 of activity plus post-testing), and the post-test was administered within one week of the end of the 16-week program. The interval from intervention start to primary outcome measurement is therefore approximately 16-17 weeks, i.e., roughly four months, which corresponds to a full university semester and satisfies the one-term minimum. Specific calendar dates are not given, but the duration itself is stated repeatedly and unambiguously.
Criterion T is met because outcomes were measured about 16 weeks after the intervention began, covering at least one full academic term.
-
D
Documented Control Group
- The control condition, its size, materials, instruction, and baseline comparability (technology familiarity, stratified gender/major, pre-test covariates) are documented, despite some internal inconsistencies in reported sample sizes.
- "Control Group (Traditional Listening Instruction) ... 2 x 45-minute sessions/week (teacher-led in classroom) ... Fixed audio from New Horizon College English (3rd ed., Book 1-2), CEFR B1 difficulty (no individual adaptation)" (Table 3, p. 12)
Relevant Quotes:
1) "the control group received traditional teacher-led listening instruction (e.g., static audio materials + post-task teacher feedback)." (p. 8)
2) "Control Group (Traditional Listening Instruction): No pre-training; baseline listening test and FLCAS anxiety assessment only. ... 2 x 45-minute sessions/week (teacher-led in classroom) ... Fixed audio from New Horizon College English (3rd ed., Book 1-2), CEFR B1 difficulty (no individual adaptation) ... Weekly 10-minute whole-class lectures on 'basic listening strategies' ... Delayed: Written feedback on homework (48 h post-session)" (Table 3, p. 12)
3) "82 Mandarin-speaking undergraduates (aged 18-22, CEFR B1-B2) from a large public university in eastern China completed all 16 weeks of intervention and data collection. The 38 excluded participants (18 from the experimental group, 20 from the control group) were removed due to excessive session absences" (p. 8)
4) "Gender balance: 41 male and 41 female participants ... Equal representation of STEM (n = 41) and Humanities (n = 41) majors" (p. 8)
5) "Baseline scores showed no significant group differences (experimental: M = 3.2 +/- 0.7; control: M = 3.1 +/- 0.8; t = 0.52, p = 0.61)" (pp. 8-9)
6) "Baseline technology literacy (Technology Literacy Index) and pre-test listening scores (CEFR levels) were included as covariates." (p. 12)
Detailed Analysis:
The D criterion requires detailed documentation of the control group: who they are, baseline characteristics, size, and what they received. The paper documents the control group size (n = 60 randomized; 20 excluded), the demographic composition of the completing cohort (age 18-22, Mandarin L1, gender and major stratification), baseline equivalence on technology familiarity with group-specific statistics, baseline listening pre-tests and FLCAS assessments, and an unusually detailed specification of exactly what instruction the control group received (Table 3: session frequency, named textbook materials, fixed CEFR B1 difficulty, strategy lectures, and feedback modality). This exceeds the typical bare "there was a control group" failure mode the criterion targets. A caveat: the paper contains internal inconsistencies (120 randomized vs. 82 completers, yet several analyses report n = 60 per group and df = 117 or n = 32 + 28 = 60 for the experimental group's HR-anxiety regression on p. 17, even though only 42 of the 60 randomized experimental participants completed all 16 weeks), and CEFR range is given as A2-B1 in the abstract but B1-B2 on p. 8. This weakens confidence in the numbers but does not eliminate the documentation itself.
Criterion D is met because the control group's size, baseline characteristics, and business-as-usual condition are described in detail, though with some internal numerical inconsistency.
-
Level 2 Criteria
-
S
School-level RCT
- Randomization was at the individual student level within a single university, not at the school or institution level.
- "A total of 120 participants were randomly assigned to two groups (n = 60 each)" (p. 8)
Relevant Quotes:
1) "A total of 120 participants were randomly assigned to two groups (n = 60 each)" (p. 8)
2) "82 Mandarin-speaking undergraduates (aged 18-22, CEFR B1-B2) from a large public university in eastern China completed all 16 weeks of intervention and data collection." (p. 8)
3) "The current experiment was conducted under controlled laboratory conditions to ensure data precision and minimize environmental noise." (p. 9)
Detailed Analysis:
The S criterion requires randomization among schools or equivalent implementing institutions. This study drew all participants from a single university and randomized individual students to conditions within that one site; no schools, campuses, or institutional units were randomized. The tutoring/personal-teaching exception applies only to the weaker class-level criterion C, not to the school-level criterion S.
Criterion S is not met because randomization occurred at the individual student level within one university, with no school-level assignment.
-
I
Independent Conduct
- The same two authors designed the AI system, designed the trial, collected the data, and analyzed the results, with no external or third-party evaluation team.
- "Yukun Liu designed the study methodology, conducted the data collection and quantitative analysis, and drafted the initial manuscript. Yan Li developed the AI-driven listening system framework, supervised the qualitative analysis" (p. 31)
Relevant Quotes:
1) "Yukun Liu designed the study methodology, conducted the data collection and quantitative analysis, and drafted the initial manuscript. Yan Li developed the AI-driven listening system framework, supervised the qualitative analysis (including interview coding and strategy log interpretation), and revised the manuscript for theoretical coherence. Both authors contributed to the design of the randomized controlled trial, validated the research instruments, and approved the final version of the manuscript." (p. 31)
2) "The system integrates commercial tools (ELSA Speak for speech recognition) with custom-developed modules for adaptive content delivery and strategy scaffolding" (p. 9)
3) "The authors declare that no funding, grants, or other financial support was received for conducting this study or preparing the manuscript." (p. 31)
Detailed Analysis:
The I criterion requires that the evaluation be conducted independently of the intervention's designers, or at minimum that an external party handle data collection and analysis. Here the author contribution statement explicitly shows the opposite: one author developed the AI-driven listening system framework and its custom modules, and the same two-person team designed the RCT, collected all data, and performed all quantitative and qualitative analyses. There is no mention of an external evaluation agency, independent enumerators, blinded assessors, or third-party oversight anywhere in the paper. This is a textbook case of the developers evaluating their own intervention.
Criterion I is not met because the intervention's own developers designed, conducted, and analyzed the trial with no independent oversight.
-
Y
Year Duration
- The interval from intervention start to final measurement was about 16-17 weeks (roughly 4 months), well short of 75% of an academic year, and no delayed follow-up was conducted.
- "The study only assessed learning outcomes via immediate post-tests (conducted within 1 week of the 16-week intervention) and failed to include delayed post-tests (e.g., 3-month or 6-month follow-ups)" (p. 24)
Relevant Quotes:
1) "it conducted a 16-week randomized controlled trial (RCT)" (Abstract, p. 1)
2) "The study only assessed learning outcomes via immediate post-tests (conducted within 1 week of the 16-week intervention) and failed to include delayed post-tests (e.g., 3-month or 6-month follow-ups)-a gap that prevents verifying the long-term retention and internalization of AI-facilitated strategies." (p. 24)
3) "To address the 16-week intervention's limitations in assessing long-term efficacy, we propose a 12-month longitudinal design" (p. 26, proposed future work only)
Detailed Analysis:
The Y criterion requires outcome measurement at least 75% of a full academic year (roughly 6.75-7.5 months of a 9-10 month year) after the intervention begins. Here the total tracking window from intervention start to final measurement was approximately 16-17 weeks (about 4 months). The authors themselves flag the absence of any delayed post-test as a limitation, and longitudinal follow-up is only proposed for future research. Four months is roughly 40-45% of an academic year, clearly below the 75% threshold.
Criterion Y is not met because outcomes were measured only about four months after the intervention started, far short of 75% of an academic year.
-
B
Balanced Control Group
- Both groups received identical instructional dosage (2 x 45-minute sessions/week over the same period) with an active, well-specified control condition, and the AI system's added technology is integral to the treatment being tested.
- "Frequency (Weeks 2-15): 2 x 45-minute sessions/week (self-paced via desktop app) [experimental] / 2 x 45-minute sessions/week (teacher-led in classroom) [control]" (Table 3, p. 12)
Relevant Quotes:
1) "Frequency (Weeks 2-15): 2 x 45-minute sessions/week (self-paced via desktop app)" [experimental] and "2 x 45-minute sessions/week (teacher-led in classroom)" [control] (Table 3, p. 12)
2) "Materials: Adaptive audio (i + 1 difficulty, interest-aligned: sports, tech, campus life) from BBC Learning English + custom clips" [experimental] vs. "Fixed audio from New Horizon College English (3rd ed., Book 1-2), CEFR B1 difficulty (no individual adaptation)" [control] (Table 3, p. 12)
3) "Strategy Training: Real-time error-driven prompts ... post-session strategy logs (10 min/session)" [experimental] vs. "Weekly 10-minute whole-class lectures on 'basic listening strategies' ... no personalized scaffolding" [control] (Table 3, p. 12)
4) "15-minute AI tool operation training" for the experimental group only, Week 1 (Table 3, p. 12)
5) "Table 3 below clarifies the control group's 'traditional instruction' (previously vague) by specifying pedagogical methods, materials, and feedback mechanisms-ensuring comparability with the experimental group's AI-driven intervention." (p. 11)
Detailed Analysis:
Following the B decision tree (per the current ERCT specification): the core instructional time is matched exactly - both groups received 2 x 45-minute listening sessions per week over Weeks 2-15, and both received strategy instruction (real-time AI prompts vs. weekly teacher strategy lectures) and feedback (instant AI feedback vs. delayed written teacher feedback). The control is an active, documented business-as-usual condition with named materials. The additional inputs unique to the experimental group are the AI hardware/software (desktop app, sensors), a one-off 15-minute operation training, and 10 minutes per session of strategy logging. The AI technology itself is the treatment variable: the study explicitly tests "the efficacy of this integrated AI system" against traditional instruction, so the devices, adaptive content, and real-time feedback are integral components of the intervention being tested, not separable confounding add-ons (analogous to the DPL-tool exception example in the standard) - this alone satisfies the "resources are the treatment" branch of the decision tree. The remaining time differences (15 minutes once; strategy logs that double as a measurement instrument) are negligible relative to roughly 21 hours of matched instruction, which also satisfies the negligible-difference branch. Therefore the allocation is balanced in the sense required by the standard, with the integral AI resources clearly noted.
Criterion B is met because instructional time was matched across groups with an active control, and the extra AI technology is integral to the treatment being tested rather than a separable resource imbalance.
-
Level 3 Criteria
-
R
Reproduced
- This December 2025 study reports no replication of itself, and no independent peer-reviewed replication of this specific AI listening system trial exists or is cited; an internet search of citing literature confirms no replication study was found.
Relevant Quotes:
1) "This study addresses a critical gap in second language acquisition (SLA) research: the lack of integration between implicit comprehensible input and explicit metacognitive strategy training in AI-driven EFL listening instruction." (Abstract, p. 1)
2) "Future research will broaden the participant base to include non-tonal L1 groups (e.g., English, Korean, Arabic) ... Additionally, longitudinal field trials in authentic classroom and mobile environments will be conducted" (p. 9)
3) "Researchers may request access through the project's corresponding author for academic replication, verification, or extension." (Appendix C, p. 31)
Detailed Analysis:
The R criterion requires independent replication of the study by a different research team in a different context, published in a peer-reviewed journal. The paper frames itself as filling a novel gap and describes replication only as a future possibility (open resources are offered "for academic replication"). It was accepted 5 December 2025 and published online 24 December 2025, leaving essentially no time for independent replication to have appeared. A citation search (Google Scholar, July 2026) identified two papers citing this article - Aristia, Pratika and Sukraini (2026) on AI-generated transcripts for EFL listening (Didaktika: Jurnal), and Azimova (2026) on AI for listening comprehension (Review of Multidisciplinary Academic) - but both are independent conceptual works that merely cite this study rather than replications of its specific AI system and trial design. The cited related studies (e.g., Li et al. 2023 in ReCALL; Zhang and Wang 2024 in Language Learning and Technology) concern adaptive AI listening instruction generally but are prior work, not replications of this specific integrated system and trial. No independent replication of this study could be identified.
Criterion R is not met because no independent, peer-reviewed replication of this specific trial exists or is referenced, confirmed via internet search.
-
A
All-subject Exams
- Only EFL listening proficiency (plus anxiety and strategy use) was measured, no other core subjects were assessed, and the prerequisite criterion E is not met.
- "Learning outcomes were measured using pre-test and post-test listening proficiency scores, the Foreign Language Classroom Anxiety Scale (FLCAS) for assessing anxiety levels, and weekly strategy logs to track metacognitive strategy use." (p. 2)
Relevant Quotes:
1) "Learning outcomes were measured using pre-test and post-test listening proficiency scores, the Foreign Language Classroom Anxiety Scale (FLCAS) for assessing anxiety levels, and weekly strategy logs to track metacognitive strategy use." (p. 2)
2) "The experimental group's post-test listening scores were 11.2 points higher (M = 79.5 vs. 68.3, d = 1.02, p < 0.001)" (Abstract, p. 1)
Detailed Analysis:
The A criterion requires standardized exam-based assessment across all main subjects taught at the educational level, and per the instructions it automatically fails when criterion E fails. Criterion E is not met here (no recognized standardized exam was used), so A cannot be met. Substantively, the study measured only English listening proficiency within a single skill of a single subject; no other university subjects (or even other English skills via standardized exams) were assessed. While the intervention is specialized, the paper offers no argument that broader assessment was unnecessary, and in any case the E prerequisite is decisive.
Criterion A is not met because criterion E fails and only a single-skill listening outcome in one subject was measured.
-
G
Graduation Tracking
- Tracking stopped at an immediate post-test within one week of the 16-week intervention, with no follow-up to graduation, and the prerequisite criterion Y is not met; no follow-up publications tracking this cohort were found online.
- "The study only assessed learning outcomes via immediate post-tests (conducted within 1 week of the 16-week intervention) and failed to include delayed post-tests" (p. 24)
Relevant Quotes:
1) "The study only assessed learning outcomes via immediate post-tests (conducted within 1 week of the 16-week intervention) and failed to include delayed post-tests (e.g., 3-month or 6-month follow-ups)" (p. 24)
2) "Preliminary pilot data from 10 experimental participants (3-month informal follow-up) suggested that summarization strategy use declined by 40% ... Without systematic delayed testing, however, these patterns cannot be validated or generalized." (p. 24)
3) "Table 6 shows the follow-up study will employ two delayed post-tests administered at three months and six months after the intervention." (p. 25, proposed future work only)
Detailed Analysis:
The G criterion requires following participants until graduation from their educational stage. Per the instructions, G automatically fails because criterion Y is not met. The paper also directly confirms the absence of any systematic follow-up beyond the immediate post-test: delayed 3- and 6-month tests are only proposed for future research, and the only longer-horizon data is an informal 3-month pilot with 10 participants that the authors themselves say cannot be validated. The undergraduate participants were not tracked to university graduation. An internet citation search (Google Scholar, July 2026) found no subsequent publication by these authors tracking this cohort further, consistent with the paper's very recent (December 2025) publication date.
Criterion G is not met because measurement ended immediately after the 16-week intervention with no tracking to graduation (and Y is not met), and no follow-up publications were found.
-
P
Pre-Registered
- The paper reports IRB approval and informed consent but contains no mention of prospective trial registration on any registry; an internet search of common registries found no matching pre-registration record.
Relevant Quotes:
1) "All procedures were reviewed and approved by the Institutional Review Board (IRB) of the host university (Approval No. EDU2024-11-03)." (p. 18)
2) "This study was approved by the Institutional Review Board of Chonnam National University (Approval No.: CNU-IRB-2024-012) and conducted in accordance with the Declaration of Helsinki" (p. 31)
3) "The datasets generated and analyzed during the current study are available from the corresponding author ... upon reasonable request." (p. 31)
Detailed Analysis:
The P criterion requires that the full study protocol (hypotheses, methods, planned analyses) be registered on a public registry before data collection began, with a verifiable registry ID and date. The paper mentions only ethics approval and consent; ethics/IRB approval is not pre-registration. No registry platform (e.g., ClinicalTrials.gov, OSF, ChiCTR, AsPredicted), registration number, or registration date appears anywhere in the manuscript. An internet search of OSF and general registries for this title, these authors, and this AI listening system found no matching pre-registration record. Notably, the paper even gives two different IRB approval numbers in different sections (EDU2024-11-03 vs. CNU-IRB-2024-012), further underscoring the absence of a single transparent, pre-registered protocol trail.
Criterion P is not met because no pre-registration statement, registry ID, or registration date is provided in the paper, and none could be located via internet search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.