Optimizing EFL vocabulary acquisition: a randomized controlled mixed-methods investigation of artificial intelligence-driven incidental, contextual, and multimodal strategies

Miao Yu

Published:
ERCT Check Date:
DOI: 10.1007/s10639-025-13803-2
  • L2 languages
  • higher education
  • China
  • gamification
  • blended learning
  • EdTech app
  • EdTech platform
  • digital assessment
  • formative assessment
0
  • C

    Randomisation was conducted at the individual student level rather than at the class or school level, and no personal-tutoring exception applies.

    "Participants were allocated to four parallel groups...using computer-generated block randomization (block size=8), stratified by gender and age to ensure demographic balance across conditions." (p. 62)

  • E

    The study used a bespoke, researcher-designed 60-item vocabulary instrument and AI-scored performance tasks rather than a widely recognised standardised exam.

    "The vocabulary mastery assessment protocol employed a 60-item instrument administered identically across pretest, posttest, and delayed posttest phases to evaluate lexical knowledge progression." (p. 65)

  • T

    The delayed posttest occurred roughly 14 weeks (about 3.2 months) after the intervention began, which falls within the one-term window, so the criterion is met, albeit marginally.

    "The pretest established baseline proficiency, the posttest measured short-term gains after the AI intervention, and the delayed posttest (6 weeks post-intervention) assessed retention." (p. 69)

  • D

    The control group's size, curriculum, materials, and baseline equivalence to the treatment arms are all explicitly documented.

    "Participants (n=98) followed a conventional curriculum devoid of AI tools, comprising textbook exercises, instructor-led drills, and peer discussions. Activities mirrored the hybrid structure of experimental groups ... to isolate AI's impact." (p. 64)

  • S

    Randomisation occurred at the individual student level across pooled universities, not at the school/institution level.

    "Participants were allocated to four parallel groups...using computer-generated block randomization (block size=8), stratified by gender and age..." (p. 62)

  • I

    The study was designed, implemented, and analysed by a single author with no independent third-party conducting the trial.

    "Miao Yu1" (sole author) and "The author declares no competing interests." (p. 53, 95)

  • Y

    The total tracked period (about 14 weeks/3.2 months) falls far short of the 75%-of-an-academic-year requirement.

    "The pretest established baseline proficiency, the posttest measured short-term gains after the AI intervention, and the delayed posttest (6 weeks post-intervention) assessed retention." (p. 69)

  • B

    Appendix Table 5 shows the AI groups (especially Multimodal, 5 sessions/week x 150 min) received substantially more weekly instructional time than the control group (2 sessions/week x 60 min), contradicting the main text's claim of matched time, and this extra time is not framed as the explicit treatment variable.

    "Frequency/Duration: 5 sessions/week (150 mins: 90 VR, 60 in-person)" [Multimodal] vs. "2 sessions/week (60 mins fully in-person)" [Control] (Table 5, p. 87)

  • R

    No independent replication of this specific study by a different research team is reported or evidenced; a Google Scholar citation check on the paper's DOI returned zero citing articles as of this check.

  • A

    Criterion E is not met, and only EFL vocabulary outcomes were assessed, with no other core subjects measured.

  • G

    Criterion Y is not met, and no follow-up beyond the 6-week delayed posttest or tracking toward graduation is reported or found in subsequent literature.

  • P

    The paper documents IRB ethics approval only; no pre-registration of the study protocol on a public registry is mentioned anywhere in the text or found externally.

Abstract

This mixed-methods study rigorously evaluates the efficacy of artificial intelligence (AI)-enhanced vocabulary learning strategies--incidental exposure, contextual priming, and multimodal scaffolding--relative to traditional instructional approaches in English as a Foreign Language (EFL) pedagogy. Employing a pretest-posttest randomized controlled trial with 383 Chinese EFL learners, the research quantifies AI's impact on both immediate lexical acquisition and long-term retention, while qualitatively elucidating learner perceptions of engagement and instructional value. Quantitative findings indicate that AI-mediated multimodal strategies yield significantly higher vocabulary gains (posttest M=137.00, SD=5.51) and sustained retention (delayed posttest M=129.00) compared to contextual (M=121.00), incidental (M=113.00), and control groups (M=76.05), with robust multivariate effects (Pillai's Trace=0.766, p<.001, partial eta squared=0.589). Thematic analysis of semi-structured interviews highlights AI's strengths in delivering personalized feedback, adaptive content, and immersive contextualization, while also surfacing potential drawbacks such as sensory overload and cultural bias in AI-generated materials.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was conducted at the individual student level rather than at the class or school level, and no personal-tutoring exception applies.
      • "Participants were allocated to four parallel groups...using computer-generated block randomization (block size=8), stratified by gender and age to ensure demographic balance across conditions." (p. 62)
      • Relevant Quotes: 1) "A total of 383 EFL learners (205 female, 178 male; M age=24.23, SD=4.5, range=19-31 years) in their third or fourth year of Bachelor of Arts in English Language Teaching (ELT) programs were recruited from three accredited universities in China." (p. 60-61) 2) "Participants were allocated to four parallel groups-Incidental AI, Contextual AI, Multimodal AI, and Conventional Instruction-using computer-generated block randomization (block size=8), stratified by gender and age to ensure demographic balance across conditions." (p. 62) 3) "Allocation concealment was rigorously implemented via sequentially numbered, opaque sealed envelopes, opened only after pretest completion." (p. 62) Detailed Analysis: The paper explicitly states that randomisation occurred at the level of the individual participant, stratified only by gender and age. No quote anywhere in the Method section describes intact classes or schools being assigned as randomisation units. Although the 383 participants were drawn from three universities, the universities themselves were not the unit of randomisation; individual students within the same institutions were split across the four conditions using block randomization, creating a clear risk of contamination between conditions on the same campus. The tutoring/personal-teaching exception does not apply: the interventions (AI apps, chatbots, VR/AR tools) were delivered as group-based, 15-session hybrid curricula (45 min in-person, 45 min online) to cohorts of 95-98 students per arm, not as one-to-one personal tutoring. Because randomisation was performed at the individual student level and no valid exception applies, criterion C is not met. Quotes verified verbatim against the source PDF; no changes required.
    • E

      Exam-based Assessment

      • The study used a bespoke, researcher-designed 60-item vocabulary instrument and AI-scored performance tasks rather than a widely recognised standardised exam.
      • "The vocabulary mastery assessment protocol employed a 60-item instrument administered identically across pretest, posttest, and delayed posttest phases to evaluate lexical knowledge progression." (p. 65)
      • Relevant Quotes: 1) "The vocabulary mastery assessment protocol employed a 60-item instrument administered identically across pretest, posttest, and delayed posttest phases to evaluate lexical knowledge progression." (p. 65) 2) "Item banks were constructed using the Academic Vocabulary List (AVL) and Corpus of Contemporary American English (COCA) frequency norms, with distractor plausibility verified via lexical decision tasks in piloting." (p. 69) 3) "To ensure transparency, reproducibility, and minimize bias in scoring, all automated scoring was conducted using GPT-4 Turbo (OpenAI, model: gpt-4-1106-preview, March 2025 release)." (p. 67) Detailed Analysis: The primary instruments (the 60-item Vocabulary Knowledge Scale and the writing/oracy/literacy performance tasks) were purpose-built by the research team for this study, drawing on word lists (AVL, COCA) rather than adopting an existing, widely recognised standardised exam such as a national curriculum test or an established international exam. Validation relied on internal expert panels, pilot samples, and psychometric modelling, and scoring of open-response items was performed by a proprietary large language model (GPT-4 Turbo) rather than a standard examination board. There is no quote indicating use of a pre-existing, externally validated standardised exam (e.g., a national English exam or an internationally recognised proficiency test) as the outcome measure. Because the assessment is a researcher-constructed instrument rather than a widely recognised standardised exam, criterion E is not met. Quotes verified verbatim against the source PDF; no changes required.
    • T

      Term Duration

      • The delayed posttest occurred roughly 14 weeks (about 3.2 months) after the intervention began, which falls within the one-term window, so the criterion is met, albeit marginally.
      • "The pretest established baseline proficiency, the posttest measured short-term gains after the AI intervention, and the delayed posttest (6 weeks post-intervention) assessed retention." (p. 69)
      • Relevant Quotes: 1) "Intervention: Over eight weeks, groups received distinct instructional modalities" (p. 63) 2) "The pretest established baseline proficiency, the posttest measured short-term gains after the AI intervention, and the delayed posttest (6 weeks post-intervention) assessed retention." (p. 69) Detailed Analysis: The intervention itself ran for eight weeks, after which an immediate posttest was administered. A delayed posttest was then conducted six weeks after the intervention ended, i.e. approximately 14 weeks (about 3.2 months) after the intervention began. This falls within the "approximately 3-4 months" window the ERCT standard treats as one academic term, so the interval from intervention start to the final (delayed) outcome measurement covers close to a full term, albeit near the lower bound rather than comfortably exceeding it. Given that the delayed-posttest interval (~14 weeks) sits at the edge of, but within, the defined term-length window, this criterion is judged met, though only marginally. Quotes verified verbatim against the source PDF; no changes required.
    • D

      Documented Control Group

      • The control group's size, curriculum, materials, and baseline equivalence to the treatment arms are all explicitly documented.
      • "Participants (n=98) followed a conventional curriculum devoid of AI tools, comprising textbook exercises, instructor-led drills, and peer discussions. Activities mirrored the hybrid structure of experimental groups ... to isolate AI's impact." (p. 64)
      • Relevant Quotes: 1) "Control group 4: traditional instruction. Participants (n=98) followed a conventional curriculum devoid of AI tools, comprising textbook exercises, instructor-led drills, and peer discussions. Activities mirrored the hybrid structure of experimental groups (e.g., 45-minute in-person quizzes, 45-minute online readings) to isolate AI's impact." (p. 64) 2) "Baseline equivalence was confirmed by independent-samples ANOVAs showing no significant pretest differences in vocabulary scores (F [3, 362]=0.93, p=.427, partial eta squared=0.01)." (p. 62) 3) Table 5 (Appendix): "Educational Content ... Textbook exercises (Cambridge English Vocabulary in Use). Printed flashcards (500 high-frequency words)." "Software/Hardware ... Software: Moodle LMS (v4.1). Hardware: Samsung Galaxy Tab S7." (p. 85-87) Detailed Analysis: The control condition's size (n=98), curriculum content (textbook exercises, drills, peer discussion), materials (named textbook, printed flashcards), session structure, and platform/hardware are all explicitly documented in the main text and in the detailed Appendix Table 5. Baseline equivalence between the control group and the three AI arms was statistically confirmed via ANOVA, showing no significant pretest differences. This level of documentation allows a reader to assess the comparability and composition of the control group. Because the control group's characteristics, activities, and baseline comparability are clearly documented, criterion D is met. Quotes verified verbatim against the source PDF; no changes required.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred at the individual student level across pooled universities, not at the school/institution level.
      • "Participants were allocated to four parallel groups...using computer-generated block randomization (block size=8), stratified by gender and age..." (p. 62)
      • Relevant Quotes: 1) "A total of 383 EFL learners ... were recruited from three accredited universities in China." (p. 60-61) 2) "Participants were allocated to four parallel groups...using computer-generated block randomization (block size=8), stratified by gender and age to ensure demographic balance across conditions." (p. 62) Detailed Analysis: Although students were drawn from three universities, the paper does not describe randomisation occurring at the level of the university, school, or institution. Instead, individual students (pooled across institutions) were assigned to conditions via block randomization stratified by gender and age. There is no quote indicating that entire universities, departments, or cohorts were randomised as units; the school/institution served only as a recruitment source, not a randomisation unit. Because randomisation was performed at the individual student level rather than the institutional/school level, criterion S is not met. Quotes verified verbatim against the source PDF; no changes required.
    • I

      Independent Conduct

      • The study was designed, implemented, and analysed by a single author with no independent third-party conducting the trial.
      • "Miao Yu1" (sole author) and "The author declares no competing interests." (p. 53, 95)
      • Relevant Quotes: 1) "Miao Yu1" (sole author byline, p. 53) 2) "1 School of Foreign Languages, Xinyang Normal University, Xinyang 464000, Henan, China" (Authors and Affiliations, p. 97) 3) "Competing interests The author declares no competing interests." (p. 95) 4) "All post-test data were evaluated by six independent assessors--three doctoral students in applied linguistics and three experienced EFL instructors (minimum five years of teaching experience)." (p. 62) Detailed Analysis: The entire study--intervention design (Table 5's technical AI configurations), group allocation logic, materials development, data analysis, and manuscript preparation--is attributed to a single author with a single institutional affiliation. No external organisation or independent research team is credited with designing or independently running the trial. The "six independent assessors" and "two independent qualitative researchers" mentioned relate only to blinded scoring of test responses and blinded coding of interviews (i.e., steps to reduce measurement/analysis bias), not to independent conduct of the trial's overall design and implementation, which remained with the sole author throughout. Because there is no evidence that the study was designed or conducted by a team independent of the researcher who designed the interventions, criterion I is not met. Quote 1's page reference was corrected from "p. 1" to "p. 53" (the article's actual first published page) to match the journal's pagination used throughout this report; no other changes required.
    • Y

      Year Duration

      • The total tracked period (about 14 weeks/3.2 months) falls far short of the 75%-of-an-academic-year requirement.
      • "The pretest established baseline proficiency, the posttest measured short-term gains after the AI intervention, and the delayed posttest (6 weeks post-intervention) assessed retention." (p. 69)
      • Relevant Quotes: 1) "Intervention: Over eight weeks, groups received distinct instructional modalities" (p. 63) 2) "The pretest established baseline proficiency, the posttest measured short-term gains after the AI intervention, and the delayed posttest (6 weeks post-intervention) assessed retention." (p. 69) Detailed Analysis: The total tracked interval, from intervention start to the final delayed posttest, is approximately 14 weeks (about 3.2 months). This is far short of 75% of an academic year (approximately 6.75-7.5 months out of a 9-10 month academic year). No quote suggests any tracking of outcomes beyond the 6-week delayed posttest. Because the study's total duration falls well short of 75% of an academic year, criterion Y is not met. Quotes verified verbatim against the source PDF; no changes required.
    • B

      Balanced Control Group

      • Appendix Table 5 shows the AI groups (especially Multimodal, 5 sessions/week x 150 min) received substantially more weekly instructional time than the control group (2 sessions/week x 60 min), contradicting the main text's claim of matched time, and this extra time is not framed as the explicit treatment variable.
      • "Frequency/Duration: 5 sessions/week (150 mins: 90 VR, 60 in-person)" [Multimodal] vs. "2 sessions/week (60 mins fully in-person)" [Control] (Table 5, p. 87)
      • Relevant Quotes: 1) "The study employed a between-subjects design... Participants (N=383) were randomly assigned to one of four groups, each receiving a 15-session hybrid curriculum (45 min in-person, 45 min online) to control for instructional format (see Table 5 for further details)." (p. 63) 2) "Participants (n=98) followed a conventional curriculum devoid of AI tools...Activities mirrored the hybrid structure of experimental groups (e.g., 45-minute in-person quizzes, 45-minute online readings) to isolate AI's impact." (p. 64) 3) Table 5, "Frequency/Duration" row: "3 sessions/week (90 mins: 45 online, 45 in-person)" [Incidental]; "4 sessions/week (120 mins: 60 online, 60 in-person)" [Contextual]; "5 sessions/week (150 mins: 90 VR, 60 in-person)" [Multimodal]; "2 sessions/week (60 mins fully in-person)" [Control]. (Table 5, p. 87) Detailed Analysis: The main narrative (Sections 2.5-2.6) states that all four groups received an identical "15-session hybrid curriculum (45 min in-person, 45 min online)" specifically "to control for instructional format" and "isolate AI's impact," which would, if accurate, satisfy Criterion B by construction. However, Appendix Table 5 directly contradicts this: it reports escalating session frequency and duration across the AI conditions (Incidental: 3 sessions/week, 90 min; Contextual: 4 sessions/week, 120 min; Multimodal: 5 sessions/week, 150 min) while the Control group receives only 2 sessions/week, 60 min, entirely in-person with no online component. This means the Multimodal group receives roughly 6.25 times the weekly instructional time of the Control group (750 vs. 120 minutes/week), a substantial resource imbalance in contact time. This extra time is not framed by the authors as the explicit treatment variable under test (the stated treatment variable is the AI-driven strategy type, not the quantity of instructional time); rather, the text explicitly (but inconsistently with Table 5) claims time was held constant to isolate the AI effect. Since the additional time/resources in the AI arms are not documented as balanced with the control group, and are not explicitly identified as the tested variable, the imbalance revealed in Table 5 constitutes a confound that the study does not resolve. Given this internal inconsistency and the clear quantitative imbalance in contact time and resources documented in Table 5, criterion B is not met.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study by a different research team is reported or evidenced; a Google Scholar citation check on the paper's DOI returned zero citing articles as of this check.
      • Relevant Quotes: 1) "Data availability The data will be made available on reasonable request." (p. 95) 2) Reference list (pp. 95-97) contains only background/related literature on AI-vocabulary tools (e.g., Feng & Ng, 2024; Crum et al., 2024; Wen et al., 2025), none of which report an independent replication of this specific four-arm Chinese EFL vocabulary study. Detailed Analysis: There is no quote anywhere in the paper indicating that this specific intervention (the incidental/contextual/ multimodal AI vocabulary strategies tested with 383 Chinese EFL undergraduate learners at three universities) has been independently replicated by a different research team in a different context and published in a peer-reviewed outlet. The cited literature discusses related but distinct AI vocabulary tools and studies, not replications of this particular design and sample. Internet re-check: the paper was published online on 27 October 2025 and, as of this verification (July 2026), a Google Scholar citation search on its DOI (10.1007/s10639-025-13803-2) returns no citing articles ("we didn't find any articles that cite" this DOI). Given the short time elapsed since publication and the absence of any citing literature, no independent replication of this specific study exists at this time. Because no independent replication of this specific study is reported or evidenced, criterion R is not met.
    • A

      All-subject Exams

      • Criterion E is not met, and only EFL vocabulary outcomes were assessed, with no other core subjects measured.
      • Relevant Quotes: 1) "Descriptive statistics for vocabulary mastery, performance tasks, and learning analytics outcomes are presented in Table 1." (p. 72) 2) "The vocabulary mastery assessment protocol employed a 60-item instrument ... to evaluate lexical knowledge progression." (p. 65) Detailed Analysis: Per the criteria-specific instruction for A, since criterion E (Exam-based Assessment) is not met--the outcome measures are researcher-constructed vocabulary/performance instruments rather than a widely recognised standardised exam--criterion A cannot be met either. Independently, the study also only assesses English vocabulary/lexical outcomes (vocabulary mastery, writing/ oracy/literacy performance tied to target lexemes, and learning analytics), with no assessment of other core subjects such as mathematics or science. Because criterion E is not met and only a single subject domain (EFL vocabulary) is assessed, criterion A is not met. Quotes verified verbatim against the source PDF; no changes required.
    • G

      Graduation Tracking

      • Criterion Y is not met, and no follow-up beyond the 6-week delayed posttest or tracking toward graduation is reported or found in subsequent literature.
      • Relevant Quotes: 1) "The pretest established baseline proficiency, the posttest measured short-term gains after the AI intervention, and the delayed posttest (6 weeks post-intervention) assessed retention." (p. 69) Detailed Analysis: Per the criteria-specific instruction for G, since criterion Y (Year Duration) is not met--the total tracked interval is only about 14 weeks--criterion G cannot be met either. Independently, there is no mention anywhere in the paper of any follow-up beyond the 6-week delayed posttest, nor any reference to a follow-up publication tracking the same cohort toward graduation. Internet re-check: a Google Scholar citation search on the paper's DOI (10.1007/s10639-025-13803-2) returned no citing articles as of this verification (July 2026), and no subsequent publication by Miao Yu tracking this cohort toward graduation could be located. Given the paper was only published online on 27 October 2025, a graduation-tracking follow-up would not yet be expected to exist. Because criterion Y is not met and no tracking toward graduation is reported or found, criterion G is not met.
    • P

      Pre-Registered

      • The paper documents IRB ethics approval only; no pre-registration of the study protocol on a public registry is mentioned anywhere in the text or found externally.
      • Relevant Quotes: 1) "Ethical considerations. The study adhered to APA Ethical Principles and received IRB approval (Ref: #2023-EFL-045)." (p. 63) 2) "Ethical approval This study was approved by the institutional review board of Xinyang Normal University and in accordance with the Declaration of Helsinki. Written consent was obtained from the participants before the study was conducted." (p. 95) Detailed Analysis: The paper documents institutional ethics (IRB) approval and informed consent, but an IRB reference number is not a pre-registration of the study protocol. No quote anywhere in the paper mentions registration of the trial's hypotheses, methods, or analysis plan on a public registry (e.g., OSF, AsPredicted, ClinicalTrials.gov, ISRCTN) prior to data collection, nor any registration date or registry ID for the protocol itself. Internet re-check: no pre-registration record for this study or its author could be located; the paper's own metadata and declarations section reference only IRB ethics approval, with no registry ID anywhere in the text or references. Because no pre-registration of the study protocol is referenced anywhere in the text or found externally, criterion P is not met.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.