Abstract
Artificial Intelligence (AI)-driven personalized language learning holds great promise for customizing content and feedback according to learners' needs and proficiency levels. This study explores the potential of AI, particularly large language models (LLMs), for enhancing personalized language learning experiences. The research examines how LLMs can be leveraged to provide dynamic, tailored content and real-time feedback, adapting to individual learners' progress and learning preferences. A group of 100 language learners participated in a 6-week experiment where they engaged with AI-powered platforms that adjusted the content's difficulty and offered personalized feedback based on their performance and proficiency levels for English grammar and vocabulary acquisition. The study used a randomized controlled trial design, assessing learners' language proficiency, motivation, engagement, and perceived effectiveness of the AI-assisted learning experience before and after the intervention. The results showed that LLMs significantly improved learners' language proficiency, with noticeable increases in their engagement and motivation. Learners reported higher satisfaction with the personalized feedback and content, appreciating the system's ability to adapt to their unique needs. However, challenges such as occasional misinterpretations by the AI and limited content variety were identified.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was at the student level, but the intervention is one-to-one personalized AI instruction akin to tutoring, so the personal-teaching exception applies.
- "Participants in the experimental group interacted with an AI-powered platform designed to adjust the content's difficulty and provide personalized feedback based on real-time performance data." (pp. 2-3)
Relevant Quotes:
1) "The participants were randomly assigned to either an experimental group (AI-assisted learning) or a control group (traditional online learning without AI adaptation)." (p. 2)
2) "A total of 100 participants, aged 18 to 35, were identified as intermediate (CEFR B1 level) learners of English as a foreign language and recruited from a university language program." (p. 2)
3) "Participants in the experimental group interacted with an AI-powered platform designed to adjust the content's difficulty and provide personalized feedback based on real-time performance data." (pp. 2-3)
Detailed Analysis:
Randomisation was performed at the individual student level: 100 individual learners were "randomly assigned" to the experimental or control condition. No classes or schools were randomised. Criterion C normally requires class-level (or stronger) randomisation, but the standard allows an exception for interventions designed for personal, one-to-one teaching such as tutoring. The intervention here is inherently individual: each learner worked alone with an AI platform that adapted content difficulty and gave personalized feedback based on that learner's own performance, functioning as a one-to-one AI tutor. The control condition was likewise individual online study with static materials. Because delivery is personal and computer-mediated rather than classroom-based, contamination between individually randomised participants (the problem criterion C guards against) is minimal, and the tutoring/personal-teaching exception reasonably applies to this personalized one-to-one AI instruction.
Criterion C is met via the personal-teaching exception because the intervention is individually delivered one-to-one adaptive AI instruction, making student-level randomisation acceptable.
-
E
Exam-based Assessment
- The outcome tests are only vaguely described as "standardized" with no test names, sources, or evidence of wide recognition, so use of a genuine standardised exam cannot be confirmed.
- "Pre- and post-intervention assessments were conducted using standardized language proficiency tests measuring grammar and vocabulary knowledge..." (p. 3)
Relevant Quotes:
1) "Pre- and post-intervention assessments were conducted using standardized language proficiency tests measuring grammar and vocabulary knowledge, self-reported motivation and engagement scales, and surveys on learners' perceptions of the learning experience." (p. 3)
2) "The CEFR (Common European Framework of Reference for Languages) is an international standard used to describe language proficiency levels, ranging from A1 (beginner) to C2 (mastery)." (p. 2)
3) "An analysis of the proficiency scores, which measured grammar and vocabulary knowledge, is summarized in Table 1." (p. 3)
Detailed Analysis:
The paper asserts that "standardized language proficiency tests" were used, but it never names the test, cites its source, or provides any evidence of its validity, reliability, or wide recognition. The ERCT standard explicitly asks to check for "Names of standardised tests used, their validity and reliability". No recognised exam (e.g., IELTS, TOEFL, Cambridge, or a national curriculum exam) is identified anywhere in the paper; the CEFR is mentioned only as the framework used to classify participants' entry proficiency (B1), not as the outcome assessment. The reported scores (e.g., 72.3, 82.7) are on an unspecified scale with no reference to any recognised testing body. An unnamed, unverifiable test that the authors merely label "standardized" cannot be confirmed as a widely recognised standardised exam, and in a brief conference paper such a description is equally consistent with a researcher-assembled grammar/vocabulary test aligned to the taught content.
Criterion E is not met because no named, widely recognised standardised exam is identified; the bare claim of "standardized" tests is unverifiable.
-
T
Term Duration
- The interval from intervention start to outcome measurement was only 6 weeks, far shorter than one academic term.
- "A group of 100 language learners participated in a 6-week experiment where they engaged with AI-powered platforms..." (p. 1)
Relevant Quotes:
1) "A group of 100 language learners participated in a 6-week experiment where they engaged with AI-powered platforms..." (p. 1)
2) "This study addresses this gap by investigating the impact of LLMs on language learners' proficiency, engagement, and motivation over a 6-week intervention." (p. 2)
3) "Pre- and post-intervention assessments were conducted using standardized language proficiency tests..." (p. 3)
4) "Future research should explore longer intervention periods..." (p. 5)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. The paper repeatedly states the experiment lasted 6 weeks, with post-tests administered immediately after the intervention ("before and after the intervention"). Six weeks is roughly half of even a short academic term and clearly below the 3-4 month threshold. No delayed follow-up measurement is reported; the authors themselves flag "longer intervention periods" as future work, confirming the short duration.
Criterion T is not met because outcomes were measured only 6 weeks after the intervention began, well short of a full academic term.
-
D
Documented Control Group
- The control group's composition, baseline pre-test scores (Table 1), and the exact condition it received are documented.
- "In contrast, the control group used standard online language learning materials with static content. They accessed a fixed library of exercises covering the same grammatical and lexical topics as the experimental group." (p. 3)
Relevant Quotes:
1) "A total of 100 participants, aged 18 to 35, were identified as intermediate (CEFR B1 level) learners of English as a foreign language and recruited from a university language program." (p. 2)
2) "The participants were randomly assigned to either an experimental group (AI-assisted learning) or a control group (traditional online learning without AI adaptation)." (p. 2)
3) "In contrast, the control group used standard online language learning materials with static content. They accessed a fixed library of exercises covering the same grammatical and lexical topics as the experimental group. Feedback was limited to pre-programmed, generic responses, such as revealing the correct answer with a brief rule explanation." (p. 3)
4) Table 1 row "Traditional learning": Pre-Test Mean (SD) 71.8 (7.9); Post-Test Mean (SD) 76.5 (7.5); Gain Score (SD) 4.7 (3.5). (Table 1, p. 3)
Detailed Analysis:
Criterion D requires documentation of who the control group is, their baseline characteristics, and the treatment they received. The paper documents: (a) the population from which both groups were drawn (100 learners aged 18-35, CEFR B1, university language program) with random assignment; (b) the control group's baseline performance (pre-test mean 71.8, SD 7.9 in Table 1); and (c) exactly what the control group received (static online materials covering the same topics, with generic pre-programmed feedback and no AI adaptation). This confirms the control condition and shows baseline comparability with the treatment group (71.8 vs 72.3). Documentation is brief - per-group sample sizes and per-group demographic breakdowns are not tabulated - but the essential elements (composition, baseline scores, and control treatment) are explicitly documented.
Criterion D is met because the control group's origin, baseline test performance, and exact condition are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Individual students within one university program were randomised, not whole schools or institutional units.
- "The participants were randomly assigned to either an experimental group (AI-assisted learning) or a control group (traditional online learning without AI adaptation)." (p. 2)
Relevant Quotes:
1) "The participants were randomly assigned to either an experimental group (AI-assisted learning) or a control group (traditional online learning without AI adaptation)." (p. 2)
2) "A total of 100 participants... were... recruited from a university language program." (p. 2)
Detailed Analysis:
Criterion S requires randomisation at the level of whole schools or equivalent institutional units. Here all 100 participants came from a single university language program and were randomised individually. No schools, sites, or institutional units were randomised, and only one institution was involved. The tutoring-style exception that can excuse student-level randomisation for criterion C does not extend to criterion S, which specifically demands school-level assignment.
Criterion S is not met because randomisation occurred at the individual student level within a single university program.
-
I
Independent Conduct
- No independent evaluators or third-party oversight are mentioned; the authors appear to have designed the platform and conducted the evaluation themselves.
Relevant Quotes:
1) "Mengkai Wang* and Yifei Gao" and "Beijing Foreign Studies University, China" (p. 1)
2) "The platform, leveraging a model based on the GPT-4 architecture, was fine-tuned on a large corpus of learner texts with expert annotations to improve its accuracy in detecting grammatical errors and generating feedback." (p. 3)
3) "The study used a randomized controlled trial design, assessing learners' language proficiency, motivation, engagement, and perceived effectiveness of the AI-assisted learning experience before and after the intervention." (p. 1)
Detailed Analysis:
Criterion I requires that the study be conducted independently of the intervention's designers, or at least that independent third-party oversight of data collection and analysis be documented. The paper contains no acknowledgment, disclosure, or methods statement indicating that any external evaluator, independent agency, or third party was involved in delivering the intervention, collecting data, or analysing results. Both authors share the same university affiliation and corresponding-author addresses ("taylericy@bfsu.edu.cn" and "smxgaoyf@126.com", p. 1), with no separate evaluation team named. The description of the fine-tuned GPT-4-based platform and the study design suggests the same two-author team configured the intervention platform and also ran and analysed the trial. With no quoted evidence of independence, the criterion fails.
Criterion I is not met because no independent third party is documented as conducting or overseeing the study, which appears to have been run entirely by the authors.
-
Y
Year Duration
- The study spanned only 6 weeks from start to final measurement, far short of 75% of an academic year.
- "A group of 100 language learners participated in a 6-week experiment..." (p. 1)
Relevant Quotes:
1) "A group of 100 language learners participated in a 6-week experiment..." (p. 1)
2) "...investigating the impact of LLMs on language learners' proficiency, engagement, and motivation over a 6-week intervention." (p. 2)
Detailed Analysis:
Criterion Y requires outcomes to be measured at least 75% of a full academic year (roughly 7+ months) after the intervention begins. The entire study, from intervention start to post- test, spanned only 6 weeks. This is far below the year- duration threshold. Additionally, per the instructions, criterion Y cannot be met when criterion T (Term Duration) is not met, and T fails here.
Criterion Y is not met because the study lasted only 6 weeks, a small fraction of an academic year.
-
B
Balanced Control Group
- Both groups received equivalent online study time and content coverage; the only difference, AI-driven personalization, is the treatment variable being tested.
- "They accessed a fixed library of exercises covering the same grammatical and lexical topics as the experimental group." (p. 3)
Relevant Quotes:
1) "Participants in the experimental group interacted with an AI-powered platform designed to adjust the content's difficulty and provide personalized feedback based on real-time performance data." (pp. 2-3)
2) "In contrast, the control group used standard online language learning materials with static content. They accessed a fixed library of exercises covering the same grammatical and lexical topics as the experimental group. Feedback was limited to pre-programmed, generic responses, such as revealing the correct answer with a brief rule explanation." (p. 3)
3) "A group of 100 language learners participated in a 6-week experiment where they engaged with AI-powered platforms that adjusted the content's difficulty and offered personalized feedback..." (p. 1)
Detailed Analysis:
Criterion B requires comparable time and resources across conditions unless the extra resource is itself the treatment variable. Here the control was an active control: both groups studied online for the same 6-week period, both completed grammar and vocabulary exercises covering "the same grammatical and lexical topics", and both received feedback. The only substantive differences - adaptive difficulty and personalized AI-generated feedback versus static content and generic pre-programmed feedback - are precisely the treatment contrast the study is designed to test (AI adaptation vs. no AI adaptation). No additional instructional time, money, human tutoring, or materials were given to the experimental group beyond the AI personalization that constitutes the intervention itself. Applying the decision tree: extra resources beyond the treatment variable are not present; the AI adaptation is integral to and identical with the tested intervention; the control received a comparable educational substitute (same topics, same modality, same duration).
Criterion B is met because the active control matched the intervention group's time, modality, and content, with the AI personalization being the treatment variable itself.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific trial is mentioned in the paper or identifiable via internet search (OpenAlex citation count = 0), as the study was only published in 2025.
Relevant Quotes:
1) "Despite the excitement surrounding these technologies, empirical research into their effectiveness in educational settings, particularly in language learning, remains limited." (p. 2)
2) "Future research should explore longer intervention periods, different age groups, and hybrid models combining AI and human feedback to further refine and optimize personalized language learning experiences." (p. 5)
Detailed Analysis:
Criterion R requires that this specific study be independently replicated by a different research team in a different context, published in a peer-reviewed venue. The paper, a 2025 conference proceedings contribution, presents itself as addressing a gap in limited empirical research and does not mention any replication of its specific trial. While other teams have run RCTs on LLM-assisted language learning (e.g., the cited Zheng et al., 2025, on LLM dialogue partners for oral proficiency), those are different interventions and designs, not replications of this specific 6-week fine-tuned GPT-4 grammar/vocabulary platform trial.
Internet verification: an OpenAlex search for this paper ("Artificial intelligence-driven personalized language learning: Customizing content and feedback to learners' needs and proficiency levels", Wang & Gao, 2025, doi.org/10.29140/97817637116240-23) returns a citation count of 0 as of the check date. Additional searches (Google Scholar, Bing, DuckDuckGo) for "Mengkai Wang" and "Yifei Gao" combined with the study's topic did not surface any independent replication, citing study, or follow-up trial by a different research team. Given the very recent publication date and zero recorded citations, no independent replication of this particular study exists or is identifiable at this time.
Criterion R is not met because no independent replication of this specific study is reported or identifiable.
-
A
All-subject Exams
- Only English grammar and vocabulary were assessed, with no standardised exams across other subjects, and prerequisite criterion E is not met.
- "An analysis of the proficiency scores, which measured grammar and vocabulary knowledge, is summarized in Table 1." (p. 3)
Relevant Quotes:
1) "An analysis of the proficiency scores, which measured grammar and vocabulary knowledge, is summarized in Table 1." (p. 3)
2) "...engaged with AI-powered platforms that adjusted the content's difficulty and offered personalized feedback based on their performance and proficiency levels for English grammar and vocabulary acquisition." (p. 1)
Detailed Analysis:
Criterion A requires standardised exam-based assessment across all main subjects, and per the instructions it automatically fails when criterion E fails. Criterion E is not met here because no named standardised exam was used. Furthermore, the study measured only English grammar and vocabulary (plus non-academic self-report measures of motivation and engagement); no other subjects were assessed. As an adult/ university language-program intervention, a specialised focus might be arguable, but with E failed and only a single subject domain measured by an unverified instrument, the criterion cannot be satisfied.
Criterion A is not met because criterion E fails and only the single domain of English grammar and vocabulary was assessed.
-
G
Graduation Tracking
- Outcomes were measured only immediately post-intervention, with no tracking of participants until graduation, and no follow-up publications were found via internet search.
- "...assessing learners' language proficiency, motivation, engagement, and perceived effectiveness of the AI-assisted learning experience before and after the intervention." (p. 1)
Relevant Quotes:
1) "The study used a randomized controlled trial design, assessing learners' language proficiency, motivation, engagement, and perceived effectiveness of the AI-assisted learning experience before and after the intervention." (p. 1)
2) "Future research should explore longer intervention periods, different age groups, and hybrid models..." (p. 5)
Detailed Analysis:
Criterion G requires tracking participants until graduation from their educational stage, and per the instructions it automatically fails when criterion Y fails (Y is not met here). Measurement occurred only immediately before and after the 6-week intervention; there is no follow-up of any kind, let alone tracking until participants completed their university language program or degree. The authors' call for "longer intervention periods" in future research confirms no long-term tracking took place.
Internet verification: searches (OpenAlex, Google Scholar, Bing, DuckDuckGo) for subsequent publications by Mengkai Wang or Yifei Gao (Beijing Foreign Studies University) tracking the same 100-learner cohort found no follow-up papers; the original paper itself has zero recorded citations, and no later work by these authors on this cohort was located.
Criterion G is not met because measurement stopped immediately after the 6-week intervention with no graduation tracking, and no follow-up publications were found.
-
P
Pre-Registered
- No pre-registration, registry ID, or protocol statement appears anywhere in the paper.
Relevant Quotes:
No quotes referencing pre-registration, a trial registry, a registration ID, or a published protocol exist anywhere in the paper.
Detailed Analysis:
Criterion P requires that the full study protocol be pre-registered on a public registry before data collection began, with quoted evidence of the registration and its timing. The paper contains no mention of any registry (e.g., ClinicalTrials.gov, OSF, AEA registry), no registration ID, no protocol paper, and no statement about pre-specified hypotheses or analysis plans. Since no registry or protocol identifier is given anywhere in the paper, there is no registration entry to look up in an external database. Absent any quoted evidence of pre-registration, the criterion fails.
Criterion P is not met because the paper contains no reference to any pre-registered protocol or registry entry.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.