Nonlinear Dynamic Language Learning Theory in AI-Mediated EFL: From Theory to Practice

Akbar Bahari

Published:
ERCT Check Date:
DOI: 10.14505/jres.v16.2(20).05
  • L2 languages
  • higher education
  • Asia
  • gamification
  • blended learning
  • EdTech app
  • EdTech platform
  • digital assessment
  • mobile learning
0
  • C

    Randomisation was performed on individual adult learners recruited by email across three universities, not on whole classes or schools, and the collaborative, instructor- delivered nature of the conditions means the tutoring exception does not apply.

    "Participants were stratified by gender, LexTALE quartile, and GPA, then randomized using covariate-adaptive minimization in REDCap by an independent statistician."

  • E

    Primary proficiency outcomes were measured with the TOEFL iBT speaking, listening and reading tests together with IELTS-aligned writing tasks scored by Criterion E-Rater, all of which are widely recognised standardised instruments.

    "Speaking used the TOEFL iBT Speaking Test (Cronbach's alpha = .92; CFA: chi-sq/df = 1.85, CFI = 0.98)"

  • T

    The intervention ran for a standardised 12 weeks with a further 8-week delayed posttest, giving roughly 20 weeks (about five months) from intervention start to final measurement, which exceeds one academic term.

    "Table 1 presents adjusted means and standard errors for 32 outcome variables across five intervention groups (N = 393) at posttest and 8-week delayed posttest."

  • D

    The static control condition's content, delivery platform, activities and baseline equivalence with the other arms are described, although group-specific sample sizes and demographics are not tabulated.

    "The Static Isomorphic Control (SIC) established a non-adaptive baseline using fixed spaced repetition and rule-based drills, validating linear models (e.g., Ebbinghaus)."

  • S

    Allocation was to individual learners via covariate-adaptive minimisation, with no randomisation of schools, campuses or any other institutional unit.

    "Participants were stratified by gender, LexTALE quartile, and GPA, then randomized using covariate-adaptive minimization in REDCap by an independent statistician."

  • I

    The sole author is the originator of NDLLT and of the CDS model being tested and also performed the conceptualisation, investigation and analysis, so the evaluation was not independent of the intervention designer.

    "Akbar Bahari: Conceptualization; Methodology; Investigation; Formal analysis; Writing - original draft; Writing - review and editing; Visualization; Supervision; Project administration; Validation."

  • Y

    The tracked period runs about 20 weeks from the start of the 12-week intervention to the 8-week delayed posttest, roughly five months, well under 75% of an academic year.

    "Additionally, the study's 8-week duration precludes conclusions about long-term retention, and the small neuroimaging subsample (n = 40) may limit the statistical power of brain-behavior analyses."

  • B

    The control arm was an active, content-matched digital curriculum, time-on-task did not differ across groups, and the extra hardware and adaptive AI given to the treatment arms is the treatment variable under test rather than an unbalanced add-on.

    "Time-on-task did not differ significantly across groups (p = .34), indicating that observed differences were not attributable to differential exposure."

  • R

    No independent replication of this trial by an unrelated research team was found; a citation-index search located six citing works, five self- or co-authored by Bahari and one unrelated systematic review, none of which is an independent RCT replication.

    "The magnitude of observed effects warrants replication across diverse contexts to confirm generalizability."

  • A

    All outcomes lie within English-as-a-foreign-language proficiency and its neurocognitive and affective correlates; no other curricular subject was assessed and no rationale for that restriction is offered.

    "Table 1 presents adjusted means and standard errors for 32 outcome variables across five intervention groups (N = 393) at posttest and 8-week delayed posttest."

  • G

    Follow-up ended with an 8-week delayed posttest, no graduation or degree-completion data were collected, criterion Y is not met (which also precludes G), and a citation search found no follow-up publication tracking the original cohort.

    "Second, longitudinal studies extending neurocognitive and proficiency tracking to 24 months would offer insights into long-term learning trajectories and critical periods."

  • P

    The paper reports only an institutional ethics approval number and names no trial registry, registration identifier, or pre-registration date preceding data collection; an internet search found no matching registration.

    "All procedures received IRB approval (#2024-NDLLT-ELT), with informed consent obtained in multiple languages and full participant rights maintained."

Abstract

Grounded in a critical-realist ontology and a pragmatic-constructivist epistemology, this study operationalizes Nonlinear Dynamic Language Learning Theory (NDLLT) in AI-mediated EFL classrooms and empirically examines motivation as a fluctuating, history-dependent system. A 12-week randomized controlled trial (N = 784; CEFR B2-C1) compared three collaborative AI conditions (AI-enhanced Socrative, team-based Kahoot!, adaptive Duolingo + collaborative production) with an active CALL control. Outcomes included TOEFL iBT skills, a 50-item NDLLS motivation scale, an 18-item feedback survey, and interviews. MANCOVA/ANCOVA tested group differences; cross-lagged structural models estimated coupling between proficiency gains and motivational change; nonlinear time-series analyses (e.g., recurrence quantification, detrended fluctuation analysis) characterized attractor strength, variability, and phase shifts. Relative to CALL, AI conditions produced larger gains in reading and writing and more time in high-engagement attractor states, moderated by emotion regulation and peer collaboration. Engagement micro-variability prospectively predicted proficiency gains, consistent with NDLLT's phase-shift hypothesis. Implementation fidelity and accessibility/fairness safeguards supported validity. Findings depict proficiency and motivation as co-evolving trajectories within learner-AI-peer ecologies and argue for proficiency-sensitive scaffolding that tunes control parameters rather than prescribing linear sequences.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was performed on individual adult learners recruited by email across three universities, not on whole classes or schools, and the collaborative, instructor- delivered nature of the conditions means the tutoring exception does not apply.
      • "Participants were stratified by gender, LexTALE quartile, and GPA, then randomized using covariate-adaptive minimization in REDCap by an independent statistician."
      • Relevant Quotes: 1) "This study employed a parallel five-arm randomized controlled trial (RCT) with a pretest-posttest design to evaluate the effectiveness of NDLLT components." (Section 3.1, p. 94) 2) "A total of 400 adult EFL learners were recruited via institutional email from three universities (September 2024-January 2025), with inclusion criteria comprising intermediate English proficiency (B1 CEFR; LexTALE >= 60, validated against TOEFL iBT, r = 0.78, Cronbach's alpha = 0.87), age 18-35, and no prior NDLLT exposure." (Section 3.1, p. 94) 3) "Participants were stratified by gender, LexTALE quartile, and GPA, then randomized using covariate-adaptive minimization in REDCap by an independent statistician." (Section 3.1, p. 94) 4) "Allocation concealment was maintained by masking group labels ("A-E"), and randomization procedures ensured balanced representation by L1 language family and region." (Section 3.1, p. 94) 5) "The standardized 12-week intervention was delivered by trained instructors with session fidelity monitored via Azure Metrics Advisor" (Section 3.2, p. 94) Detailed Analysis: Criterion C requires random allocation of intact classes (or a stronger unit such as schools), so that treatment and control learners are not mixed within the same teaching group. Every statement about allocation in this paper refers to individual participants: learners were recruited one by one via institutional email, stratified on individual attributes (gender, LexTALE quartile, GPA), and then individually assigned by covariate-adaptive minimisation in REDCap. No quote anywhere in the paper describes classes, cohorts, sites, or institutions as the unit of randomisation, and the phrase "balanced representation by L1 language family and region" again concerns individual-level balance. The tutoring exception in the standard permits student- level randomisation when the intervention is inherently one-to-one personal teaching. That exception does not cleanly apply here. The paper frames the work as taking place in "AI-mediated EFL classrooms," the intervention was "delivered by trained instructors," and several arms are explicitly collaborative and group-based (the SFN arm uses "Decentralized peer-AI collaboration" and AR group tasks, and the CDS arm includes "Neuro-synchronized teamwork (EEG-HRV coherence)"; Appendix A). The abstract likewise describes "three collaborative AI conditions." With five arms mixed inside shared classrooms at the same universities, contamination between conditions is exactly the risk criterion C is designed to exclude. Verification note (verified checked, quotes and paper re-read in full): a genuine, verbatim internal contradiction is confirmed. The abstract reports a 12-week trial with "N = 784" and "three collaborative AI conditions ... with an active CALL control" (four groups), while Section 3.1 (Method) reports "400 adult EFL learners" recruited, 7 excluded as outliers, and a final analysed sample of "N = 393" (Section 4.1) across a "parallel five-arm" design (five groups: SIC, AHT, SFN, NCS, CDS). 784 is neither the recruited (400) nor analysed (393) total, and it is unclear how a four-condition abstract reconciles with the five-arm Method. This sample-size/ design mismatch further undermines confidence in the allocation description but does not itself change the criterion decision, since every available description of the randomisation unit (individual-level, REDCap, covariate-adaptive) is internally consistent even though the headline N is not. Criterion C is not met because randomisation was carried out on individual students rather than on classes or schools, and the classroom-based, partly collaborative intervention does not qualify for the personal-tutoring exception.
    • E

      Exam-based Assessment

      • Primary proficiency outcomes were measured with the TOEFL iBT speaking, listening and reading tests together with IELTS-aligned writing tasks scored by Criterion E-Rater, all of which are widely recognised standardised instruments.
      • "Speaking used the TOEFL iBT Speaking Test (Cronbach's alpha = .92; CFA: chi-sq/df = 1.85, CFI = 0.98)"
      • Relevant Quotes: 1) "Speaking used the TOEFL iBT Speaking Test (Cronbach's alpha = .92; CFA: chi-sq/df = 1.85, CFI = 0.98) with AI-enhanced pronunciation analysis (Speechify, Eloquence AI; r = .85, p < .001) and BERT-based grammar assessments (RMSEA = 0.04)." (Section 3.4, p. 95) 2) "Writing utilized IELTS-aligned tasks (alpha = .89; CFA: CFI = 0.95), Criterion E-Rater diagnostics (Phi = .89), Lexical Complexity Analyzer (RMSEA = 0.038), and Coh-Metrix 3.0 (alpha = .93)." (Section 3.4, p. 95) 3) "Listening (TOEFL iBT, alpha = .91) and reading (Praat metrics, kappa = .91) employed automated protocols." (Section 3.4, pp. 95-96) 4) "Outcomes included TOEFL iBT skills, a 50-item NDLLS motivation scale, an 18-item feedback survey, and interviews." (Abstract, p. 89) 5) "1.1 Speaking Proficiency Test (TOEFL iBT) ... Standardized ETS protocols; digital recording; inter-rater calibration; task counterbalancing" (Appendix B, p. 111) 6) "4.1 Reading Comprehension ... TOEFL iBT LTT algorithms; standardized monitor calibration; Delphi panel validation" (Appendix B, p. 112) Detailed Analysis: Criterion E asks whether the study's outcome measurement rests on standardised, externally recognised examinations rather than instruments built by the researchers to match their own intervention. The paper's headline linguistic outcomes are anchored on the TOEFL iBT (speaking, listening, reading), administered under "Standardized ETS protocols" with ETS-certified raters, and on IELTS-aligned writing tasks scored with ETS Criterion E-Rater. TOEFL iBT and IELTS are among the best-known standardised English proficiency examinations, and their use here is documented with reliability and convergent-validity evidence against IELTS bands. The study additionally reports many bespoke or author-assembled measures (AI-derived pronunciation and fluency indices, Coh-Metrix and LIWC analytics, a 50-item NDLLS motivation scale, EEG/fMRI markers). These are supplementary; they do not displace the standardised core. Because the standard asks only that a standard exam-based assessment be used, the presence of extra custom analytics does not defeat the criterion. Verification note: the previous ranking's quotes dropped several Greek-letter symbols (alpha, chi-squared, kappa, Phi) during PDF text extraction. These have been restored against the source PDF above; the substance of the quotes and the "met" decision are unchanged. Criterion E is met because proficiency outcomes were assessed with recognised standardised examinations (TOEFL iBT, IELTS-aligned tasks scored by Criterion E-Rater) rather than solely with instruments invented for this study.
    • T

      Term Duration

      • The intervention ran for a standardised 12 weeks with a further 8-week delayed posttest, giving roughly 20 weeks (about five months) from intervention start to final measurement, which exceeds one academic term.
      • "Table 1 presents adjusted means and standard errors for 32 outcome variables across five intervention groups (N = 393) at posttest and 8-week delayed posttest."
      • Relevant Quotes: 1) "The standardized 12-week intervention was delivered by trained instructors with session fidelity monitored via Azure Metrics Advisor; a CONSORT flow diagram is provided in the supplementary materials (Figure 2)." (Section 3.2, p. 94) 2) "A 12-week randomized controlled trial (N = 784; CEFR B2-C1) compared three collaborative AI conditions" (Abstract, p. 89) 3) "Table 1 presents adjusted means and standard errors for 32 outcome variables across five intervention groups (N = 393) at posttest and 8-week delayed posttest." (Section 4.1, p. 96) 4) "Eight-week delayed posttest assessments evaluated intervention durability (Table 3)." (Section 4.1, p. 97) 5) "A total of 400 adult EFL learners were recruited via institutional email from three universities (September 2024-January 2025)" (Section 3.1, p. 94) 6) "Additionally, the study's 8-week duration precludes conclusions about long-term retention, and the small neuroimaging subsample (n = 40) may limit the statistical power of brain-behavior analyses." (Section 8, p. 103) Detailed Analysis: Criterion T requires at least one full academic term (roughly 3-4 months) between the start of the intervention and the primary outcome measurement. Both the abstract and the Method describe a 12-week intervention, which is itself approximately one academic term. Outcomes were then measured again at a delayed posttest eight weeks after the immediate posttest, so the interval from intervention start to the final measurement point is approximately 20 weeks, or close to five months. That comfortably exceeds a single term. Verification note (flagged for manual review): quote 6 is a confirmed, verbatim internal contradiction. Section 8 (Limitations) literally states "the study's 8-week duration," which directly conflicts with the Method's repeated, more detailed statement of a "standardized 12-week intervention" (with an accompanying CONSORT flow diagram) plus an "8-week delayed posttest" on top of that. The Method's description is treated as the more authoritative source because it is specific, appears twice independently (Section 3.2 and Table 1/3), and is tied to a CONSORT diagram; the Limitations sentence most plausibly mis-states the delayed-posttest window as the whole study duration, but the paper never resolves this explicitly. Under the most conservative possible reading (intervention = 8 weeks only, ~1.8 months), criterion T would be in doubt even though 8 weeks itself is stated as a duration, not a start-to-measurement interval, so this reading is not self-consistent either. Given the unresolved contradiction, this criterion is decided on the balance of the more detailed and repeated 12-week description, but the discrepancy itself should be treated as a documentation defect warranting manual review rather than an incidental slip. Criterion T is met because, on the best-supported reading of the paper's own figures, the interval from the start of the 12-week intervention to the 8-week delayed posttest is approximately five months, which is at least one full academic term; however, the paper's own Limitations section contradicts its stated duration and this should be flagged for manual review.
    • D

      Documented Control Group

      • The static control condition's content, delivery platform, activities and baseline equivalence with the other arms are described, although group-specific sample sizes and demographics are not tabulated.
      • "The Static Isomorphic Control (SIC) established a non-adaptive baseline using fixed spaced repetition and rule-based drills, validating linear models (e.g., Ebbinghaus)."
      • Relevant Quotes: 1) "The Static Isomorphic Control (SIC) established a non-adaptive baseline using fixed spaced repetition and rule-based drills, validating linear models (e.g., Ebbinghaus)." (Section 3.3, p. 95) 2) "Group 1: SIC (Static Control) ... Non-adaptive baseline (linear curriculum validated against Memrise) ... 15 CEFR modules (A1-B1); Leitner system (24h/7d/30d intervals); explicit SVO drills ... React Native v0.72.4 & Firebase v9.23.0; MediaPipe Gaze v0.10.3 (30Hz iris tracking) ... Personalization: None (fixed curriculum validated via Nation, 2006)." (Appendix A, pp. 110-111) 3) "Baseline proficiency equivalence was confirmed across groups using MANOVA (Pillai's Trace = 0.02, F(16,1556) = 1.08, p = .41; Cohen's d < 0.20 for all key variables)." (Section 3.1, p. 94) 4) "At pretest, all groups demonstrated comparable baseline performance (M range: 21.3-23.8), with no statistically significant differences (p = .214)." (Section 4.1, p. 99) 5) "Domain/Variable G1: SIC ... Speaking Proficiency 19.27 (1.03) ... Pronunciation Accuracy 67.90 (1.82) ... Lexical Complexity 65.59 (3.67) ... Motivation Scale 99.10 (2.70)" (Table 1, pp. 96-97) 6) "inclusion criteria comprising intermediate English proficiency (B1 CEFR; LexTALE 60 ...), age 18-35, and no prior NDLLT exposure" (Section 3.1, p. 94) Detailed Analysis: Criterion D asks whether the control condition is documented well enough for a reader to judge comparability: who the control participants were, how they performed at baseline, and what they actually received. On what the controls received, the paper is unusually detailed. Appendix A devotes a full column to Group 1 (SIC), specifying the curriculum (15 CEFR A1-B1 modules), the delivery schedule (Leitner spaced repetition at 24h, 7d and 30d intervals), the activity types (vocabulary grids, explicit SVO grammar rules), the software stack, the absence of personalisation, and even the observed outcomes (65% lexical retention at 7-day delay, engagement decay over 15% per week). This is a far richer description of a control arm than most trials supply, and it confirms the control was an active business-as-usual style condition rather than a no-treatment group. On baseline comparability, the paper reports a formal equivalence test across arms (MANOVA, Pillai's Trace = 0.02, p = .41, all Cohen's d < 0.20) and states that pretest means were comparable across groups (M range 21.3-23.8, p = .214). Table 1 reports the SIC group's values on the key outcome variables alongside the other arms. The weakness is that no per-arm sample size or per-arm demographic breakdown (gender, L1, age) is given; only whole-sample inclusion criteria and stratification variables are stated, and the total N itself is reported inconsistently (400 recruited, 393 analysed, 784 in the abstract - see the verified sample-size contradiction documented under criterion C). This is a documentation defect, but the substance the criterion asks for - control group composition rules, baseline performance, and the treatment the controls received - is present. Criterion D is met because the control condition's curriculum, delivery, platform and baseline equivalence are documented in the Method and Appendix A, notwithstanding the absence of a per-arm demographic table and the unresolved N discrepancy noted elsewhere in this report.
  • Level 2 Criteria

    • S

      School-level RCT

      • Allocation was to individual learners via covariate-adaptive minimisation, with no randomisation of schools, campuses or any other institutional unit.
      • "Participants were stratified by gender, LexTALE quartile, and GPA, then randomized using covariate-adaptive minimization in REDCap by an independent statistician."
      • Relevant Quotes: 1) "Participants were stratified by gender, LexTALE quartile, and GPA, then randomized using covariate-adaptive minimization in REDCap by an independent statistician." (Section 3.1, p. 94) 2) "A total of 400 adult EFL learners were recruited via institutional email from three universities (September 2024-January 2025)" (Section 3.1, p. 94) 3) "The sample was limited to a single East Asian university (N = 393) with relatively homogeneous L1 backgrounds and uniform access to technology, which restricts the generalizability of results across different linguistic, cultural, and socioeconomic contexts." (Section 8, p. 103) 4) "First, cross-linguistic validation is needed through cluster-randomized trials involving typologically diverse language pairs and multilingual contexts to test the broader applicability of the NDLLT framework." (Section 8, p. 103) Detailed Analysis: Criterion S requires that whole educational institutions or implementation units be the randomised unit. The paper randomised individual participants; the only unit named anywhere in the allocation description is the participant. Verification note: quotes 2 and 3 confirm a genuine, verbatim site-count contradiction, separate from the duration contradiction under T and the sample-size contradiction under C. Section 3.1 (Method) states participants were "recruited via institutional email from three universities," while Section 8 (Limitations) states "the sample was limited to a single East Asian university (N = 393)." These cannot both be literally true; either the three-university recruitment description or the single-university limitation description is in error. This does not change the criterion decision (both readings still describe individual-level, not institution-level, randomisation - three universities as parallel recruitment pools is still not the same as randomising at the institution level), but it further undermines confidence in the paper's site/sample reporting and should be flagged for manual review alongside the T and C discrepancies. Decisively, the authors themselves list cluster-randomised designs as future work rather than as something already done: "cross-linguistic validation is needed through cluster-randomized trials." This confirms the present trial was not cluster- or school-randomised regardless of which site count is accurate. Criterion S is not met because randomisation was performed on individual learners, and the paper explicitly positions cluster-randomised trials as a future step.
    • I

      Independent Conduct

      • The sole author is the originator of NDLLT and of the CDS model being tested and also performed the conceptualisation, investigation and analysis, so the evaluation was not independent of the intervention designer.
      • "Akbar Bahari: Conceptualization; Methodology; Investigation; Formal analysis; Writing - original draft; Writing - review and editing; Visualization; Supervision; Project administration; Validation."
      • Relevant Quotes: 1) "Akbar Bahari: Conceptualization; Methodology; Investigation; Formal analysis; Writing - original draft; Writing - review and editing; Visualization; Supervision; Project administration; Validation." (Credit Authorship Contribution Statement, p. 104) 2) "This study addresses these gaps by proposing and testing Nonlinear Dynamic Language Learning Theory (NDLLT) as a unifying framework for AI-mediated language learning." (Introduction, p. 90) 3) "Central to this research is the Comprehensive Dynamic System (CDS) model, an instructional framework specifically developed to operationalize NDLLT's core principles in classroom contexts." (Section 7, p. 102) 4) "I thank the participating EFL students and instructors at the partner universities for their time and insight; the independent statistician who prepared the allocation sequence and envelopes; and the linguistics/anthropology panel who audited prompts for cultural fairness." (Acknowledgments, p. 104) 5) "A modified triple-blind protocol minimized bias: participants were masked to allocation (with sham EEG for non-NCS arms), outcome assessors were blinded (70% automated scoring via e-rater, 30% by trained raters, kappa = 0.87), and statistical analyses were conducted by blinded analysts on anonymized data." (Section 3.2, p. 94) 6) "I also appreciate the Socrative, Kahoot!, and Duolingo teams for granting research access to educational features without influencing study design, analysis, or reporting." (Acknowledgments, p. 104) Detailed Analysis: Criterion I requires that the trial be conducted independently of whoever designed the intervention, or at minimum that third-party oversight of measurement and analysis be documented. Here the paper is explicit that the author both proposed the theory ("proposing and testing Nonlinear Dynamic Language Learning Theory") and developed the instructional model under test (the CDS model, "specifically developed to operationalize NDLLT's core principles"). The CRediT statement then assigns to that same single author the conceptualisation, methodology, investigation, formal analysis, validation, supervision and project administration. Designer and evaluator are one person. There are partial mitigations. An independent statistician prepared the allocation sequence, blinded analysts are said to have run the statistics on anonymised data, 70% of scoring was automated via e-rater, and the platform vendors (Socrative, Kahoot!, Duolingo) are stated not to have influenced design, analysis or reporting. Vendor independence is genuinely relevant, but the conflict this criterion targets is not vendor influence - it is the author's own authorship of the intervention. Unlike the exception examples in the standard (where an external agency ran data collection, or a government led the trial), no external body led or evaluated this trial; the blinding claims are unverifiable assertions within a single-author paper, and no named third-party evaluator is identified. Verification note: this pattern is corroborated by an internet citation-index check performed for criterion R below - every empirical follow-up study citing this paper that was located is solely or jointly authored by Bahari himself, consistent with a single-investigator research programme rather than one subject to independent evaluation. Criterion I is not met because the sole author designed the NDLLT/CDS intervention and simultaneously ran the investigation, analysis and reporting, with no external evaluation body responsible for the trial.
    • Y

      Year Duration

      • The tracked period runs about 20 weeks from the start of the 12-week intervention to the 8-week delayed posttest, roughly five months, well under 75% of an academic year.
      • "Additionally, the study's 8-week duration precludes conclusions about long-term retention, and the small neuroimaging subsample (n = 40) may limit the statistical power of brain-behavior analyses."
      • Relevant Quotes: 1) "The standardized 12-week intervention was delivered by trained instructors with session fidelity monitored via Azure Metrics Advisor" (Section 3.2, p. 94) 2) "Table 1 presents adjusted means and standard errors for 32 outcome variables across five intervention groups (N = 393) at posttest and 8-week delayed posttest." (Section 4.1, p. 96) 3) "Additionally, the study's 8-week duration precludes conclusions about long-term retention, and the small neuroimaging subsample (n = 40) may limit the statistical power of brain-behavior analyses." (Section 8, p. 103) 4) "Second, longitudinal studies extending neurocognitive and proficiency tracking to 24 months would offer insights into long-term learning trajectories and critical periods." (Section 8, p. 103) 5) "A total of 400 adult EFL learners were recruited via institutional email from three universities (September 2024-January 2025)" (Section 3.1, p. 94) Detailed Analysis: Criterion Y requires outcomes to be measured at least 75% of an academic year (roughly 7 months of a 9-10 month year) after the intervention begins. Taking the most generous reading of the paper's own figures (see the T-criterion contradiction note above regarding the disputed 8-week vs. 12-week duration), the intervention ran 12 weeks and the final measurement was an 8-week delayed posttest, giving about 20 weeks or 4.6 months of tracking from intervention start. That is well short of 75% of an academic year under either possible reading of the paper's contradictory duration statements. The authors confirm the shortfall in their own Limitations section, conceding that the duration "precludes conclusions about long-term retention" and calling for future studies extending tracking to 24 months. Even the recruitment window (September 2024 to January 2025) spans only about five months, leaving no room for a year-long follow-up. Criterion Y is not met because the maximum interval from intervention start to final measurement is roughly five months under any reading of the paper's duration statements, far below the 75%-of-an-academic-year threshold, as the authors themselves acknowledge.
    • B

      Balanced Control Group

      • The control arm was an active, content-matched digital curriculum, time-on-task did not differ across groups, and the extra hardware and adaptive AI given to the treatment arms is the treatment variable under test rather than an unbalanced add-on.
      • "Time-on-task did not differ significantly across groups (p = .34), indicating that observed differences were not attributable to differential exposure."
      • Relevant Quotes: 1) "Time-on-task did not differ significantly across groups (p = .34), indicating that observed differences were not attributable to differential exposure." (Section 4.1, p. 100) 2) "The Static Isomorphic Control (SIC) established a non-adaptive baseline using fixed spaced repetition and rule-based drills, validating linear models (e.g., Ebbinghaus)." (Section 3.3, p. 95) 3) "This study systematically tested NDLLT through five experimental groups contrasting linear and nonlinear L2 acquisition dynamics (see appendix A)." (Section 3.3, p. 95) 4) "Group 1: SIC ... 15 CEFR modules (A1-B1); Leitner system (24h/7d/30d intervals); explicit SVO drills ... React Native v0.72.4 & Firebase v9.23.0; MediaPipe Gaze v0.10.3 (30Hz iris tracking)." (Appendix A, pp. 110-111) 5) "Neuro-Crossmodal Scaffolding (NCS) integrated biosensors (Muse 2 EEG, Apple Watch HRV) for embodied AR tasks modulated by LSTM/PPO, aligning with cross-modal plasticity." (Section 3.3, p. 95) 6) "Hardware cost ($847/learner); sensor drift (NeuroKit2 ICA artifact removal)." (Appendix A, Group 4 NCS, p. 110) 7) "Implementation fidelity was high: task completion rates exceeded 95% for all groups (CDS: M = 98.2%, SD = 1.1%)." (Section 4.1, p. 100) Detailed Analysis (re-applied against the updated criterion B decision tree, since the ERCT specification's wording for this criterion has changed since the prior check): Step 0 - EXTRA_RESOURCES_PRESENT: yes. The NCS arm alone carries a stated hardware cost of $847 per learner (Muse 2 EEG, Apple Watch), the SFN arm uses Meta Quest 3 headsets and federated infrastructure, and the AHT and CDS arms consume GPT-4 API and Kubernetes resources that the control does not. Step 1a - NEGLIGIBLE_DIFF: no, the budget/technology gap is substantial ($847/learner and dedicated hardware), so this branch does not apply. Step 2 - RESOURCES_ARE_TREATMENT: yes. The whole design is explicitly framed as a contrast of "linear and nonlinear L2 acquisition dynamics," where the adaptive AI, biosensor and federated architectures are precisely the manipulated variables; the SIC arm exists specifically as the "non-adaptive baseline." Removing the sensors and adaptive models from the treatment arms would remove the treatment itself, so under the decision tree this branch alone is sufficient to return "met," provided the control receives a comparable business-as-usual level of core instructional time - which is documented below. Time, separately, is explicitly balanced regardless of the resource-integrality finding: all arms underwent the same standardized 12-week programme, task completion rates exceeded 95% in every group, and the authors report a direct statistical test showing "Time-on-task did not differ significantly across groups (p = .34)." The control was not an idle or no-treatment group: it received a full 15-module CEFR A1-B1 curriculum with Leitner spaced repetition, explicit grammar drills and its own mobile app platform, i.e. an active, content-matched comparison condition. This matches the standard's DPL-tool exception example, where added devices integral to the tested intervention did not defeat criterion B. Criterion B is met because equal instructional time was documented across arms with an active content-matched control, and the extra hardware and adaptive AI budget in the treatment arms constitutes the treatment variable itself (per the updated decision tree's Step 2) rather than an unmatched supplementary resource.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this trial by an unrelated research team was found; a citation-index search located six citing works, five self- or co-authored by Bahari and one unrelated systematic review, none of which is an independent RCT replication.
      • "The magnitude of observed effects warrants replication across diverse contexts to confirm generalizability."
      • Relevant Quotes (from the paper itself): 1) "The magnitude of observed effects warrants replication across diverse contexts to confirm generalizability." (Section 4.1, p. 100) 2) "This study addresses these gaps by proposing and testing Nonlinear Dynamic Language Learning Theory (NDLLT) as a unifying framework for AI-mediated language learning." (Introduction, p. 90) 3) "First, cross-linguistic validation is needed through cluster-randomized trials involving typologically diverse language pairs and multilingual contexts to test the broader applicability of the NDLLT framework." (Section 8, p. 103) 4) "Article info: Received 25 August 2025; Revised 8 September 2025; Accepted 20 September 2025; Published 30 December 2025." (p. 89) Internet Research (OpenAlex/Crossref citation-index search for works citing this paper, DOI 10.14505/jres.v16.2(20).05, OpenAlex ID W4417329462, performed 2026-07-27; abstract text below is reconstructed from indexed metadata, not confirmed against the original full texts, since full-text access was not available for these citing works): a) Bahari, A. (2026). "Evaluating the Effectiveness of AI-Driven Approaches on EFL Learners' Expository Writing Skills." TESOL Journal. Sole-authored by the same author as the present study; a new sample (291 postgraduate students) testing commercial writing tools, not a replication of the present NDLLT/CDS trial. b) Li, Q., & Bahari, A. (2026). "AI-mediated scaffolding in academic writing: a longitudinal mixed-methods study..." Interactive Learning Environments. Co-authored with Bahari; a new 16-week trial with 445 postgraduate TEFL students on writing scaffolding, not a replication of this study's design or population. c) Bahari, A. (2026). "Leveraging AI-enhanced interventions for targeted EFL teacher development..." Interactive Learning Environments. Sole-authored by Bahari; different population (187 EFL teachers), different intervention. d) Xu, X., & Bahari, A. (2026). "The impact of AI-enhanced interactive learning environments on EFL teachers' computational thinking pedagogical knowledge..." Interactive Learning Environments. Co-authored with Bahari; different population (125 EFL teachers), different construct. e) Li, Z., Gu, C., & Bahari, A. (2026). "Comparing AI-enhanced instructional designs for EFL postgraduate proficiency: a multisite randomized controlled trial..." Computer Assisted Language Learning. Co-authored with Bahari; a new multisite RCT (N = 245) with different interventions and a different outcome design, not an independent replication of the present study. f) Wu, M., & Liu, M. (2026). "Synthesizing AI-enhanced language learning through the lens of complex dynamic systems theory: a systematic literature review." Journal for EuroCALL (jccall). The only citing work with no Bahari co-authorship; however it is a systematic literature review of 35 studies through a Complex Dynamic Systems Theory lens, not an empirical RCT replication of this specific study's design or findings. Detailed Analysis: Criterion R requires that this specific study, or its central experimental claim, has been independently replicated by a different research team and published in a peer-reviewed outlet. Nothing in the paper itself reports such a replication, and the author repeatedly frames replication and validation as work still to be done. The internet search corroborates this: of six works citing this paper as of 2026-07-27, five are authored or co-authored by Bahari himself (i.e., self-citation / the author's own follow-up research programme, not independent replication), and the sixth is an unrelated systematic literature review rather than an empirical replication attempt. No independent research team was found to have reproduced this trial's design, population, or findings. Criterion R is not met because there is no independent peer-reviewed replication of this trial; all located empirical follow-up work is by the same author, and the one unrelated citing work is a literature review, not a replication.
    • A

      All-subject Exams

      • All outcomes lie within English-as-a-foreign-language proficiency and its neurocognitive and affective correlates; no other curricular subject was assessed and no rationale for that restriction is offered.
      • "Table 1 presents adjusted means and standard errors for 32 outcome variables across five intervention groups (N = 393) at posttest and 8-week delayed posttest."
      • Relevant Quotes: 1) "Outcomes included TOEFL iBT skills, a 50-item NDLLS motivation scale, an 18-item feedback survey, and interviews." (Abstract, p. 89) 2) "RQ1 (Quantitative): To what extent do NDLLT-aligned, adaptive AI interventions improve L2 proficiency (speaking fluency, writing complexity, reading accuracy, listening comprehension) relative to non-adaptive controls, and how do changes in neural connectivity and efficiency correlate with these gains?" (Introduction, p. 90) 3) "Domain/Variable ... Speaking Proficiency ... Pronunciation Accuracy ... Lexical Complexity ... Frontal Theta Power ... White Matter Connectivity ... Cognitive Load Scale ... Motivation Scale" (Table 1, pp. 96-97) 4) "The study employed theory-driven instruments triangulating behavioral, neurocognitive, and systemic metrics across all language domains to minimize confounds (e.g., placebo effects)." (Section 3.4, p. 95) 5) "The present study robustly demonstrates that NDLLT interventions significantly enhance L2 proficiency across fluency, complexity, accuracy, and comprehension domains." (Section 5, p. 101) Detailed Analysis: Criterion A requires impact to be measured with standardised exams across all the main subjects the participants study, so that gains in the target subject are not achieved at the expense of others. The 32 outcome variables here are extensive but entirely confined to one domain. They cover English subskills (speaking, pronunciation, grammar, fluency, writing, listening, reading, vocabulary, interactional competence) plus neurocognitive markers (frontal theta, fMRI, DTI) and affective measures (motivation, anxiety, cognitive load). The authors describe this breadth as spanning "all language domains" - which is coverage within English, not coverage of the participants' wider curriculum. These are university students who study other subjects; no mathematics, science, or other academic attainment measure was collected, and there is no evidence about whether substantial time spent in AR headsets and biosensor sessions affected performance elsewhere. The specification's exception applies to highly specialised interventions in upper secondary or vocational education where a clear rationale is given for measuring only related outcomes. No such rationale is offered here, and the exception is not framed for general higher-education populations of this kind. Criterion A is not met because outcome measurement was restricted to English proficiency and its correlates, with no assessment of other main subjects and no stated justification for that restriction.
    • G

      Graduation Tracking

      • Follow-up ended with an 8-week delayed posttest, no graduation or degree-completion data were collected, criterion Y is not met (which also precludes G), and a citation search found no follow-up publication tracking the original cohort.
      • "Second, longitudinal studies extending neurocognitive and proficiency tracking to 24 months would offer insights into long-term learning trajectories and critical periods."
      • Relevant Quotes (from the paper itself): 1) "Eight-week delayed posttest assessments evaluated intervention durability (Table 3)." (Section 4.1, p. 97) 2) "Additionally, the study's 8-week duration precludes conclusions about long-term retention" (Section 8, p. 103) 3) "Second, longitudinal studies extending neurocognitive and proficiency tracking to 24 months would offer insights into long-term learning trajectories and critical periods." (Section 8, p. 103) 4) "17 Delayed Posttesting Temporal Stability ... 8-week delayed roleplay (same academic scenarios) ... Excellent 8-week stability; negligible practice effects" (Appendix C, p. 115) Internet Research (OpenAlex/Crossref citation-index search for works citing this paper, performed 2026-07-27; see full list of six citing works under criterion R above): none of the located follow-up publications by the same author (Bahari, solo or co-authored, 2026) describe continued tracking of the original 393-participant CEFR B2-C1 adult EFL cohort from this trial. Instead, each uses an entirely new sample and population: 291 postgraduate students (expository writing study), 445 postgraduate TEFL students (academic-writing scaffolding study), 187 EFL teachers (teacher-development study), 125 EFL teachers (computational thinking study), and 245 postgraduate students at four East Asian universities (multisite instructional-design RCT). No degree-completion, programme-exit, or graduation-cohort follow-up of the original participants was found in any indexed source. Detailed Analysis: Criterion G requires participants to be tracked through to graduation from their educational stage, so that long-term consequences of the intervention can be observed. The final data collection point in this study is an 8-week delayed posttest following a 12-week programme. There is no mention of degree completion, programme exit, credential attainment, or any administrative follow-up of the cohort, and the participants (adults aged 18-35 recruited by institutional email) are not even described in terms of a shared graduation cohort. The author explicitly identifies extended tracking as unfinished business, proposing future "longitudinal studies extending neurocognitive and proficiency tracking to 24 months." The citation-index search confirms no follow-up publication tracking this specific cohort to graduation exists as of 2026-07-27. Additionally, the standard's chaining rule applies: because criterion Y is not met, criterion G cannot be met. Criterion G is not met because tracking stopped at an 8-week delayed posttest with no graduation data, no indexed follow-up study of the original cohort was found, and criterion Y was not satisfied.
    • P

      Pre-Registered

      • The paper reports only an institutional ethics approval number and names no trial registry, registration identifier, or pre-registration date preceding data collection; an internet search found no matching registration.
      • "All procedures received IRB approval (#2024-NDLLT-ELT), with informed consent obtained in multiple languages and full participant rights maintained."
      • Relevant Quotes: 1) "All procedures received IRB approval (#2024-NDLLT-ELT), with informed consent obtained in multiple languages and full participant rights maintained." (Section 3.2, p. 94) 2) "A modified triple-blind protocol minimized bias: participants were masked to allocation (with sham EEG for non-NCS arms), outcome assessors were blinded ... and statistical analyses were conducted by blinded analysts on anonymized data." (Section 3.2, p. 94) 3) "The standardized 12-week intervention was delivered by trained instructors with session fidelity monitored via Azure Metrics Advisor; a CONSORT flow diagram is provided in the supplementary materials (Figure 2)." (Section 3.2, p. 94) 4) "An a priori power analysis using G*Power 3.1 (MANCOVA, f = 0.15, alpha = 0.05, 1-beta = 0.90) determined that N = 350 was required for adequate power" (Section 3.1, p. 94) 5) "5.1 Dialogue Performance ... OSF repository workflows; dual Shure microphone setup; AI-human triangulation protocols" (Appendix B, p. 112) Internet Research (performed 2026-07-27): searches of the OSF API and a general web search for "NDLLT" or "Akbar Bahari" combined with trial-registry terms (ClinicalTrials. gov, OSF, AsPredicted, ISRCTN, AEA RCT registry) returned no matching pre-registration record. This corroborates the complete absence of any registry citation, identifier, or registration date inside the paper itself. Detailed Analysis: Criterion P requires a pre-registered study protocol - identified registry, registration identifier, and a registration date demonstrably before data collection began. The paper supplies none of these. The only formal approval cited is an institutional review board number (#2024-NDLLT-ELT), which is an ethics clearance, not a public pre-registration of hypotheses and analysis plans. The word "protocol" appears only in the sense of a blinding procedure or an administration procedure, never as a registered study protocol. There is no ClinicalTrials.gov, ISRCTN, OSF, AEA or AsPredicted identifier anywhere in the article, and no registration date is given against the September 2024 - January 2025 data collection window. Two items come closest but fall short: the a priori power analysis shows some advance planning but is not a public registration, and a passing mention of "OSF repository workflows" in an Appendix B "Replicability Measures" column refers to a data/material workflow aspiration for one instrument, with no repository link, project identifier, or date. Criterion P is not met because no trial registry entry, registration identifier, or pre-registration date is reported anywhere in the paper, and an independent internet search for such a registration returned no results.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.