Level 1 Criteria
-
C Class-level RCT
- Randomisation was performed on individual adult learners recruited by email across three universities, not on whole classes or schools, and the collaborative, instructor- delivered nature of the conditions means the tutoring exception does not apply.
- "Participants were stratified by gender, LexTALE quartile, and GPA, then randomized using covariate-adaptive minimization in REDCap by an independent statistician."
- Relevant Quotes: 1) "This study employed a parallel five-arm randomized controlled trial (RCT) with a pretest-posttest design to evaluate the effectiveness of NDLLT components." (Section 3.1, p. 94) 2) "A total of 400 adult EFL learners were recruited via institutional email from three universities (September 2024-January 2025), with inclusion criteria comprising intermediate English proficiency (B1 CEFR; LexTALE >= 60, validated against TOEFL iBT, r = 0.78, Cronbach's alpha = 0.87), age 18-35, and no prior NDLLT exposure." (Section 3.1, p. 94) 3) "Participants were stratified by gender, LexTALE quartile, and GPA, then randomized using covariate-adaptive minimization in REDCap by an independent statistician." (Section 3.1, p. 94) 4) "Allocation concealment was maintained by masking group labels ("A-E"), and randomization procedures ensured balanced representation by L1 language family and region." (Section 3.1, p. 94) 5) "The standardized 12-week intervention was delivered by trained instructors with session fidelity monitored via Azure Metrics Advisor" (Section 3.2, p. 94) Detailed Analysis: Criterion C requires random allocation of intact classes (or a stronger unit such as schools), so that treatment and control learners are not mixed within the same teaching group. Every statement about allocation in this paper refers to individual participants: learners were recruited one by one via institutional email, stratified on individual attributes (gender, LexTALE quartile, GPA), and then individually assigned by covariate-adaptive minimisation in REDCap. No quote anywhere in the paper describes classes, cohorts, sites, or institutions as the unit of randomisation, and the phrase "balanced representation by L1 language family and region" again concerns individual-level balance. The tutoring exception in the standard permits student- level randomisation when the intervention is inherently one-to-one personal teaching. That exception does not cleanly apply here. The paper frames the work as taking place in "AI-mediated EFL classrooms," the intervention was "delivered by trained instructors," and several arms are explicitly collaborative and group-based (the SFN arm uses "Decentralized peer-AI collaboration" and AR group tasks, and the CDS arm includes "Neuro-synchronized teamwork (EEG-HRV coherence)"; Appendix A). The abstract likewise describes "three collaborative AI conditions." With five arms mixed inside shared classrooms at the same universities, contamination between conditions is exactly the risk criterion C is designed to exclude. Verification note (verified checked, quotes and paper re-read in full): a genuine, verbatim internal contradiction is confirmed. The abstract reports a 12-week trial with "N = 784" and "three collaborative AI conditions ... with an active CALL control" (four groups), while Section 3.1 (Method) reports "400 adult EFL learners" recruited, 7 excluded as outliers, and a final analysed sample of "N = 393" (Section 4.1) across a "parallel five-arm" design (five groups: SIC, AHT, SFN, NCS, CDS). 784 is neither the recruited (400) nor analysed (393) total, and it is unclear how a four-condition abstract reconciles with the five-arm Method. This sample-size/ design mismatch further undermines confidence in the allocation description but does not itself change the criterion decision, since every available description of the randomisation unit (individual-level, REDCap, covariate-adaptive) is internally consistent even though the headline N is not. Criterion C is not met because randomisation was carried out on individual students rather than on classes or schools, and the classroom-based, partly collaborative intervention does not qualify for the personal-tutoring exception.
-
E Exam-based Assessment
- Primary proficiency outcomes were measured with the TOEFL iBT speaking, listening and reading tests together with IELTS-aligned writing tasks scored by Criterion E-Rater, all of which are widely recognised standardised instruments.
- "Speaking used the TOEFL iBT Speaking Test (Cronbach's alpha = .92; CFA: chi-sq/df = 1.85, CFI = 0.98)"
- Relevant Quotes: 1) "Speaking used the TOEFL iBT Speaking Test (Cronbach's alpha = .92; CFA: chi-sq/df = 1.85, CFI = 0.98) with AI-enhanced pronunciation analysis (Speechify, Eloquence AI; r = .85, p < .001) and BERT-based grammar assessments (RMSEA = 0.04)." (Section 3.4, p. 95) 2) "Writing utilized IELTS-aligned tasks (alpha = .89; CFA: CFI = 0.95), Criterion E-Rater diagnostics (Phi = .89), Lexical Complexity Analyzer (RMSEA = 0.038), and Coh-Metrix 3.0 (alpha = .93)." (Section 3.4, p. 95) 3) "Listening (TOEFL iBT, alpha = .91) and reading (Praat metrics, kappa = .91) employed automated protocols." (Section 3.4, pp. 95-96) 4) "Outcomes included TOEFL iBT skills, a 50-item NDLLS motivation scale, an 18-item feedback survey, and interviews." (Abstract, p. 89) 5) "1.1 Speaking Proficiency Test (TOEFL iBT) ... Standardized ETS protocols; digital recording; inter-rater calibration; task counterbalancing" (Appendix B, p. 111) 6) "4.1 Reading Comprehension ... TOEFL iBT LTT algorithms; standardized monitor calibration; Delphi panel validation" (Appendix B, p. 112) Detailed Analysis: Criterion E asks whether the study's outcome measurement rests on standardised, externally recognised examinations rather than instruments built by the researchers to match their own intervention. The paper's headline linguistic outcomes are anchored on the TOEFL iBT (speaking, listening, reading), administered under "Standardized ETS protocols" with ETS-certified raters, and on IELTS-aligned writing tasks scored with ETS Criterion E-Rater. TOEFL iBT and IELTS are among the best-known standardised English proficiency examinations, and their use here is documented with reliability and convergent-validity evidence against IELTS bands. The study additionally reports many bespoke or author-assembled measures (AI-derived pronunciation and fluency indices, Coh-Metrix and LIWC analytics, a 50-item NDLLS motivation scale, EEG/fMRI markers). These are supplementary; they do not displace the standardised core. Because the standard asks only that a standard exam-based assessment be used, the presence of extra custom analytics does not defeat the criterion. Verification note: the previous ranking's quotes dropped several Greek-letter symbols (alpha, chi-squared, kappa, Phi) during PDF text extraction. These have been restored against the source PDF above; the substance of the quotes and the "met" decision are unchanged. Criterion E is met because proficiency outcomes were assessed with recognised standardised examinations (TOEFL iBT, IELTS-aligned tasks scored by Criterion E-Rater) rather than solely with instruments invented for this study.
-
T Term Duration
- The intervention ran for a standardised 12 weeks with a further 8-week delayed posttest, giving roughly 20 weeks (about five months) from intervention start to final measurement, which exceeds one academic term.
- "Table 1 presents adjusted means and standard errors for 32 outcome variables across five intervention groups (N = 393) at posttest and 8-week delayed posttest."
- Relevant Quotes: 1) "The standardized 12-week intervention was delivered by trained instructors with session fidelity monitored via Azure Metrics Advisor; a CONSORT flow diagram is provided in the supplementary materials (Figure 2)." (Section 3.2, p. 94) 2) "A 12-week randomized controlled trial (N = 784; CEFR B2-C1) compared three collaborative AI conditions" (Abstract, p. 89) 3) "Table 1 presents adjusted means and standard errors for 32 outcome variables across five intervention groups (N = 393) at posttest and 8-week delayed posttest." (Section 4.1, p. 96) 4) "Eight-week delayed posttest assessments evaluated intervention durability (Table 3)." (Section 4.1, p. 97) 5) "A total of 400 adult EFL learners were recruited via institutional email from three universities (September 2024-January 2025)" (Section 3.1, p. 94) 6) "Additionally, the study's 8-week duration precludes conclusions about long-term retention, and the small neuroimaging subsample (n = 40) may limit the statistical power of brain-behavior analyses." (Section 8, p. 103) Detailed Analysis: Criterion T requires at least one full academic term (roughly 3-4 months) between the start of the intervention and the primary outcome measurement. Both the abstract and the Method describe a 12-week intervention, which is itself approximately one academic term. Outcomes were then measured again at a delayed posttest eight weeks after the immediate posttest, so the interval from intervention start to the final measurement point is approximately 20 weeks, or close to five months. That comfortably exceeds a single term. Verification note (flagged for manual review): quote 6 is a confirmed, verbatim internal contradiction. Section 8 (Limitations) literally states "the study's 8-week duration," which directly conflicts with the Method's repeated, more detailed statement of a "standardized 12-week intervention" (with an accompanying CONSORT flow diagram) plus an "8-week delayed posttest" on top of that. The Method's description is treated as the more authoritative source because it is specific, appears twice independently (Section 3.2 and Table 1/3), and is tied to a CONSORT diagram; the Limitations sentence most plausibly mis-states the delayed-posttest window as the whole study duration, but the paper never resolves this explicitly. Under the most conservative possible reading (intervention = 8 weeks only, ~1.8 months), criterion T would be in doubt even though 8 weeks itself is stated as a duration, not a start-to-measurement interval, so this reading is not self-consistent either. Given the unresolved contradiction, this criterion is decided on the balance of the more detailed and repeated 12-week description, but the discrepancy itself should be treated as a documentation defect warranting manual review rather than an incidental slip. Criterion T is met because, on the best-supported reading of the paper's own figures, the interval from the start of the 12-week intervention to the 8-week delayed posttest is approximately five months, which is at least one full academic term; however, the paper's own Limitations section contradicts its stated duration and this should be flagged for manual review.
-
D Documented Control Group
- The static control condition's content, delivery platform, activities and baseline equivalence with the other arms are described, although group-specific sample sizes and demographics are not tabulated.
- "The Static Isomorphic Control (SIC) established a non-adaptive baseline using fixed spaced repetition and rule-based drills, validating linear models (e.g., Ebbinghaus)."
- Relevant Quotes: 1) "The Static Isomorphic Control (SIC) established a non-adaptive baseline using fixed spaced repetition and rule-based drills, validating linear models (e.g., Ebbinghaus)." (Section 3.3, p. 95) 2) "Group 1: SIC (Static Control) ... Non-adaptive baseline (linear curriculum validated against Memrise) ... 15 CEFR modules (A1-B1); Leitner system (24h/7d/30d intervals); explicit SVO drills ... React Native v0.72.4 & Firebase v9.23.0; MediaPipe Gaze v0.10.3 (30Hz iris tracking) ... Personalization: None (fixed curriculum validated via Nation, 2006)." (Appendix A, pp. 110-111) 3) "Baseline proficiency equivalence was confirmed across groups using MANOVA (Pillai's Trace = 0.02, F(16,1556) = 1.08, p = .41; Cohen's d < 0.20 for all key variables)." (Section 3.1, p. 94) 4) "At pretest, all groups demonstrated comparable baseline performance (M range: 21.3-23.8), with no statistically significant differences (p = .214)." (Section 4.1, p. 99) 5) "Domain/Variable G1: SIC ... Speaking Proficiency 19.27 (1.03) ... Pronunciation Accuracy 67.90 (1.82) ... Lexical Complexity 65.59 (3.67) ... Motivation Scale 99.10 (2.70)" (Table 1, pp. 96-97) 6) "inclusion criteria comprising intermediate English proficiency (B1 CEFR; LexTALE 60 ...), age 18-35, and no prior NDLLT exposure" (Section 3.1, p. 94) Detailed Analysis: Criterion D asks whether the control condition is documented well enough for a reader to judge comparability: who the control participants were, how they performed at baseline, and what they actually received. On what the controls received, the paper is unusually detailed. Appendix A devotes a full column to Group 1 (SIC), specifying the curriculum (15 CEFR A1-B1 modules), the delivery schedule (Leitner spaced repetition at 24h, 7d and 30d intervals), the activity types (vocabulary grids, explicit SVO grammar rules), the software stack, the absence of personalisation, and even the observed outcomes (65% lexical retention at 7-day delay, engagement decay over 15% per week). This is a far richer description of a control arm than most trials supply, and it confirms the control was an active business-as-usual style condition rather than a no-treatment group. On baseline comparability, the paper reports a formal equivalence test across arms (MANOVA, Pillai's Trace = 0.02, p = .41, all Cohen's d < 0.20) and states that pretest means were comparable across groups (M range 21.3-23.8, p = .214). Table 1 reports the SIC group's values on the key outcome variables alongside the other arms. The weakness is that no per-arm sample size or per-arm demographic breakdown (gender, L1, age) is given; only whole-sample inclusion criteria and stratification variables are stated, and the total N itself is reported inconsistently (400 recruited, 393 analysed, 784 in the abstract - see the verified sample-size contradiction documented under criterion C). This is a documentation defect, but the substance the criterion asks for - control group composition rules, baseline performance, and the treatment the controls received - is present. Criterion D is met because the control condition's curriculum, delivery, platform and baseline equivalence are documented in the Method and Appendix A, notwithstanding the absence of a per-arm demographic table and the unresolved N discrepancy noted elsewhere in this report.