Abstract
Purpose: The purpose of this study is to investigate the transformative role of artificial intelligence (AI) tools in enhancing academic writing proficiency among English as a Foreign Language (EFL) learners. By focusing on the balance between syntactic complexity and clarity, the research evaluates the effectiveness of AI-enhanced educational tools such as Grammarly, ProWritingAid, Hemingway Editor, Quillbot, Writefull and Turnitin Revision Assistant. Utilizing a pretest-posttest randomized controlled trial, the study aims to measure improvements in clarity, precision, syntactic complexity, argumentation and adherence to academic standards, providing insights into AI's potential in educational practices. Design/methodology/approach: This study employs a pretest-posttest randomized controlled trial design to evaluate the impact of AI-enhanced educational tools on academic writing skills among 466 EFL postgraduate students. Participants are randomly assigned to six experimental groups, each using a different AI tool and a control group using traditional computer-assisted language learning methods. Writing proficiency is assessed using the IELTS writing test, syntactic complexity analysis, readability tests and a rubric for academic standards. Quantitative data are analyzed using ANOVA, while qualitative data from interviews and surveys provide insights into learners' perceptions of AI tools' effectiveness. Findings: The study finds that AI-enhanced educational tools significantly improve academic writing proficiency among EFL postgraduate students compared to traditional methods. Notably, the Turnitin Revision Assistant demonstrates remarkable effectiveness across multiple dimensions, including clarity, precision, syntactic complexity and adherence to academic standards. Other tools like Grammarly and ProWritingAid also show substantial improvements in writing skills. Qualitative feedback reveals that learners perceive AI tools as beneficial, though challenges such as over-reliance and maintaining personal voice are noted.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was performed at the individual student level into six AI-tool groups and one control group, not at the class level, and no tutoring exception is invoked.
- "A rigorously powered sample of 459 EFL postgraduate students (stratified by age/gender) was randomized into six experimental groups (AI tools: e.g., Grammarly, Hemingway Editor) and one control group (traditional CALL)."
Relevant Quotes:
1) "This study employs a pretest-posttest randomized controlled trial design to evaluate the impact of AI-enhanced educational tools on academic writing skills among 466 EFL postgraduate students." (Abstract, Design/methodology/approach)
2) "A rigorously powered sample of 459 EFL postgraduate students (stratified by age/gender) was randomized into six experimental groups (AI tools: e.g., Grammarly, Hemingway Editor) and one control group (traditional CALL)." (Section 4, Methodology)
3) "Participants are randomly assigned to six experimental groups, each using a different AI tool and a control group using traditional computer-assisted language learning methods." (Abstract)
Detailed Analysis:
The paper only describes randomisation at the level of individual students ("459 EFL postgraduate students...was randomized into six experimental groups"). There is no mention of classes, sections, or schools being the unit of randomisation, and no indication that entire cohorts or classrooms were assigned as a block. The ERCT standard requires class-level (or stronger) randomisation to avoid contamination, unless the intervention is a one-to-one personal tutoring intervention, in which case student-level randomisation is acceptable. Although the AI tools here are used individually by each student on their own writing, the paper never frames the study as a personal-tutoring intervention, nor does it discuss any steps taken to prevent contamination between students who may attend the same courses/institution and could share tool access or compare experiences. Absent such a quote or framing, the exception cannot be applied.
Note: the "459" randomised figure quoted above is itself inconsistent with the abstract, Table 1 (which sums to 466 participants across the seven groups), and the pretest/ posttest tables (Table 2, Table 3), all of which use N = 466 throughout, with no attrition or exclusion of 7 participants explained anywhere in the paper. This is a documented internal inconsistency in the paper's sample size reporting (see also criterion D below).
Because randomisation occurred at the student level without an explicit, quoted tutoring-exception rationale, criterion C is not met.
-
E
Exam-based Assessment
- The primary academic-writing outcome was assessed with the IELTS Academic Writing Test, a widely recognised international standardised exam.
- "Writing proficiency was measured using the IELTS Academic Writing Test (pretest-posttest), which evaluated task achievement, coherence, lexical resource, and grammatical accuracy through two tasks: a descriptive report and an argumentative essay, scored on a 0-9 scale with high reliability (Cronbach's α = 0.88) and construct validity (r = 0.85, p < 0.001)."
Relevant Quotes:
1) "Writing proficiency was measured using the IELTS Academic Writing Test (pretest-posttest), which evaluated task achievement, coherence, lexical resource, and grammatical accuracy through two tasks: a descriptive report and an argumentative essay, scored on a 0-9 scale with high reliability (Cronbach's α = 0.88) and construct validity (r = 0.85, p < 0.001)." (Section 4.1, Instruments)
2) "Syntactic complexity was analyzed using a specialized test based on (Atak et al., 2021) and (Qian et al., 2021), focusing on sentence length and subordinate clauses, with scores ranging from 0 to 10..." (Section 4.1)
3) "An academic writing rubric (Appendix C) evaluated writing quality across ten criteria, with scores ranging from 0 to 50..." (Section 4.1)
4) "Argumentation and persuasion skills were assessed using a 20-item rubric-based test (Appendix D)..." (Section 4.1)
Detailed Analysis:
Criterion E requires that the study's academic outcome be measured with a widely recognised standardised exam rather than a bespoke instrument. The paper's headline outcome, "academic writing proficiency," was assessed using the IELTS Academic Writing Test, an internationally recognised, standardised English-language assessment used well beyond this single study. This satisfies the core requirement for a standardised exam-based assessment.
Several supplementary measures used in the study (the syntactic-complexity test, the researcher-built academic writing rubric, the argumentation/persuasion rubric, and the Flesch-Kincaid readability score) are custom or adapted instruments rather than standardised exams, and would not on their own satisfy criterion E. However, since the primary academic-outcome instrument (IELTS) is a genuine standardised exam, criterion E is met.
-
T
Term Duration
- No start date, end date, or duration in weeks/months is given anywhere for the intervention, so a term-length follow-up cannot be verified.
Relevant Quotes:
1) "This study employs a pretest-posttest randomized controlled trial (RCT) grounded in three interlocking theoretical frameworks..." (Section 4, Methodology)
2) "Interventions targeted discrete writing skills (Appendix A), with standardized duration and assessment via validated pretest-posttest rubrics measuring clarity, argumentation, vocabulary diversity, and academic conventions." (Section 4)
3) "Second, the study focused on short-term improvements in academic writing, leaving the long-term impacts of AI-enhanced tools unexplored." (Section 6, Discussion, Limitations)
Detailed Analysis:
Criterion T requires a clearly documented interval of at least one academic term between intervention start and outcome measurement. The paper states that intervention duration was "standardized" across groups but never specifies what that duration actually was (no dates, no number of weeks or sessions, no mention of a semester or term). No timeline for when the pretest or posttest was administered relative to the start of the AI-tool practice is provided anywhere in the methodology, results, or appendices referenced in the text.
The authors themselves acknowledge in the limitations section that "the study focused on short-term improvements in academic writing, leaving the long-term impacts of AI-enhanced tools unexplored," which strongly suggests the intervention window was brief and not a full term, but even this does not give a quantifiable duration.
Because no quote establishes the length of the intervention-to-measurement interval, and the authors characterise the study as measuring only short-term effects, criterion T is not met.
-
D
Documented Control Group
- The control group is identified only by a group label, sample size, and a generic descriptor ("Regular CALL Course"), without demographic or baseline detail specific to that group.
- "Control Group (Group 7) 68 Regular CALL Course Baseline for comparative analysis"
Relevant Quotes:
1) "Control Group (Group 7) 68 Regular CALL Course Baseline for comparative analysis" (Table 1)
2) "Participants are randomly assigned to six experimental groups, each using a different AI tool and a control group using traditional computer-assisted language learning methods." (Abstract)
3) "A rigorously powered sample of 459 EFL postgraduate students (stratified by age/gender) was randomized into six experimental groups...and one control group (traditional CALL)." (Section 4)
4) "4. Baseline Equivalence: Pretest group comparisons confirmed baseline equivalence, with no significant differences across groups." (Section 5.1)
Detailed Analysis:
Criterion D requires detailed documentation of the control group's demographics, baseline performance, and the conditions it received. The paper reports only the control group's sample size (N=68) and a one-line descriptor ("Regular CALL Course") in a summary table, without describing what that regular course actually consists of (content, contact hours, materials), and without reporting any control-group-specific demographic or baseline-score breakdown; all baseline statistics reported (Tables 2-3, the "baseline equivalence" statement) are pooled across the full sample of 466/459 students rather than broken out by group.
This assessment is reinforced by a broader data-integrity concern: Table 2 is titled "Descriptive statistics of academic writing performance across groups" but its body reports only pooled totals (N = 466) for the whole sample, not the seven per-group breakdowns the title and surrounding narrative imply. The narrative text separately cites per-group means for Group 6 (e.g. IELTS posttest M = 5.85, pretest M = 2.32; syntactic complexity posttest M = 6.54, pretest M = 1.76) and for the control group (e.g. syntactic complexity posttest M = 2.89) that cannot be cross-checked against any published per-group table, and the total sample size itself is reported inconsistently (466 in the abstract and Tables 1-3 vs 459 in the methodology text). Combined, this makes it impossible to independently verify the control group's actual documented characteristics.
Because there is no group-specific demographic profile, verifiable baseline performance data, or detailed description of what the control condition entailed, the control group is not documented to the standard required by criterion D.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred at the individual student level within a single participant pool, not among schools.
- "A rigorously powered sample of 459 EFL postgraduate students (stratified by age/gender) was randomized into six experimental groups (AI tools: e.g., Grammarly, Hemingway Editor) and one control group (traditional CALL)."
Relevant Quotes:
1) "A rigorously powered sample of 459 EFL postgraduate students (stratified by age/gender) was randomized into six experimental groups (AI tools: e.g., Grammarly, Hemingway Editor) and one control group (traditional CALL)." (Section 4)
2) "Participants are randomly assigned to six experimental groups, each using a different AI tool and a control group using traditional computer-assisted language learning methods." (Abstract)
Detailed Analysis:
Criterion S requires randomisation at the school (or equivalent institutional) level. The paper provides no indication that multiple schools/universities were involved or that randomisation occurred above the individual-student level; only one institution (Urmia University) is named as the author's affiliation, and randomisation is described exclusively as assigning individual students to one of seven groups.
Since there is no quote indicating school-level (or higher) randomisation, and the stronger requirement subsumes the already-unmet class-level criterion, criterion S is not met.
-
I
Independent Conduct
- The study has a single author who both designed the AI writing-tool intervention and conducted the entire evaluation, with no independent third-party evaluator.
- "No other individuals have made significant contributions to the present study."
Relevant Quotes:
1) "Akbar Bahari" (sole listed author, title page)
2) "No other individuals have made significant contributions to the present study." (Author contributions statement)
3) "1 Urmia University, Urmia" and "Corresponding Author: Akbar Bahari (bahariakbar2020@gmail.com)" (title page, affiliation and corresponding-author lines)
Detailed Analysis:
Criterion I requires that the study be conducted independently of those who designed or have an interest in the intervention, typically via an external evaluation team or third-party data collectors. Here, a single author is credited with the entire study, and the paper explicitly states no other individuals contributed significantly. There is no mention of an external evaluation agency, blinded test administrators, or independent data-collection team.
Because the same sole author designed, implemented, and evaluated the intervention with no documented independent oversight, criterion I is not met.
-
Y
Year Duration
- Since criterion T (Term Duration) is not met and no duration is documented at all, criterion Y cannot be met either.
Relevant Quotes:
1) "Second, the study focused on short-term improvements in academic writing, leaving the long-term impacts of AI-enhanced tools unexplored. Longitudinal studies are needed to assess whether the observed gains are sustained over time..." (Section 6, Discussion, Limitations)
Detailed Analysis:
Criterion Y requires outcomes to be measured after at least 75% of a full academic year, and per the standard's dependency rule, if criterion T is not met, Y is automatically not met. As documented under criterion T, no start/end dates or duration are given for the intervention, and the authors explicitly describe their own study as focused on "short-term improvements," with longitudinal follow-up flagged as future work rather than something already conducted.
Criterion Y is not met.
-
B
Balanced Control Group
- Duration and assessment were standardized across all groups, and the tool used (AI vs. regular CALL) is itself the treatment variable being tested, so the control received a comparable "business as usual" course.
- "Interventions targeted discrete writing skills (Appendix A), with standardized duration and assessment via validated pretest-posttest rubrics measuring clarity, argumentation, vocabulary diversity, and academic conventions."
Relevant Quotes:
1) "Participants are randomly assigned to six experimental groups, each using a different AI tool and a control group using traditional computer-assisted language learning methods." (Abstract)
2) "Control Group (Group 7) 68 Regular CALL Course Baseline for comparative analysis" (Table 1)
3) "Interventions targeted discrete writing skills (Appendix A), with standardized duration and assessment via validated pretest-posttest rubrics measuring clarity, argumentation, vocabulary diversity, and academic conventions." (Section 4)
Detailed Analysis:
Applying the criterion-B decision tree: extra resources are present in the sense that experimental groups used specialised AI software, but the paper explicitly states that "standardized duration" applied across all seven arms, indicating equal instructional time regardless of group. The study's core research question is precisely "does using AI tool X for the same amount of course time improve writing more than the regular CALL course," i.e., which writing-support tool is used is itself the treatment variable under test (analogous to the B exception for resources framed as the primary treatment variable). The control group's "Regular CALL Course" represents the standard business-as-usual language-learning activity that would have occurred regardless, matched in time to the experimental arms.
There is no quote indicating the AI-tool groups received materially more class time, budget, or contact hours than the control group; the difference is the specific software tool used during equivalent course time. On this basis, criterion B is met.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study is mentioned or discoverable; the study was only published in 2025.
Relevant Quotes:
1) "This study offers a novel contribution by empirically evaluating the specific impacts of various AI tools on distinct writing competencies among EFL learners, addressing existing gaps in the literature." (Abstract, Originality/ value)
Detailed Analysis:
Criterion R requires independent replication of this specific study by a different research team, published in a peer-reviewed outlet. The paper explicitly frames itself as addressing a "novel" gap, and no prior or concurrent replication of this exact design (six-arm AI-tool comparison for EFL postgraduate academic writing) is referenced.
An internet search (July 2026) for independent replications of this specific six-arm AI-writing-tool RCT found none. Bahari has several closely related, more recent studies (e.g. "AI-mediated scaffolding in academic writing" and "Evaluating the Effectiveness of AI-Driven Approaches on EFL Learners' Expository Writing Skills," TESOL Journal, 2026), but these are by the same author and are not independent replications of this specific study's design, sample, or outcome measures. Other located material (e.g. Turnitin's own internal pilot-study reports on Revision Assistant) is not a peer-reviewed independent replication of this RCT. Given the paper was published in August 2025, there has been essentially no time for an independent replication to appear in the literature.
Criterion R is not met.
-
A
All-subject Exams
- Only writing/language outcomes were assessed; no other core academic subjects (e.g. mathematics, science) were measured, and no specialised-subject exception is argued.
- "Writing proficiency was measured using the IELTS Academic Writing Test (pretest-posttest)..."
Relevant Quotes:
1) "Writing proficiency was measured using the IELTS Academic Writing Test (pretest-posttest)..." (Section 4.1)
2) "Syntactic complexity was analyzed using a specialized test...Readability was assessed using the Flesch-Kincaid Readability Test...An academic writing rubric...Argumentation and persuasion skills were assessed using a 20-item rubric-based test..." (Section 4.1)
Detailed Analysis:
Criterion A requires that all main subjects taught at the relevant educational level be assessed, unless the intervention is a highly specialised one (e.g. vocational training) with an explicit rationale for a narrow scope. All outcome measures in this study (IELTS writing, syntactic complexity, readability, academic writing rubric, argumentation) are dimensions of a single subject area: academic/EFL writing. No other core subjects (mathematics, science, social studies, etc.) are assessed, and the paper does not offer an explicit vocational/specialised-education rationale of the kind the exception requires (this is a general postgraduate EFL writing course, not a narrowly vocational program).
Criterion A is not met.
-
G
Graduation Tracking
- Since criterion Y (Year Duration) is not met, criterion G is automatically not met; the study also reports only an immediate posttest with no follow-up.
- "Second, the study focused on short-term improvements in academic writing, leaving the long-term impacts of AI-enhanced tools unexplored."
Relevant Quotes:
1) "Second, the study focused on short-term improvements in academic writing, leaving the long-term impacts of AI-enhanced tools unexplored. Longitudinal studies are needed to assess whether the observed gains are sustained over time..." (Section 6, Discussion, Limitations)
Detailed Analysis:
Criterion G requires tracking of participants until graduation, and by the standard's dependency rule, if criterion Y is not met, G cannot be met either. As established above, Y is not met because no duration data exists and the authors themselves label the design as short-term only, with no follow-up conducted after the immediate posttest and no mention of any planned or published follow-up study tracking the same cohort.
An internet search (July 2026) for subsequent papers by Akbar Bahari tracking this same cohort of 466/459 EFL postgraduate students through to programme completion found no such follow-up publication; later Bahari studies located (e.g. the 2026 TESOL Journal and Interactive Learning Environments papers) describe different samples and designs, not a graduation-tracking extension of this trial.
Criterion G is not met.
-
P
Pre-Registered
- No pre-registration platform, registry ID, or registration date is mentioned anywhere in the paper.
Relevant Quotes:
1) (No statement referencing a trial registry, protocol pre-registration, or registration date was found anywhere in the Methodology, Instruments, Results, Discussion, or supplementary-material sections of the paper.)
Detailed Analysis:
Criterion P requires a quoted reference to a pre-registered study protocol (hypotheses, methods, planned analyses) filed before data collection began, typically via a named registry with a date. The paper describes its hypotheses (H1, H2) and statistical approach within the article itself but never states that this plan was filed with a registry in advance of data collection, nor gives any registration identifier or date.
An internet search (July 2026) for a pre-registration of this study (by title, DOI, and author) on OSF, AsPredicted, ClinicalTrials.gov, and general web search found no matching registry entry.
Because no pre-registration evidence is present, criterion P is not met.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.