Abstract
In the context of the burgeoning field of second language (L2) education, where proficient writing plays an integral role in effective language acquisition and communication, the ever-increasing technology development has influenced the trajectory of L2 writing development. To address the need for enhanced writing skills among English as a Foreign Language (EFL) learners, this study investigates the efficacy of Automated Writing Evaluation (AWE) training. A randomized controlled trial employing repeated measures was conducted, involving a participant pool of 190 Chinese EFL students. The study comprehensively assessed the effects of AWE training, utilizing the Grammarly platform, an AI-driven program, on various dimensions of writing skills, encompassing task achievement, coherence and cohesion, lexicon, and grammatical accuracy. Control variables included writing self-efficacy and global English proficiency. Writing skills were evaluated through the administration of an International English Language Testing System (IELTS) writing sample test. The results unequivocally demonstrate that the experimental group consistently exhibited superior performance across all facets of writing skills compared to the control group.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomization was performed at the individual student level within institutes, not at the class or school level, and the group-based intervention does not qualify for the one-to-one tutoring exception.
- "Post-enrollment, a blocked randomization technique was employed, utilizing computer-generated random numbers to allocate students to either the control or experimental groups." (p. 5)
Relevant Quotes:
1) "The randomization process was facilitated by bundling the two course options (AWE-based and conventional) under a single course-tandem, aptly named the English Writing Course. Enrollment into the course-tandem was exclusive, thus ensuring a controlled environment for the study. Post-enrollment, a blocked randomization technique was employed, utilizing computer-generated random numbers to allocate students to either the control or experimental groups." (p. 5)
2) "In total, 95 students were randomly assigned to the AWE-based writing intervention group (average age: M = 21.6, SD = 2.9; 60% female), while another 95 students were assigned to the traditional writing course group (average age: M = 21.4, SD = 2.7; 40% female)." (p. 5)
3) "Through this approach, an equitable distribution of students was achieved across all participating institutes." (p. 5)
Detailed Analysis:
The ERCT C criterion requires randomisation of entire classes (or schools) to prevent contamination between treatment and control participants who share the same learning environment. Here the unit of randomisation was clearly the individual student: computer-generated random numbers allocated enrolled students to the experimental or control condition within each institute. Both conditions then ran simultaneously within the same institutes ("Within each institute, a control group was established, participating in a conventional writing course", p. 5), so treated and untreated students coexisted in the same settings, which is exactly the contamination scenario the criterion is designed to avoid. The exception for personal one-to-one tutoring does not apply: the intervention was delivered through a software tool plus weekly group writing workshops, not individual tutoring.
Criterion C is not met because students, not classes or schools, were the unit of randomisation and no valid exception applies.
-
E
Exam-based Assessment
- Writing outcomes were measured with sample tasks from the IELTS, a widely recognised standardised examination, scored with the standard IELTS analytic rubric by two independent raters with good inter-rater reliability.
- "In this study, two sample tasks from the International English Language Testing System (IELTS) were used to measure the writing skills of the participants." (p. 5)
Relevant Quotes:
1) "In this study, two sample tasks from the International English Language Testing System (IELTS) were used to measure the writing skills of the participants." (p. 5)
2) "The writing performance of the participants was assessed using an analytic essay scoring scale based on the IELTS rubric." (p. 5)
3) "The IELTS rubric, renowned for its reliability, is extensively employed for assessing writing abilities within second language contexts." (p. 5)
4) "The selection of the IELTS rubric for the analytic essay scoring scale was based on its comprehensive nature and established reliability and validity in assessing writing skills." (p. 5)
5) "To ensure the consistency of the scoring process, two independent raters were recruited, and inter-rater reliability was calculated using Cohen's Kappa, which was reported to be 0.82." (p. 5)
6) "To evaluate the participants' general English language proficiency and ensure their comparability, the Oxford Placement Test (OPT) developed by Allan (2004) was employed." (p. 5)
Detailed Analysis:
Criterion E requires that outcomes be measured with a standard, widely recognised exam rather than a test custom built for the study. The primary outcome (writing skill in task achievement, coherence and cohesion, lexicon, and grammatical accuracy) was measured with authentic sample writing tasks from the IELTS, an internationally recognised standardised English examination, and scored with the official IELTS analytic band rubric (1-9 scale). The authors did not construct their own test items; they drew both the tasks and the scoring scale from an established standardised assessment with documented validity and reliability, and additionally verified scoring consistency with two independent raters (Cohen's Kappa = 0.82). Baseline comparability was checked with another standardised instrument, the Oxford Placement Test. Although the IELTS tasks were administered locally rather than as an official IELTS sitting, the instrument itself is a standard, widely recognised exam, which satisfies the intent of the criterion.
Criterion E is met because outcomes were assessed with sample tasks and the rubric of the standardised IELTS examination rather than a researcher-made test.
-
T
Term Duration
- Outcomes were measured at the end of the 12-week instructional period, an interval of approximately three months from the intervention start, which corresponds to one academic term.
- "Both groups received 12 weeks of instruction, during which the treatment group underwent AWE-based instruction." (p. 4)
Relevant Quotes:
1) "Both groups received 12 weeks of instruction, during which the treatment group underwent AWE-based instruction." (p. 4)
2) "Initial measurements were conducted as pretests, seamlessly integrated into the first two sessions of the respective course. Subsequent posttest measurements were conducted during the final session of the course." (p. 4)
3) "Subsequently, a post-test task was given to both the experimental and control groups after the completion of the 12-week instructional period, which included the AWE-based instruction and traditional writing instruction without AWE, respectively." (p. 5)
4) "The AWE tool was provided by Grammarly and was used by the students to submit a written essay in English every week for a period of 12 weeks." (p. 5)
Detailed Analysis:
Criterion T requires the interval from intervention start to the primary outcome measurement to be at least one full academic term (approximately 3-4 months, i.e., a semester or equivalent). The intervention ran for 12 weeks, with pretests in the first two sessions and the posttest in the final session of the course, so the start-to-measurement interval is roughly 12 weeks, about three months. Twelve weeks is the standard length of a semester/term of instruction in this context (the intervention was delivered as a complete writing course), and thus sits at the boundary the standard defines as one term. No exact calendar dates are given, which slightly weakens the documentation, but the quoted 12-week course length is explicit and consistent throughout the paper.
Criterion T is met because outcomes were measured 12 weeks (about one full academic term) after the intervention began.
-
D
Documented Control Group
- The control group's size, demographics, baseline scores, and the instruction it received are documented in detail, including baseline equivalence tests and Table 1 statistics.
- "In total, 95 students were randomly assigned to the AWE-based writing intervention group (average age: M = 21.6, SD = 2.9; 60% female), while another 95 students were assigned to the traditional writing course group (average age: M = 21.4, SD = 2.7; 40% female)." (p. 5)
Relevant Quotes:
1) "In total, 95 students were randomly assigned to the AWE-based writing intervention group (average age: M = 21.6, SD = 2.9; 60% female), while another 95 students were assigned to the traditional writing course group (average age: M = 21.4, SD = 2.7; 40% female)." (p. 5)
2) "On the other hand, the control group in this study received traditional writing instruction without the use of an AWE tool. The students in the control group were asked to write an essay in English every week for a period of 8 weeks, which were graded by the instructor based on a rubric that evaluated various aspects of writing, including grammar, spelling, vocabulary, and organization." (pp. 5-6)
3) "To ensure that the training conditions did not differ significantly at the outset of the research, two-tailed t-tests were performed to examine the pretest measures for all dependent and control variables. The baseline equivalence was examined for key characteristics." (p. 6)
4) "The pretest means for both groups were similar for all measures, and no significant differences were found." (p. 6)
5) "TABLE 1 Means and standard deviations for each group in pre- and post-tests." (p. 7)
Detailed Analysis:
Criterion D requires detailed documentation of the control group: who they are, their baseline performance, and what they received. The paper reports the control group's size (n = 95), mean age and standard deviation, gender composition, and the instruction it received (traditional writing instruction with weekly instructor-graded essays and writing workshops, without AWE). Table 1 provides pretest and posttest means and standard deviations for the control group on all four writing outcome measures, plus global English proficiency (OPT) and writing self-efficacy, and baseline equivalence was formally tested with t-tests showing no significant pretest differences. This is sufficient documentation to judge comparability of the control group.
Criterion D is met because the control group's composition, baseline performance, and treatment conditions are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred among individual students within institutes, not among schools or institutions, so the school-level RCT requirement is not satisfied.
- "Through this approach, an equitable distribution of students was achieved across all participating institutes." (p. 5)
Relevant Quotes:
1) "Within each institute, a control group was established, participating in a conventional writing course." (p. 5)
2) "Post-enrollment, a blocked randomization technique was employed, utilizing computer-generated random numbers to allocate students to either the control or experimental groups. Through this approach, an equitable distribution of students was achieved across all participating institutes." (p. 5)
3) "The study cohort comprised 190 intermediate EFL students (60% female), all of whom were enrolled in one of four distinct writing courses hosted by different institutes offering the writing intervention." (p. 4)
Detailed Analysis:
Criterion S requires randomisation at the level of the school or implementing institution. In this study four institutes hosted the courses, but the institutes themselves were not randomised to conditions; instead, both experimental and control conditions ran inside every institute, and individual students were allocated to conditions by computer-generated random numbers. This is explicitly student-level, not school-level (or even class-level), randomisation.
Criterion S is not met because randomisation was conducted among individual students within each institute rather than among schools or institutions.
-
I
Independent Conduct
- The same research team that designed the AWE-based intervention also delivered the training and conducted the pretests and analyses, with no external or third-party evaluation reported.
- "In addition, the pretest and all trainings were conducted by the research team to ensure consistency and fidelity to the experimental design." (p. 6)
Relevant Quotes:
1) "The implementation of the AWE-based writing evaluation was overseen by a team of researchers collaborating with two proficient English teachers." (p. 5)
2) "In addition, the pretest and all trainings were conducted by the research team to ensure consistency and fidelity to the experimental design." (p. 6)
3) "All the authors equally contributed to completing this project." (p. 9, Author contributions)
4) "To ensure the consistency of the scoring process, two independent raters were recruited, and inter-rater reliability was calculated using Cohen's Kappa, which was reported to be 0.82." (p. 5)
5) "The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest." (p. 9)
Detailed Analysis:
Criterion I requires that the study be conducted independently of those who designed the intervention, or at least that third-party oversight of data collection and analysis be documented. Here the authors designed the AWE-based instructional program, oversaw its implementation, and explicitly state that the pretest and all trainings were conducted by the research team itself. There is no mention of an external evaluation agency, an independent data-collection team, or blinded outcome administration. The recruitment of two independent raters for essay scoring is a useful reliability measure but concerns only the scoring step; it does not make the overall conduct of the trial independent of the intervention designers. The fact that the intervention uses a commercial third-party tool (Grammarly) does not confer independence either, since the evaluation itself was run entirely by the authors.
Criterion I is not met because the intervention designers themselves conducted the study without documented independent oversight.
-
Y
Year Duration
- The interval from intervention start to outcome measurement was only 12 weeks, far short of 75% of an academic year.
- "Both groups received 12 weeks of instruction, during which the treatment group underwent AWE-based instruction." (p. 4)
Relevant Quotes:
1) "Both groups received 12 weeks of instruction, during which the treatment group underwent AWE-based instruction." (p. 4)
2) "Subsequently, a post-test task was given to both the experimental and control groups after the completion of the 12-week instructional period..." (p. 5)
3) "Another limitation is that the study only examined short-term effects of AWE-based instruction, and it is unclear whether these effects would persist over time." (p. 9)
Detailed Analysis:
Criterion Y requires that outcomes be measured at least 75% of a full academic year (roughly 9-10 months, so at least about 7 months) after the intervention begins. In this study the entire tracking window from intervention start to posttest was 12 weeks (about 3 months), and the authors themselves describe the study as examining only short-term effects with no later follow-up measurement. Three months is well below 75% of an academic year.
Criterion Y is not met because the outcome measurement occurred only 12 weeks after the intervention started, far less than 75% of an academic year.
-
B
Balanced Control Group
- The experimental group had 12 weeks of weekly essay practice versus 8 weeks for controls, and the paper is contradictory about whether controls received the matching workshops, so time and resources were not verifiably balanced.
- "The students in the control group were asked to write an essay in English every week for a period of 8 weeks, which were graded by the instructor based on a rubric..." (pp. 5-6)
Relevant Quotes:
1) "The AWE tool was provided by Grammarly and was used by the students to submit a written essay in English every week for a period of 12 weeks." (p. 5)
2) "In addition to the AWE tool, the students in the experimental group received a weekly one-hour writing workshop that focused on developing their writing skills and providing additional opportunities for practice." (p. 5)
3) "In contrast, the control group received traditional writing instruction without the integration of AWE or the additional writing workshops." (p. 4)
4) "The students in the control group were asked to write an essay in English every week for a period of 8 weeks, which were graded by the instructor based on a rubric..." (pp. 5-6)
5) "The students in the control group also received a weekly one-hour writing workshop that was similar in content and structure to the workshops provided to the experimental group. However, the writing workshops in the control group did not include the use of an AWE tool." (p. 6)
Detailed Analysis:
The treatment variable is the AWE (Grammarly) feedback, which is integral to the intervention being tested, so the presence of automated feedback itself does not violate criterion B. However, the paper documents additional, non-integral resource imbalances. First, the experimental group submitted weekly essays for 12 weeks while the control group wrote weekly essays "for a period of 8 weeks", implying the intervention group received roughly 50% more structured writing practice; this extra practice time is not the treatment variable and is not matched. Second, the paper is internally contradictory about the workshops: the overview (p. 4) states the control group did not receive "the additional writing workshops", while the procedure section (p. 6) states the control group received a weekly workshop "similar in content and structure". Given this contradiction, it cannot be verified from the paper that instructional time was actually balanced between the two conditions. Applying the decision tree: extra resources are present (additional weeks of essay practice, and possibly workshops), the difference is not negligible, the extra practice time is not framed as the treatment variable, and the control group is not clearly documented as matching it.
Criterion B is not met because the intervention group received unmatched additional writing practice (12 vs 8 weeks of weekly essays) and the paper contradicts itself on whether control students received the matching workshops, so balanced time on task cannot be verified.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific trial is reported in the paper, and an internet search of the citing literature found no independent peer-reviewed replication of this study.
- "To address this limitation, future research should replicate the study with different populations." (p. 9)
Relevant Quotes:
1) "One limitation is that the study was conducted with a sample of Chinese EFL learners only, which limits the generalizability of the findings to other EFL contexts. To address this limitation, future research should replicate the study with different populations." (p. 9)
2) "As no similar studies were found, the widely accepted classification of effect sizes was employed..." (p. 6)
3) "Despite this trend, there is little empirical research investigating the effectiveness of AWE on EFL writing skills in China. This study aims to address this gap..." (p. 2)
Detailed Analysis:
Criterion R requires that the specific study (its central experimental claim, design, and context) be independently replicated by a different research team and published in a peer-reviewed journal. The paper itself reports no replication; on the contrary, the authors state that no similar studies were found when selecting effect-size benchmarks and explicitly call for future replication. Prior related work cited in the paper (e.g., Liao 2016; Wang et al. 2013; Waer 2023; Parra and Calero 2019) consists of earlier, differently designed studies of AWE tools in other contexts; these predate or parallel this trial and are not replications of this specific randomized controlled trial of Grammarly-based instruction with Chinese EFL learners.
Internet Search (2026): the OpenAlex and Semantic Scholar records for this paper (DOI 10.3389/fpsyg.2023.1249991) list approximately 87-89 citing works, spanning studies on AI writing feedback (ChatGPT, Grammarly, DeepSeek, Gemini) across Indonesian, Thai, Malaysian, Iranian, Spanish-speaking, and other Chinese EFL populations. None of these citing works describes an independent replication of this specific trial's design, population, and intervention (a 12-week Grammarly-AWE RCT with Chinese EFL learners at Tangshan Normal University); they cite it as background or a methodological comparator rather than reproducing it. No published independent replication of this particular 2023 trial was identified.
Criterion R is not met because no independent, peer-reviewed replication of this specific study is reported or found.
-
A
All-subject Exams
- Only English writing skill was assessed, with no measurement of other main subjects and no explicit rationale offered, so the all-subject requirement is not satisfied.
- "To evaluate the effectiveness of the AWE-based instruction, four measures including task achievement, coherence and cohesion, lexicon, and grammatical accuracy were used." (p. 6)
Relevant Quotes:
1) "To evaluate the effectiveness of the AWE-based instruction, four measures including task achievement, coherence and cohesion, lexicon, and grammatical accuracy were used." (p. 6)
2) "In this study, two sample tasks from the International English Language Testing System (IELTS) were used to measure the writing skills of the participants." (p. 5)
3) "Nonetheless, this study's limitations include using a single measure for writing skills and the need for further research on the long-term effectiveness of AWE on language learning outcomes." (p. 9)
4) "Lastly, the study did not examine the effects of AWE-based instruction on other aspects of writing, such as discourse organization and rhetorical strategies." (p. 9)
Detailed Analysis:
Criterion A requires standardised exam-based assessment across all main subjects taught at the relevant educational level, to detect possible negative spillovers on non-target subjects. This study measured only English L2 writing (four sub-scores of a single IELTS writing task); no other subject, and indeed no other English skill (reading, listening, speaking), was assessed as an outcome, and the authors acknowledge relying on a single measure of writing. The context is an extracurricular EFL writing course for adults, which narrows the plausible subject range, but the paper offers no explicit rationale of the kind the standard's specialised-intervention exception requires, and even within English instruction only the writing skill was examined. With such a narrow outcome set, possible trade-offs with other learning could not be detected.
Criterion A is not met because only one skill in one subject (English writing) was assessed, without a justified exception covering all main subjects.
-
G
Graduation Tracking
- Measurement stopped at the 12-week posttest with no follow-up tracking to graduation, no follow-up publications were found, and the prerequisite Year Duration criterion is also unmet.
- "Another limitation is that the study only examined short-term effects of AWE-based instruction, and it is unclear whether these effects would persist over time." (p. 9)
Relevant Quotes:
1) "Subsequent posttest measurements were conducted during the final session of the course." (p. 4)
2) "Another limitation is that the study only examined short-term effects of AWE-based instruction, and it is unclear whether these effects would persist over time. Therefore, future studies should investigate the long-term effects of AWE-based instructional programs on second language writing skills." (p. 9)
3) "Following the study's completion, all students were invited to engage in the alternate course as a continuation of their learning process." (p. 5)
Detailed Analysis:
Criterion G requires tracking participants until graduation from their educational stage. Here all measurement ended at the final session of the 12-week course; the authors describe the study as short-term only and call for future long-term research. No follow-up assessments, cohort tracking, or follow-up publications on this sample are mentioned in the paper.
Internet Search (2026): a search of the paper's citing literature (OpenAlex/Semantic Scholar, ~87-89 citing works as of 2026) and of the authors' names (Ping Wei, Xiaosai Wang, Hui Dong, Tangshan Normal University) did not surface any follow-up or longitudinal publication by these authors that tracks the same 190-student cohort toward graduation. No such papers were found, so none are cited here.
In addition, per the ranking instructions, criterion G cannot be met when criterion Y (Year Duration) is not met, and Y is not met for this study.
Criterion G is not met because tracking stopped at the immediate posttest with no graduation follow-up found in the paper or in subsequent literature, and the prerequisite Y criterion also fails.
-
P
Pre-Registered
- The paper contains no mention of pre-registration on any registry platform or of a protocol published before data collection began, and no external pre-registration record was found online.
Relevant Quotes:
1) "The studies involving humans were approved by the School of Foreign Languages, Tangshan Normal University, Tangshan, Hebei, China." (p. 9, Ethics statement)
2) "To rigorously assess the efficacy of the writing intervention, we adopted a randomized controlled trial (RCT) design with repeated measures (Friedman et al., 2010)." (p. 4)
Detailed Analysis:
Criterion P requires that the full study protocol, including hypotheses, methods, and planned analyses, be registered on a public registry before data collection began, with the registration identifiable and datable. The paper reports institutional ethics approval but contains no reference to any trial registry (e.g., ClinicalTrials.gov, OSF, AEA registry, Chinese Clinical Trial Registry), no registration number, and no statement that a protocol was published before data collection. Ethics approval is not a substitute for pre-registration.
Internet Search (2026): searches for a pre-registration record associated with this study, its authors, or its institution (Tangshan Normal University) did not locate any entry on OSF, AsPredicted, ChiCTR, or similar registries. No pre-registration record was found.
Criterion P is not met because no pre-registration statement, registry link, or registration date appears anywhere in the paper, and none was found through independent search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.