Abstract
The affordances of ChatGPT in language learning and teaching have gained increasing traction. While studies began to investigate the potential of ChatGPT as a feedback provider, little attention was given to ChatGPT's potential impact on students' writing performance and the ideal L2 writing self vis-a-vis the established automated writing evaluation systems (AWE). To address these gaps, a sequential explanatory mixed methods design was adopted. One hundred and fifty second-year university students from three writing classes in a Chinese public university were recruited and randomly divided into a ChatGPT group, an AWE group, and a control group. After an eleven-week intervention, the ANCOVA results showed that while the ChatGPT group scored significantly higher than the AWE group and the control group in post-writing performance as measured by their writing score, in terms of students' ideal L2 writing self, the ChatGPT group performed significantly lower than the AWE group with a medium effect size. Qualitative analysis of students' reflection papers revealed students' (over)reliance on the tool and the accompanying loss of creativity and agency. Pedagogical implications as well as directions for future research are also discussed.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Random assignment was described as being carried out at the level of individual student participants rather than at the level of entire classes.
- "In total, 150 participants from three writing classes agreed to take part in the research and were randomly assigned to a ChatGPT group (n=50), an AWE group (n=50), and a control group (n=50)." (p. 1154-1155)
Relevant Quotes:
1) "To mitigate the potential impact of confounding variables, three out of five writing classes, all taught by the same instructor, were randomly selected for this study." (p. 1154)
2) "In total, 150 participants from three writing classes agreed to take part in the research and were randomly assigned to a ChatGPT group (n=50), an AWE group (n=50), and a control group (n=50)." (p. 1154-1155)
3) "They are second-year non-English major students coming from disciplines including engineering, science, and arts." (p. 1154-1155)
Detailed Analysis:
The paper describes a two-stage selection process: first, three intact writing classes (out of five) taught by the same instructor were randomly selected to participate; second, the 150 individual student participants pooled from these three classes "were randomly assigned" to the ChatGPT, AWE, and control conditions. The wording used ("150 participants ... were randomly assigned") explicitly frames randomization as occurring at the participant level, not at the level of the whole class. Although each condition happened to end up with n=50, matching the typical size of one class, the paper never states that each intact class was assigned wholesale to a single condition; instead it explicitly frames assignment as being of individual participants. No further quote clarifies that condition assignment preserved intact class membership. This creates a realistic risk that students in different conditions attended the same physical class sessions (e.g. during the shared peer feedback activities described in week two), which is exactly the contamination risk the class-level criterion is designed to guard against. The intervention here (comparing AI feedback tools during out-of-class, individual writing revision) is not a one-to-one personal tutoring intervention, so the tutoring exception does not apply.
Because the explicit textual evidence describes random assignment at the student level rather than at the class (or school) level, and no exception applies, criterion C is not met.
-
E
Exam-based Assessment
- Writing performance was scored using the rating criteria of CET-4, a nationally recognized standardized English test, based on prompts adapted from that same test.
- "The writing for both the pre-test and post-test of all three groups was scored by the first researcher and an experienced rater of CET-4, using the well-established rating criteria of CET-4 (See Appendix C), which is one of the most widely recognized English tests throughout the country" (p. 1158)
Relevant Quotes:
1) "To ensure the reliability of the writing prompts, the genre of all four writing prompts employed in this study was argumentative and was adapted from a pool of writing sections in College English Test Band Four (See Appendix B)." (p. 1156)
2) "The writing for both the pre-test and post-test of all three groups was scored by the first researcher and an experienced rater of CET-4, using the well-established rating criteria of CET-4 (See Appendix C), which is one of the most widely recognized English tests throughout the country (Huang, 2023)." (p. 1158)
3) "Pearson's correlation coefficient for the two raters' independent ratings was 0.80 (p<0.01), demonstrating satisfactory inter-rater reliability." (p. 1158)
Detailed Analysis:
The primary academic outcome, writing performance, was assessed using writing prompts directly adapted from the College English Test Band Four (CET-4), and scoring was conducted using the CET-4's own established rating criteria (reproduced in Appendix C), rather than a custom rubric invented solely for this study. CET-4 is explicitly described as "one of the most widely recognized English tests throughout the country," i.e. a genuine, widely used national standardized examination, not a bespoke instrument created to flatter the intervention. Inter-rater reliability was also reported, adding to the assessment's credibility. The secondary outcome (ideal L2 writing self) is a validated psychological questionnaire rather than an exam, but it is not the primary academic-outcome measure; the writing score is.
Because the writing outcome was scored using CET-4's established, nationally recognized rating criteria rather than a custom-built measure, criterion E is met.
-
T
Term Duration
- Outcomes were measured at the end of an eleven-week intervention, which is shorter than the roughly one-term (3-4 month) interval required, with no further follow-up.
- "The experiment started in week one and lasted for eleven weeks." ... "In week eleven, the post-test of the same questionnaire and a writing prompt were conducted among all three groups in class." (p. 1156-1157)
Relevant Quotes:
1) "The whole procedure for this study is presented in Figure 1. The experiment started in week one and lasted for eleven weeks." (p. 1156)
2) "Prior to the intervention, in week one, the same pre-test writing prompt and questionnaires were administered to all three groups to investigate their baseline writing performance and ideal L2 writing self." (p. 1156)
3) "In week eleven, the post-test of the same questionnaire and a writing prompt were conducted among all three groups in class." (p. 1157)
4) "In addition, the duration of our study was relatively short, and to gain deeper insights into sustained outcomes, future research may consider conducting experiments with longer intervention periods." (p. 1168, Limitations)
Detailed Analysis:
The intervention and its primary outcome measurement both occurred within an eleven-week window: baseline data were collected in week one and the post-test was administered in week eleven, i.e. roughly ten weeks after the intervention began. Eleven weeks (about 2.5 months) is shorter than the "approximately 3-4 months" that the ERCT standard treats as a typical academic term, and the paper does not provide a quote defining a shorter local "term" that this interval would satisfy. There is no subsequent or delayed measurement beyond the week-eleven post-test; the authors themselves acknowledge the short duration as a limitation of the study and call for "longer intervention periods" in future work, confirming that no term-length follow-up tracking was conducted.
Because the interval from intervention start to outcome measurement (about ten to eleven weeks) falls short of a full academic term and no further follow-up was conducted, criterion T is not met.
-
D
Documented Control Group
- The control group's size, demographics, baseline scores, and business-as-usual treatment (peer feedback only, no AI/AWE tool) are clearly documented.
- "31 (mean age = 20.07, SD=0.83; 45.16% female) for the control group." (p. 1155)
Relevant Quotes:
1) "The final dataset consisted of data from 35 participants (mean age = 20.29, SD=0.71; 48.57% female) for the ChatGPT group, 42 participants (mean age = 20.00, SD=0.88; 59.52% female) for the AWE groups, and 31 (mean age = 20.07, SD=0.83; 45.16% female) for the control group." (p. 1155)
2) "Table 1. One-way ANOVA results of pre-tests ... Control 31 3.84 0.74 0.13 [ideal L2 writing self] ... Control 31 8.26 1.39 0.25 [writing score]" (Table 1, p. 1159-1160)
3) "meanwhile, the AWE group could use AWE systems to assist with their revision, while the ChatGPT group could revise with assistance offered by ChatGPT." (p. 1157, implying the control group revised using peer feedback alone, without an AI/AWE tool)
4) "Although the mean score and the standard deviation for each group across the three variables were slightly different, ANOVA analysis showed that there was no statistically significant difference among the three groups in both variables." (p. 1159)
Detailed Analysis:
The paper documents the control group's sample size (31 participants after attrition/screening), demographics (mean age, gender split), and baseline scores for both dependent variables (writing performance and ideal L2 writing self), with statistical confirmation that the three groups did not differ significantly at baseline. The nature of the control condition (business-as-usual: standard course instruction and peer feedback, but no AWE or ChatGPT tool access) is also described, by contrast with the explicit descriptions of what the AWE and ChatGPT groups received.
Because the control group's size, demographics, baseline performance, and treatment are clearly documented, criterion D is met.
-
Level 2 Criteria
-
S
School-level RCT
- Randomization/selection occurred among individual students from three classes at a single university, not among multiple schools.
- "This research took place in a College English Writing course at a public university in Eastern China." (p. 1154)
Relevant Quotes:
1) "This research took place in a College English Writing course at a public university in Eastern China." (p. 1154)
2) "three out of five writing classes, all taught by the same instructor, were randomly selected for this study." (p. 1154)
3) "150 participants from three writing classes agreed to take part in the research and were randomly assigned to a ChatGPT group (n=50), an AWE group (n=50), and a control group (n=50)." (p. 1154-1155)
Detailed Analysis:
The entire study was conducted within a single public university in Eastern China, using three classes taught by one instructor. There is no mention of multiple schools or institutions being involved, nor of randomization occurring at the school/institution level. This is a single-site, within-institution study with (at best) class-level selection and individual-level condition assignment.
Because the study did not involve randomization across multiple schools/institutions, criterion S is not met.
-
I
Independent Conduct
- The study's own authors (the "first researcher") scored the writing outcomes themselves; no independent third party conducted data collection or analysis.
- "The writing for both the pre-test and post-test of all three groups was scored by the first researcher and an experienced rater of CET-4" (p. 1158)
Relevant Quotes:
1) "The writing for both the pre-test and post-test of all three groups was scored by the first researcher and an experienced rater of CET-4, using the well-established rating criteria of CET-4" (p. 1158)
2) "Prior to the experiment, we conducted a briefing session with the instructor about our research purpose, research design, and research procedure, and the instructor agreed to participate in our experiment." (p. 1154)
3) "Two researchers coded the data independently using open coding and focused coding in the first two rounds. Then two researchers discussed together the disagreements in codes ... and resolved the disagreements." (p. 1159)
Detailed Analysis:
Data collection, scoring, and analysis were conducted by the study's own research team: the "first researcher" (one of the paper's authors) personally scored the writing outcome measure alongside one external CET-4 rater, and the same two researchers who designed the study also carried out the qualitative coding. There is no statement of an independent, third-party organization conducting data collection or analysis free of involvement in the study's design. The class instructor cooperated with the research team but was not described as an independent evaluator of outcomes.
Because the authors who designed the study were also centrally involved in scoring and analyzing the outcome data, with no independent evaluator described, criterion I is not met.
-
Y
Year Duration
- Since criterion T (Term Duration) was not met, and the study's total duration of eleven weeks is far short of an academic year, criterion Y cannot be met.
- "The experiment started in week one and lasted for eleven weeks." (p. 1156)
Relevant Quotes:
1) "The experiment started in week one and lasted for eleven weeks." (p. 1156)
2) "In week eleven, the post-test of the same questionnaire and a writing prompt were conducted among all three groups in class." (p. 1157)
Detailed Analysis:
Per the ERCT standard's specific rule, if criterion T (Term Duration) is not met, criterion Y (Year Duration) is automatically not met. Independently, the study's total span of eleven weeks (roughly 2.5 months) is far shorter than the 75% of an academic year (roughly 9-10 months) required by criterion Y, with no extended follow-up reported.
Criterion Y is not met both because criterion T was not met and because the study duration is far shorter than a year.
-
B
Balanced Control Group
- Access to the AWE/ChatGPT tools is the explicit treatment variable being tested against a business-as-usual (peer feedback only) control, so the resource asymmetry is integral to the study design.
- "meanwhile, the AWE group could use AWE systems to assist with their revision, while the ChatGPT group could revise with assistance offered by ChatGPT." (p. 1157)
Relevant Quotes:
1) "In week two, as part of the course curriculum, students in all groups were asked to form a group of two or three and give peer feedback on each other's first writing in class." (p. 1156, applying equally to all three groups)
2) "For the AWE and ChatGPT groups, a brief training session was also given in week two; students were shown how to submit their writing and receive feedback using Youdao in the AWE group ... and how to use ChatGPT ... in the ChatGPT group." (p. 1156)
3) "After class, participants of all groups were asked to revise their first writing based on the peer feedback they received; meanwhile, the AWE group could use AWE systems to assist with their revision, while the ChatGPT group could revise with assistance offered by ChatGPT." (p. 1156-1157)
4) "The same procedure was repeated for prompt two and prompt three; the only difference was that the instructor assessment was given on writing of prompt two for all three groups as required in the course curriculum." (p. 1157)
Detailed Analysis:
Applying the ERCT criterion B decision tree: extra resources are present, since the ChatGPT and AWE groups received access to AI-based feedback tools (plus a brief training session on how to use them) that the control group did not receive. However, these additional resources are precisely the treatment variable under investigation: the entire purpose of the study, as stated in the research questions, is to compare the effects of ChatGPT feedback versus AWE feedback versus no automated feedback tool (the control, business-as-usual condition) on writing performance and the ideal L2 writing self. The control group was not deprived of some incidental extra resource; rather, it represents the standard, business-as-usual condition (regular course instruction and in-class peer feedback, which all three groups equally received) against which the explicit treatment (tool access) is compared. The brief tool-training sessions given to the AWE/ChatGPT groups are a necessary component of delivering that treatment, not a separate confounding resource. All three groups were otherwise treated identically (same writing prompts, same in-class peer feedback activity, same instructor assessment on prompt two, same pre/post-test schedule).
Because the additional resource (AI/AWE tool access) is the explicit, primary treatment variable being tested against a business-as-usual control, and all other course activities were held constant across groups, criterion B is met.
-
Level 3 Criteria
-
R
Reproduced
- No evidence of independent replication was found; the study is newly published (online Feb 2025) and internet searches found no other research team reporting a reproduction of this specific three-arm ChatGPT/AWE/ control comparison.
Detailed Analysis:
The paper does not reference any prior or concurrent independent replication of this specific study design (a three-arm comparison of ChatGPT feedback, AWE feedback, and a no-tool control on writing performance and ideal L2 writing self). The article was published online on 03 Feb 2025 and appears in the 2026 volume (39:4, 1148-1175), making independent peer-reviewed replication of this exact study design unlikely to exist yet.
Internet searches (web search plus citation-tracking for Shi, Chai, Zhou & Aubrey 2025/2026) turned up citing and related items such as a "Best Evidence in Brief" summary of this same paper (cuspbeb.com), a related but distinct study on ChatGPT as an AWE tool (published in Education and Information Technologies, 2025), and a three-level meta-analysis of AI-assisted feedback on L2 writing (ScienceDirect, 2026) that cites this and other primary studies. None of these is an independent replication of this specific three-arm comparison by a different research team; they are secondary discussions, adjacent studies with different designs, or reviews/meta-analyses that aggregate many primary studies rather than reproduce this one. The broader literature on ChatGPT and AWE effects on writing includes many related but non-replicating studies (e.g. Song & Song, 2023; Niloy et al., 2023; Zou & Huang, 2023), none of which is described as an independent replication of this specific three-arm comparison and its findings.
Because no independent replication of this specific study was found after an internet search, criterion R is not met.
-
A
All-subject Exams
- Only English writing performance was assessed; no other core subjects were measured.
- "the ChatGPT group scored significantly higher than the AWE group and the control group in post-writing performance as measured by their writing score" (Abstract)
Relevant Quotes:
1) "the ChatGPT group scored significantly higher than the AWE group and the control group in post-writing performance as measured by their writing score" (Abstract)
2) "The independent variable was the type of intervention ... and the dependent variables were the corresponding post-test data including students' writing scores and their responses to the questionnaire of the ideal L2 writing self." (p. 1159)
Detailed Analysis:
The study's only academic-outcome measure is English writing performance (via CET-4-based scoring), scored against the pre-defined dependent variables of writing score and the ideal L2 writing self questionnaire (a psychological/motivational measure, not an exam). No other core school subjects (e.g. mathematics, science) were assessed, and the paper offers no explicit justification of a highly-specialized, upper-secondary/vocational rationale that would invoke the stated exception for narrowly-focused interventions.
Because only a single subject (English writing) was assessed via a standardized measure, with no other core subjects tested and no exception invoked, criterion A is not met.
-
G
Graduation Tracking
- Since criterion Y (Year Duration) was not met, and there was no tracking of participants beyond the eleven-week post-test, criterion G cannot be met.
- "the duration of our study was relatively short, and to gain deeper insights into sustained outcomes, future research may consider conducting experiments with longer intervention periods." (p. 1168)
Relevant Quotes:
1) "In week eleven, the post-test of the same questionnaire and a writing prompt were conducted among all three groups in class." (p. 1157)
2) "In addition, the duration of our study was relatively short, and to gain deeper insights into sustained outcomes, future research may consider conducting experiments with longer intervention periods." (p. 1168)
Detailed Analysis:
Per the ERCT standard's specific rule, since criterion Y (Year Duration) was not met, criterion G (Graduation Tracking) is automatically not met. Independently, the paper's data collection ended with the week-eleven post-test and reflection papers; the authors explicitly list the short study duration as a limitation and call for future research with longer intervention periods, with no mention of any plan or attempt to track participants through to course, degree, or program graduation.
An internet search for follow-up publications by the same author team (Shi, Chai, Zhou, Aubrey) tracking this same cohort of 150 students toward graduation found none; the only related items identified were a third-party summary of this same paper and unrelated studies by other author teams on ChatGPT/AWE writing feedback.
Because there is no follow-up tracking toward graduation (and criterion Y was not met), criterion G is not met.
-
P
Pre-Registered
- No pre-registration of the study protocol, hypotheses, or analysis plan is mentioned anywhere in the paper, and no registry entry was found via internet search.
Detailed Analysis:
A review of the methods, procedures, and disclosure sections of the paper finds no reference to a registry platform (e.g. OSF, AsPredicted, ClinicalTrials.gov, or an equivalent), no registration ID, and no registration date. The "Disclosure statement" section only states competing interests information ("The authors report there are no competing interests to declare."), with no mention of pre-registration.
An internet search for a pre-registration record associated with this study (by title, authors, and topic) did not locate any entry on common registries (OSF, AsPredicted, ClinicalTrials.gov, or similar trial registries).
Because no evidence of pre-registration is present anywhere in the paper or in an internet search, criterion P is not met.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.