ChatGPT as an automated writing evaluation tool: how students perceive it and how it affects their writing

Linqian Ding, Di Zou, Lucas Kohnke

Published:
ERCT Check Date:
DOI: 10.1007/s10639-025-13775-3
  • L2 languages
  • higher education
  • China
  • Asia
  • EdTech platform
  • formative assessment
0
  • C

    The paper is an explicitly quasi-experimental study with only two intact classes, one flipped into each condition, which is not a properly implemented class-level RCT.

    "The study employed a quasi-experimental design with two intact groups (Groups A and B)." (p. 26783)

  • E

    Writing outcomes were scored with a researcher-adopted six-aspect rubric rather than any standardised, widely recognised exam.

    "To assess writing quality, we employed a comprehensive writing rubric (see Appendix 1). It assessed six aspects of academic writing: mechanics, tone, grammar, organisation, content, and APA formatting." (p. 26784)

  • T

    The whole intervention and outcome measurement occurred within roughly two weeks (Weeks 5-7 of the course), far short of one academic term.

    "In Week 6, the students in both groups completed their Literature review writing task in class within 60 min." (p. 26783)

  • D

    There is no genuine control group, and group-specific documentation is limited to baseline draft scores, with demographics reported only for the pooled sample.

    "One class was randomly assigned to be Group A (n=30) and the other to be Group B (n=47)." (p. 26782)

  • S

    Assignment involved two intact classes within a single university; no schools or institutional units were randomised.

    "All of the participants were enrolled in two classes of Academic English Writing taught by the same instructor." (p. 26782)

  • I

    The authors designed, delivered, rated, and analysed the intervention themselves, with a member of the research team even providing the Group B feedback; no independent evaluator was involved.

    "In contrast, the ChatGPT feedback from Group B was collected and forwarded to a human rater – a third researcher involved in the study" (p. 26784)

  • Y

    Since the tracking interval was only about two weeks, the 75%-of-an-academic-year requirement is necessarily unmet.

    "the experiment compared the first and final drafts of a single assignment, which allowed the identification of improvements but could not provide evidence of long-term writing development." (p. 26794)

  • B

    The extra instructor feedback given to Group B is the explicit treatment variable being tested (combined feedback versus ChatGPT-only feedback), so the resource difference between arms is integral to the design.

    "It also compares the effectiveness of ChatGPT with ChatGPT plus teacher feedback in improving student writing." (p. 26778)

  • R

    No independent replication of this specific study exists; the paper cites related but distinct ChatGPT-feedback studies, not reproductions of this trial, and a targeted internet search found no later replication.

  • A

    Only academic English writing was assessed, with a custom rubric, so neither the all-subject coverage nor the standardised-exam prerequisite (criterion E) is satisfied.

    "It assessed six aspects of academic writing: mechanics, tone, grammar, organisation, content, and APA formatting." (p. 26784)

  • G

    Measurement ended with the revised draft and a Week 7 questionnaire; no participant was tracked to graduation, no follow-up publication was found, and criterion Y is unmet.

    "the experiment compared the first and final drafts of a single assignment, which allowed the identification of improvements but could not provide evidence of long-term writing development." (p. 26794)

  • P

    The paper contains no mention of any pre-registered protocol, registry, or registration date, and no matching registration was found in registries such as OSF via internet search.

Abstract

Generative artificial intelligence (GAI) language models, exemplified by ChatGPT, are significantly helping language learners in writing practice by providing immediate formative feedback. This study investigates whether ChatGPT's writing feedback influences graduate students' academic writing abilities and compares it with combined instructor and ChatGPT feedback. The students receiving only ChatGPT feedback (Group A) showed significant improvements in mechanics, tone, grammar, APA formatting, and overall writing quality, with lesser gains in organisation and no significant change in content. In contrast, the students who received combined feedback (Group B) exhibited significant enhancements in all of the parameters that were evaluated, including the areas in which Group A lagged. However, the only significant between-group difference was in grammar, although there was a marginal difference in organisation that suggested a trend towards greater improvement in Group B. The participants expressed positive attitudes about using ChatGPT as an automated writing evaluation tool, especially for correcting grammar and enriching vocabulary.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • The paper is an explicitly quasi-experimental study with only two intact classes, one flipped into each condition, which is not a properly implemented class-level RCT.
      • "The study employed a quasi-experimental design with two intact groups (Groups A and B)." (p. 26783)
      • Relevant Quotes: 1) "The study employed a quasi-experimental design with two intact groups (Groups A and B)." (p. 26783) 2) "All of the participants were enrolled in two classes of Academic English Writing taught by the same instructor. ... One class was randomly assigned to be Group A (n=30) and the other to be Group B (n=47)." (p. 26782) 3) "This study employed a quasi-experimental design and a mixed-methods approach to investigate the impact of ChatGPT as an AWE tool on students' writing quality." (p. 26794) Detailed Analysis: Criterion C requires a Randomised Controlled Trial conducted at the class level, with randomisation clearly described and properly implemented. The authors themselves label the study "quasi-experimental" with "two intact groups". Although one of the two intact classes was "randomly assigned" to each condition, randomising only two pre-existing clusters is effectively a single coin flip: it cannot balance class-level confounds (different class compositions, n=30 vs n=47), and the paper provides no description of a randomisation procedure (no method, no allocation concealment). Moreover, neither arm is a control in the RCT sense: both groups received feedback interventions (ChatGPT only vs ChatGPT plus instructor feedback), so there is no randomised untreated comparison. The intervention is also not one-to-one tutoring, so the personal-teaching exception does not apply. Because the paper explicitly self-identifies as quasi-experimental rather than an RCT and the two-cluster assignment cannot constitute properly implemented class-level randomisation, the criterion fails. Criterion C is not met because the study is a quasi-experimental design with two intact classes rather than a properly randomised class-level controlled trial.
    • E

      Exam-based Assessment

      • Writing outcomes were scored with a researcher-adopted six-aspect rubric rather than any standardised, widely recognised exam.
      • "To assess writing quality, we employed a comprehensive writing rubric (see Appendix 1). It assessed six aspects of academic writing: mechanics, tone, grammar, organisation, content, and APA formatting." (p. 26784)
      • Relevant Quotes: 1) "To assess writing quality, we employed a comprehensive writing rubric (see Appendix 1). It assessed six aspects of academic writing: mechanics, tone, grammar, organisation, content, and APA formatting. These features of writing are commonly considered in university-level writing assessments (e.g. Brigham Young University, 2024; Utica University, 2024). Each aspect was scored on a five-point scale, with 1 being the lowest and 5 being the highest." (p. 26784) 2) "the first drafts and revisions from both groups were collected and evaluated blindly using the same writing rubric. Both researchers independently assessed six randomly selected compositions to ensure reliability and achieved strong inter-rater agreement." (p. 26785) 3) "all participants were native Mandarin speakers from mainland China and had taken the International English Language Testing System (IELTS) examination as their application for graduate study." (p. 26782) Detailed Analysis: Criterion E requires that outcomes be measured with a standard, widely recognised standardised exam rather than an instrument assembled for the study. The outcome here is a score on a single course assignment (a literature review draft and its revision) rated by the researchers using a rubric adapted from university writing-rubric examples. This is a researcher-scored custom rubric applied to a course task, not a recognised standardised examination such as IELTS or TOEFL. IELTS is mentioned only as background to describe participants' entry proficiency, not as an outcome measure. Because the outcome instrument is study-specific and researcher-scored, it does not satisfy the standardisation requirement. Criterion E is not met because outcomes were measured with a custom rubric on a single course assignment rather than a standardised exam.
    • T

      Term Duration

      • The whole intervention and outcome measurement occurred within roughly two weeks (Weeks 5-7 of the course), far short of one academic term.
      • "In Week 6, the students in both groups completed their Literature review writing task in class within 60 min." (p. 26783)
      • Relevant Quotes: 1) "In Week 5 of the course, both groups attended a 60-minute training session to learn how to use ChatGPT to obtain writing feedback." (p. 26783) 2) "In Week 6, the students in both groups completed their Literature review writing task in class within 60 min. ... During the second half of the Week 6 class, the students used ChatGPT independently to obtain feedback on their drafts." (p. 26783) 3) "In Week 7, students in both groups completed a perception questionnaire designed to gather their views on the feedback they received." (p. 26784) 4) "The study had two major Limitations. First, the experiment compared the first and final drafts of a single assignment, which allowed the identification of improvements but could not provide evidence of long-term writing development." (p. 26794) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Here the intervention (ChatGPT feedback on one literature-review draft) began in Week 6, the revised final drafts were submitted immediately after the feedback cycle, and perceptions were collected in Week 7. The entire start-to-measurement interval is therefore on the order of one to two weeks. The authors themselves acknowledge that the single-assignment pre/post comparison "could not provide evidence of long-term writing development" and recommend a future longitudinal design. There is no term-long follow-up of outcomes. Criterion T is not met because outcomes were measured within about one to two weeks of the intervention start, far less than one academic term.
    • D

      Documented Control Group

      • There is no genuine control group, and group-specific documentation is limited to baseline draft scores, with demographics reported only for the pooled sample.
      • "One class was randomly assigned to be Group A (n=30) and the other to be Group B (n=47)." (p. 26782)
      • Relevant Quotes: 1) "This study included 77 graduate students at a university in Hong Kong. They were all enrolled in the Faculty of Education, with a specific focus on English Language Education. The majority of participants were female (n=68) ... Most of them aged between 18 and 34 (n=71), and six aged 35 or above." (p. 26782) 2) "One class was randomly assigned to be Group A (n=30) and the other to be Group B (n=47)." (p. 26782) 3) "The results of the independent t-test showed that there were no significant differences between the first drafts of the students in Groups A and B in mechanics (p =.52), tone (p =.58), grammar (p =.63), organisation (p =.25), APA formatting (p =.179), and overall (p =.25) scores." (p. 26785) 4) "the students who received combined feedback (Group B) exhibited significant enhancements in all of the parameters that were evaluated" (p. 26777) Detailed Analysis: Criterion D requires a well-documented control group, including demographics, baseline performance, and the treatment it received. This study has no control condition at all: both arms received an intervention (Group A ChatGPT feedback only; Group B ChatGPT plus refined instructor feedback), and no group experienced business-as-usual writing instruction without AI feedback. Even treating Group A as the comparison arm, its documentation is partial: baseline draft scores are reported per group (Tables 1-2) and baseline equivalence was tested, but demographic characteristics (gender, age, language background, IELTS range) are reported only for the pooled 77 participants, never separately for each group. Without a documented untreated control condition, the criterion's core requirement of detailed control group data for proper comparison is not satisfied. Criterion D is not met because the study contains no untreated control group and per-group documentation is limited to baseline writing scores.
  • Level 2 Criteria

    • S

      School-level RCT

      • Assignment involved two intact classes within a single university; no schools or institutional units were randomised.
      • "All of the participants were enrolled in two classes of Academic English Writing taught by the same instructor." (p. 26782)
      • Relevant Quotes: 1) "This study included 77 graduate students at a university in Hong Kong." (p. 26782) 2) "All of the participants were enrolled in two classes of Academic English Writing taught by the same instructor. ... One class was randomly assigned to be Group A (n=30) and the other to be Group B (n=47)." (p. 26782) Detailed Analysis: Criterion S requires randomisation at the level of schools or comparable implementing institutions. This study took place within one university, in two classes taught by the same instructor, and the unit of assignment was the intact class (and only two of them). No schools, campuses, or institutional sites were randomised, and the design is self-described as quasi-experimental rather than a cluster-randomised trial. Criterion S is not met because assignment occurred between two classes inside a single institution, not at the school level.
    • I

      Independent Conduct

      • The authors designed, delivered, rated, and analysed the intervention themselves, with a member of the research team even providing the Group B feedback; no independent evaluator was involved.
      • "In contrast, the ChatGPT feedback from Group B was collected and forwarded to a human rater – a third researcher involved in the study" (p. 26784)
      • Relevant Quotes: 1) "In contrast, the ChatGPT feedback from Group B was collected and forwarded to a human rater – a third researcher involved in the study – who was a native English speaker with over eight years of experience in teaching academic writing." (p. 26784) 2) "They uploaded their drafts to the university platform, which the instructor and the research team could access." (p. 26783) 3) "Both researchers independently assessed six randomly selected compositions to ensure reliability and achieved strong inter-rater agreement." (p. 26785) 4) "We modified the original scale by replacing references to 'Pigai' with 'ChatGPT' and recalculated Cronbach's alpha using our data set." (p. 26784) Detailed Analysis: Criterion I requires that the study be conducted independently of those who designed the intervention, e.g., by an external evaluation team. Here the same research team designed the intervention procedure, adapted the prompt and the questionnaire, delivered the feedback conditions (with "a third researcher involved in the study" personally producing the refined Group B feedback), scored the writing with their own rubric, and performed all analyses. Although draft scoring was described as blind, there is no external agency, third-party evaluator, or independent oversight mentioned anywhere in the paper, and the Conflict of interest statement ("None", p. 26796) does not establish independence of conduct. Criterion I is not met because the intervention was designed, implemented, rated, and analysed entirely by the same research team with no independent evaluator.
    • Y

      Year Duration

      • Since the tracking interval was only about two weeks, the 75%-of-an-academic-year requirement is necessarily unmet.
      • "the experiment compared the first and final drafts of a single assignment, which allowed the identification of improvements but could not provide evidence of long-term writing development." (p. 26794)
      • Relevant Quotes: 1) "Each course ran for 16 weeks, with one 120-minute session per week." (p. 26782) 2) "In Week 5 of the course, both groups attended a 60-minute training session ... In Week 6, the students in both groups completed their Literature review writing task ... In Week 7, students in both groups completed a perception questionnaire." (pp. 26783-26784) 3) "Future research could adopt a longitudinal design to examine whether ChatGPT usage can contribute to long-term writing progress. For example, researchers might follow students over a 16-week course and track their writing across multiple assignments" (p. 26794) Detailed Analysis: Criterion Y requires outcome measurement at least 75% of an academic year after intervention start, and per the prompt rules it automatically fails when criterion T fails. The intervention-to-measurement window here spans roughly Weeks 5/6 to Week 7 of a one-semester course, i.e., about one to two weeks. The authors explicitly flag the absence of longitudinal tracking as a limitation and propose even a 16-week follow-up only as future work. This is far below 75% of an academic year. Criterion Y is not met because outcomes were measured about two weeks after the intervention began, and criterion T is already unmet.
    • B

      Balanced Control Group

      • The extra instructor feedback given to Group B is the explicit treatment variable being tested (combined feedback versus ChatGPT-only feedback), so the resource difference between arms is integral to the design.
      • "It also compares the effectiveness of ChatGPT with ChatGPT plus teacher feedback in improving student writing." (p. 26778)
      • Relevant Quotes: 1) "It also compares the effectiveness of ChatGPT with ChatGPT plus teacher feedback in improving student writing." (p. 26778) 2) "How does combining ChatGPT and teacher feedback affect student writing compared to ChatGPT feedback alone?" (p. 26782) 3) "In Week 5 of the course, both groups attended a 60-minute training session to learn how to use ChatGPT to obtain writing feedback." (p. 26783) 4) "After receiving ChatGPT feedback, the Group A participants were instructed to revise their drafts based solely on this feedback and then upload their final versions." (p. 26784) 5) "In contrast, the ChatGPT feedback from Group B was collected and forwarded to a human rater ... He then refined it with a focus on structural coherence and topic relevance. In addition, he provided targeted revision strategies and personalised explanations tailored to each student's writing topic." (p. 26784) Detailed Analysis: Following the decision tree: (1) Do the arms differ in resources? Yes - Group B received an added layer of expert human feedback (a rater with 8+ years of academic-writing teaching experience refining and personalising the ChatGPT feedback), which is extra adult input relative to Group A. (2) Is this extra resource the treatment variable? Yes - the study's second research question is explicitly whether "combining ChatGPT and teacher feedback" outperforms "ChatGPT feedback alone", so the additional human feedback is the integral treatment contrast being tested, not an accidental confound. Otherwise the two arms were treated identically: both classes had the same instructor, the same 16-week course, the same 60-minute ChatGPT training, the same in-class writing task, and the same standardised prompt. The added human feedback for Group B is therefore clearly framed as the primary treatment variable, which the standard's exception explicitly permits against a ChatGPT-only comparison condition. Criterion B is met because the only resource difference between arms (added instructor feedback) is the explicit treatment variable under test, with all other inputs matched.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study exists; the paper cites related but distinct ChatGPT-feedback studies, not reproductions of this trial, and a targeted internet search found no later replication.
      • Relevant Quotes: 1) "Zhang et al. (2025) compared improvement in different aspects of writing between students who received only ChatGPT feedback with those who received hybrid feedback (ChatGPT and teacher feedback)." (p. 26781) 2) "This aligned with findings from Yang et al. (2025) and Zhang et al. (2025), who also reported that ChatGPT helped students improve surface-level features but had limited effect on higher-order writing skills." (p. 26791) Detailed Analysis: Criterion R requires independent replication of the specific study by a different team, published in a peer-reviewed journal. This paper was published online in October 2025 and reports a novel comparison in its particular context (Hong Kong graduate students, a single literature-review assignment, a six-aspect rubric). The related studies it cites (e.g., Zhang et al., 2025; Yang et al., 2025; Shi et al., 2025) are prior, methodologically different investigations of ChatGPT feedback, not replications of this study's design and findings, and one cited team overlaps in affiliation context. Internet searches (Google Scholar-style web search, Semantic Scholar, and citation lookups) as of July 2026 found no later paper that reproduces this specific quasi-experiment (two-intact-class comparison of ChatGPT-only versus ChatGPT-plus-teacher feedback on the same six-aspect rubric) by an independent team; the article is too recent (published October 2025) for an independent replication to plausibly have appeared and been peer-reviewed yet. Criterion R is not met because no independent peer-reviewed replication of this specific study has been identified.
    • A

      All-subject Exams

      • Only academic English writing was assessed, with a custom rubric, so neither the all-subject coverage nor the standardised-exam prerequisite (criterion E) is satisfied.
      • "It assessed six aspects of academic writing: mechanics, tone, grammar, organisation, content, and APA formatting." (p. 26784)
      • Relevant Quotes: 1) "To assess writing quality, we employed a comprehensive writing rubric (see Appendix 1). It assessed six aspects of academic writing: mechanics, tone, grammar, organisation, content, and APA formatting." (p. 26784) 2) "This study investigates whether ChatGPT's writing feedback influences graduate students' academic writing abilities and compares it with combined instructor and ChatGPT feedback." (p. 26777) Detailed Analysis: Criterion A requires standardised exam-based assessment across all main subjects, and by rule it automatically fails when criterion E fails. Criterion E is not met here (custom rubric on one assignment), so A cannot be met. Substantively, the study measures only one domain - academic English writing on a single literature-review task - and no other subjects or courses in the graduate programme were assessed. While a writing-focused scope is understandable for a graduate EAP course, the paper offers no justification framed as a specialised-intervention exception, and in any case the E prerequisite already fails. Criterion A is not met because only one subject (academic writing) was assessed and the exam-based prerequisite (criterion E) is unmet.
    • G

      Graduation Tracking

      • Measurement ended with the revised draft and a Week 7 questionnaire; no participant was tracked to graduation, no follow-up publication was found, and criterion Y is unmet.
      • "the experiment compared the first and final drafts of a single assignment, which allowed the identification of improvements but could not provide evidence of long-term writing development." (p. 26794)
      • Relevant Quotes: 1) "The classes were part of the second semester of their one-year graduate program." (p. 26782) 2) "In Week 7, students in both groups completed a perception questionnaire designed to gather their views on the feedback they received." (p. 26784) 3) "the experiment compared the first and final drafts of a single assignment, which allowed the identification of improvements but could not provide evidence of long-term writing development. Future research could adopt a longitudinal design" (p. 26794) Detailed Analysis: Criterion G requires tracking participants until graduation from their educational stage, and by rule it fails when criterion Y fails. Data collection ended with the revised final drafts and the Week 7 perception questionnaire, mid way through the second semester of a one-year programme. There is no mention of following students to programme completion, and the authors explicitly state the design cannot speak to long-term development, proposing longitudinal follow-up only as future research. An internet search for later publications by Ding, Zou, or Kohnke tracking this same cohort of 77 graduate students (e.g., through subsequent semesters or to programme completion) found no such follow-up paper as of July 2026. Criterion G is not met because tracking stopped immediately after the single-assignment revision cycle, well before graduation, no follow-up publication was located, and criterion Y is already unmet.
    • P

      Pre-Registered

      • The paper contains no mention of any pre-registered protocol, registry, or registration date, and no matching registration was found in registries such as OSF via internet search.
      • Relevant Quotes: 1) "Ethics declarations I confirm that all the research meets ethical guidelines and adheres to the legal requirements of the study country." (p. 26796) 2) "Data Availability The datasets generated during and/or analysed during the current study are available from the corresponding author on reasonable request." (p. 26796) Detailed Analysis: Criterion P requires that the full study protocol be registered on a public registry before data collection began, with a verifiable link and date. The paper's declarations cover only ethics compliance, funding, data availability, and conflicts of interest. There is no mention of ClinicalTrials.gov, OSF, AsPredicted, a trial registry ID, or any pre-specified analysis plan anywhere in the manuscript. An internet search for a pre-registration record tied to this DOI or to this study's title/authors turned up no matching entry in OSF or other registries. With no registration statement at all and none found externally, timing relative to data collection cannot be verified and the criterion fails. Criterion P is not met because the paper contains no reference to any pre-registration of hypotheses, methods, or analyses, and none was found via external search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.