Abstract
Nowadays, Chinese EFL learners are increasingly using automated writing evaluation (AWE) to provide feedback on their writing. However, AWE is relatively inadequate for providing elaborate feedback on the global aspects of writing that require human judgment. Thus, how to combine AWE with human feedback is a significant issue worth exploring. Regarding the effects of AWE+human feedback, cohesion and coherence are rarely studied. Thus, this study aimed to address that gap. This research employed a quasi-experimental study with a quantitative method to address AWE+human feedback. It also aimed to compare the effects of traditional feedback with AWE+human feedback modes, that is, teacher-only, AWE+teacher (AT), and AWE+peer+teacher (APT), on EFL learners' writing quality in terms of holistic score, cohesion, and coherence. A total of 90 EFL learners from three intact classes of English major were randomly assigned to the control and experimental groups. The control group received only teacher feedback. The experimental groups received AT feedback and APT feedback in their writing process, respectively. Two instruments, iWrite and Coh-Metrix, were used to collect the data. Results showed that all feedback types improved holistic scores, coherence, and cohesion, with the APT model producing the most significant improvements across these dimensions. The APT group demonstrated particularly high holistic scores. It enhanced sentence-level coherence and lexical cohesion, suggesting that integrating AWE, peer, and teacher feedback provides a comprehensive and effective approach to developing writing proficiency.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was carried out at the class level (three intact classes randomly assigned to conditions), meeting the class-level RCT requirement.
- "Three intact classes were allocated randomly to serve as experimental and control groups." (p. 902)
Relevant Quotes:
1) "A total of 90 EFL learners from three intact classes of English major were randomly assigned to the control and experimental groups." (Abstract, p. 899)
2) "This study employed a quasi-experimental research design. Three intact classes were randomly assigned as one control group and two experimental groups." (p. 902)
3) "The participants of this study were 90 sophomores from three intact classes of Business English Major. Three intact classes were allocated randomly to serve as experimental and control groups." (p. 902)
4) "The same instructor will teach three classes, which was necessary to prevent the instructor's experience and manner from functioning as an extraneous variable affecting the study's validity." (p. 902)
Detailed Analysis:
The unit of assignment was the intact class, not individual students within a class: three whole classes (n=30 each) were randomly allocated to the teacher-only (TF), AWE+teacher (AT), and AWE+peer+teacher (APT) conditions. This satisfies the ERCT requirement that entire classes, rather than students within a single classroom, be randomised, which prevents cross-group contamination of feedback conditions. The randomisation procedure itself is only briefly described ("allocated randomly", with no detail on the mechanism), the design is self-described as "quasi-experimental" because intact classes were used, and with only one class per condition class-level effects are confounded with treatment; these are weaknesses in rigour, but the stated unit of randomisation is the class, which is what criterion C checks. Baseline equivalence was verified via ANOVA on pretest scores (F=.587, p>.05).
Criterion C is met because entire intact classes, not individual students within a class, were randomly assigned to the feedback conditions.
-
E
Exam-based Assessment
- Outcomes were single researcher-administered essay tasks scored by iWrite and Coh-Metrix rather than an official standardised exam administration, so the exam-based assessment requirement is not satisfied.
- "These tests were selected from the College English Test Band 6 (CET6) writing exercises, and they have the same difficulty level and similar topics." (p. 904)
Relevant Quotes:
1) "Two instruments, iWrite and Coh-Metrix, were used to collect the data." (Abstract, p. 899)
2) "The AWE tool used in the present study was iWrite 3.0. iWrite (https://iwrite.unipus.cn/) is an AWE system developed explicitly for Chinese English learners by the Foreign Language Teaching and Research Press and the National Research Center of Foreign Language Education in 2015." (p. 902)
3) "Writing pretests and posttests were used to examine the effects of the treatments of the study. These tests were selected from the College English Test Band 6 (CET6) writing exercises, and they have the same difficulty level and similar topics. The pretest topic was 'My View on the Use of PowerPoint (PPT),' while the posttest topic was 'On a Harmonious Dormitory Life'." (p. 904)
4) "For cohesion and coherence, six indices of data were generated from the pretest and posttest texts of three groups using Coh-Metrix." (p. 904)
Detailed Analysis:
Criterion E requires that outcomes be measured with a widely recognised standardised exam, not a researcher- arranged assessment. Here the outcome measures were (a) holistic scores generated by iWrite 3.0, the very AWE system that formed part of the intervention, and (b) six Coh-Metrix computational indices of cohesion and coherence. Although the two writing prompts were "selected from the College English Test Band 6 (CET6) writing exercises", the students did not sit an actual administration of CET6 or any other official standardised examination; the essays were single researcher-administered tasks scored by automated research tools rather than by the standardised exam's official scoring protocol. Using prompts borrowed from a standardised test does not make the assessment a standardised exam, and scoring outcomes with the intervention tool itself (iWrite) further departs from an independent standardised assessment.
Criterion E is not met because outcomes were measured with researcher-administered writing tasks scored by iWrite and Coh-Metrix, not by an actual widely recognised standardised examination.
-
T
Term Duration
- The intervention ran from week 3 to week 14 with the posttest in week 15 of a semester-long study, so outcomes were tracked for roughly one full academic term after the intervention began.
- "This study lasted 15 weeks. ... Weeks 3-14. These weeks were treatment weeks. ... In week 15, the posttest was implemented." (p. 904)
Relevant Quotes:
1) "This study lasted 15 weeks. Students had a writing lesson once a week." (p. 904)
2) "Weeks 3-14. These weeks were treatment weeks." (p. 904)
3) "In week 15, the posttest was implemented to obtain data on the effectiveness of the treatments throughout the program." (p. 904)
4) "Data for the current study were collected in the third semester when the students began to learn basic writing skills and to improve English writing progressively." (p. 902)
5) "First, it was conducted on a relatively small sample of EFL learners from a single local university in China over a single academic term." (p. 907)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. The study ran across a full 15-week semester: pretest in week 1, treatment across weeks 3-14 (four three-week writing cycles), and posttest in week 15. From the start of the intervention (week 3) to outcome measurement (week 15) is about 12 weeks, roughly three months, and the authors themselves describe the study as covering "a single academic term" within the third semester. The interval from intervention start to measurement therefore spans approximately one full academic term.
Criterion T is met because the treatment began in week 3 and outcomes were measured in week 15 of a semester-long (15-week) study, an interval of about one academic term.
-
D
Documented Control Group
- The control class (n=30) is documented with baseline pretest scores, population details, and a clear description of the teacher-only feedback condition it received.
- "The control group received only teacher feedback. Students submitted their work through the iWrite system without seeing any automated feedback." (p. 904)
Relevant Quotes:
1) "The participants of this study were 90 sophomores from three intact classes of Business English Major." (p. 902)
2) "The control group received only teacher (TF) feedback, while the first experimental group was given AWE+Teacher (AT) feedback, and the second experimental group was given AWE+Peer+Teacher (APT) feedback." (p. 902)
3) "The results of the ANOVA test (see Table1) showed that there was no significant difference (F=.587 p>.05) among the pretest scores of the experimental groups and the control group." (p. 902)
4) "The control group received only teacher feedback. Students submitted their work through the iWrite system without seeing any automated feedback. After the teacher had reviewed the work, students revised it based on the teacher's comments and resubmitted it for final grading." (p. 904)
5) "In other words, these three classes received the same teaching and writing tasks except for different feedback modes." (p. 904)
6) Table 3 documents "Group1(n=30)" (teacher-only feedback) alongside the two treatment groups, and Table 4 reports the TF group's pretest mean holistic score of 63.127. (pp. 904-905)
Detailed Analysis:
Criterion D requires clear documentation of the control group's composition, baseline performance, and the conditions it experienced. The paper documents the control group's size (one intact class, n=30), its population (sophomore Business English majors at a local Chinese university in their third semester), its baseline writing performance (pretest means reported in Tables 1 and 4, with an ANOVA confirming no significant baseline differences), and precisely what it received during the study (the same lectures, teacher, and writing tasks as the treatment groups, with teacher-only feedback and no visible automated feedback). Demographic detail beyond major, year of study, and institution is limited, but the documentation of the control condition, its size, baseline scores, and treatment is sufficient for proper comparison.
Criterion D is met because the control group's size, baseline performance, and exact conditions are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation was at the class level within one university department, not at the school or institution level.
- "Three intact classes were allocated randomly to serve as experimental and control groups." (p. 902)
Relevant Quotes:
1) "This study was conducted in a Foreign Language Department at a local Chinese university." (p. 902)
2) "Three intact classes were allocated randomly to serve as experimental and control groups." (p. 902)
3) "First, it was conducted on a relatively small sample of EFL learners from a single local university in China over a single academic term." (p. 907)
Detailed Analysis:
Criterion S requires randomisation at the level of whole schools or comparable institutional units. This study took place within a single department of one Chinese university, and the unit of randomisation was the intact class (three classes within that department). No schools, campuses, or other institutional units were randomised, and the authors explicitly note the single-university setting as a limitation.
Criterion S is not met because randomisation occurred among three classes within a single university department, not among schools or institutions.
-
I
Independent Conduct
- The intervention designers themselves conducted, collected, and analysed the study with no independent evaluators or third-party oversight.
- "To comprehensively analyze the writing quality regarding coherence and cohesion, the author will employ six indices (see Table 2) automatically generated from Coh-Metrix 3.0." (p. 903)
Relevant Quotes:
1) "Yan Zhang is a Ph.D. candidate at Universiti Putra Malaysia and a lecturer in the Department of Foreign Languages at Yuncheng University, China." (p. 909)
2) "The same instructor will teach three classes, which was necessary to prevent the instructor's experience and manner from functioning as an extraneous variable affecting the study's validity." (p. 902)
3) "This study used quantitative data analysis with SPSS 25 and Coh-Metrix to address RQ1..." (p. 904)
4) "To comprehensively analyze the writing quality regarding coherence and cohesion, the author will employ six indices (see Table 2) automatically generated from Coh-Metrix 3.0." (p. 903)
Detailed Analysis:
Criterion I requires the study to be conducted independently of those who designed the intervention. Here the authors themselves designed the feedback intervention (the AT and APT feedback procedures), organised its delivery in classes at the first author's own institution (Yuncheng University), and performed the data collection and analysis ("the author will employ six indices"). There is no external evaluation team, no third-party oversight, and no independence statement anywhere in the paper. The only mitigating design element is that iWrite itself was developed by an unrelated organisation, but the study design, conduct, and analysis were all done by the intervention designers.
Criterion I is not met because the same research team designed the feedback intervention and also conducted, collected, and analysed the study without any independent third-party involvement.
-
Y
Year Duration
- The whole study, from pretest to posttest, spanned only one 15-week semester, far short of 75% of an academic year.
- "This study lasted 15 weeks." (p. 904)
Relevant Quotes:
1) "This study lasted 15 weeks." (p. 904)
2) "First, it was conducted on a relatively small sample of EFL learners from a single local university in China over a single academic term. Extending the study across a longer period and involving a larger sample size could yield more robust results and further significant indicators." (p. 907)
Detailed Analysis:
Criterion Y requires that outcomes be measured at least 75% of an academic year (roughly 9-10 months, so at least ~7 months) after the intervention begins. This study spanned a single 15-week semester, with the posttest at week 15, about 12-13 weeks after the intervention started. The authors themselves flag the single-term duration as a limitation and call for a longer study period. A 15-week window falls far short of 75% of an academic year.
Criterion Y is not met because the study covered only one 15-week semester, well below 75% of an academic year.
-
B
Balanced Control Group
- All groups had the same teacher, lectures, tasks, and iWrite training, and the extra AWE/peer feedback in the experimental groups was the explicit treatment variable being tested, so the design is balanced.
- "In other words, these three classes received the same teaching and writing tasks except for different feedback modes. Only feedback modes affected their final writing quality." (p. 904)
Relevant Quotes:
1) "It also aimed to compare the effects of traditional feedback with AWE+human feedback modes, that is, teacher-only, AWE+teacher (AT), and AWE+peer+teacher (APT), on EFL learners' writing quality in terms of holistic score, cohesion, and coherence." (Abstract, p. 899)
2) "During these treatment weeks, the teacher delivered her planned lectures to all classes during scheduled class hours. Then, after class, Writing Task 1 was released to students through iWrite, and different classes received different feedback modes. In other words, these three classes received the same teaching and writing tasks except for different feedback modes. Only feedback modes affected their final writing quality." (p. 904)
3) "All students, including the control group, were trained on how to use iWrite 3.0 and submit writing tasks. However, only experimental groups received automated feedback." (p. 904)
4) "The control group received only teacher feedback. Students submitted their work through the iWrite system without seeing any automated feedback." (p. 904)
5) "Experimental Group 1 combined AWE with teacher feedback. Students first submitted their work through the iWrite system, received automated feedback with scores and specific criteria ratings, and revised based on this feedback as many times as desired before resubmission." (p. 904)
6) "The same instructor will teach three classes..." (p. 902)
Detailed Analysis:
Following the criterion B decision tree: the intervention groups did receive additional inputs relative to the control - automated iWrite feedback (AT and APT) and peer feedback (APT), plus the extra revision cycles these enable. However, these additional feedback layers are precisely the treatment variable the study was designed to test: the research questions explicitly compare teacher-only versus AWE+teacher versus AWE+peer+teacher feedback modes, so the extra feedback is integral to the intervention rather than a separable confounding add-on. All other educational inputs were deliberately held constant: the same instructor taught all three classes, all classes received the same lectures during the same scheduled hours, all wrote the same four writing tasks on the same schedule, all submitted through iWrite, and all (including the control) received identical iWrite training. The control group also received genuine teacher feedback and revision opportunities, making it an active comparison rather than a no-treatment group. The additional automated and peer feedback (with associated extra revision activity) in the experimental groups is the explicit treatment contrast being tested against the teacher-feedback baseline.
Criterion B is met because instruction, tasks, teacher, and platform training were equalised across groups, and the additional AWE and peer feedback layers were the explicit treatment variable being tested rather than an unbalanced supplementary resource.
-
Level 3 Criteria
-
R
Reproduced
- The study is framed as novel gap-filling research, was published very recently (May 2025), and a July 2026 internet search found no independent peer-reviewed replication of this specific three-arm feedback design.
Relevant Quotes:
1) "Regarding the effects of AWE+human feedback, cohesion and coherence are rarely studied. Thus, this study aimed to address that gap." (Abstract, p. 899)
2) "These studies didn't explore the effect of AWE+peer+teacher on writing quality. Therefore, there is a need to continue to study the effect of this multi-feedback mode." (p. 901)
Detailed Analysis:
Criterion R requires that this specific study be independently replicated by a different research team in a different context and published in a peer-reviewed journal. The paper positions itself as filling a gap (the effect of AWE+peer+teacher feedback on cohesion and coherence had not been studied), so no prior replication exists by definition. The paper was published in May 2025. An internet search conducted in July 2026 for independent replications found related but distinct studies by other teams (e.g., other AWE or AWE+peer feedback trials with Chinese EFL learners such as Chen and Cui, 2022, and earlier feasibility work such as Huang and Zhang, 2014 and Chen and Gong, 2022 cited in the paper), but none of these is an independent replication of this specific three-arm TF/AT/APT design measuring holistic scores, cohesion, and coherence; they are predecessors or different intervention contrasts. Given the paper's recent publication date (May 2025), an independent peer-reviewed replication would not plausibly exist yet in any case. No independent replication of this study was found.
Criterion R is not met because no independent, peer- reviewed replication of this specific study by a different research team was found.
-
A
All-subject Exams
- Criterion E is unmet and outcomes covered only English writing quality, not all main subjects, so the all-subject exams requirement fails.
- "This research employed a quasi-experimental study with a quantitative method to address AWE+human feedback ... on EFL learners' writing quality in terms of holistic score, cohesion, and coherence." (Abstract, p. 899)
Relevant Quotes:
1) "This research employed a quasi-experimental study with a quantitative method to address AWE+human feedback ... on EFL learners' writing quality in terms of holistic score, cohesion, and coherence." (Abstract, p. 899)
2) "Second, while this research focused on holistic quality and coherence, it did not explore other dimensions of writing quality, such as accuracy, fluency, linguistic diversity, and syntactic complexity." (p. 907)
Detailed Analysis:
Criterion A requires standardised exam-based assessment across all main subjects taught at the education level, and explicitly requires criterion E as a prerequisite. Here criterion E is not met (no standardised exam was used), so criterion A automatically fails. Furthermore, the study measured only English argumentative writing quality; no other subjects in the students' degree programme were assessed, and no justification for a specialised- intervention exception is offered beyond the study's focus.
Criterion A is not met because criterion E fails and only English writing was assessed, with no other subjects measured.
-
G
Graduation Tracking
- Measurement ended at the week-15 posttest in the students' sophomore year with no follow-up to graduation, prerequisite criterion Y also fails, and no follow-up publications tracking this cohort were found online.
- "In week 15, the posttest was implemented to obtain data on the effectiveness of the treatments throughout the program." (p. 904)
Relevant Quotes:
1) "In week 15, the posttest was implemented to obtain data on the effectiveness of the treatments throughout the program." (p. 904)
2) "Extending the study across a longer period and involving a larger sample size could yield more robust results and further significant indicators." (p. 907)
Detailed Analysis:
Criterion G requires participants to be tracked until graduation from their educational stage, and per the instructions it cannot be met when criterion Y is not met, as is the case here. Measurement ended with the week-15 posttest during the participants' second (sophomore) year; there is no follow-up of any kind, let alone tracking to university graduation, and the authors' own limitations section calls for a longer study period. An internet search conducted in July 2026 for subsequent publications by the same author team (Zhang, Jalaluddin, Mamat, Zhao) tracking this cohort of sophomores toward graduation found no such follow-up papers; only the original May 2025 article was located.
Criterion G is not met because data collection stopped at the week-15 posttest with no tracking of participants to graduation, and no follow-up publications were found.
-
P
Pre-Registered
- The paper contains no reference to any pre-registration, registry ID, or published protocol, and none was found through an external registry search.
Relevant Quotes:
No quotes mentioning pre-registration, a trial registry, a registration ID, or a published protocol appear anywhere in the paper. The acknowledgements mention only a funding project: "We would like to extend our gratitude to the Shanxi Provincial Department of Education, China, for their funding and support through the project ... (SZH-230017), which has made this research possible." (p. 907)
Detailed Analysis:
Criterion P requires the full study protocol (hypotheses, methods, planned analyses) to be registered in a public registry before data collection began, with quoted evidence of the registration and its timing. The paper contains no reference to any registry (e.g., ClinicalTrials.gov, OSF, AsPredicted, ISRCTN), no registration number, and no statement that a protocol was published before the study. The only identifier given (SZH-230017) is a provincial funding project code, not a pre-registration. An internet search conducted in July 2026 for a pre-registered protocol by this author team on major registries (OSF Registries, AsPredicted, ClinicalTrials.gov, ISRCTN) found no matching entry.
Criterion P is not met because there is no mention of any pre-registered protocol or registry entry in the paper, and none was found through external search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.