Abstract
This study examines the impact of AI-assisted writing feedback on the writing skills of secondary-level EFL students. A sample of 60 Turkish high school students was divided into an experimental group receiving feedback from an AI writing assistant and a control group receiving traditional teacher feedback. Over an 8 week period, students wrote multiple essays; pre-test and post-test writing assessments were administered. Results indicated that the experimental group showed significantly greater improvement in overall writing performance compared to the control group. The AI-assisted feedback group exhibited notable gains in grammar, vocabulary, and coherence, outperforming their peers with teacher-only feedback. These findings suggest that AI-driven feedback can effectively supplement writing instruction, enhancing writing proficiency in an EFL context. The study highlights the pedagogical potential of integrating AI tools into writing curricula to improve student outcomes.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was done at the individual student level within a single school, not at the class or school level, and the classroom-based feedback intervention does not qualify for the one-to-one tutoring exception.
- "The students were randomly assigned into two groups of equal size: an experimental group (n = 30) and a control group (n = 30)." (p. 3)
Relevant Quotes:
1) "The students were randomly assigned into two groups of equal size: an experimental group (n = 30) and a control group (n = 30)." (p. 3)
2) "We employed a quasi-experimental pre-test/post-test control group design." (p. 3)
3) "Aside from the feedback mechanism, both groups followed the same writing curriculum and were taught by the same instructor to control for instructional differences." (pp. 3-4)
4) "An independent-samples t-test confirmed no statistically significant difference between the groups at pre-test, t(58) = 0.45, p = 0.654, indicating successful random assignment." (p. 7)
Detailed Analysis:
The paper describes random assignment of 60 individual students within one public high school in Türkiye into experimental and control groups (quotes 1 and 4), although the authors confusingly also label the design "quasi-experimental" (quote 2). Taking the randomisation statements at face value, the unit of randomisation is the individual student, not intact classes or schools. Both groups were taught by the same instructor (quote 3), which is precisely the contamination scenario the C criterion is designed to prevent: the teacher delivering feedback to both conditions could transfer intervention practices to control students. The exception for personal teaching/tutoring does not apply: the intervention is a classroom-based writing programme in which drafts were written during regular class time in a shared computer-lab/classroom environment; it is not framed by the authors as a one-to-one tutoring intervention, and no quote in the paper claims the tutoring exception.
Criterion C is not met because students were randomised individually within a single school rather than by class or school, with no applicable tutoring exception.
-
E
Exam-based Assessment
- Outcomes were measured with researcher-administered essay tests scored on a rubric adapted from the Cambridge B2 scale, not with a widely recognised standardised exam.
- "To evaluate the students' writing, we used an analytic scoring rubric adapted from the Cambridge English B2 Writing Assessment Scale..." (p. 5)
Relevant Quotes:
1) "The pre-test consisted of a timed persuasive essay (approximately 300 words) on a common prompt appropriate for their level (e.g., arguing a position on a school-related issue)." (p. 4)
2) "The pre-test and post-test essays were the primary instruments for measuring writing performance. To evaluate the students' writing, we used an analytic scoring rubric adapted from the Cambridge English B2 Writing Assessment Scale, covering four dimensions: Content..., Organization..., Language Use..., and Mechanics..." (p. 5)
3) "The post-test prompt was a new topic of similar difficulty and genre to the pre-test (for example, another persuasive essay on a different issue)." (p. 5)
4) "Two experienced EFL instructors, who were not otherwise involved in the study, scored the essays independently." (p. 5)
Detailed Analysis:
The E criterion requires a standard, widely recognised standardised exam, not an assessment created for the study. Here the outcome instruments were essay prompts devised by the researchers for the study ("a common prompt appropriate for their level", "another persuasive essay on a different issue"), administered in class and scored with a rubric that was only "adapted from" the Cambridge English B2 Writing Assessment Scale. Adapting a rubric from a recognised scale does not make the assessment itself a standardised exam: no named national or international standardised test (e.g., Cambridge B2 First itself, TOEFL, a national exam) was actually administered. The instrument is therefore a custom, study-specific assessment, even though the scoring procedure was careful (independent raters, high inter-rater reliability).
Criterion E is not met because outcomes were measured with custom essay tests scored via an adapted rubric rather than a recognised standardised exam.
-
T
Term Duration
- The interval from intervention start to outcome measurement was only 8 weeks, which is shorter than one full academic term (approximately 3-4 months).
- "Both groups completed a writing pre-test at the beginning of the study and a writing post-test at the end of the 8-week intervention period." (p. 3)
Relevant Quotes:
1) "Both groups completed a writing pre-test at the beginning of the study and a writing post-test at the end of the 8-week intervention period." (p. 3)
2) "Following the pre-test, the intervention was implemented over 8 weeks during regular class time." (p. 4)
3) "At the end of Week 8, all participants completed a writing post-test." (p. 5)
4) "Second, the duration of the intervention was relatively short (8 weeks). We were able to observe short-term improvements in writing skills, but it remains unclear whether these gains would be sustained over a longer period." (p. 13)
Detailed Analysis:
The T criterion requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. Here the intervention began in Week 2 (after a Week 1 pre-test) and the post-test was administered at the end of Week 8, so the start-to-measurement interval is at most 8 weeks (roughly 2 months). The authors themselves acknowledge the duration was "relatively short (8 weeks)" and that only short-term improvements were observed. There is no delayed follow-up measurement extending the tracking window to a full term.
Criterion T is not met because outcomes were measured only 8 weeks after intervention start, short of a full academic term.
-
D
Documented Control Group
- The control group is clearly documented with its size (n = 30), comparable demographics and baseline proficiency, baseline test scores, and a description of the teacher-only feedback it received.
- "Both groups had similar demographic profiles and baseline English proficiency." (p. 3)
Relevant Quotes:
1) "The students were randomly assigned into two groups of equal size: an experimental group (n = 30) and a control group (n = 30). Both groups had similar demographic profiles and baseline English proficiency." (p. 3)
2) "The participants were 60 secondary school students (aged 15-17) enrolled in an EFL program at a public high school in Türkiye. All participants were intermediate-level English learners, as determined by their school's placement test and their previous semester English grades." (p. 3)
3) "Control Group (Teacher-Only Feedback): Students in the control group followed a traditional writing process. After writing the first draft of each weekly essay, they submitted it to the teacher for feedback." (p. 4)
4) "the mean pre-test score was 25.8 (SD = 5.4) for the experimental group and 25.2 (SD = 5.1) for the control group. An independent-samples t-test confirmed no statistically significant difference between the groups at pre-test, t(58) = 0.45, p = 0.654" (p. 7)
5) "Aside from the feedback mechanism, both groups followed the same writing curriculum and were taught by the same instructor to control for instructional differences." (pp. 3-4)
Detailed Analysis:
The D criterion requires clear documentation of the control group's composition, baseline performance, and the conditions it experienced. The paper specifies the control group's size (n = 30), age range, proficiency level (intermediate, per school placement test and grades), demographic comparability, baseline writing scores with a statistical baseline-equivalence test (Table 1 and quote 4), and a detailed description of exactly what the control group received (weekly essays with delayed written teacher feedback, same curriculum and same instructor as the experimental group). This is sufficient to judge comparability and to confirm the control condition was business-as-usual teacher feedback.
Criterion D is met because the control group's size, baseline characteristics, baseline scores, and treatment conditions are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- The study took place in a single school with randomisation at the individual student level, so there was no school-level randomisation.
- "The participants were 60 secondary school students (aged 15-17) enrolled in an EFL program at a public high school in Türkiye." (p. 3)
Relevant Quotes:
1) "The participants were 60 secondary school students (aged 15-17) enrolled in an EFL program at a public high school in Türkiye." (p. 3)
2) "The students were randomly assigned into two groups of equal size: an experimental group (n = 30) and a control group (n = 30)." (p. 3)
3) "First, the sample size (N = 60) and single-school context limit the generalizability of the findings. All participants were from one public high school in Turkey..." (p. 13)
Detailed Analysis:
The S criterion requires randomisation among schools (or equivalent implementing units). This study was conducted entirely within one public high school, and assignment to conditions was made at the individual student level. The authors explicitly acknowledge the "single-school context" as a limitation. With only one institution involved, school-level randomisation was impossible by design.
Criterion S is not met because randomisation occurred at the student level within a single school, not among schools.
-
I
Independent Conduct
- The authors themselves designed the study, conducted the data collection, and performed the analyses, with no independent third-party evaluation team beyond the two external essay raters.
- "M.E: was responsible for conceptualizing the study, designing the research methodology, conducting the data collection, and performing the statistical analyses." (p. 14)
Relevant Quotes:
1) "M.E: was responsible for conceptualizing the study, designing the research methodology, conducting the data collection, and performing the statistical analyses. A.N.D: contributed to the literature review, the interpretation of findings, and the refinement of the discussion section." (p. 14)
2) "Two experienced EFL instructors, who were not otherwise involved in the study, scored the essays independently." (p. 5)
3) "This study did not receive any external funding. All expenses related to the research were internally covered by the corresponding author." (p. 15)
Detailed Analysis:
The I criterion requires that the study be conducted independently of those who designed the intervention. Here the same author (M.E.) conceptualised the study, designed the methodology (including the AI-assisted feedback intervention protocol), collected the data, and ran the statistical analyses. There is no external evaluation agency, no independent oversight body, and no statement of third-party conduct of the trial. The only element of independence is that two EFL instructors not otherwise involved in the study scored the essays blind of the study operations (quote 2); while this reduces scoring bias, it falls well short of the criterion's requirement that the evaluation itself (data collection, analysis, conclusions) be independent of the intervention designers. The intervention tool (Grammarly) is commercial and not authored by the researchers, but the study implementation and analysis were entirely in the designers' hands.
Criterion I is not met because the intervention designers themselves conducted the data collection and analysis with no independent evaluation team.
-
Y
Year Duration
- The study lasted only 8 weeks, far short of 75% of an academic year, and since criterion T is not met this criterion automatically fails as well.
- "Second, the duration of the intervention was relatively short (8 weeks)." (p. 13)
Relevant Quotes:
1) "Following the pre-test, the intervention was implemented over 8 weeks during regular class time." (p. 4)
2) "At the end of Week 8, all participants completed a writing post-test." (p. 5)
3) "Second, the duration of the intervention was relatively short (8 weeks)." (p. 13)
4) "Longitudinal research would be particularly valuable: for instance, following students for a semester or an entire school year to examine whether the improvements from AI feedback persist, increase, or plateau over time." (p. 14)
Detailed Analysis:
The Y criterion requires outcome measurement at least 75% of an academic year (roughly 7-9 months) after intervention start. The tracking interval here is 8 weeks from intervention start to post-test, about 20-25% of a typical academic year. The authors themselves recommend future studies follow students "for a semester or an entire school year", confirming that this study did not. Additionally, the weaker T (Term Duration) criterion is not met, and per the instructions, if T is not met then Y cannot be met.
Criterion Y is not met because the 8-week duration falls far short of 75% of an academic year.
-
B
Balanced Control Group
- The extra input (immediate AI-generated feedback via Grammarly, within a hybrid AI-plus-teacher feedback package) is the explicit treatment variable tested against business-as-usual teacher feedback, with in-class time, curriculum, and teacher feedback time deliberately equalised across groups.
- "We attempted to balance the total amount of feedback each group received by ensuring the teacher spent roughly equivalent time on each control student's essay as the experimental students spent interacting with AI feedback." (pp. 4-5)
Relevant Quotes:
1) "The independent variable was the type of feedback on writing (AI-assisted feedback vs. traditional teacher feedback), and the dependent variable was the students' writing performance..." (p. 3)
2) "Throughout the 8-week period, both groups wrote on similar topics and had the same amount of in-class writing time for each task." (p. 5)
3) "We attempted to balance the total amount of feedback each group received by ensuring the teacher spent roughly equivalent time on each control student's essay as the experimental students spent interacting with AI feedback. However, the depth and immediacy of feedback naturally differed between the conditions." (pp. 4-5)
4) "In addition to the automated feedback, the teacher provided a brief follow-up focusing on higher-order concerns (e.g., content relevance, argument strength) after students revised their drafts. This hybrid approach ensured that AI feedback was used for detailed language-level corrections while the teacher addressed content and organization as needed." (p. 4)
5) "Additionally, the experimental group effectively received more total feedback (AI feedback on every draft plus some teacher comments) than the control group (teacher comments only). This difference reflects a realistic scenario in which AI enables a greater volume of feedback without extra teacher effort, but it also means that part of the performance gain might be attributed to feedback quantity." (p. 14)
6) "To ensure fairness, the control-group students were given access to the AI writing tool (and a brief tutorial on its use) after the conclusion of the experiment..." (p. 6)
Detailed Analysis:
Applying the decision tree: extra resources ARE present in the experimental condition — access to Grammarly (Education version) accounts in the computer lab, immediate AI feedback on every draft, plus brief teacher follow-up comments, so the experimental group received a greater total volume of feedback (quote 5). The next question is whether these extra resources are the treatment variable itself. They are: the stated independent variable is "the type of feedback on writing (AI-assisted feedback vs. traditional teacher feedback)" (quote 1), and the AI tool, its immediacy, and the hybrid AI-plus-teacher structure are explicitly framed as the intervention package being tested against a business-as-usual teacher-feedback control (quotes 1, 4). The paper also documents genuine balancing efforts for non-integral inputs: identical curriculum, same instructor, same in-class writing time (quote 2), and teacher time per control essay matched to experimental students' AI interaction time (quote 3). The authors candidly acknowledge that feedback quantity remains a partial confound (quote 5), but under the ERCT standard this greater feedback volume is integral to what AI-assisted feedback is (a tool that "enables a greater volume of feedback without extra teacher effort") rather than a separable add-on. This parallels the specification's DPL example, where the device and its use were integral to the intervention tested against teaching-as-usual, and criterion B was met.
Criterion B is met because the additional AI feedback resource is the explicit, integral treatment variable tested against a business-as-usual control, with other educational inputs (time, curriculum, teacher) deliberately balanced.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study exists; earlier similar trials by other teams (e.g., Wei et al. 2023) predate this study and are prior related work, not replications of it.
- "For instance, Wei et al. [22], in a randomized controlled trial, demonstrated that EFL students who received feedback via Grammarly showed significantly higher post-test scores across multiple writing sub-domains..." (p. 2)
Relevant Quotes:
1) "For instance, Wei et al. [22], in a randomized controlled trial, demonstrated that EFL students who received feedback via Grammarly showed significantly higher post-test scores across multiple writing sub-domains (e.g., grammar, mechanics, organization) than peers who received only teacher feedback." (p. 2)
2) "These findings are echoed by Liu et al. [11], whose quasi-experimental study revealed that students receiving AI-supported instruction not only outperformed the control group in holistic writing scores but also exhibited higher motivation and engagement in writing tasks." (p. 2)
3) "Studies with larger and more diverse samples are needed to see if our results hold true across different educational contexts and student populations." (p. 14)
Detailed Analysis:
The R criterion requires that this study be independently replicated by a different team, in a different context, in a peer-reviewed journal. The paper was accepted in October 2025 and is very recent; it cites earlier related studies of Grammarly/AI feedback by other teams (Wei et al. 2023 in Chinese EFL learners; Liu et al. 2021), but these were published before the present study and are prior literature that the present study builds on — the present study is closer to being a replication of them than the reverse. Following the standard's guidance (the PAX-GBG example), the existence of other trials of similar interventions in other contexts does not constitute independent reproduction of this particular study of Turkish secondary students.
Internet research update (2026-07-27): a web search identified a preprint by the same authors, "AI-Assisted Writing Feedback for Enhancing Secondary Students' Writing Skills: An Experimental Study" (Ekizoğlu & Demir, Research Square, preprint doi 10.21203/rs.3.rs-6430737/v1), which is an earlier version of this same study rather than an independent replication by a different team. No subsequent publication by a different research team attempting to replicate this specific Turkish secondary-school Grammarly-vs-teacher-feedback trial was located in the sources checked (Google-style web search, Springer/Discover Education citing articles, Research Square, ResearchGate). The authors themselves call for future studies to test whether the results hold (quote 3), and no independent replication of this specific study was found.
Criterion R is not met because no independent replication of this specific study exists; earlier similar trials are prior work and the only closely related item found (a preprint by the same authors) is not an independent replication.
-
A
All-subject Exams
- Only EFL writing was assessed, no other school subjects were measured, and since criterion E is not met this criterion automatically fails as well.
- "The pre-test and post-test essays were the primary instruments for measuring writing performance." (p. 5)
Relevant Quotes:
1) "The pre-test and post-test essays were the primary instruments for measuring writing performance." (p. 5)
2) "the dependent variable was the students' writing performance, measured by analytic scores on writing assessments." (p. 3)
3) "Third, our measurement of improvement was focused on writing test scores and rubric sub-components, which capture tangible gains in written product quality." (p. 13)
Detailed Analysis:
The A criterion requires standardised exam-based assessment across all main subjects taught at the students' level. This study measured only EFL writing performance; no outcomes in mathematics, science, Turkish language, or any other core high school subject were assessed. There is no stated specialised- vocational rationale that would justify the single-subject exception for these general secondary students. Additionally, criterion E (a prerequisite for A) is not met because the writing assessments themselves were custom, non-standardised instruments, so A fails on that ground as well.
Criterion A is not met because only a single custom writing assessment was used and no other main subjects were measured.
-
G
Graduation Tracking
- Measurement ended at the Week 8 post-test with no follow-up toward graduation, and since criterion Y is not met this criterion automatically fails as well.
- "We were able to observe short-term improvements in writing skills, but it remains unclear whether these gains would be sustained over a longer period." (p. 13)
Relevant Quotes:
1) "At the end of Week 8, all participants completed a writing post-test." (p. 5)
2) "We were able to observe short-term improvements in writing skills, but it remains unclear whether these gains would be sustained over a longer period. It is possible that continued use of AI feedback is needed to maintain the advantage... questions that only a longitudinal study could answer." (p. 13)
3) "Longitudinal research would be particularly valuable: for instance, following students for a semester or an entire school year..." (p. 14)
Detailed Analysis:
The G criterion requires tracking participants until graduation from their educational stage. The final measurement here was the post-test at the end of Week 8; the participants were aged 15-17 and were not followed to the end of high school. The authors explicitly frame longitudinal follow-up as unanswered future work, confirming no graduation tracking occurred.
Internet research update (2026-07-27): a web search for follow-up or longitudinal publications by Ekizoğlu and/or Demir tracking this same cohort (e.g., toward high-school graduation) did not locate any such paper; only the original Discover Education article and its earlier Research Square preprint of the same 8-week study were found. No follow-up publication tracking this cohort was identified. Furthermore, since criterion Y is not met, this criterion is automatically not met per the instructions.
Criterion G is not met because data collection stopped at the 8-week post-test with no tracking to graduation, and no follow-up publication tracking this cohort was found.
-
P
Pre-Registered
- The paper reports ethics approval but contains no mention of any pre-registered protocol on a trial registry before data collection began, and no registration for this study was found via internet search.
Relevant Quotes:
1) "The research was approved by the relevant institutional review board (IRB) and the school administration prior to the start of the study." (p. 6)
2) "Ethical approval for this study was obtained from the Scientific Research and Publication Ethics Committee of Çağ University (Approval Number: 2025/06; Reference Code: E-81570533-604.01.01-2500004112)." (p. 15)
Detailed Analysis:
The P criterion requires the full study protocol (hypotheses, methods, planned analyses) to be pre-registered on a public registry before data collection began. The paper mentions only IRB/ethics-committee approval, which is not pre-registration. There is no reference anywhere in the paper to a registry such as ClinicalTrials.gov, ISRCTN, OSF, or AEA RCT Registry, no registration ID, and no registration date. Ethics approval alone does not satisfy the transparency purpose of pre-registration (preventing selective reporting).
Internet research update (2026-07-27): a web search for a pre-registration record by Ekizoğlu and/or Demir for this study (checking for mentions of OSF, AsPredicted, ClinicalTrials.gov, or ISRCTN in connection with the paper) found no evidence of any pre-registration; only the published article, its earlier preprint, and citing works were found.
Criterion P is not met because no pre-registration of the study protocol is mentioned in the paper or found via internet search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.