Abstract
Effective feedback plays a critical role in enhancing the writing skills of English as a Foreign Language (EFL) learners. This study examines the comparative effectiveness of three feedback approaches--Teacher e-feedback, AI-based feedback, and a hybrid model--in enhancing the writing performance of Iranian intermediate-level EFL learners. A randomized controlled trial was conducted with 88 intermediate-level EFL learners, who were randomly assigned to one of three feedback groups: (1) teacher e-feedback, (2) AI-generated feedback using tools such as ChatGPT and Grammarly, and (3) a hybrid approach that combined both feedback types. Writing proficiency was assessed using IELTS writing tasks and the Oxford Placement Test before and after the intervention. Significant differences were observed between the groups, with the hybrid feedback group showing the most substantial improvements, particularly in task achievement, coherence, and grammatical accuracy. AI feedback was most effective in enhancing lexical resources. Qualitative reflections supported the quantitative findings, with participants in the hybrid group reporting increased confidence, reduced anxiety, and appreciation for the balanced, dual-source feedback. These results highlight the pedagogical potential of integrating human and AI feedback to enhance EFL writing instruction.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Intact classes were the unit of random assignment to the three feedback conditions, satisfying class-level randomisation.
- "For the experiment, the three selected classes were randomly assigned to one of the three feedback groups." (p. 4)
Relevant Quotes:
1) "A randomized controlled trial was conducted with 88 intermediate-level EFL learners, who were randomly assigned to one of three feedback groups: (1) teacher e-feedback, (2) AI-generated feedback using tools such as ChatGPT and Grammarly, and (3) a hybrid approach that combined both feedback types." (p. 1, Abstract)
2) "From the six IELTS writing classes offered at these two branches, three classes were randomly selected to participate in the study." (p. 4)
3) "For the experiment, the three selected classes were randomly assigned to one of the three feedback groups." (p. 4)
4) "This random assignment ensured that each group had a balanced number of participants, allowing for a fair comparison of the different feedback methods." (p. 5)
Detailed Analysis:
Although the abstract loosely says that learners "were randomly assigned," the Methods section clarifies the actual unit of randomisation: three intact IELTS writing classes were randomly selected from six, and then those whole classes were randomly assigned to the three feedback conditions. Students were not individually randomised within a shared classroom, so each condition was delivered to a separate intact class, preventing within-class contamination between treatment and control students. This satisfies the class-level requirement of criterion C. A caveat is that with only three classes (one per arm) the randomisation is minimal and class effects cannot be separated from treatment effects, but the criterion as defined checks the unit of randomisation, which here is the class.
Criterion C is met because entire classes, not individual students within one classroom, were randomly assigned to the feedback conditions.
-
E
Exam-based Assessment
- Outcomes were assessed with standard IELTS writing tasks scored via the official IELTS Writing Band Descriptors, a widely recognised standardised assessment.
- "The IELTS Writing Band Descriptors were used to assess the writing improvement of participants in this study, both in the pre-test and post-test." (p. 5)
Relevant Quotes:
1) "Writing proficiency was assessed using IELTS writing tasks and the Oxford Placement Test before and after the intervention." (p. 1, Abstract)
2) "In this study, all groups of L2 learners were asked to write a 200-word essay on a randomly selected topic from Cambridge Book 3 for Task 2, both before the intervention (pre-test) and after it (post-test)." (p. 5)
3) "The IELTS Writing Band Descriptors were used to assess the writing improvement of participants in this study, both in the pre-test and post-test." (p. 5)
4) "The homogeneity of the participants was assessed using the Oxford Placement Test (OPT) to ensure that there were no significant differences in English proficiency levels among the three groups at the beginning of the study." (p. 5)
5) "Both raters have over 5 years of experience teaching IELTS writing and were trained on the use of the IELTS Writing Band Descriptors for this study." (p. 6)
Detailed Analysis:
The outcome measure was not a custom test designed by the researchers. The writing tasks were taken from official Cambridge IELTS practice materials (Cambridge Book 3, Task 2), and performance was scored against the official IELTS Writing Band Descriptors, an internationally recognised, standardised assessment framework (Task Achievement, Coherence and Cohesion, Lexical Resource, Grammatical Range and Accuracy). Baseline proficiency was measured with the Oxford Placement Test, another widely recognised standardised instrument. Scoring was double-rated, blind, with weighted Cohen's kappa of 0.85, supporting reliability. Although the essays were administered by the researchers rather than in an official IELTS sitting (and were 200 words rather than the official 250-word Task 2 minimum), the tasks and scoring rubric are standard, widely recognised instruments rather than study-specific tests.
Criterion E is met because outcomes were measured with standard IELTS writing tasks scored using the official IELTS Writing Band Descriptors, not a custom-made assessment.
-
T
Term Duration
- The intervention ran for the language center's entire two-and-a-half-month semester with the post-test at its end, covering one full term in this context.
- "Over the two-and-a-half-month semester, the three groups attended regular IELTS writing classes, which were consistent with the other classes offered at the language center." (p. 10)
Relevant Quotes:
1) "Students in these classes were informed about the nature of the study at the beginning of their two-and-a-half-month semester." (p. 4)
2) "A week later, writing instructors administered the writing pre-test to all three groups." (p. 9)
3) "Over the two-and-a-half-month semester, the three groups attended regular IELTS writing classes, which were consistent with the other classes offered at the language center." (p. 10)
4) "Upon completion of the educational sessions, the writing post-test was administered to all three groups." (p. 10)
Detailed Analysis:
The intervention began at the start of the language center's semester and outcomes were measured at the end of that same semester, so the interval from intervention start to measurement is the full semester, about 2.5 months. The ERCT standard defines a term as "a semester or equivalent (approximately 3-4 months)" and instructs evaluators to check the paper for the local definition of a term when the academic calendar differs. In this private language center context, the full academic cycle is explicitly a "two-and-a-half-month semester," and the feedback intervention ran continuously for that entire semester with measurement at its end. The tracking therefore covers one complete term as defined in the study's own context, though it sits slightly below the typical 3-4 month range.
Criterion T is met because outcomes were measured a full semester (the term unit in this context, 2.5 months) after the intervention began.
-
D
Documented Control Group
- The comparison (teacher e-feedback) group's size, conditions, and baseline equivalence via OPT and writing pre-test are clearly documented.
- "The ANOVA test results indicated that there were differences in the mean scores of the Oxford Placement Test (OPT) among the three groups; however, these differences were not statistically significant (F = 0.105, p = 0.901)." (p. 5)
Relevant Quotes:
1) "The participants in this study were 88 adult intermediate students of ESL enrolled in IELTS writing courses at two branches of a single language center in Shiraz, Iran." (p. 4)
2) "The study included participants aged between 21 and 35 (M = 28, SD = 4.2) with an intermediate level of English proficiency, as assessed by a placement test conducted by the language center." (p. 4)
3) "...resulting in final class sizes ranging from 28 students (Group 3) to 30 students (Groups 1 and 2)." (p. 4)
4) "Group 1 (Teacher E-Feedback): Received electronic feedback from teachers." (p. 4)
5) "The ANOVA test results indicated that there were differences in the mean scores of the Oxford Placement Test (OPT) among the three groups; however, these differences were not statistically significant (F = 0.105, p = 0.901)." (p. 5)
6) "The pre-test consisted of a 200-word essay, which was collected and analyzed to gauge participants' initial writing abilities and provide a baseline for comparison." (p. 9)
Detailed Analysis:
This is a three-arm comparative trial with no untreated arm; the Teacher E-Feedback group (Group 1) functions as the comparison/control condition, representing the conventional feedback practice against which the AI and hybrid conditions are compared. The paper documents this group's size (n = 30), the exact treatment it received (video e-feedback via Snagit and Telegram following the researcher-provided guidelines in Table 3), participant demographics (adult learners aged 21-35, intermediate proficiency), and baseline equivalence via both the Oxford Placement Test homogeneity ANOVA and a writing pre-test used as ANCOVA covariate. All groups attended the same regular IELTS classes with the same instructor, so the comparison group's conditions are clearly specified. Demographics are reported for the sample as a whole rather than per group, which is a minor limitation, but baseline comparability is verified statistically.
Criterion D is met because the comparison group's size, composition, baseline performance, and treatment conditions are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation was at the class level within a single language center, not among schools or sites.
- "For the experiment, the three selected classes were randomly assigned to one of the three feedback groups." (p. 4)
Relevant Quotes:
1) "The participants in this study were 88 adult intermediate students of ESL enrolled in IELTS writing courses at two branches of a single language center in Shiraz, Iran." (p. 4)
2) "From the six IELTS writing classes offered at these two branches, three classes were randomly selected to participate in the study." (p. 4)
3) "For the experiment, the three selected classes were randomly assigned to one of the three feedback groups." (p. 4)
Detailed Analysis:
Criterion S requires randomisation among schools or equivalent educational institutions/sites. Here, all participants came from a single language center (two branches), and the unit of randomisation was the class, not the institution or branch. Three classes drawn from across the two branches were assigned to conditions; no schools, centers, or sites were randomised. There is no quote indicating institution-level assignment.
Criterion S is not met because randomisation occurred at the class level within one language center, not at the school or site level.
-
I
Independent Conduct
- The authors designed the feedback protocols and ran and analysed the trial themselves; only essay scoring was independent and blinded, so independent conduct is not established.
- "Table 3 outlines the feedback guidelines provided by the researchers to the writing instructor for offering supplementary e-feedback to students in Group 1." (p. 7)
Relevant Quotes:
1) "To ensure instructional consistency and control for variability in teaching style, a single, experienced writing instructor conducted the classes for all three experimental groups." (p. 7)
2) "Table 3 outlines the feedback guidelines provided by the researchers to the writing instructor for offering supplementary e-feedback to students in Group 1." (p. 7)
3) "The scoring was performed independently by two experienced EFL writing instructors who were not involved in the study's interventions." (p. 6)
4) "All pre-test and post-test essays were anonymized, assigned a random code, and mixed so that raters were blind to the participant's identity, experimental group assignment (Teacher, AI, or Hybrid), and the time of the test (pre-test vs. post-test)." (p. 6)
5) "Qualitative data from the reflection prompts were analyzed using thematic analysis by two independent raters..." (p. 10)
Detailed Analysis:
Criterion I requires that the study be conducted independently of the intervention designers, or at least that data collection and analysis be explicitly independent of them. Here the authors themselves designed the feedback protocols (the guidelines in Tables 3-5 were "provided by the researchers"), organised the interventions, and carried out the statistical analysis; one author's own prior e-feedback work is cited as the basis for the teacher e-feedback approach. The only independent elements are the two blinded essay raters and the second coder for the qualitative data, which mitigates scoring bias but does not amount to an external evaluation team independent of the intervention designers for data collection, analysis, and conclusions. No statement of third-party oversight or an external evaluator is present.
Criterion I is not met because the same research team designed the feedback interventions and conducted and analysed the study, with independence limited to blinded essay scoring.
-
Y
Year Duration
- The full interval from intervention start to final measurement was only about 2.5 months, well below 75% of an academic year.
- "Over the two-and-a-half-month semester, the three groups attended regular IELTS writing classes..." (p. 10)
Relevant Quotes:
1) "Students in these classes were informed about the nature of the study at the beginning of their two-and-a-half-month semester." (p. 4)
2) "Over the two-and-a-half-month semester, the three groups attended regular IELTS writing classes..." (p. 10)
3) "Upon completion of the educational sessions, the writing post-test was administered to all three groups." (p. 10)
Detailed Analysis:
Criterion Y requires that outcomes be measured at least 75% of a full academic year (roughly 9-10 months, so at least about 7 months) after the intervention begins. Here the entire interval from intervention start to the post-test was one two-and-a-half-month semester, with no delayed follow-up measurement afterwards. 2.5 months is roughly 25-28% of a standard academic year, far below the 75% threshold, even allowing for contextual variation in academic calendars.
Criterion Y is not met because outcomes were measured only 2.5 months after the intervention began, far short of 75% of an academic year.
-
B
Balanced Control Group
- All three active arms had matched class time, curriculum, and instructor, and the hybrid group's additional dual-source feedback is integral to the treatment being tested.
- "The instructor followed standardized curricula and feedback protocols for each group (as detailed in Tables 3, 4, and 5) to ensure the primary difference between the groups was the feedback modality itself." (p. 7)
Relevant Quotes:
1) "To ensure instructional consistency and control for variability in teaching style, a single, experienced writing instructor conducted the classes for all three experimental groups. The instructor followed standardized curricula and feedback protocols for each group (as detailed in Tables 3, 4, and 5) to ensure the primary difference between the groups was the feedback modality itself." (p. 7)
2) "Over the two-and-a-half-month semester, the three groups attended regular IELTS writing classes, which were consistent with the other classes offered at the language center. However, as detailed in the intervention section, the mode of feedback received by L2 learners varied throughout the semester." (p. 10)
3) "For Group 3, a hybrid method was utilized that integrated both AI feedback and instructor e-feedback." (p. 8)
4) "In addition to the AI feedback, the instructor provided supplementary e-feedback on the revised submissions." (p. 9)
5) "Throughout the semester, students' main source of feedback was AI tools, specifically ChatGPT-4 (OpenAI) and Grammarly (Premium version)." (pp. 7-8)
Detailed Analysis:
Applying the decision tree: all three arms were active conditions receiving the same class time, the same regular IELTS curriculum, and the same single instructor, so core educational inputs (time and instruction) were matched across groups. Each arm received a feedback treatment of comparable structure governed by parallel researcher-provided guidelines (Tables 3, 4, 5) covering the same feedback categories. The hybrid group did receive two feedback sources (AI plus teacher e-feedback), which constitutes somewhat more total feedback input than the single-source arms; however, this dual-source combination is precisely the treatment variable under investigation - the study's explicit purpose is to compare teacher-only, AI-only, and combined feedback - so the extra feedback in the hybrid arm is integral to the intervention being tested rather than a separable, confounding add-on. Similarly, the AI group's tool access (ChatGPT-4, Grammarly Premium) is the defined treatment for that arm. No arm received unmatched additional class time or budget outside the feedback modality itself.
Criterion B is met because all groups received equal class time, curriculum, and instructor, and the differing feedback inputs (including the hybrid group's dual-source feedback) are integral to the treatment contrast being tested.
-
Level 3 Criteria
-
R
Reproduced
- No independent peer-reviewed replication of this specific three-arm feedback RCT was found, and the paper itself frames such evidence as scarce.
- "The hybrid model is theorized as the optimal solution to resolve this trade-off, yet robust empirical evidence from direct, three-way comparisons is scarce." (p. 4)
Relevant Quotes:
1) "However, relatively few studies have directly compared these approaches, particularly in EFL contexts, to determine their relative effectiveness. This gap informs the rationale for our study." (p. 2)
2) "The hybrid model is theorized as the optimal solution to resolve this trade-off, yet robust empirical evidence from direct, three-way comparisons is scarce." (p. 4)
3) "Future research should aim to validate these findings through multi-institutional studies across varied educational landscapes to establish broader applicability." (p. 16)
Detailed Analysis:
The paper positions itself as filling a gap ("robust empirical evidence from direct, three-way comparisons is scarce") and calls for future validation studies, indicating no replication of this specific trial existed at publication (August 2025). An OpenAlex citation-graph search for works citing this paper (12 citing works as of July 2026: e.g., Zhao & Li 2026, "A Systematic Comparative Review of Generative AI vs. Teacher Feedback in EFL Writing"; Xie et al. 2026, "Developing Lexical Richness: A Longitudinal Learner Corpus Analysis"; Askari & Yazdani 2026, "ChatGPT and Adult English Language Learning: A Scoping Literature Review"; among others) found no independent peer-reviewed replication of this specific three-arm RCT (teacher e-feedback vs. AI feedback vs. hybrid, IELTS writing, Iranian EFL context) by an independent team; the citing works are literature reviews, scoping studies, or differently-scoped empirical work rather than direct methodological replications.
Criterion R is not met because no independent peer-reviewed replication of this specific study was found.
-
A
All-subject Exams
- Only EFL writing was assessed; no other subjects or language skills were measured and no explicit exception rationale is provided.
- "ANCOVA was used to analyze overall writing performance, as well as specific components such as Task Achievement/Response, Coherence and Cohesion, Lexical Resource, and Grammatical Range and Accuracy." (p. 10)
Relevant Quotes:
1) "Writing proficiency was assessed using IELTS writing tasks and the Oxford Placement Test before and after the intervention." (p. 1, Abstract)
2) "ANCOVA was used to analyze overall writing performance, as well as specific components such as Task Achievement/Response, Coherence and Cohesion, Lexical Resource, and Grammatical Range and Accuracy." (p. 10)
3) "The participants in this study were 88 adult intermediate students of ESL enrolled in IELTS writing courses at two branches of a single language center in Shiraz, Iran." (p. 4)
Detailed Analysis:
Criterion A requires standardised exam-based measurement across all main subjects taught, to detect negative spillover onto non-target areas. This study measured only English writing (and its four IELTS sub-components); no other subject or even other English skills (reading, listening, speaking) were assessed. The context is a specialised adult IELTS writing course at a private language center, where writing is essentially the sole subject of instruction, which resembles the standard's exception for highly specialised interventions; however, the exception applies to upper secondary or vocational education and requires that the rationale for measuring only related outcomes be clearly explained, and the paper offers no such explicit justification for restricting measurement to writing alone. Under a strict reading, the single-domain outcome measurement without an explicitly justified exception does not satisfy the criterion.
Criterion A is not met because only writing outcomes were assessed, with no measurement of other subjects or skills and no explicit exception rationale given in the paper.
-
G
Graduation Tracking
- Measurement ended at the end-of-semester post-test with no follow-up tracking to graduation, no subsequent cohort-tracking publication was found, and prerequisite Y fails.
- "Upon completion of the educational sessions, the writing post-test was administered to all three groups." (p. 10)
Relevant Quotes:
1) "Upon completion of the educational sessions, the writing post-test was administered to all three groups. This post-test included another 200-word essay to assess the participants' progress in writing improvement." (p. 10)
2) "Future research should aim to validate these findings through multi-institutional studies across varied educational landscapes to establish broader applicability." (p. 16)
Detailed Analysis:
Criterion G requires tracking participants until graduation from their educational stage. Measurement here ended with the post-test at the end of the 2.5-month semester; there is no delayed follow-up, no tracking of whether learners completed their IELTS preparation or achieved target band scores in the real exam, and no reference to planned or published follow-up studies on this cohort. An OpenAlex citation-graph search for subsequent publications by Soori, Khojasteh, or Javed citing or building on this paper found no follow-up study tracking this cohort; the 12 works citing this paper are by unrelated authors and do not track these participants. In addition, per the ranking instructions, criterion G cannot be met when criterion Y (Year Duration) is not met, and Y fails here.
Criterion G is not met because tracking stopped at the end-of-semester post-test with no follow-up through any graduation point, no subsequent cohort-tracking publication by the authors was found, and prerequisite criterion Y is not met.
-
P
Pre-Registered
- The paper provides no pre-registration reference, registry ID, or registration date, and no registration record was found in IRCT or elsewhere.
Relevant Quotes:
1) "A randomized controlled trial with a pre-test and post-test design was chosen for this study." (p. 4)
2) "Notably, all participants provided informed consent by signing a paper consent form, and their participation was voluntary." (p. 5)
3) "Per the study's explanatory sequential design, the placeholder 'x' was defined as the 'Hybrid Feedback approach' after quantitative results identified this intervention as the most effective." (p. 6, Table 2 note)
Detailed Analysis:
The paper contains no mention of pre-registration on any registry platform (e.g., ClinicalTrials.gov, OSF, AEA, IRCT), no registration ID, and no registration date. A direct check of the publisher's article landing page found no registration details, and a search of the Iranian Registry of Clinical Trials (IRCT), the registry most relevant to this Iran-based study, returned zero matching results for this trial. Moreover, the study design itself included elements finalised only after the data were analysed (the qualitative reflection prompts were tailored to whichever group performed best), which is inconsistent with a fully pre-specified, pre-registered protocol. There is no quoted evidence that hypotheses, methods, and analysis plans were published before data collection began.
Criterion P is not met because no pre-registration statement, registry ID, or registration date appears in the paper, on the publisher's page, or in the IRCT registry.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.