Effects of AI generated and teacher feedback on EFL learners writing performance and emotional experience

Alireza Maleki

Published:
ERCT Check Date:
DOI: 10.1007/s44163-026-00935-8
  • L2 languages
  • adult education
  • Asia
  • formative assessment
0
  • C

    Randomisation was at the individual student level within one course at a single institute, not at class or school level, and the one-to-one tutoring exception does not apply.

    "They were randomly assigned to two groups: an experimental group (n = 11) receiving AI-generated feedback and a control group (n = 11) receiving teacher feedback." (p. 4)

  • E

    Outcomes were measured with custom-designed parallel writing tasks scored on a rubric adapted from IELTS descriptors, not with a widely recognised standardised exam.

    "Writing achievement was measured using two parallel writing tasks designed for the pre- and post-test phases." (p. 4)

  • T

    The interval from intervention start to outcome measurement was only two weeks (Week 1 to Week 3), far short of the required full academic term.

    "Two weeks later (Week 3), both groups completed the post-test writing task under the same conditions." (p. 5)

  • D

    The control group's size, baseline scores, participant characteristics, and the exact feedback condition it received are clearly documented in the text and Tables 1-2.

    "The control group received teacher-written feedback from the same instructor using a pre-planned template mirroring the structure and approximate length of the AI feedback (one positive remark, three improvement points, two suggestions)." (p. 5)

  • S

    The study randomised individual students within one private language institute; no schools or institutional units were randomised.

    "They were randomly assigned to two groups: an experimental group (n = 11) receiving AI-generated feedback and a control group (n = 11) receiving teacher feedback." (p. 4)

  • I

    The single author designed the intervention, ran the study, and analysed the data himself, with only blinded essay raters and no independent evaluation team.

    "A.M was responsible for writing and supervising this paper." (p. 9)

  • Y

    The full study interval was about two weeks, far below 75% of an academic year, and prerequisite criterion T is also not met.

    "Two weeks later (Week 3), both groups completed the post-test writing task under the same conditions." (p. 5)

  • B

    Both groups received structurally matched feedback of the same approximate length under identical task conditions, so time and resources were balanced and only the feedback source differed.

    "This parallel structure was used to control for feedback quantity and organization, allowing the comparison to focus on feedback source rather than format." (p. 5)

  • R

    The recently published study presents itself as a novel contribution and no independent peer-reviewed replication of this specific trial exists.

    "The study makes a distinct contribution by (a) directly comparing LLM-based AI feedback with teacher feedback ... (c) situating these findings within an underrepresented EFL institute context." (p. 3)

  • A

    Only EFL paragraph writing was assessed with a custom instrument, and since prerequisite criterion E fails, A fails as well.

    "Writing achievement was measured using two parallel writing tasks designed for the pre- and post-test phases." (p. 4)

  • G

    Measurement ended at the two-week post-test with no follow-up or graduation tracking, no subsequent tracking paper was found, and prerequisite criterion Y is not met.

    "In addition, longitudinal studies could investigate how sustained exposure to AI feedback influences learners' attitudes, self-efficacy, and writing development over time." (p. 9)

  • P

    No pre-registration, registry ID, or registration date is reported anywhere in the paper, and no external registry record was found.

    "Clinical trial number: Not applicable." (p. 10)

Abstract

Although considerable research has explored the role of feedback in second language writing, limited studies have compared the effects of AI-generated feedback and teacher-written feedback on both academic performance and emotional experience, particularly in EFL contexts. This mixed-methods brief report investigated the impact of AI-mediated feedback on Iranian EFL learners' writing achievement and perceptions. Twenty-two upper-intermediate institute students were randomly assigned to either an AI feedback group (n=11) or a teacher feedback group (n=11). Both groups completed two parallel opinion paragraph writing tasks and received feedback accordingly. Quantitative analysis showed that the AI feedback group achieved significantly greater improvement in post-test scores than the teacher feedback group, highlighting the role of automated formative feedback in enhancing writing performance. Qualitative thematic analysis revealed that AI feedback promoted clarity, autonomy, and motivation, whereas teacher feedback provided emotional reassurance and interpersonal support. These findings suggest that AI systems, when designed and implemented thoughtfully, can complement human feedback by fostering both linguistic progress and learner autonomy in EFL writing instruction.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was at the individual student level within one course at a single institute, not at class or school level, and the one-to-one tutoring exception does not apply.
      • "They were randomly assigned to two groups: an experimental group (n = 11) receiving AI-generated feedback and a control group (n = 11) receiving teacher feedback." (p. 4)
      • Relevant Quotes: 1) "Twenty-two upper-intermediate institute students were randomly assigned to either an AI feedback group (n=11) or a teacher feedback group (n=11)." (p. 1, Abstract) 2) "They were randomly assigned to two groups: an experimental group (n = 11) receiving AI-generated feedback and a control group (n = 11) receiving teacher feedback." (p. 4) 3) "In Week 1, all participants completed the pre-test writing task under identical classroom conditions. Their drafts were collected and randomly assigned to feedback treatments." (p. 5) 4) "The participants were 22 Iranian EFL learners (11 males and 11 females) aged 18–22, enrolled in an upper-intermediate writing course at a private language institute in Mashhad, Iran." (p. 3) Detailed Analysis: The ERCT C criterion requires randomisation at the class level (or stronger, school level) to prevent contamination between treatment and control participants. The quotes show that individual students (or their individual drafts) drawn from a single upper-intermediate writing course at one institute were randomly assigned to the two feedback conditions. This is student-level randomisation within the same course, exactly the design the criterion is meant to guard against, since classmates in the AI and teacher feedback groups share the same classroom environment and could share or discuss the feedback they received. Exception check: the exception applies to interventions that are inherently personal teaching such as one-to-one tutoring. The intervention here is written formative feedback on paragraph drafts, delivered within a normal group-based writing course; the paper never frames it as one-to-one tutoring or personal teaching, so the exception does not apply. Criterion C is not met because randomisation was performed at the individual student level within a single course, not at the class or school level, and no tutoring exception applies.
    • E

      Exam-based Assessment

      • Outcomes were measured with custom-designed parallel writing tasks scored on a rubric adapted from IELTS descriptors, not with a widely recognised standardised exam.
      • "Writing achievement was measured using two parallel writing tasks designed for the pre- and post-test phases." (p. 4)
      • Relevant Quotes: 1) "Writing achievement was measured using two parallel writing tasks designed for the pre- and post-test phases. Each task required learners to compose a 200–250-word opinion paragraph within 30 min on familiar, school-related topics." (p. 4) 2) "Essays were assessed by two independent EFL instructors using an analytic rubric adapted from the IELTS writing descriptors. The rubric comprised five criteria—task achievement, organization, grammar, vocabulary, and mechanics—each scored on a five-point scale (0–4), with a total possible score of 20." (p. 4) 3) "The prompts were matched in structure, lexical range, and cognitive demand to reduce practice effects." (p. 4) Detailed Analysis: Criterion E requires that outcomes be measured with a widely recognised standardised exam rather than an instrument created for the study. Here the outcome measure consists of two writing prompts designed by the researcher specifically for the pre- and post-test phases of this study. Although the scoring rubric was "adapted from the IELTS writing descriptors", the assessment itself is not the IELTS exam or any other recognised standardised test; it is a custom pair of 200-250-word opinion-paragraph tasks with a researcher-adapted 20-point rubric. Adapting descriptors from a standardised exam does not make the assessment itself a standardised exam, since neither the tasks, administration, nor scoring correspond to an externally validated, widely recognised testing protocol. Criterion E is not met because outcomes were measured with custom writing tasks designed for the study and scored with a researcher-adapted rubric, not a recognised standardised exam.
    • T

      Term Duration

      • The interval from intervention start to outcome measurement was only two weeks (Week 1 to Week 3), far short of the required full academic term.
      • "Two weeks later (Week 3), both groups completed the post-test writing task under the same conditions." (p. 5)
      • Relevant Quotes: 1) "In Week 1, all participants completed the pre-test writing task under identical classroom conditions." (p. 5) 2) "Two weeks later (Week 3), both groups completed the post-test writing task under the same conditions." (p. 5) 3) "After the two-week intervention, the AI feedback group seemed to achiev a notably higher post-test mean (M = 13.91, SD = 1.27) than the teacher feedback group (M = 12.00, SD = 1.25)." (p. 5) 4) "The small sample size and short intervention period limit generalizability." (p. 9) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (about 3-4 months) after the intervention begins. The quotes clearly establish the timeline: the pre-test and start of the feedback treatment occurred in Week 1 and the post-test was administered in Week 3, i.e. the entire interval from intervention start to outcome measurement was approximately two weeks. The authors themselves describe it as a "two-week intervention" and acknowledge the "short intervention period" as a limitation. Two weeks falls far short of a full academic term, and there was no delayed follow-up measurement of any kind. Criterion T is not met because outcomes were measured only two weeks after the intervention began, far less than one full academic term.
    • D

      Documented Control Group

      • The control group's size, baseline scores, participant characteristics, and the exact feedback condition it received are clearly documented in the text and Tables 1-2.
      • "The control group received teacher-written feedback from the same instructor using a pre-planned template mirroring the structure and approximate length of the AI feedback (one positive remark, three improvement points, two suggestions)." (p. 5)
      • Relevant Quotes: 1) "The participants were 22 Iranian EFL learners (11 males and 11 females) aged 18–22, enrolled in an upper-intermediate writing course at a private language institute in Mashhad, Iran. Their proficiency level was determined by their scores on the institute's standardized placement test, corresponding to B2 on the CEFR scale." (p. 3) 2) "All participants had at least three years of prior formal English instruction and similar exposure to writing practice. They were randomly assigned to two groups: an experimental group (n = 11) receiving AI-generated feedback and a control group (n = 11) receiving teacher feedback." (p. 4) 3) "The control group received teacher-written feedback from the same instructor using a pre-planned template mirroring the structure and approximate length of the AI feedback (one positive remark, three improvement points, two suggestions)." (p. 5) 4) "Both groups started at nearly the same mean level (AI group M = 10.23, SD = 1.18; Teacher group M = 10.41, SD = 1.11)." (p. 5) 5) "Table 1 presents the individual pre-test and post-test writing scores of all participants in the AI feedback and teacher feedback groups." (p. 5) Detailed Analysis: Criterion D requires clear documentation of the control group: who they are, their size, baseline performance, and what treatment they received. The paper documents the control group's size (n = 11), the shared demographic profile of the sample (age 18-22, gender balance, B2 proficiency confirmed by placement test, at least three years of prior English instruction), baseline writing performance at both group level (Table 2: M=10.41, SD=1.11) and individual level (Table 1), and precisely what the control condition consisted of (teacher-written feedback from the same instructor using a template matched in structure and length to the AI feedback). Testing and classroom conditions were identical for both groups. This constitutes adequate documentation for comparison, even though group-by-group demographic breakdowns are not separately tabulated. Criterion D is met because the control group's size, baseline performance, participant characteristics, and exact treatment are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • The study randomised individual students within one private language institute; no schools or institutional units were randomised.
      • "They were randomly assigned to two groups: an experimental group (n = 11) receiving AI-generated feedback and a control group (n = 11) receiving teacher feedback." (p. 4)
      • Relevant Quotes: 1) "They were randomly assigned to two groups: an experimental group (n = 11) receiving AI-generated feedback and a control group (n = 11) receiving teacher feedback." (p. 4) 2) "The participants were 22 Iranian EFL learners ... enrolled in an upper-intermediate writing course at a private language institute in Mashhad, Iran." (p. 3) Detailed Analysis: Criterion S requires randomisation at the level of whole schools or equivalent educational units. This study took place within a single private language institute in Mashhad, and randomisation was performed across individual students in one writing course. No schools, institutes, sites, or other institutional units were randomised; there was only one institution involved in total. Consequently the study cannot satisfy school-level randomisation. Criterion S is not met because randomisation occurred at the individual student level within a single institute, with no school-level assignment.
    • I

      Independent Conduct

      • The single author designed the intervention, ran the study, and analysed the data himself, with only blinded essay raters and no independent evaluation team.
      • "A.M was responsible for writing and supervising this paper." (p. 9)
      • Relevant Quotes: 1) "A.M was responsible for writing and supervising this paper." (p. 9, Author contributions) 2) "The AI prompt was developed based on formative feedback principles (clarity, balance, and actionability) and aligned with the analytic writing rubric used for assessment." (p. 5) 3) "The control group received teacher-written feedback from the same instructor using a pre-planned template mirroring the structure and approximate length of the AI feedback." (p. 5) 4) "To ensure scoring reliability, both raters independently evaluated all scripts without access to group assignment." (p. 4) 5) "Essays were assessed by two independent EFL instructors using an analytic rubric adapted from the IELTS writing descriptors." (p. 4) Detailed Analysis: Criterion I requires that the study be conducted independently of the intervention's designers, e.g. by an external evaluation team. Here a single author designed the study, developed the AI feedback prompt, arranged the teacher feedback template, and conducted the analysis and write-up; there is no external evaluation agency or third-party oversight of data collection, analysis, or conclusions. The only element of independence is that two EFL instructors rated the essays blind to group assignment, which is a rater-blinding step rather than independent conduct of the trial: the raters worked within the study run by the author, and the interviews, thematic analysis, and statistical analysis were all performed by the same researcher who designed the intervention. This does not match the standard's examples of independence (external evaluation teams, government-led trials, independent enumerators covering the full data pipeline). Criterion I is not met because the same single author designed the intervention and conducted the study and analysis, with no independent third-party evaluation beyond blinded essay raters.
    • Y

      Year Duration

      • The full study interval was about two weeks, far below 75% of an academic year, and prerequisite criterion T is also not met.
      • "Two weeks later (Week 3), both groups completed the post-test writing task under the same conditions." (p. 5)
      • Relevant Quotes: 1) "In Week 1, all participants completed the pre-test writing task under identical classroom conditions." (p. 5) 2) "Two weeks later (Week 3), both groups completed the post-test writing task under the same conditions." (p. 5) 3) "After the two-week intervention, the AI feedback group seemed to achiev a notably higher post-test mean..." (p. 5) 4) "The small sample size and short intervention period limit generalizability." (p. 9) Detailed Analysis: Criterion Y requires outcomes to be tracked for at least 75% of a full academic year (roughly 9-10 months) from intervention start. The entire study, from pre-test to post-test, spanned about two weeks, and no longer-term follow-up was conducted; the author explicitly lists the short intervention period as a limitation and calls for "longer instructional periods" in future research. Two weeks is a tiny fraction of an academic year. In addition, the prerequisite criterion T (term duration) is not met, which by the instructions means Y cannot be met. Criterion Y is not met because the study lasted only about two weeks from intervention start to final measurement, nowhere near 75% of an academic year.
    • B

      Balanced Control Group

      • Both groups received structurally matched feedback of the same approximate length under identical task conditions, so time and resources were balanced and only the feedback source differed.
      • "This parallel structure was used to control for feedback quantity and organization, allowing the comparison to focus on feedback source rather than format." (p. 5)
      • Relevant Quotes: 1) "The experimental group received AI-generated written feedback using ChatGPT (OpenAI, GPT-4-turbo, April 2025 update)." (p. 5) 2) "Feedback was copied directly from the ChatGPT interface without modification and printed for each student." (p. 5) 3) "The control group received teacher-written feedback from the same instructor using a pre-planned template mirroring the structure and approximate length of the AI feedback (one positive remark, three improvement points, two suggestions). This parallel structure was used to control for feedback quantity and organization, allowing the comparison to focus on feedback source rather than format." (p. 5) 4) "In Week 1, all participants completed the pre-test writing task under identical classroom conditions." (p. 5) 5) "Two weeks later (Week 3), both groups completed the post-test writing task under the same conditions. ... The use of parallel tasks, controlled feedback structure, and random group assignment helped minimize confounding variables." (p. 5) Detailed Analysis: Criterion B requires that the control group receive comparable time and resources so the intervention's specific effect is isolated. Following the decision tree: did the intervention add extra time or budget for the treatment group? No. Both groups completed the same two writing tasks under identical classroom conditions on the same schedule, and both groups received written feedback of deliberately matched quantity and structure (one positive remark, three improvement points, two suggestions), differing only in whether the feedback source was ChatGPT or the instructor. This is an active, structurally matched control: the control group received an equivalent educational input (teacher feedback of the same length and format) rather than nothing. No additional instructional time, materials, or budget were given to the AI group; the printed AI feedback is a direct substitute for the printed teacher feedback. Since no meaningful extra resources were present in the treatment condition, the balance requirement is satisfied. Criterion B is met because both groups received feedback of matched quantity, structure, and delivery under identical conditions, with the feedback source being the only difference.
  • Level 3 Criteria

    • R

      Reproduced

      • The recently published study presents itself as a novel contribution and no independent peer-reviewed replication of this specific trial exists.
      • "The study makes a distinct contribution by (a) directly comparing LLM-based AI feedback with teacher feedback ... (c) situating these findings within an underrepresented EFL institute context." (p. 3)
      • Relevant Quotes: 1) "This mixed-methods brief report investigated the impact of AI-mediated feedback on Iranian EFL learners' writing achievement and perceptions." (p. 1) 2) "While recent studies have documented the effectiveness, usability, and accuracy of AI-generated feedback ... particularly from LLM-based tools, fewer studies have systematically compared such feedback with teacher-written feedback while simultaneously examining both writing performance and learners' emotional or affective experiences." (p. 2) 3) "The study makes a distinct contribution by (a) directly comparing LLM-based AI feedback with teacher feedback, (b) integrating objective writing performance data with qualitative evidence of learners' emotional and motivational experiences, and (c) situating these findings within an underrepresented EFL institute context." (p. 3) Detailed Analysis: Criterion R requires that this specific study be independently replicated by a different research team, in a different context, in a peer-reviewed journal. The paper itself claims novelty ("distinct contribution", "underrepresented EFL institute context"), which indicates no prior replication of this design existed at publication. While there is a broader literature comparing AI/automated feedback with teacher feedback (e.g. Banihashem et al. 2024, Er et al. 2025, Zhao 2025 cited in the references), those are related studies with different designs, populations, and outcome sets, not replications of this particular RCT contrasting ChatGPT feedback with structurally matched teacher feedback on writing performance plus emotional experience in an Iranian institute. The article was published online on 8 February 2026, and no independent replication of this specific study by a different team has been identified. Internet searches for citing or follow-up publications referencing this article (via Crossref, Springer, and general web search as of July 2026) returned no independent replication study by another research team; the article is too recently published for a replication to plausibly exist yet. Criterion R is not met because no independent published replication of this specific study exists; the paper explicitly frames itself as novel.
    • A

      All-subject Exams

      • Only EFL paragraph writing was assessed with a custom instrument, and since prerequisite criterion E fails, A fails as well.
      • "Writing achievement was measured using two parallel writing tasks designed for the pre- and post-test phases." (p. 4)
      • Relevant Quotes: 1) "Writing achievement was measured using two parallel writing tasks designed for the pre- and post-test phases." (p. 4) 2) "Quantitative analysis showed that the AI feedback group achieved significantly greater improvement in post-test scores than the teacher feedback group." (p. 1) 3) "Moreover, only one AI tool and one writing genre were tested." (p. 9) Detailed Analysis: Criterion A requires standardised exam-based assessment across all main subjects, and by the instructions it cannot be met when criterion E is not met. Criterion E fails here because the outcome measure was a custom writing task, so A automatically fails as well. Substantively, the study measured only EFL opinion-paragraph writing, a single skill within a single subject; no other subjects (or even other English skills or genres, as the author acknowledges) were assessed. As a private language-institute EFL course, the setting could arguably justify a narrow English focus, but even that focus was measured with a non-standardised custom instrument, so the prerequisite is not satisfied. Criterion A is not met because criterion E is not met and only a single custom writing measure in one subject was used.
    • G

      Graduation Tracking

      • Measurement ended at the two-week post-test with no follow-up or graduation tracking, no subsequent tracking paper was found, and prerequisite criterion Y is not met.
      • "In addition, longitudinal studies could investigate how sustained exposure to AI feedback influences learners' attitudes, self-efficacy, and writing development over time." (p. 9)
      • Relevant Quotes: 1) "Two weeks later (Week 3), both groups completed the post-test writing task under the same conditions. ... Immediately after post-testing, semi-structured interviews were conducted to capture learners' perceptions and emotional experiences." (p. 5) 2) "In addition, longitudinal studies could investigate how sustained exposure to AI feedback influences learners' attitudes, self-efficacy, and writing development over time." (p. 9) Detailed Analysis: Criterion G requires tracking participants until graduation from their educational stage, and by the instructions it cannot be met when criterion Y is not met. Y fails here, so G automatically fails. Substantively, measurement ended at the Week 3 post-test and immediate interviews; there was no follow-up of any kind, let alone tracking to course completion or graduation, and the author explicitly defers longitudinal investigation to future research. Internet searches for subsequent publications by the same author (Alireza Maleki) that might track this cohort further (via Crossref, Springer, and general web search as of July 2026) found no such follow-up paper; his only other located publication (Maleki, Curr Psychol, 2025, on growth mindset feedback) is an unrelated study with a different sample. No follow-up publications tracking this specific cohort exist. Criterion G is not met because data collection stopped two weeks after the intervention began with no graduation tracking, no follow-up publication was found, and prerequisite criterion Y also fails.
    • P

      Pre-Registered

      • No pre-registration, registry ID, or registration date is reported anywhere in the paper, and no external registry record was found.
      • "Clinical trial number: Not applicable." (p. 10)
      • Relevant Quotes: 1) "This study was reviewed and approved by the institutional review board (IRB) at Kashmar Institute of Higher Education, Iran." (p. 10) 2) "Clinical trial number: Not applicable." (p. 10) Detailed Analysis: Criterion P requires pre-registration of the full study protocol (hypotheses, methods, planned analyses) on a public registry before data collection began. The paper contains no mention of any registry (e.g. ClinicalTrials.gov, OSF, AsPredicted, AEA registry), no registration ID, and no registration date; the declarations section explicitly states "Clinical trial number: Not applicable." IRB approval is an ethics review, not a pre-registered protocol. A targeted internet check for a registered protocol under the author's name or the study title (as of July 2026) found no matching registration record on any public trial or study registry, consistent with the paper's own declaration. There is therefore no evidence that the hypotheses and analysis plan were registered before data collection. Criterion P is not met because the study was not pre-registered on any registry and the paper explicitly lists no trial number.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.