Rethinking AI-assisted writing instruction: feedback literacy scripts, calibration training, and student writing development

Zhe Dai

Published:
ERCT Check Date:
DOI: 10.3389/fpsyg.2026.1829268
  • L2 languages
  • higher education
  • China
  • formative assessment
0
  • C

    Randomisation was conducted at the individual student level within one institution, not at the class or school level, and the intervention is not one-to-one tutoring.

    "Participants were assigned to four groups of 30 through stratified randomization based on baseline writing ability, English proficiency, and prior AI use experience." (p. 5)

  • E

    Writing outcomes were scored with a researcher-designed 100-point analytic rubric by two teachers, not with a widely recognised standardised exam.

    "Writing quality was rated independently by two experienced English teachers using a 100-point analytic rubric comprising four dimensions..." (p. 7)

  • T

    The whole study, from baseline to the final retention task, lasted only six weeks, which is well short of one academic term.

    "The study lasted six weeks and consisted of four time points." (p. 5)

  • D

    The control (regular AI) group's size, condition, baseline writing scores, and comparability checks are clearly documented.

    "The regular AI group received only standard AI feedback and did not receive additional metacognitive training." (p. 5)

  • S

    Randomisation occurred among individual students within a single institution, not among schools or comparable implementation units.

    "The study included 120 undergraduate English majors from the host institution... Participants were assigned to four groups of 30 through stratified randomization..." (p. 5)

  • I

    A single author designed the interventions and also conducted, analysed, and reported the study with no external or third-party evaluation.

    "ZD: Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing." (p. 12)

  • Y

    The entire study lasted six weeks, far short of 75% of an academic year, and criterion T is already not met.

    "The study lasted six weeks and consisted of four time points." (p. 5)

  • B

    All groups received the same AI feedback and writing tasks; the modest extra training time (FRAC/APCA activities) is the explicit treatment variable being tested against regular AI use.

    "All students had access to AI feedback during T1 and T2, but the way they engaged with that feedback differed across groups." (p. 5)

  • R

    The paper, published in June 2026, reports no independent replication of this specific FRAC/APCA trial and none could be identified via internet search.

  • A

    Only English argumentative writing was assessed, with no standardised exams and no measurement of other core subjects, and criterion E is not met.

    "All writing tasks were academic argumentative essays to ensure comparability across tasks." (p. 6)

  • G

    Tracking ended at the six-week retention task with no follow-up to graduation, and prerequisite criterion Y is not met.

    "At T3, AI support was withdrawn again, and students were required to complete the writing task independently..." (p. 5)

  • P

    The paper reports ethics approval but no pre-registration of the study protocol on any registry.

Abstract

Introduction: As generative AI becomes increasingly integrated into writing instruction, the central educational challenge is not only to provide feedback but also to help students interpret, evaluate, and use such feedback critically. This study examined whether two metacognitive interventions--a Feedback Literacy Script (FRAC) and an Assessment-Performance Calibration Activity (APCA)--could improve students' writing quality and self-assessment accuracy in AI-assisted writing. Methods: A 2 x 2 mixed factorial design was employed with 120 undergraduate English majors assigned to four conditions: regular AI use, FRAC only, APCA only, and FRAC+APCA. Across four writing task points, the study collected data on writing quality gain, self-assessment accuracy, overconfidence, effective feedback uptake, and revision depth. Results: The findings showed differentiated effects of the two interventions. FRAC had a stronger direct effect on writing quality, effective feedback uptake, and deep revision, suggesting that feedback literacy training primarily improved how students processed and enacted AI feedback. APCA showed the clearest effect on self-assessment accuracy and overconfidence reduction, indicating that calibration training more directly strengthened students' internal judgment. The combined intervention produced the highest gains in overall writing quality and the strongest retention after AI support was removed, but it did not outperform APCA alone in self-assessment accuracy.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was conducted at the individual student level within one institution, not at the class or school level, and the intervention is not one-to-one tutoring.
      • "Participants were assigned to four groups of 30 through stratified randomization based on baseline writing ability, English proficiency, and prior AI use experience." (p. 5)
      • Relevant Quotes: 1) "Participants were assigned to four groups of 30 through stratified randomization based on baseline writing ability, English proficiency, and prior AI use experience." (p. 5) 2) "The study included 120 undergraduate English majors from the host institution, aged 19 to 22 years (M = 20.3, SD = 0.9)." (p. 5) 3) "Based on the presence or absence of the two intervention factors, four groups were established: a regular AI group, a FRAC-only group, an APCA-only group, and a combined FRAC + APCA group." (p. 5) Detailed Analysis: The unit of randomisation was the individual student: 120 undergraduates from a single host institution were stratified and randomly assigned to four groups of 30. There is no indication that intact classes or schools were randomised. The ERCT exception for personal teaching (e.g., one-to-one tutoring) does not apply: the interventions (FRAC feedback-literacy scripts and APCA calibration activities) were delivered as structured training within a shared course context, not as personal tutoring. Because students in different conditions came from the same institution and programme, contamination between conditions cannot be ruled out, which is exactly the problem the class-level requirement is designed to avoid. Criterion C is not met because randomisation was at the student level within one institution and no valid tutoring exception applies.
    • E

      Exam-based Assessment

      • Writing outcomes were scored with a researcher-designed 100-point analytic rubric by two teachers, not with a widely recognised standardised exam.
      • "Writing quality was rated independently by two experienced English teachers using a 100-point analytic rubric comprising four dimensions..." (p. 7)
      • Relevant Quotes: 1) "Writing quality was rated independently by two experienced English teachers using a 100-point analytic rubric comprising four dimensions: content and argumentation, organization and structure, language and expression, and evidence and citation, with each dimension worth 25 points." (p. 7) 2) "All writing tasks were academic argumentative essays to ensure comparability across tasks. At each of the four time points, students wrote 800 to 1000 words within 90 min." (p. 6) 3) "All participants had passed CET-6 or had attained an equivalent level of English proficiency and possessed basic academic writing skills." (p. 5) Detailed Analysis: The primary outcomes (writing quality, writing quality gain) were measured with study-specific argumentative essay tasks scored on a custom 100-point analytic rubric developed and calibrated by the research team, with an ICC of 0.92 between two raters. Good inter-rater reliability does not make the instrument a standardised exam: the tasks and rubric were designed for this study, not drawn from a widely recognised standardised testing programme. The only standardised test mentioned (CET-6) was used solely as an eligibility criterion, not as an outcome measure. Secondary measures (self-assessment accuracy, adoption rates, revision depth) are likewise bespoke process measures. Criterion E is not met because outcomes were measured with custom essay tasks and a researcher-designed rubric rather than a recognised standardised exam.
    • T

      Term Duration

      • The whole study, from baseline to the final retention task, lasted only six weeks, which is well short of one academic term.
      • "The study lasted six weeks and consisted of four time points." (p. 5)
      • Relevant Quotes: 1) "The study lasted six weeks and consisted of four time points." (p. 5) 2) "...be able to complete all tasks during the six-week study period." (p. 5) 3) "At T3, AI support was withdrawn again, and students were required to complete the writing task independently, allowing us to observe whether the earlier intervention effects could be retained without assistance." (p. 5) Detailed Analysis: The intervention and all outcome measurement occurred within a six-week window spanning four writing task points (T0 baseline through T3 retention). A term under the ERCT standard is approximately 3-4 months; the interval from intervention start to the final outcome measurement here is at most about five to six weeks. There is no longer-term follow-up measurement reported. Criterion T is not met because the interval from intervention start to final measurement was about six weeks, far less than one full academic term.
    • D

      Documented Control Group

      • The control (regular AI) group's size, condition, baseline writing scores, and comparability checks are clearly documented.
      • "The regular AI group received only standard AI feedback and did not receive additional metacognitive training." (p. 5)
      • Relevant Quotes: 1) "The regular AI group received only standard AI feedback and did not receive additional metacognitive training." (p. 5) 2) "Participants were assigned to four groups of 30 through stratified randomization based on baseline writing ability, English proficiency, and prior AI use experience. Baseline equivalence tests showed no significant group differences in T0 initial draft scores, F(3, 116) = 0.82, p = 0.485, or in English proficiency and AI use experience, indicating good comparability across groups." (p. 5) 3) "The baseline means were 68.48 (SD = 7.43) for the regular AI group, 68.79 (SD = 7.95) for the FRAC group, 67.98 (SD = 9.11) for the APCA group, and 68.82 (SD = 9.40) for the combined group." (p. 5) 4) "The study included 120 undergraduate English majors from the host institution, aged 19 to 22 years (M = 20.3, SD = 0.9). Female students comprised 78% of the sample..." (p. 5) 5) "Students in the regular AI group could use AI feedback freely but received no metacognitive training." (p. 5) Detailed Analysis: The control condition (regular AI group, n = 30) is clearly defined: these students received the same AI feedback access at T1 and T2 as all other groups but no metacognitive training. Its baseline writing performance is reported (M = 68.48, SD = 7.43), baseline equivalence across groups is formally tested, sample demographics and eligibility criteria are described, and Table 1 reports the control group's outcome statistics (writing gain, SAA error, overconfidence rate, EAR, revision depth). This provides enough documentation to judge comparability and what the control group experienced. Criterion D is met because the control group's size, condition, baseline performance, and comparability are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred among individual students within a single institution, not among schools or comparable implementation units.
      • "The study included 120 undergraduate English majors from the host institution... Participants were assigned to four groups of 30 through stratified randomization..." (p. 5)
      • Relevant Quotes: 1) "The study included 120 undergraduate English majors from the host institution, aged 19 to 22 years (M = 20.3, SD = 0.9)." (p. 5) 2) "Participants were assigned to four groups of 30 through stratified randomization based on baseline writing ability, English proficiency, and prior AI use experience." (p. 5) Detailed Analysis: The study was conducted at one host institution (Zhejiang Institute of Communications) and randomised individual students to conditions. No schools, campuses, sites, or other institution-level units were randomised. Since only one institution participated, school-level randomisation was impossible by design. Criterion S is not met because randomisation was at the individual student level within a single institution, not at the school level.
    • I

      Independent Conduct

      • A single author designed the interventions and also conducted, analysed, and reported the study with no external or third-party evaluation.
      • "ZD: Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing." (p. 12)
      • Relevant Quotes: 1) "ZD: Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing." (p. 12) 2) "The researcher reviewed these logs according to a unified feedback protocol and provided brief written comments for students to use in the next task." (p. 6) 3) "These records were included in the analysis as process data, and implementation was verified through submission rates, completeness checks, and classroom observations by the researcher." (p. 6) 4) "Electronic files were stored on password-protected devices accessible only to the researcher." (p. 6) Detailed Analysis: This is a single-author study in which the same person designed the FRAC and APCA interventions, implemented them, collected the data, monitored implementation via classroom observation, analysed the results, and wrote the paper. Two English teachers rated essays and independent coders double-coded a portion of process data, but there is no statement of an external evaluation team, third-party oversight, or independence between intervention design and trial conduct. This is the paradigm case of developer-led evaluation that criterion I is designed to guard against. Criterion I is not met because the intervention designer also conducted, analysed, and reported the study without documented independent oversight.
    • Y

      Year Duration

      • The entire study lasted six weeks, far short of 75% of an academic year, and criterion T is already not met.
      • "The study lasted six weeks and consisted of four time points." (p. 5)
      • Relevant Quotes: 1) "The study lasted six weeks and consisted of four time points." (p. 5) 2) "...be able to complete all tasks during the six-week study period." (p. 5) Detailed Analysis: Criterion Y requires tracking of outcomes for at least 75% of a full academic year (roughly 9-10 months) from the intervention start. Here the full study, including the baseline, both AI-supported tasks, the transfer task, and the unsupported retention task, was completed within six weeks. Additionally, per the ranking rules, criterion Y cannot be met when criterion T (term duration) is not met, and T fails here. Criterion Y is not met because the six-week study duration is far below 75% of an academic year.
    • B

      Balanced Control Group

      • All groups received the same AI feedback and writing tasks; the modest extra training time (FRAC/APCA activities) is the explicit treatment variable being tested against regular AI use.
      • "All students had access to AI feedback during T1 and T2, but the way they engaged with that feedback differed across groups." (p. 5)
      • Relevant Quotes: 1) "All students had access to AI feedback during T1 and T2, but the way they engaged with that feedback differed across groups. Students in the regular AI group could use AI feedback freely but received no metacognitive training." (p. 5) 2) "FRAC training was completed before T1 and lasted approximately 15 to 25 min, with brief reviews conducted before T1 and T2." (p. 6) 3) "APCA was implemented continuously from T0 to T3, with each round lasting approximately 10 to 15 min." (p. 6) 4) "This study used a 2 × 2 mixed factorial design to examine the effects of two metacognitive interventions—the Feedback Literacy Script (FRAC) and the Assessment-Performance Calibration Activity (APCA)—on students' writing quality and self-assessment accuracy." (p. 4-5) 5) "The researcher reviewed these logs according to a unified feedback protocol and provided brief written comments for students to use in the next task." (p. 6) Detailed Analysis: Applying the criterion B decision tree: the core educational inputs were held constant across all four groups - the same writing tasks, the same time limits, and identical access to standardised Claude-generated AI feedback at T1 and T2. The additional resources in the intervention arms consist of the metacognitive training activities themselves: a one-off 15-25 min FRAC script training with brief reviews, and 10-15 min APCA calibration rounds with researcher comments on calibration logs. These additional activities ARE the treatment variables in the 2 x 2 design - the study explicitly tests whether adding FRAC and/or APCA on top of regular AI use improves outcomes, with the regular AI group serving as the business-as-usual comparison by design (RESOURCES_ARE_TREATMENT = true). The extra time involved is also modest (tens of minutes over six weeks) relative to the four 90-min writing tasks common to all groups. It should be noted that APCA students additionally received brief written researcher comments and teacher-rating comparisons, but these are integral components of the calibration intervention being tested rather than a separable, confounding add-on. Criterion B is met because the added metacognitive training is the explicit treatment variable tested against a regular AI-use control that otherwise received identical tasks and AI feedback.
  • Level 3 Criteria

    • R

      Reproduced

      • The paper, published in June 2026, reports no independent replication of this specific FRAC/APCA trial and none could be identified via internet search.
      • Relevant Quotes: 1) "Future research should test the intervention in more diverse contexts, collect comparable process data across all groups, and incorporate variables such as trust in AI feedback, perceived effort..." (p. 11) 2) "This study discussed the role of structured metacognitive support in students' use of AI feedback for writing revision." (p. 11) Detailed Analysis: Criterion R requires independent replication of the specific study by a different research team, published in a peer-reviewed journal. This paper was published on 12 June 2026, only about six weeks before this verification, and the author explicitly frames testing "the intervention in more diverse contexts" as future research. Internet searches (Frontiers article page, general web search engines) were conducted for independent replications of this specific FRAC/APCA 2x2 trial by other authors; no such replication was found. Given the very recent publication date and the novelty of the FRAC/APCA instruments, this is expected and no quotes from a replication study can be provided. Criterion R is not met because no independent peer-reviewed replication of this specific study exists or could be found.
    • A

      All-subject Exams

      • Only English argumentative writing was assessed, with no standardised exams and no measurement of other core subjects, and criterion E is not met.
      • "All writing tasks were academic argumentative essays to ensure comparability across tasks." (p. 6)
      • Relevant Quotes: 1) "All writing tasks were academic argumentative essays to ensure comparability across tasks." (p. 6) 2) "Writing quality was rated independently by two experienced English teachers using a 100-point analytic rubric..." (p. 7) Detailed Analysis: Criterion A requires standardised exam-based assessment across all main subjects, and it cannot be met when criterion E fails, which is the case here. The study measured only English academic writing quality plus process variables (self-assessment accuracy, feedback uptake, revision depth); no other subjects were assessed. While a narrow focus might be argued for a specialised undergraduate English-major context, the prerequisite criterion E is unmet because no standardised exams were used at all. Criterion A is not met because criterion E fails and only a single subject (English writing) was measured with custom instruments.
    • G

      Graduation Tracking

      • Tracking ended at the six-week retention task with no follow-up to graduation, and prerequisite criterion Y is not met.
      • "At T3, AI support was withdrawn again, and students were required to complete the writing task independently..." (p. 5)
      • Relevant Quotes: 1) "At T3, AI support was withdrawn again, and students were required to complete the writing task independently, allowing us to observe whether the earlier intervention effects could be retained without assistance." (p. 5) 2) "The study lasted six weeks and consisted of four time points." (p. 5) Detailed Analysis: Criterion G requires participants to be tracked until graduation from their educational stage. The undergraduate participants (aged 19-22) were followed only through the final retention task at the end of the six-week study; no longer-term follow-up or planned follow-up publication is mentioned. Internet searches for follow-up publications by Zhe Dai tracking this same cohort toward graduation were conducted; none were found, consistent with the paper being too recent (published 12 June 2026) for follow-up studies to plausibly exist yet. In addition, per the ranking rules, criterion G cannot be met when criterion Y is not met, and Y fails here. Criterion G is not met because measurement stopped six weeks into the study with no graduation tracking, and no follow-up publication could be found.
    • P

      Pre-Registered

      • The paper reports ethics approval but no pre-registration of the study protocol on any registry.
      • Relevant Quotes: 1) "This study involved human participants and was approved by the university ethics committee of the host institution under Approval No. ZJIC-IRB-2025-EDU-371." (p. 6) 2) "The studies involving humans were approved by the Ethics Committee of Zhejiang Institute of Communications, Hangzhou, China under Approval No. ZJIC-IRB-2025-EDU-371." (p. 11) Detailed Analysis: Criterion P requires that the full study protocol, including hypotheses, methods, and planned analyses, be registered on a public registry before data collection began. The paper mentions only institutional ethics approval (an IRB approval number), which is not pre-registration of a protocol. No registry name, registration ID, or registration date (e.g., ClinicalTrials.gov, OSF, AsPredicted, Chinese Clinical Trial Registry) appears anywhere in the paper. Internet searches for a pre-registration record tied to this study or its IRB approval number were conducted; none were found. Criterion P is not met because no pre-registration of the study protocol is reported or discoverable.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.