The Impact of Multi-Level Anonymity in Asynchronous Online Peer Feedback for EFL Writing in Higher Education

Piansin Pinchai

Published:
ERCT Check Date:
DOI: 10.69598/hasss.25.2.273433
  • L2 languages
  • higher education
  • Asia
  • digital assessment
  • formative assessment
0
  • C

    Randomization was at the individual student level (and the two groups were in fact separate academic-year cohorts), not at the class or school level, and the group-based peer feedback intervention does not qualify for the one-to-one tutoring exception.

    "62 first-year English major students from a mid-sized university in Thailand were randomly assigned to a control group of 31" (p. 271)

  • E

    Outcomes were measured with a course-specific narrative paragraph writing task scored by a researcher-made rubric, not a widely recognised standardised exam.

    "a narrative paragraph of 240–260 words was used as the first and final draft to test the two groups of participants before (pretest) and after (posttest)" (p. 274)

  • T

    The intervention and outcome measurement were completed within a single 3-week instructional unit, far short of one academic term.

    "This learning unit spanned 3 weeks with 1 3-hour meeting per week." (p. 274)

  • D

    The control group's size, cohort, condition (OPF plus teacher feedback), and baseline and final writing scores are clearly documented, with sample demographics reported.

    "The control group of 31 participants enrolled in an English reading and writing course in the second semester of academic year 2022 and received multi-level anonymity in asynchronous OPF peer feedback with teacher intervention." (p. 274)

  • S

    The study took place in a single university with allocation at the student/cohort level; no school-level randomisation occurred.

    "The site of this study was Mae Fah Luang University. The target sample was 62 first-year English major students" (p. 273)

  • I

    The single author designed, taught, administered, and scored the intervention with no external or independent evaluation team mentioned.

    "On the teacher's side, the rubric was used and applied by only 1 instructor in order to ensure fairness, consistency, and reliability for all 62 participants." (p. 274)

  • Y

    The full study lasted only 3 weeks, nowhere near 75% of an academic year, so the year-duration requirement fails (as does its prerequisite T).

    "This learning unit spanned 3 weeks with 1 3-hour meeting per week." (p. 274)

  • B

    Both groups received the identical 3-week OPF activity and class time, and the only input difference (presence or absence of teacher feedback) is the explicit treatment variable being tested.

    "1) to compare the effectiveness of multi-level anonymity in asynchronous OPF with and without teacher feedback intervention" (p. 271)

  • R

    No independent replication of this specific study exists; it was published in May 2025, has zero recorded citations as of the verification date, and no external search identifies any replication by another team.

  • A

    Only EFL narrative paragraph writing was assessed, with a custom rubric rather than standardised exams, so all-subject coverage fails (and prerequisite E is unmet).

    "All participants completed a narrative paragraph writing task, which was used for data collection and analysis." (p. 271)

  • G

    Measurement ended with the final draft at week 3 of the unit; no tracking of participants to graduation is reported, and no follow-up publication tracking this cohort was found.

    "Once this feedback process was completed by week 3, the teacher shared the review forms with the file attached for all participants to access and review before the final draft production stage." (p. 275)

  • P

    The paper contains no mention of any pre-registered protocol, registry, or registration date, and no external registry record for this study could be found.

Abstract

Online Peer Feedback (OPF) is a proven and effective peer editing tool in EFL writing classrooms. However, the level of effectiveness varies depending the relationship between peer editors and on the available peer editing tools. This study investigates the impact of multi-level anonymity in asynchronous OPF by addressing two key objectives 1) to compare the effectiveness of multi-level anonymity in asynchronous OPF with and without teacher feedback intervention; and 2) to evaluate the quality of multi-level anonymity in asynchronous OPF compared to traditional teacher feedback. The study utilized a randomized pretest-posttest control group design. 62 first-year English major students from a mid-sized university in Thailand were randomly assigned to a control group of 31, and they received triple-anonymity asynchronous OPF together with teacher feedback. The experimental group of 31 received triple-anonymity asynchronous OPF without any teacher feedback. All participants completed a narrative paragraph writing task, which was used for data collection and analysis. Quantitative data were analyzed using descriptive statistics and bivariate correlations. Two key findings emerged: 1) There was no significant difference in writing skill improvement between the experimental and control groups. Triple-anonymity asynchronous OPF with or without teacher feedback intervention are both practical peer assessment tools in EFL writing classrooms; and moreover, 2) there was a moderate positive correlation between peers and the teacher feedback in the experimental group. This indicates reliability of peer feedback in the triple-anonymity asynchronous OPF group without teacher intervention. These results suggest that the incorporation of triple-anonymity asynchronous OPF into writing instruction can develop students' writing skills and can enhance assessment methods in higher education EFL classrooms.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomization was at the individual student level (and the two groups were in fact separate academic-year cohorts), not at the class or school level, and the group-based peer feedback intervention does not qualify for the one-to-one tutoring exception.
      • "62 first-year English major students from a mid-sized university in Thailand were randomly assigned to a control group of 31" (p. 271)
      • Relevant Quotes: 1) "The study utilized a randomized pretest-posttest control group design. 62 first-year English major students from a mid-sized university in Thailand were randomly assigned to a control group of 31, and they received triple-anonymity asynchronous OPF together with teacher feedback." (Abstract, p. 271) 2) "The control group of 31 participants enrolled in an English reading and writing course in the second semester of academic year 2022 and received multi-level anonymity in asynchronous OPF peer feedback with teacher intervention. The experimental group of 31 participants enrolled in an English reading and writing course in the second semester of academic year 2023 and received multi-level anonymity in asynchronous OPF without the teacher's intervention." (p. 274) 3) "In order to set a baseline for evaluation and analysis, a pretest-posttest randomized experimental design was implemented." (p. 274) Detailed Analysis: Criterion C requires randomisation of entire classes (or stronger, entire schools). The abstract claims students "were randomly assigned" to conditions, i.e., allocation at the individual student level, not at the class level. Moreover, the Method section reveals that the control group was drawn from the 2022 academic year cohort and the experimental group from the 2023 cohort, meaning the two conditions were two different intact course cohorts in different years rather than concurrently randomised classes; no description of a class- or school-level randomisation procedure is given anywhere in the paper. The intervention (anonymous small-group online peer feedback within a writing class) is not personal one-to-one tutoring, so the tutoring exception does not apply. Criterion C is not met because allocation was described at the student level (and in practice by year cohort), with no class-level or school-level randomisation.
    • E

      Exam-based Assessment

      • Outcomes were measured with a course-specific narrative paragraph writing task scored by a researcher-made rubric, not a widely recognised standardised exam.
      • "a narrative paragraph of 240–260 words was used as the first and final draft to test the two groups of participants before (pretest) and after (posttest)" (p. 274)
      • Relevant Quotes: 1) "To compare the improvement in paragraph writing skills between the experimental and control groups (RQ1), a narrative paragraph of 240–260 words was used as the first and final draft to test the two groups of participants before (pretest) and after (posttest) the administration of the multi-level anonymity in asynchronous OPF." (p. 274) 2) "Before administering the task, three experts on the teaching team validated the task's face and construct validity, ensuring the appropriateness and consistency of the test instruction, content, and course objectives." (p. 274) 3) "A detailed rubric with three quality levels focusing on four aspects - topic sentence, supporting details, concluding sentence, and mechanics was used to evaluate the task." (p. 274) 4) "On the teacher's side, the rubric was used and applied by only 1 instructor in order to ensure fairness, consistency, and reliability for all 62 participants." (p. 274) Detailed Analysis: Criterion E requires the outcome to be measured with a standard, widely recognised standardised exam. Here, the outcome was a narrative paragraph writing assignment created for the course, validated informally by "three experts on the teaching team," and scored with a custom rubric applied by a single instructor. No national, state, or internationally recognised standardised test (e.g., IELTS, TOEFL, national curriculum exam) was used. The paper mentions CEFR levels (A2+ to B1) only to describe baseline proficiency of the sample, not as an outcome measure. Criterion E is not met because the outcome measure was a custom, course-specific writing task with a researcher-made rubric rather than a standardised exam.
    • T

      Term Duration

      • The intervention and outcome measurement were completed within a single 3-week instructional unit, far short of one academic term.
      • "This learning unit spanned 3 weeks with 1 3-hour meeting per week." (p. 274)
      • Relevant Quotes: 1) "This 3-week activity started with a first draft submission, peer feedback production via Google Forms, peer feedback evaluation via Google Forms, and then a final draft submission." (p. 274) 2) "This learning unit spanned 3 weeks with 1 3-hour meeting per week." (p. 274) 3) "In week 1, the teacher took measures to ensure fairness in the process by using a pseudonym for each participant on the first draft... After 1.30 hours of in-class writing, the teacher systematically assigned three anonymous peers to provide peer feedback via Google Forms" (p. 274) 4) "Once this feedback process was completed by week 3, the teacher shared the review forms with the file attached for all participants to access and review before the final draft production stage." (p. 275) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. In this study the entire cycle - first draft (pretest), peer feedback, and final draft (posttest) - was completed within a 3-week learning unit. The posttest (final draft) was collected at the end of week 3, only about three weeks after the intervention started. Three weeks is far shorter than an academic term, and no longer-term follow-up measurement is reported. Criterion T is not met because the interval from intervention start to outcome measurement was only about 3 weeks, well under one academic term.
    • D

      Documented Control Group

      • The control group's size, cohort, condition (OPF plus teacher feedback), and baseline and final writing scores are clearly documented, with sample demographics reported.
      • "The control group of 31 participants enrolled in an English reading and writing course in the second semester of academic year 2022 and received multi-level anonymity in asynchronous OPF peer feedback with teacher intervention." (p. 274)
      • Relevant Quotes: 1) "The target sample was 62 first-year English major students (58.06% female, 41.94% male) enrolled in an English reading and writing course in the second semester of the academic years 2022 and 2023. The participants generally obtained basic to intermediate English proficiency, ranging from A2+ to B1, according to the Common European Framework of Reference (CEFR)." (pp. 273-274) 2) "The control group of 31 participants enrolled in an English reading and writing course in the second semester of academic year 2022 and received multi-level anonymity in asynchronous OPF peer feedback with teacher intervention." (p. 274) 3) "Control group (n = 31) First draft 6.63 1.53 ... Final draft 8.12 1.25" (Table 1, p. 276) Detailed Analysis: Criterion D requires that the control group be well documented: who they are, their size, baseline performance, and what they received. The paper specifies the control group size (n = 31), their cohort and course (first-year English majors, second semester of academic year 2022), their general proficiency range (A2+ to B1 CEFR), the overall gender composition of the sample, and exactly what treatment they received (triple-anonymity asynchronous OPF plus teacher feedback). Table 1 reports the control group's baseline (first draft) and outcome (final draft) means and standard deviations. This is comparable in detail to accepted examples of criterion D. A minor weakness is that demographics are reported for the pooled sample rather than per group, but the essential documentation (size, condition, baseline scores) is present. Criterion D is met because the control group's size, composition, baseline performance, and condition are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • The study took place in a single university with allocation at the student/cohort level; no school-level randomisation occurred.
      • "The site of this study was Mae Fah Luang University. The target sample was 62 first-year English major students" (p. 273)
      • Relevant Quotes: 1) "The site of this study was Mae Fah Luang University." (p. 273) 2) "62 first-year English major students from a mid-sized university in Thailand were randomly assigned to a control group of 31" (Abstract, p. 271) Detailed Analysis: Criterion S requires randomisation among schools or equivalent institutional units. This study was conducted at one single university (Mae Fah Luang University), and allocation was described at the individual student level (in practice, two course cohorts in successive academic years). No multiple institutions were involved and no school-level randomisation was performed or described. Criterion S is not met because the trial involved a single institution with student/cohort-level allocation, not school-level randomisation.
    • I

      Independent Conduct

      • The single author designed, taught, administered, and scored the intervention with no external or independent evaluation team mentioned.
      • "On the teacher's side, the rubric was used and applied by only 1 instructor in order to ensure fairness, consistency, and reliability for all 62 participants." (p. 274)
      • Relevant Quotes: 1) "Piansin Pinchai. School of Liberal Arts, Mae Fah Luang University, Thailand. Corresponding author: Piansin Pinchai" (p. 271) 2) "On the teacher's side, the rubric was used and applied by only 1 instructor in order to ensure fairness, consistency, and reliability for all 62 participants." (p. 274) 3) "The teacher explained the task and provided samples of components of the narrative paragraph before assigning participants to then get into small groups of 2–3 to practice writing paragraph outlines and producing the writing." (p. 274) 4) "In week 1, the teacher took measures to ensure fairness in the process by using a pseudonym for each participant on the first draft in order to maintain anonymity." (p. 274) Detailed Analysis: Criterion I requires the study to be conducted independently of the intervention's designers, or at least documented third-party oversight of data collection and analysis. This is a single-author study in which the same person appears to have designed the multi-level anonymity OPF procedure, taught the course, administered the intervention, scored all drafts as the sole instructor-rater, and analysed the data. There is no mention of an external evaluation team, independent test administrators, blinded raters other than the pseudonym procedure for peers, or any third-party oversight. Criterion I is not met because the intervention designer also conducted, scored, and analysed the study without documented independent oversight.
    • Y

      Year Duration

      • The full study lasted only 3 weeks, nowhere near 75% of an academic year, so the year-duration requirement fails (as does its prerequisite T).
      • "This learning unit spanned 3 weeks with 1 3-hour meeting per week." (p. 274)
      • Relevant Quotes: 1) "This learning unit spanned 3 weeks with 1 3-hour meeting per week." (p. 274) 2) "This 3-week activity started with a first draft submission, peer feedback production via Google Forms, peer feedback evaluation via Google Forms, and then a final draft submission." (p. 274) Detailed Analysis: Criterion Y requires outcomes to be measured at least 75% of a full academic year (~9-10 months) after the intervention begins. Here the entire intervention and measurement cycle lasted 3 weeks within one semester unit, with no follow-up beyond the final draft in week 3. Additionally, per the criteria-specific instructions, criterion Y cannot be met when criterion T (Term Duration) is not met, and T failed. Criterion Y is not met because the tracking interval was only 3 weeks, far below 75% of an academic year.
    • B

      Balanced Control Group

      • Both groups received the identical 3-week OPF activity and class time, and the only input difference (presence or absence of teacher feedback) is the explicit treatment variable being tested.
      • "1) to compare the effectiveness of multi-level anonymity in asynchronous OPF with and without teacher feedback intervention" (p. 271)
      • Relevant Quotes: 1) "This study investigates the impact of multi-level anonymity in asynchronous OPF by addressing two key objectives 1) to compare the effectiveness of multi-level anonymity in asynchronous OPF with and without teacher feedback intervention" (Abstract, p. 271) 2) "The control group of 31 participants... received multi-level anonymity in asynchronous OPF peer feedback with teacher intervention. The experimental group of 31 participants... received multi-level anonymity in asynchronous OPF without the teacher's intervention." (p. 274) 3) "To prepare participants in both the experimental and control group, they were all introduced to a narrative paragraph writing lesson." (p. 274) 4) "This learning unit spanned 3 weeks with 1 3-hour meeting per week." (p. 274) 5) "In this study, the participants in both the control and experimental groups were 31 each, with three hours of contact time each week" (p. 274) Detailed Analysis: Applying the criterion B decision tree: both conditions received identical class time, materials, procedure, and rubric (the same 3-week unit, 3 hours of weekly contact time, the same preparatory writing lesson, and the same triple-anonymity asynchronous OPF procedure via Google Forms). The only differing input is that the control group additionally received teacher feedback while the experimental group did not. This asymmetry is not an unplanned confound but precisely the explicit treatment variable stated in the study's first research question ("with and without teacher feedback intervention"), so EXTRA_RESOURCES_PRESENT and RESOURCES_ARE_TREATMENT both hold, which per the decision tree yields "met" regardless of which arm received the extra input. Notably here the *control* arm (not the nominal "experimental" arm) is the one receiving the additional resource (teacher feedback), while the "experimental" arm received strictly fewer inputs (peer feedback alone) - this is the reverse of the typical failure pattern (intervention group gaining an unmatched extra resource) and does not create the bias problem criterion B is designed to catch, since no group received an unaccounted-for advantage favouring a positive result for the tested intervention. A caveat is that the two arms were taught in different academic years (2022 vs 2023), which is a comparability risk for other criteria, but in terms of documented time and resources the conditions were structured identically apart from the tested variable. Criterion B is met because both groups received identical time, structure, and materials, and the sole input difference (teacher feedback) is the explicit treatment variable under test.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study exists; it was published in May 2025, has zero recorded citations as of the verification date, and no external search identifies any replication by another team.
      • Relevant Quotes: 1) "Future studies should confirm the effectiveness of the multi-level anonymity in asynchronous OPF in EFL across diverse classrooms and learners by exploring broader demographic variations." (p. 278) 2) "Second, this study collected data from only one type of writing task, specifically a narrative paragraph. Future studies should include a variety of writing tasks to further explore its effectiveness." (p. 278) Detailed Analysis: Criterion R requires that the study be independently replicated by a different research team in a different context, published in a peer-reviewed journal. The paper itself, published 14 May 2025, explicitly calls for future studies to confirm the effectiveness of the approach, indicating no replication existed at publication. Related literature cited in the paper (e.g., Rod & Nubdal, 2022 on double-blind multiple peer reviews; Awada & Diab, 2023 on OPF versus face-to-face) studies related but distinct interventions and does not replicate this specific triple-anonymity asynchronous OPF trial. As part of this verification, the paper's DOI (10.69598/hasss.25.2.273433) was checked against OpenAlex and Semantic Scholar; both databases report a citation count of 0 as of July 2026, with no citing or related works identified. Given the paper's very recent publication date (14 May 2025), only about 14 months before this check, no independent replication could plausibly have been conducted and published yet. Criterion R is not met because no independent published replication of this specific study could be identified.
    • A

      All-subject Exams

      • Only EFL narrative paragraph writing was assessed, with a custom rubric rather than standardised exams, so all-subject coverage fails (and prerequisite E is unmet).
      • "All participants completed a narrative paragraph writing task, which was used for data collection and analysis." (p. 271)
      • Relevant Quotes: 1) "All participants completed a narrative paragraph writing task, which was used for data collection and analysis." (Abstract, p. 271) 2) "A detailed rubric with three quality levels focusing on four aspects - topic sentence, supporting details, concluding sentence, and mechanics was used to evaluate the task." (p. 274) Detailed Analysis: Criterion A requires standardised exam-based assessment across all main subjects. This study measured only one narrow outcome - English narrative paragraph writing - within a single university course, using a custom rubric. No other subjects were assessed. Furthermore, per the criteria-specific instructions, criterion A cannot be met when criterion E is not met, and E failed because the assessment was custom-made. No justified specialised-intervention exception applies, since even the single assessed subject was not measured with a standardised exam. Criterion A is not met because only one custom writing measure was used and criterion E is unmet.
    • G

      Graduation Tracking

      • Measurement ended with the final draft at week 3 of the unit; no tracking of participants to graduation is reported, and no follow-up publication tracking this cohort was found.
      • "Once this feedback process was completed by week 3, the teacher shared the review forms with the file attached for all participants to access and review before the final draft production stage." (p. 275)
      • Relevant Quotes: 1) "This 3-week activity started with a first draft submission, peer feedback production via Google Forms, peer feedback evaluation via Google Forms, and then a final draft submission." (p. 274) 2) "Once this feedback process was completed by week 3, the teacher shared the review forms with the file attached for all participants to access and review before the final draft production stage." (p. 275) Detailed Analysis: Criterion G requires following participants until graduation from their educational stage. The participants were first-year university students, and all measurement concluded with the final draft at the end of the 3-week unit. The paper reports no follow-up beyond this point. As part of this verification, OpenAlex and Semantic Scholar were checked for later works by the same author (Piansin Pinchai) that might track this cohort toward graduation; the source paper itself has a citation count of 0 and no follow-up publication by this author tracking the same 2022/2023 cohorts could be identified. Additionally, per the criteria-specific instructions, criterion G cannot be met when criterion Y is not met, and Y failed. Criterion G is not met because tracking ended with the final draft in week 3, and no graduation follow-up study exists.
    • P

      Pre-Registered

      • The paper contains no mention of any pre-registered protocol, registry, or registration date, and no external registry record for this study could be found.
      • Relevant Quotes: 1) "Received: 21 October 2024, Revised: 21 March 2025, Accepted: 18 April 2025, Published: 14 May 2025" (p. 271) 2) "In order to set a baseline for evaluation and analysis, a pretest-posttest randomized experimental design was implemented." (p. 274) Detailed Analysis: Criterion P requires that the full study protocol be registered on a public registry before data collection began. The paper contains no reference to any registry (e.g., ClinicalTrials.gov, OSF, AsPredicted, ISRCTN), no registration ID, and no registration date. Data collection occurred during the 2022 and 2023 academic years, and the manuscript was only received in October 2024; nothing indicates any protocol was registered beforehand. A search for a pre-registration record by this author or for this study title did not surface any registry entry, consistent with the paper's silence on this point. Criterion P is not met because no pre-registration statement, registry, or ID appears anywhere in the paper, and no external registry record was found.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.