The Effects of the Guided Dialogic Peer Feedback-Based Writing Instruction on Chinese EFL Students' Writing Performance in an Integrated Blended Learning Environment

Yin Deng, Pragasit Sitthitikul

Published:
ERCT Check Date:
DOI: 10.61508/refl.v32i1.277804
  • L2 languages
  • higher education
  • China
  • blended learning
  • EdTech platform
  • formative assessment
1
  • C

    Two intact classes were randomly assigned as experimental and control groups, so randomisation occurred at the class level rather than among individual students within one class.

    "The participants from two intact classes were randomly assigned as an experimental group (31 students) and a control group (32 students)." (Abstract, p. 1)

  • E

    The pre- and post-tests were taken from the writing sections of past TEM-4, a standardised national English proficiency test for English majors in Chinese universities, rather than a test specially designed for the study.

    "Pre- and post-tests: Those tests were selected from the writing sections of the past TEM-4." (p. 8)

  • T

    The intervention began in Week 1 and outcomes were measured with a post-test in Week 18, an interval of roughly 4.5 months that exceeds one full academic term.

    "Over a period of 18 weeks, the experimental group engaged in guided dialogic peer feedback instruction, whereas the control group was given traditional teaching." (Abstract, p. 1)

  • D

    The control group's size, gender composition, baseline proficiency (Gaokao and TEM-4 scores), and business-as-usual condition are all clearly documented.

    "To be more specific, a class of 31 students (4 males and 27 females) was assigned as the experimental group while a class of 32 (6 males and 26 females) students as the control group." (p. 8)

  • S

    Randomisation involved only two intact classes within a single independent college; no schools or institutions were randomised.

    "The study took place at an independent college in southwestern China, which ranked in the top 50% of such institutions." (p. 7)

  • I

    The authors designed the GDPF instruction, delivered it, collected the data, and the researcher himself served as one of the essay raters, with no independent third-party evaluation reported.

    "Therefore, the current study analyzed scores from two other writing teachers and the researcher himself." (p. 8)

  • Y

    Outcomes were measured 18 weeks (about 4.5 months) after the intervention began, which is well short of 75% of a full academic year.

    "On Week 18, 60-minute post-tests were administered to both groups and the scores were collected." (p. 11)

  • B

    Both groups received identical in-class instruction, wrote the same essays on the same schedule, and the additional online dialogic peer feedback activity in the experimental group is itself the treatment variable being tested rather than a separable resource add-on.

    "To sum up, all students from both groups received the same in-class instruction. In terms of roles, students in the control group functioned as writers only while students in the experimental group functioned as both writers and reviewers." (p. 11)

  • R

    No independent replication of this specific GDPF study exists; the paper was published in late 2024/early 2025 and an external search found no peer-reviewed replication by a different team.

  • A

    Only argumentative writing in English was assessed; no other main subjects were measured and the paper offers no explicit rationale for a specialised-intervention exception.

    "Both groups were assessed through writing assignments on two different topics, administered as a pre-test and a post-test." (Abstract, p. 1)

  • G

    Measurement stopped at the Week 18 post-test with no follow-up tracking of students to graduation, and criterion Y is not met, which also rules out G.

    "On Week 18, 60-minute post-tests were administered to both groups and the scores were collected." (p. 11)

  • P

    The paper contains no mention of any trial registry, pre-registration ID, or pre-registered protocol, and no registration for this study was found in an internet search.

Abstract

While extensive research exists on peer feedback and its effects on writing, there are few experimental studies that rigorously investigate the effects of guided dialogic peer feedback on students' argumentative writing performance. This study, adopting a mixed-methods approach, examined the influence of guided dialogic peer feedback within a blended learning setting on the writing performance of 63 university students studying English as a second language (L2). The participants from two intact classes were randomly assigned as an experimental group (31 students) and a control group (32 students). Over a period of 18 weeks, the experimental group engaged in guided dialogic peer feedback instruction, whereas the control group was given traditional teaching. Both groups were assessed through writing assignments on two different topics, administered as a pre-test and a post-test. Additionally, students from the experimental group filled out a questionnaire and some engaged in semi-structured interviews after the treatment. The study's findings, derived from a series of t-tests and ANCOVA, revealed that dialogic peer feedback significantly enhanced the students' writing performance. Furthermore, the questionnaire responses indicated a positive student perception towards this instructional approach, a sentiment echoed in the interview analyses.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Two intact classes were randomly assigned as experimental and control groups, so randomisation occurred at the class level rather than among individual students within one class.
      • "The participants from two intact classes were randomly assigned as an experimental group (31 students) and a control group (32 students)." (Abstract, p. 1)
      • Relevant Quotes: 1) "The participants from two intact classes were randomly assigned as an experimental group (31 students) and a control group (32 students)." (Abstract, p. 1) 2) "The participants in the study came from two intact classes of the English Department, aging from 19 to 21." (p. 8) 3) "To be more specific, a class of 31 students (4 males and 27 females) was assigned as the experimental group while a class of 32 (6 males and 26 females) students as the control group." (p. 8) Detailed Analysis: Criterion C requires that randomisation be conducted at the class level (or stronger) so that treatment and control groups are properly isolated. The abstract explicitly states that the two intact classes were "randomly assigned" to the experimental and control conditions, meaning the unit of assignment was the whole class, not individual students within a single classroom. This design avoids within-class contamination between treated and untreated students. The description of the randomisation mechanism is brief (no detail on how the random allocation of the two classes was performed), and with only two clusters the randomisation is minimal, but the stated unit of randomisation satisfies the class-level requirement of the standard. Criterion C is met because entire intact classes, not individual students within one class, were randomly assigned to the experimental and control conditions.
    • E

      Exam-based Assessment

      • The pre- and post-tests were taken from the writing sections of past TEM-4, a standardised national English proficiency test for English majors in Chinese universities, rather than a test specially designed for the study.
      • "Pre- and post-tests: Those tests were selected from the writing sections of the past TEM-4." (p. 8)
      • Relevant Quotes: 1) "Pre- and post-tests: Those tests were selected from the writing sections of the past TEM-4." (p. 8) 2) "...their scores of Gaokao and TEM-4 (Test for English Majors Band 4, a standardized English proficiency test for English majors in Chinese universities)..." (p. 8) 3) "Both writing topics, online education and music downloads, were not covered during the writing course." (p. 8) 4) "Their essays were graded with the analytic scoring rubric developed by McDonough et al. (2018) and scores were collected as data." (p. 9) 5) "The results showed that ICC values for pre- and post-tests were 0.806 and 0.895, indicating the scores given by the research were reliable and could be collected as data for later analysis." (p. 8) Detailed Analysis: Criterion E requires that outcomes be measured with a standard, widely recognised exam rather than a custom test built for the study. The outcome tasks were drawn directly from the writing sections of past TEM-4 papers, and the paper itself identifies TEM-4 as "a standardized English proficiency test for English majors in Chinese universities." The topics were deliberately not covered in the course, reducing alignment bias. A caveat is that scoring was done by the researcher and two other writing teachers using a published analytic rubric (McDonough et al., 2018) rather than through official TEM-4 grading, though inter-rater reliability was verified (ICC 0.806 and 0.895). Because the assessment instrument itself originates from a recognised standardised national exam and was not specially constructed for the study, the core requirement of the criterion is satisfied. Criterion E is met because the writing tests were taken from the writing sections of the standardised national TEM-4 exam rather than being custom-designed for the study.
    • T

      Term Duration

      • The intervention began in Week 1 and outcomes were measured with a post-test in Week 18, an interval of roughly 4.5 months that exceeds one full academic term.
      • "Over a period of 18 weeks, the experimental group engaged in guided dialogic peer feedback instruction, whereas the control group was given traditional teaching." (Abstract, p. 1)
      • Relevant Quotes: 1) "Over a period of 18 weeks, the experimental group engaged in guided dialogic peer feedback instruction, whereas the control group was given traditional teaching." (Abstract, p. 1) 2) "On Week 1, a 30-minute brief introduction to the study was given to both groups... A 60-minute pre-test was then conducted to both groups." (p. 9) 3) "On Week 18, 60-minute post-tests were administered to both groups and the scores were collected." (p. 11) 4) "A treatment spanning 18 weeks was carried out." (p. 3) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. Here the intervention started at the beginning of the semester (Week 1 introduction and pre-test, followed by the writing cycles) and the post-test was administered in Week 18. Eighteen weeks corresponds to about 4.5 months, which covers a full university semester/term in the Chinese higher education context. The interval from intervention start to outcome measurement therefore meets or exceeds the one-term minimum. Criterion T is met because the interval from intervention start (Week 1) to outcome measurement (Week 18) spans a full 18-week semester, at least one academic term.
    • D

      Documented Control Group

      • The control group's size, gender composition, baseline proficiency (Gaokao and TEM-4 scores), and business-as-usual condition are all clearly documented.
      • "To be more specific, a class of 31 students (4 males and 27 females) was assigned as the experimental group while a class of 32 (6 males and 26 females) students as the control group." (p. 8)
      • Relevant Quotes: 1) "To be more specific, a class of 31 students (4 males and 27 females) was assigned as the experimental group while a class of 32 (6 males and 26 females) students as the control group." (p. 8) 2) "Students from both groups were considered to have similar English proficiency as no significant differences were found in terms of their scores of Gaokao and TEM-4 (Test for English Majors Band 4...) with p values of .486 and .111 respectively." (p. 8) 3) "The participants in the study came from two intact classes of the English Department, aging from 19 to 21. None of them had received training in peer feedback." (p. 8) 4) "During the treatment, all participants from both groups were required to follow the same in-class activities and write 4 argumentative essays on 4 controversial topics..." (p. 9) 5) "To sum up, all students from both groups received the same in-class instruction. In terms of roles, students in the control group functioned as writers only while students in the experimental group functioned as both writers and reviewers." (p. 11) 6) "An independent t-test was conducted to compare the pre-test scores between the experimental and control groups. The analysis revealed a statistically significant difference (p = 0.001, t = 3.643, mean difference = 0.405)..." (p. 13) Detailed Analysis: Criterion D requires clear documentation of the control group's composition, baseline performance, and the treatment it received. The paper reports the control group's size (32), gender breakdown (6 males, 26 females), age range (19-21), lack of prior peer feedback training, and baseline English proficiency evidence (comparable Gaokao and TEM-4 scores with p values reported). Pre-test writing scores for the control group were also measured and reported, and the paper is explicit that the control group followed the same in-class activities and traditional teaching, functioning as "writers only" with no peer feedback component. A baseline difference in pre-test writing scores was detected and transparently reported, then adjusted for with ANCOVA. This constitutes adequate documentation for comparison purposes. Criterion D is met because the control group's size, demographics, baseline proficiency, and business-as-usual condition are described in detail.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation involved only two intact classes within a single independent college; no schools or institutions were randomised.
      • "The study took place at an independent college in southwestern China, which ranked in the top 50% of such institutions." (p. 7)
      • Relevant Quotes: 1) "The study took place at an independent college in southwestern China, which ranked in the top 50% of such institutions." (p. 7) 2) "The participants from two intact classes were randomly assigned as an experimental group (31 students) and a control group (32 students)." (Abstract, p. 1) 3) "The participants in the study came from two intact classes of the English Department..." (p. 8) Detailed Analysis: Criterion S requires that randomisation occur at the level of schools or comparable institutional units. This study was conducted entirely within one independent college in southwestern China, and the unit of random assignment was the intact class (two classes of second-year English majors). No multiple institutions were involved and no school-level randomisation took place, so the design does not reach the school-level standard. Criterion S is not met because randomisation was performed on two classes within a single college, not among schools or institutions.
    • I

      Independent Conduct

      • The authors designed the GDPF instruction, delivered it, collected the data, and the researcher himself served as one of the essay raters, with no independent third-party evaluation reported.
      • "Therefore, the current study analyzed scores from two other writing teachers and the researcher himself." (p. 8)
      • Relevant Quotes: 1) "The treatment, namely the GDPF, was adapted from the model proposed by Yoon (2011)." (p. 8) 2) "Therefore, the current study analyzed scores from two other writing teachers and the researcher himself." (p. 8) 3) "Students were requied to exchange feedback online following strict, step-by-step guidance provided by the instructor." (p. 3) 4) "By observing relevant studies, the researcher found a gap in research." (p. 7) 5) "In terms of the guided dialogic peer feedback, students from the experimental group were asked to provide feedback to each draft under the guidance framework (see Table 1) adapted by Nelson and Schunn (2009) and Er et al. (2021)." (p. 11) Detailed Analysis: Criterion I requires that the study be conducted independently from the designers of the intervention, or at least that data collection and analysis involve documented third-party oversight. In this study the same author/researcher adapted the GDPF instruction, implemented it in his own teaching context, administered the tests, conducted the interviews, and participated directly in scoring the essays ("the researcher himself" was one of the three raters). Although two other writing teachers also rated the essays and inter-rater reliability was checked, this is internal quality control, not independent conduct by an external evaluation team. No statement of independent or third-party oversight of data collection, analysis, or conclusions appears anywhere in the paper. Criterion I is not met because the intervention designers themselves conducted the study and the researcher participated in scoring, with no independent evaluation team.
    • Y

      Year Duration

      • Outcomes were measured 18 weeks (about 4.5 months) after the intervention began, which is well short of 75% of a full academic year.
      • "On Week 18, 60-minute post-tests were administered to both groups and the scores were collected." (p. 11)
      • Relevant Quotes: 1) "A treatment spanning 18 weeks was carried out." (p. 3) 2) "On Week 18, 60-minute post-tests were administered to both groups and the scores were collected." (p. 11) 3) "Secondly, the research design implemented in this study facilitated the observation of the effects of guided dialogic peer feedback within a constrained time frame. Future research could extend this temporal scope to assess the sustained effects of dialogic peer feedback on L2 writing proficiency over a longer duration." (p. 20) Detailed Analysis: Criterion Y requires outcome measurement at least 75% of a full academic year (roughly 9-10 months, i.e., about 7 or more months) after the intervention begins. Here the entire study, from the Week 1 introduction and pre-test to the Week 18 post-test, covered a single 18-week semester, about 4.5 months. This is only about half of a typical academic year and below the 75% threshold. The authors themselves acknowledge the "constrained time frame" as a limitation and call for longer-duration follow-up in future research. Criterion Y is not met because the 18-week tracking period falls well short of 75% of an academic year.
    • B

      Balanced Control Group

      • Both groups received identical in-class instruction, wrote the same essays on the same schedule, and the additional online dialogic peer feedback activity in the experimental group is itself the treatment variable being tested rather than a separable resource add-on.
      • "To sum up, all students from both groups received the same in-class instruction. In terms of roles, students in the control group functioned as writers only while students in the experimental group functioned as both writers and reviewers." (p. 11)
      • Relevant Quotes: 1) "During the treatment, all participants from both groups were required to follow the same in-class activities and write 4 argumentative essays on 4 controversial topics..." (p. 9) 2) "A total of 63 Students from both groups produced 3 drafts over a writing cycle which consisted of 4 weeks with one class of 90 minutes per week, except that students in the experimental group needed to upload their writings online and conduct dialogic peer feedback in the process as well." (p. 10) 3) "To sum up, all students from both groups received the same in-class instruction. In terms of roles, students in the control group functioned as writers only while students in the experimental group functioned as both writers and reviewers." (p. 11) 4) "On Week 1, a 30-minute brief introduction to the study was given to both groups and the experimental group also received additional training on the dialogic peer feedback as well as the scoring rubric on the writing in order to ensure consensus among all participants." (p. 9) 5) "The current study set out to investigate the effects of a teaching instruction that incorporates blended learning and guided dialogic peer feedback (GDPF)." (p. 3) Detailed Analysis: Criterion B asks whether time and resources were balanced between conditions, unless the additional inputs are integral to the treatment being tested. Both groups had identical in-class time (one 90-minute class per week), identical writing cycles, identical topics, and the same number of drafts and teacher evaluations. The extra elements received by the experimental group were: (a) initial training on dialogic peer feedback and the scoring rubric, (b) access to the Kdocs online platform, and (c) the time spent giving and receiving dialogic peer feedback online (and in follow-up offline discussions). Applying the decision tree: extra engagement time exists, but these inputs are exactly the guided dialogic peer feedback intervention that the study is explicitly designed to test - the research question is whether GDPF itself improves writing performance. The feedback activity, its training, and the platform are integral components of the treatment package rather than separable, confounding add-ons, and the control group received the full business-as-usual instruction with the same classroom dosage. Per the standard's exception, when the additional activity is the primary treatment variable, a business-as-usual control is acceptable. It should be clearly noted, however, that experimental students did spend additional out-of-class time on reviewing peers' work, which is inherent to peer feedback interventions. This assessment was re-checked against the current ERCT specification's criterion B decision tree (extra resources present, but they are the treatment variable being tested), and the conclusion is unchanged. Criterion B is met because in-class instruction and writing tasks were identical across groups and the additional peer feedback activity, training, and platform are integral to the GDPF treatment being tested.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific GDPF study exists; the paper was published in late 2024/early 2025 and an external search found no peer-reviewed replication by a different team.
      • Relevant Quotes: 1) "While extensive research exists on peer feedback and its effects on writing, there are few experimental studies that rigorously investigate the effects of guided dialogic peer feedback on students' argumentative writing performance." (Abstract, p. 1) 2) "This assertion opens up avenues for future research to explore the enduring effects and potential adaptability of the GDPF in diverse educational settings." (p. 20) 3) "...there seems to be an absence of empirical studies that justify the employment of guided dialogic peer feedback in blended courses in teaching of writing within the Chinese context. To fill this research gap..." (p. 3) Detailed Analysis: Criterion R requires independent replication of the study by a different research team in a different context, published in a peer-reviewed journal. The paper positions itself as filling a research gap, explicitly stating there is an absence of comparable empirical studies, and invites future research to test the GDPF in other settings. Related prior studies cited in the paper (e.g., Wood, 2021; Ng & Yu, 2021; Noroozi et al., 2020) investigate dialogic or online peer feedback generally but are not replications of this specific GDPF intervention and design. A fresh internet search (repeated during verification) for independent replications of this study, its authors, or the specific GDPF design found only the original publication, its abstract mirrored on ResearchGate, and general peer feedback literature; no peer-reviewed independent replication of this trial was identified, which is unsurprising given its very recent publication (available online 27 December 2024, published in rEFLections 32(1), January-April 2025). Criterion R is not met because no independent peer-reviewed replication of this specific study could be found.
    • A

      All-subject Exams

      • Only argumentative writing in English was assessed; no other main subjects were measured and the paper offers no explicit rationale for a specialised-intervention exception.
      • "Both groups were assessed through writing assignments on two different topics, administered as a pre-test and a post-test." (Abstract, p. 1)
      • Relevant Quotes: 1) "Both groups were assessed through writing assignments on two different topics, administered as a pre-test and a post-test." (Abstract, p. 1) 2) "The first research question concerning the effects of the instruction on students' writing performance were responded by analyzing pre- and post-tests results..." (p. 8) 3) "The English program covers comprehensive language skills, Western cultures, and foundational translation knowledge." (p. 7) Detailed Analysis: Criterion A requires standardised exam-based assessment across all main subjects taught at the educational level, or a clearly justified exception for highly specialised interventions. This study measured only English argumentative writing performance. The English major curriculum described in the paper includes comprehensive language skills, Western cultures, and translation, none of which were assessed. Even within English proficiency, only the writing skill was measured. Although one could argue the intervention targets a specialised university writing course, the paper provides no explicit rationale of the kind the exception requires, and the exception is framed for upper secondary or vocational education. Effects on other subjects or skills (positive or negative) therefore remain unmeasured. Criterion A is not met because only writing in English was assessed, with no coverage of other main subjects and no quoted justification for an exception.
    • G

      Graduation Tracking

      • Measurement stopped at the Week 18 post-test with no follow-up tracking of students to graduation, and criterion Y is not met, which also rules out G.
      • "On Week 18, 60-minute post-tests were administered to both groups and the scores were collected." (p. 11)
      • Relevant Quotes: 1) "On Week 18, 60-minute post-tests were administered to both groups and the scores were collected." (p. 11) 2) "Future research could extend this temporal scope to assess the sustained effects of dialogic peer feedback on L2 writing proficiency over a longer duration." (p. 20) 3) "This assertion opens up avenues for future research to explore the enduring effects and potential adaptability of the GDPF in diverse educational settings." (p. 20) Detailed Analysis: Criterion G requires that participants be tracked until graduation from their educational stage. The participants were second-year university students, and all outcome measurement ended with the post-test in Week 18 of the same semester. The authors explicitly identify the constrained time frame as a limitation and defer longitudinal follow-up to future research, indicating no continued tracking of this cohort through to graduation. A search for follow-up publications by the same authors tracking this cohort (e.g., by Yin Deng or Pragasit Sitthitikul) found no such papers, as this study appears to be a recent, standalone publication. Additionally, per the ranking rules, criterion G cannot be met when criterion Y is not met, and Y fails here. Criterion G is not met because tracking ended at the Week 18 post-test, far short of graduation, no follow-up publications were found, and criterion Y is not met.
    • P

      Pre-Registered

      • The paper contains no mention of any trial registry, pre-registration ID, or pre-registered protocol, and no registration for this study was found in an internet search.
      • Relevant Quotes: 1) "This was a mixed-methods approach which employed an embedded experimental design (Creswell & Plano Clark, 2011)." (p. 7) 2) "Article history: Received: 1 Jan 2024, Accepted: 21 Dec 2024, Available online: 27 Dec 2024" (p. 1) Detailed Analysis: Criterion P requires that the full study protocol, including hypotheses, methods, and planned analyses, be registered on a public registry before data collection began. The paper describes its mixed-methods design and analysis procedures but contains no reference to any registry platform (e.g., ClinicalTrials.gov, OSF, AsPredicted, ISRCTN), no registration ID, and no statement about a pre-registered protocol or analysis plan anywhere in the text. An internet search for a pre-registration record by these authors or for this specific GDPF study found no matching registration on OSF, AsPredicted, or similar registries. Without any quoted evidence of pre-registration and its timing, and no external registration record found, the criterion cannot be satisfied. Criterion P is not met because no pre-registration statement, registry ID, or protocol reference appears in the paper, and no external registration record was found online.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.