Integrating Structured Reflection into Process Writing: Enhancing Metacognitive Engagement and Writing Performance in Collaborative and Individual ESL Contexts

Kuldip Kaur Maktiar Singh, Irene Yoke Chu Leong, Chun Keat Yeap

Published:
ERCT Check Date:
DOI: 10.51584/IJRIAS.2025.100600103
  • language arts
  • L2 languages
  • higher education
  • Asia
0
  • C

    Randomisation was carried out at the individual student level within a single writing course, not at the class or school level, and no personal-tutoring exception applies.

    "Participants were randomly assigned to either a collaborative writing group (n = 30) or an individual writing group (n = 30)."

  • E

    Writing quality and reflection depth were scored by the researchers' own trained raters using researcher-adapted rubrics, not a widely recognised standardised exam.

    "Writing Quality was assessed using CEFR-aligned rubrics covering five traits: content, coherence, range, accuracy, and mechanics."

  • T

    The intervention spanned only a short, explicitly acknowledged "short intervention period" covering a single three-stage writing cycle, far below one academic term.

    "The relatively small sample size and the short intervention period limit the generalizability of the results."

  • D

    The comparison (individual) group's size, structure, and baseline (prewriting) performance are explicitly documented and shown to be statistically comparable to the collaborative group.

    "Results showed comparable mean scores between the groups (Collaborative: M = 64.06, SD = 4.50; Individual: M = 64.40, SD = 4.66), supporting valid comparisons in subsequent stages."

  • S

    The study was conducted with one cohort in one writing course at a single university, with no school-level randomisation.

    "The study involved 60 undergraduate ESL students enrolled in a writing course at a Malaysian university."

  • I

    There is no indication of an independent, third-party team conducting or overseeing the study; the same institution's researchers appear to have designed, run, and evaluated the intervention.

    "Two trained raters, both experienced English language educators, independently scored all reflections."

  • Y

    Because criterion T (Term Duration) is not met, and the study is explicitly a short, single writing-cycle intervention, the year-long duration requirement is not met.

    "The relatively small sample size and the short intervention period limit the generalizability of the results."

  • B

    Both arms received identical structured-reflection scaffolding at each writing stage; the only manipulated variable was the collaborative-versus- individual structure itself, which is the intended treatment contrast, so no extra time or budget was given to either arm alone.

    "At each stage, students were provided with structured reflection prompts grounded in Moon's Reflective Thinking Model."

  • R

    No independent replication of this study is reported or plausible, given its very recent publication date, and a dedicated internet search found no such replication.

  • A

    Because criterion E is not met, criterion A cannot be met; additionally, only ESL writing performance was assessed, not all core subjects.

    "Writing Quality was assessed using CEFR-aligned rubrics covering five traits: content, coherence, range, accuracy, and mechanics."

  • G

    Because criterion Y is not met, criterion G cannot be met; the study also explicitly stops at the end of the short intervention with no further follow-up, and no subsequent tracking papers by these authors were found.

    "Future studies could adopt a longitudinal design and expand to different ESL populations to validate these outcomes over time."

  • P

    The paper describes institutional ethical approval but makes no mention of pre-registering the study protocol, hypotheses, or analysis plan before data collection, and no registry entry was found via internet search.

Abstract

This study investigates how structured reflection embedded within process-oriented writing improves ESL learners' writing quality and metacognitive engagement in both collaborative and individual contexts. Drawing upon Murray's Process Writing Model and Moon's Reflective Thinking Model, a mixed-methods study involving 60 Malaysian undergraduates revealed that collaborative learners consistently outperformed their individual peers in writing traits and reflective depth. CEFR-aligned assessments and thematic analysis of journals and surveys underscored that structured reflection heightened writing awareness, purposeful revision, and metacognitive growth. Collaborative settings further amplified dialogic reflection through peer interaction. The findings support embedding structured reflection within the process writing cycle to enhance ESL writing outcomes and recommend its pedagogical adoption.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was carried out at the individual student level within a single writing course, not at the class or school level, and no personal-tutoring exception applies.
      • "Participants were randomly assigned to either a collaborative writing group (n = 30) or an individual writing group (n = 30)."
      • Relevant Quotes: 1) "The study involved 60 undergraduate ESL students enrolled in a writing course at a Malaysian university." (p. 1350) 2) "Participants were randomly assigned to either a collaborative writing group (n = 30) or an individual writing group (n = 30)." (p. 1350) 3) "The collaborative group worked in triads, jointly composing texts and submitting group reflections at each stage." (p. 1350) Detailed Analysis: The paper describes a single cohort of 60 undergraduates in one writing course at one university, with individual students randomly assigned to either the collaborative (then grouped into triads) or individual condition. There is no description of randomising intact classes or schools; the unit of randomisation is the student, and both conditions appear to run concurrently within the same course context, creating a risk of contamination between students in the two arms. The intervention is a process-writing/reflection pedagogy, not one-to-one tutoring, so the tutoring exception does not apply. All three quotes above were checked against the source PDF and are verbatim. Criterion C is not met because randomisation occurred at the individual student level within one course, not at the class or school level.
    • E

      Exam-based Assessment

      • Writing quality and reflection depth were scored by the researchers' own trained raters using researcher-adapted rubrics, not a widely recognised standardised exam.
      • "Writing Quality was assessed using CEFR-aligned rubrics covering five traits: content, coherence, range, accuracy, and mechanics."
      • Relevant Quotes: 1) "Writing Quality was assessed using CEFR-aligned rubrics covering five traits: content, coherence, range, accuracy, and mechanics. These rubrics ensured objective and standardized evaluation of learners' writing performance." (p. 1350) 2) "Reflection Depth was evaluated using a 5-point rubric adapted from Moon's Reflective Thinking Model. Scores ranged from 1 (descriptive recall) to 5 (critical reflection), with criteria based on indicators such as self-awareness, reasoning, and goal-setting." (p. 1351) 3) "Two trained raters, both experienced English language educators, independently scored all reflections. Inter-rater reliability was confirmed using Cohen's Kappa (κ = 0.87), indicating high consistency." (p. 1351) Detailed Analysis: Both outcome instruments are researcher-constructed for this study: writing quality was rated by two trained raters against a "CEFR-aligned" analytic rubric applied to students' own in-course essays, and reflection depth was scored using a bespoke 5-point rubric "adapted from Moon's Reflective Thinking Model." Neither instrument is a widely recognised, externally administered standardised exam (e.g., a national curriculum test or an established international writing test); rather, they are rater-applied scoring frameworks built for this research. Referencing the CEFR framework does not by itself make the rubric a standardised exam, since the actual instrument and its administration were designed and applied by the study team. All quotes above were checked and are verbatim from the source. Criterion E is not met because the study relies on researcher-adapted rubrics scored by the authors' own raters, not a standardised exam.
    • T

      Term Duration

      • The intervention spanned only a short, explicitly acknowledged "short intervention period" covering a single three-stage writing cycle, far below one academic term.
      • "The relatively small sample size and the short intervention period limit the generalizability of the results."
      • Relevant Quotes: 1) "The study design was based on Murray's Process Writing Model, which guided learners through three iterative stages: prewriting, drafting, and revising." (p. 1350) 2) "Despite the robust findings, this study has several limitations. The relatively small sample size and the short intervention period limit the generalizability of the results." (p. 1354) 3) "Future studies could adopt a longitudinal design and expand to different ESL populations to validate these outcomes over time." (p. 1354) Detailed Analysis: No explicit calendar start/end dates are given, but the intervention is described only as a single writing-course cycle through prewriting, drafting, and revising, with post-intervention writing scores, reflection scores, and a survey collected once at the end. The authors themselves characterise this as a "short intervention period" and call for future "longitudinal design," confirming the tracking window was well short of a full academic term. All quotes above were checked and are verbatim from the source. Criterion T is not met because the study explicitly describes a short, single-cycle intervention rather than term-long tracking.
    • D

      Documented Control Group

      • The comparison (individual) group's size, structure, and baseline (prewriting) performance are explicitly documented and shown to be statistically comparable to the collaborative group.
      • "Results showed comparable mean scores between the groups (Collaborative: M = 64.06, SD = 4.50; Individual: M = 64.40, SD = 4.66), supporting valid comparisons in subsequent stages."
      • Relevant Quotes: 1) "To ensure baseline equivalence, all students completed a prewriting task assessed using CEFR-aligned criteria. Results showed comparable mean scores between the groups (Collaborative: M = 64.06, SD = 4.50; Individual: M = 64.40, SD = 4.66), supporting valid comparisons in subsequent stages." (p. 1350) 2) "Table 2. Prewriting Scores of Collaborative and Individual Groups" listing Group, Number of Participants, Structure, Mean Score, SD: "Individual 30 Individual participants 64.40 4.66" (Table 2, p. 1350) 3) "The individual group completed all writing and reflections independently." (p. 1350) Detailed Analysis: Although the design compares two active conditions rather than an untreated control, the paper clearly documents the comparison (individual) arm's sample size (n = 30), its structure ("individual participants," each completing writing and reflections independently), and its baseline prewriting performance (mean and SD), explicitly confirming statistical equivalence with the collaborative arm prior to the intervention. This level of detail satisfies the intent of criterion D to document the comparison group's characteristics and baseline for proper comparison. All quotes above were checked and are verbatim from the source. Criterion D is met because the comparison group's size, structure, and baseline scores are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • The study was conducted with one cohort in one writing course at a single university, with no school-level randomisation.
      • "The study involved 60 undergraduate ESL students enrolled in a writing course at a Malaysian university."
      • Relevant Quotes: 1) "The study involved 60 undergraduate ESL students enrolled in a writing course at a Malaysian university." (p. 1350) 2) "Participants were randomly assigned to either a collaborative writing group (n = 30) or an individual writing group (n = 30)." (p. 1350) Detailed Analysis: There is no mention of multiple schools or institutions being sampled or randomised; the entire study took place within one writing course at one Malaysian university, with individual students (not schools) as the unit of assignment. Both quotes above were checked and are verbatim from the source. Criterion S is not met because randomisation occurred among individual students in a single course at one institution, not among schools.
    • I

      Independent Conduct

      • There is no indication of an independent, third-party team conducting or overseeing the study; the same institution's researchers appear to have designed, run, and evaluated the intervention.
      • "Two trained raters, both experienced English language educators, independently scored all reflections."
      • Relevant Quotes: 1) "Two trained raters, both experienced English language educators, independently scored all reflections. Inter-rater reliability was confirmed using Cohen's Kappa (κ = 0.87), indicating high consistency." (p. 1351) 2) Biodata: "Irene Leong is currently an Associate Professor at the Academy of Language Studies in University Technology MARA, Melaka, Malaysia... Yeap Chun Keat is a Senior Lecturer in the Academy of Language Studies, Universiti Teknologi MARA..." (p. 1356) 3) No statement anywhere in the Methodology, Ethical Approval, or Acknowledgements sections describes an external evaluation agency conducting data collection or analysis independent of the authors. Detailed Analysis: The "two trained raters" quote refers only to inter-rater reliability of scoring reflections, not to organisational independence from the intervention's designers. All three authors are faculty at the same Academy of Language Studies, Universiti Teknologi MARA, and there is no mention of an external or third-party team conducting the data collection or analysis independently of the people who designed the study. All quotes above were checked and are verbatim from the source (biodata quote combines two adjacent biodata entries with an ellipsis). Criterion I is not met because no independent, third-party conduct of the study is documented.
    • Y

      Year Duration

      • Because criterion T (Term Duration) is not met, and the study is explicitly a short, single writing-cycle intervention, the year-long duration requirement is not met.
      • "The relatively small sample size and the short intervention period limit the generalizability of the results."
      • Relevant Quotes: 1) "The relatively small sample size and the short intervention period limit the generalizability of the results." (p. 1354) 2) "Future studies could adopt a longitudinal design and expand to different ESL populations to validate these outcomes over time." (p. 1354) Detailed Analysis: Per the ERCT specific instruction for criterion Y, if criterion T is not met, Y cannot be met either. The study also provides no evidence of tracking spanning anywhere near 75% of an academic year; it is explicitly described as a short intervention with only a single post-intervention measurement point, and the authors call for future longitudinal designs, confirming none was used here. Both quotes above were checked and are verbatim from the source. Criterion Y is not met because the study duration was short and criterion T was not satisfied.
    • B

      Balanced Control Group

      • Both arms received identical structured-reflection scaffolding at each writing stage; the only manipulated variable was the collaborative-versus- individual structure itself, which is the intended treatment contrast, so no extra time or budget was given to either arm alone.
      • "At each stage, students were provided with structured reflection prompts grounded in Moon's Reflective Thinking Model."
      • Relevant Quotes: 1) "At each stage, students were provided with structured reflection prompts grounded in Moon's Reflective Thinking Model. These prompts encouraged engagement across multiple levels of reflection — descriptive, dialogic, and critical." (p. 1350) 2) "The collaborative group worked in triads, jointly composing texts and submitting group reflections at each stage." (p. 1350) 3) "The individual group completed all writing and reflections independently." (p. 1350) 4) "To ensure baseline equivalence, all students completed a prewriting task assessed using CEFR-aligned criteria. Results showed comparable mean scores between the groups..." (p. 1350) Detailed Analysis (re-checked against the current criterion B decision tree): Step 0-1, EXTRA_RESOURCES_PRESENT: both arms move through the same three writing stages (prewriting, drafting, revising) with the same structured reflection prompts derived from the same theoretical models, and neither arm is described as receiving additional class time, budget, or materials that the other lacked. Since no extra time or budget is present in either condition relative to the other, this branch resolves to met() without needing to reach the "resources are the treatment" or "control matches resources" branches. The only difference between arms is the social structure of the task (working in triads versus working alone), which is the explicit collaborative-versus-individual contrast the study was designed to test, not a separable resource imbalance. All quotes above were checked and are verbatim from the source. Criterion B is met because the two conditions receive identical reflective scaffolding, time, and materials, differing only in the collaborative-versus-individual structure that is the explicit object of study.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this study is reported or plausible, given its very recent publication date, and a dedicated internet search found no such replication.
      • Relevant Quotes: 1) No quotes in the paper reference a prior or independent replication of this specific intervention or design. 2) "Received: 17 June 2025; Accepted: 19 June 2025; Published: 15 July 2025" (p. 1346) Detailed Analysis: The paper contains no mention of any prior or external attempt to replicate its specific collaborative-versus-individual structured-reflection design. Additionally, given the paper's very recent publication (July 2025), it is not plausible that an independent replication by a different research team has yet appeared in a peer-reviewed outlet. An internet search (web search on the paper title, authors, and design terms) was conducted specifically to check for independent replications; no such papers were found, and none are claimed to exist. Both quotes above were checked and are verbatim from the source. Criterion R is not met because no independent replication of this study exists or was found.
    • A

      All-subject Exams

      • Because criterion E is not met, criterion A cannot be met; additionally, only ESL writing performance was assessed, not all core subjects.
      • "Writing Quality was assessed using CEFR-aligned rubrics covering five traits: content, coherence, range, accuracy, and mechanics."
      • Relevant Quotes: 1) "Writing Quality was assessed using CEFR-aligned rubrics covering five traits: content, coherence, range, accuracy, and mechanics." (p. 1350) 2) No quotes anywhere in the paper describe assessment of any subject other than ESL writing (e.g., mathematics, science, or other core subjects). Detailed Analysis: Per the ERCT specific instruction for criterion A, since criterion E (Exam-based Assessment) is not met, criterion A cannot be met. Independently, the study measures only ESL writing quality and reflective depth; it does not assess outcomes in any other main subject area, nor does it provide a justification for restricting outcomes to a single specialised domain in the way the standard's exception envisions. The quote above was checked and is verbatim from the source. Criterion A is not met because criterion E is not met and only a single subject (ESL writing) was assessed.
    • G

      Graduation Tracking

      • Because criterion Y is not met, criterion G cannot be met; the study also explicitly stops at the end of the short intervention with no further follow-up, and no subsequent tracking papers by these authors were found.
      • "Future studies could adopt a longitudinal design and expand to different ESL populations to validate these outcomes over time."
      • Relevant Quotes: 1) "Despite the robust findings, this study has several limitations. The relatively small sample size and the short intervention period limit the generalizability of the results." (p. 1354) 2) "Future studies could adopt a longitudinal design and expand to different ESL populations to validate these outcomes over time." (p. 1354) 3) No mention anywhere in the paper of tracking participants beyond the immediate post-intervention writing scores, reflection scores, and survey. Detailed Analysis: Per the ERCT specific instruction for criterion G, since criterion Y is not met, G cannot be met. Additionally, all data collection (writing quality, reflection depth, post-intervention survey) occurred once, shortly after the single writing cycle ended, with no indication of any follow-up toward course completion or graduation, and the authors explicitly recommend a "longitudinal design" as future work, confirming none occurred here. An internet search for subsequent papers by Kuldip Kaur Maktiar Singh, Irene Yoke Chu Leong, or Chun Keat Yeap tracking this same cohort toward graduation was conducted; no such follow-up paper was found, only other unrelated studies by these authors. All quotes above were checked and are verbatim from the source. Criterion G is not met because criterion Y is not met and no follow-up beyond the immediate post- intervention measures is reported or found.
    • P

      Pre-Registered

      • The paper describes institutional ethical approval but makes no mention of pre-registering the study protocol, hypotheses, or analysis plan before data collection, and no registry entry was found via internet search.
      • Relevant Quotes: 1) "This study was conducted in accordance with institutional ethical guidelines. Ethical approval was obtained from the Research Ethics Committee of the Academy of Language Studies, Universiti Teknologi MARA, Melaka, Malaysia. All participants provided informed consent prior to participation." (p. 1355) 2) No quotes anywhere in the paper reference a public registry (e.g., OSF, ISRCTN, AEA RCT Registry) or a pre-registration date for the study's hypotheses, methods, or analysis plan. Detailed Analysis: The only transparency-related disclosure in the paper concerns institutional research-ethics approval and informed consent, which is distinct from pre-registration of the study's protocol, hypotheses, and planned analyses on a public registry before data collection began. No registry name, identifier, or registration date is provided anywhere in the text. An internet search for a pre-registration record under this title or these authors' names (e.g., on OSF or similar registries) found no matching entry. Both quotes above were checked and are verbatim from the source. Criterion P is not met because no pre-registration of the study protocol is documented.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.