The Effect of Reflection-Supported Process-Based Writing Teaching on Iraqi EFL Students' Writing Performance and Attitude

Salam Hamid Abbas

Published:
ERCT Check Date:
DOI: 10.24093/awej/vol7no4.4
  • L2 languages
  • higher education
  • Asia
0
  • C

    The paper explicitly adopts a non-randomized design with two intact sections and never describes any randomisation procedure, so class-level randomisation is not properly documented or implemented.

    "The non-randomized pretest/posttest control group design is adopted in this study." (p. 49)

  • E

    Outcomes were measured with a researcher-administered composition test scored on an adopted rubric and a researcher-made attitude scale, not with any widely recognised standardised exam.

    "To evaluate participants' writing performance, a writing test is administered in which participants are asked to write a composition of at least three paragraphs with about 225 words on one of three topics that are selected by participants themselves." (p. 50)

  • T

    The intervention ran for 15 weeks (a full university semester) with outcomes measured at the end of the experiment, satisfying the one-term minimum interval.

    "The experiment is inaugurated in the second semester of the academic year 2015/2016 and lasted for 15 weeks." (p. 52)

  • D

    The control group's size, condition, and baseline writing, intelligence, and attitude data are documented and shown to be statistically equivalent to the experimental group.

    "During the experiment, both groups, the experimental and control, are assigned two writing lesson periods per week and given the same number and topics of writing assignments." (pp. 52-53)

  • S

    The study involved two class sections within a single college of one university, with no school-level randomisation of any kind.

    "While the sample is restricted to 88 second year students in the Department of English – College of Education/Ibn Rushd for Human Sciences of the University of Baghdad during the academic year 2015-2016." (p. 49)

  • I

    The single author designed the intervention, prepared the instruments, ran the experiment, and scored the tests himself, with no independent third-party conduct or oversight.

    "The intrascorer method in which the researcher himself score the test papers of the sample in the two administration" (p. 52)

  • Y

    The study lasted only 15 weeks (one semester), well short of 75% of an academic year, with no longer follow-up tracking.

    "The experiment is inaugurated in the second semester of the academic year 2015/2016 and lasted for 15 weeks." (p. 52)

  • B

    Both groups received identical process-based teaching, the same two weekly lessons, and the same assignments, with the only difference being the reflection sheets that constitute the treatment variable itself.

    "During the experiment, both groups, the experimental and control, are assigned two writing lesson periods per week and given the same number and topics of writing assignments." (pp. 52-53)

  • R

    No independent published replication of this specific reflection-supported process-writing study was found; later similar studies elsewhere are thematically related but are not framed as replications of this trial.

  • A

    Only EFL writing performance and writing attitude were measured, criterion E is not met, and no other core subjects were assessed.

    "At the end of the experiment, the two instruments of the study, i.e., a writing performance test and attitude toward writing scale are administered on both groups." (Abstract, p. 42)

  • G

    Measurement stopped at the end of the 15-week experiment with no tracking of students to graduation, criterion Y is not met, and no follow-up publication tracking the same cohort was found.

    "The attitude scale and writing performance test are administered to the sample of the study in one session at the end of the experiment." (p. 53)

  • P

    The paper contains no mention of any pre-registered protocol, registry, or registration date, and no external pre-registration record was found.

Abstract

The study aims at finding out the effect of process-based writing teaching supported by students' reflection on their performance in, and attitude toward writing. It hypothesizes that there is no statistically significant difference between the mean score of the experimental group taught writing according to the reflection-supported process-based approach and the control group taught writing according to the process-based approach in the writing performance test and writing attitude scale. To achieve the aims of the study, two second year sections in the Department of English of the College of Education/Ibn Rushd for Human Sciences are randomly assigned as the experimental and control groups with 43 and 45 students respectively. The experiment in this study lasts for 15 weeks during which both groups are taught writing according to the process approach and given one writing assignment per week. Only the experimental group students are required to reflect on their writing performance in every writing assignment by using a reflection sheet prepared for this purpose. At the end of the experiment, the two instruments of the study, i.e., a writing performance test and attitude toward writing scale are administered on both groups. The statistical manipulation of the results achieved shows that supporting the process orientation to writing teaching with a phase of students reflection on their writing performance is effective in developing their writing performance and helping them formulate positive attitude toward writing. In the light of the results and conclusions achieved, a set of recommendations is put forward.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • The paper explicitly adopts a non-randomized design with two intact sections and never describes any randomisation procedure, so class-level randomisation is not properly documented or implemented.
      • "The non-randomized pretest/posttest control group design is adopted in this study." (p. 49)
      • Relevant Quotes: 1) "To achieve the aims of the study, two second year sections in the Department of English of the College of Education/ Ibn Rushd for Human Sciences are randomly assigned as the experimental and control groups with 43 and 45 students respectively." (Abstract, p. 42) 2) "The non-randomized pretest/posttest control group design is adopted in this study. In this design two or more matched -on-the relevant- variables groups are involved. One of these groups is assigned as a control group and the other(s) is assigned as experimental." (p. 49) 3) "While the sample is restricted to 88 second year students in the Department of English – College of Education/Ibn Rushd for Human Sciences of the University of Baghdad during the academic year 2015-2016. The sample is divided into a control group with 45 participants and an experimental group with 43 participants." (p. 49) Detailed Analysis: Criterion C requires a genuine RCT with randomisation at the class level (or stronger) that is clearly described and properly implemented. The abstract briefly claims that the two intact sections "are randomly assigned as the experimental and control groups", which would nominally be class-level assignment. However, the Methodology section directly contradicts this by stating that "the non-randomized pretest/posttest control group design is adopted in this study", i.e., the authors themselves classify the study as a quasi-experiment with matched intact groups rather than a randomised trial. No randomisation procedure (method of allocation, who performed it, how) is described anywhere in the paper, and only two pre-existing class sections were involved, so even under the most charitable reading the "randomisation" is a single coin flip between two intact groups, which cannot be verified from the text. The intervention is whole-class instruction, not one-to-one tutoring, so the tutoring exception does not apply. Criterion C is not met because the paper self-describes its design as non-randomized and provides no description of any properly implemented class-level randomisation.
    • E

      Exam-based Assessment

      • Outcomes were measured with a researcher-administered composition test scored on an adopted rubric and a researcher-made attitude scale, not with any widely recognised standardised exam.
      • "To evaluate participants' writing performance, a writing test is administered in which participants are asked to write a composition of at least three paragraphs with about 225 words on one of three topics that are selected by participants themselves." (p. 50)
      • Relevant Quotes: 1) "To evaluate participants' writing performance, a writing test is administered in which participants are asked to write a composition of at least three paragraphs with about 225 words on one of three topics that are selected by participants themselves. Participants' written compositions are scores analytically according to a rubric adopted from O'Malley & Pierce (1996: 145)." (p. 50) 2) "The scale is prepared by the researcher based on four different scales of attitude toward writing proposed by Wollcot & Buhr 1987, DeMent 2008, Erkan & Saban 2011, and Scott 2012." (p. 50) 3) "The final version of the scale includes 30 items to be responded to according to a rating of five Likert points (strongly agree, agree, somewhat agree, disagree, and strongly disagree)." (p. 51) Detailed Analysis: Criterion E requires that outcomes be measured with standard, widely recognised standardised exams rather than instruments created for the study. Here the primary outcome, writing performance, was measured with a composition task devised and administered by the researcher and scored with a rubric adopted from a textbook (O'Malley & Pierce, 1996); this is a study-specific assessment, not a national or otherwise standardised examination. The second outcome, attitude toward writing, was measured with a 30-item Likert scale explicitly "prepared by the researcher" by selecting and adjusting items from four earlier scales; an attitude scale is in any case not an exam. The only standardised instrument mentioned, Raven's Progressive Matrices, was used solely for baseline equalization of intelligence, not as an outcome measure. Criterion E is not met because all outcome measures were custom, researcher-made instruments rather than recognised standardised exams.
    • T

      Term Duration

      • The intervention ran for 15 weeks (a full university semester) with outcomes measured at the end of the experiment, satisfying the one-term minimum interval.
      • "The experiment is inaugurated in the second semester of the academic year 2015/2016 and lasted for 15 weeks." (p. 52)
      • Relevant Quotes: 1) "The experiment is inaugurated in the second semester of the academic year 2015/2016 and lasted for 15 weeks." (p. 52) 2) "The experiment in this study lasts for 15 weeks during which both groups are taught writing according to the process approach and given one writing assignment per week." (Abstract, p. 42) 3) "At the end of the experiment, the writing performance post-test and attitude toward writing scale are administered on participants in both groups." (p. 53) 4) "The attitude scale and writing performance test are administered to the sample of the study in one session at the end of the experiment." (p. 53) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. The intervention started at the beginning of the second semester of the 2015/2016 academic year and ran for 15 weeks, with the writing post-test and attitude scale administered at the end of the experiment. A 15-week interval corresponds to roughly 3.5 months and constitutes a full university semester (term) in the Iraqi higher-education context, so the interval from intervention start to outcome measurement meets the one-term threshold. Criterion T is met because the interval from intervention start to outcome measurement was 15 weeks, covering one full academic term.
    • D

      Documented Control Group

      • The control group's size, condition, and baseline writing, intelligence, and attitude data are documented and shown to be statistically equivalent to the experimental group.
      • "During the experiment, both groups, the experimental and control, are assigned two writing lesson periods per week and given the same number and topics of writing assignments." (pp. 52-53)
      • Relevant Quotes: 1) "The sample is divided into a control group with 45 participants and an experimental group with 43 participants." (p. 49) 2) "This equalization checking involves a writing pre-test scores, intelligence, and the attitude toward writing. The data collection instruments in the equalization checking are a writing performance pre-test, Raven's progressive matrices intelligence test (RPM), and an attitude toward writing scale prepared by the researcher." (pp. 49-50) 3) "T-test results indicate no statistically significant difference between the two groups as the computed values of the writing pre-test, intelligence test, and attitude scale are found 0.94, 1.037, and 0.62 respectively at 0.05 level of significance and 86 degree of freedom." (p. 50) 4) "During the experiment, both groups, the experimental and control, are assigned two writing lesson periods per week and given the same number and topics of writing assignments." (pp. 52-53) 5) "Table 2: Checking the equalization of the experimental and control groups in the writing pre-test, intelligence test, and attitude toward writing" (p. 50) Detailed Analysis: Criterion D requires clear documentation of the control group: who they are, their size, baseline characteristics, and what treatment they received. The paper specifies the control group size (45 second-year English majors at the same college), reports baseline means and standard deviations for writing pre-test, Raven's Progressive Matrices intelligence, and attitude toward writing in Table 2, and statistically verifies equivalence between groups on all three variables. The control condition is explicitly described: the same process-based writing teaching (Tompkins model), two writing lessons per week, and identical assignments, with only the reflection component withheld. Broader demographic detail (e.g., age, gender composition) is not reported, but the documentation of composition, baseline performance, and treatment received is sufficient for meaningful comparison. Criterion D is met because the control group's size, baseline equivalence data, and condition are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • The study involved two class sections within a single college of one university, with no school-level randomisation of any kind.
      • "While the sample is restricted to 88 second year students in the Department of English – College of Education/Ibn Rushd for Human Sciences of the University of Baghdad during the academic year 2015-2016." (p. 49)
      • Relevant Quotes: 1) "While the sample is restricted to 88 second year students in the Department of English – College of Education/Ibn Rushd for Human Sciences of the University of Baghdad during the academic year 2015-2016." (p. 49) 2) "To achieve the aims of the study, two second year sections in the Department of English of the College of Education/ Ibn Rushd for Human Sciences are randomly assigned as the experimental and control groups with 43 and 45 students respectively." (Abstract, p. 42) Detailed Analysis: Criterion S requires randomisation among schools or equivalent implementing institutions. This study took place entirely within one department of one college at the University of Baghdad, and the units involved were two class sections, not separate institutions. No schools, colleges, or sites were randomised, and the paper additionally self-describes the design as non-randomized. Criterion S is not met because the study involved a single institution with assignment (at best) of two class sections, not school-level randomisation.
    • I

      Independent Conduct

      • The single author designed the intervention, prepared the instruments, ran the experiment, and scored the tests himself, with no independent third-party conduct or oversight.
      • "The intrascorer method in which the researcher himself score the test papers of the sample in the two administration" (p. 52)
      • Relevant Quotes: 1) "only participants in the experimental group are asked and encouraged to regularly reflect on every written performance they submit by using a reflection sheet prepared by the researcher for this purpose (Appendix C)." (p. 53) 2) "The scale is prepared by the researcher based on four different scales of attitude toward writing" (p. 50) 3) "The intrascorer method in which the researcher himself score the test papers of the sample in the two administration, and the interscorer method in which another university EFL instructor is asked to score the sample test papers of the second administration." (p. 52) 4) "The present study is intended to experiment engaging Iraqi EFL students in practicing reflection on their writing performance ... as a technique that may help them promote their writing skills" (p. 43) Detailed Analysis: Criterion I requires that the study be conducted independently from those who designed the intervention. Here a single author (the researcher) conceived the reflection-supported intervention, prepared the reflection sheet and the attitude scale, ran the 15-week experiment in his own institution, and scored the writing tests himself (the "intrascorer" reliability check explicitly states "the researcher himself score the test papers"). The only outside involvement is a second EFL instructor double-scoring a pilot subsample for reliability and expert juries validating the instruments; neither constitutes independent conduct of data collection, implementation, or analysis, and no external evaluation team or third-party oversight is mentioned. Criterion I is not met because the intervention designer personally conducted, administered, and scored the study without independent evaluation.
    • Y

      Year Duration

      • The study lasted only 15 weeks (one semester), well short of 75% of an academic year, with no longer follow-up tracking.
      • "The experiment is inaugurated in the second semester of the academic year 2015/2016 and lasted for 15 weeks." (p. 52)
      • Relevant Quotes: 1) "The experiment is inaugurated in the second semester of the academic year 2015/2016 and lasted for 15 weeks." (p. 52) 2) "At the end of the experiment, the writing performance post-test and attitude toward writing scale are administered on participants in both groups." (p. 53) Detailed Analysis: Criterion Y requires that outcomes be measured at least 75% of a full academic year (roughly 7+ months of a 9-10 month year) after the intervention begins. The entire experiment spanned 15 weeks within the second semester of 2015/2016, about 3.5 months, and outcomes were collected in a single session at the end of that period. There was no follow-up beyond the end of the semester. Fifteen weeks is roughly 35-40% of an academic year, far below the 75% threshold. Since criterion T's weaker one-term threshold was met but the stronger year-long threshold was not, criterion Y correctly remains unmet independently of T. Criterion Y is not met because tracking from intervention start to measurement covered only about 15 weeks, well under 75% of an academic year.
    • B

      Balanced Control Group

      • Both groups received identical process-based teaching, the same two weekly lessons, and the same assignments, with the only difference being the reflection sheets that constitute the treatment variable itself.
      • "During the experiment, both groups, the experimental and control, are assigned two writing lesson periods per week and given the same number and topics of writing assignments." (pp. 52-53)
      • Relevant Quotes: 1) "During the experiment, both groups, the experimental and control, are assigned two writing lesson periods per week and given the same number and topics of writing assignments." (pp. 52-53) 2) "Tompkins Model (2004) to process writing is used in teaching and training participants in both groups." (p. 53) 3) "However, only participants in the experimental group are asked and encouraged to regularly reflect on every written performance they submit by using a reflection sheet prepared by the researcher for this purpose (Appendix C)." (p. 53) 4) "Only the experimental group students are required to reflect on their writing performance in every writing assignment by using a reflection sheet prepared for this purpose." (Abstract, p. 42) Detailed Analysis: Criterion B requires that the control group receive balanced educational time and resources unless the extra resource is itself the treatment variable being tested. Applying the current decision tree: both groups received the same instructional programme (process-based writing teaching following the Tompkins model), the same two writing lesson periods per week, and the same number and topics of weekly assignments, so class time and materials were equal (CONTROL_MATCHES_RESOURCES / no meaningful EXTRA_RESOURCES_PRESENT beyond the reflection sheet itself). The sole difference is that experimental students completed a self-directed reflection sheet after each assignment, and the paper explicitly frames "reflection-supported" writing teaching (i.e., reflection practice itself) as the independent variable under test, so even if this is treated as an "extra resource" it is RESOURCES_ARE_TREATMENT: integral to, and the explicit object of, the experimental manipulation, not a separable add-on requiring matching in the control group. No additional teacher time, materials, or budget beyond the one-page reflection sheet is reported for the experimental group. Criterion B is met because instructional time, teaching approach, and assignments were identical across groups, and the only added element (student reflection) is the treatment variable itself, satisfying the decision tree's "resources are the treatment" branch.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent published replication of this specific reflection-supported process-writing study was found; later similar studies elsewhere are thematically related but are not framed as replications of this trial.
      • Relevant Quotes: 1) "Moreover, related literature is not conclusive regarding the effect of students reflection on their performance and attitudes. Accordingly, the present study is intended to contribute to the literature in this respect." (p. 44) 2) "Unfortunately, although reflection plays a significant role in learning, little research has been done on the effect of reflection type or amount on the outcomes achieved by students (Bringle & Hatcher, 1999:181)." (p. 46) Detailed Analysis: Criterion R requires independent replication of the study by a different research team in a different context, published in a peer-reviewed journal. The paper itself positions the work as filling a gap ("related literature is not conclusive"), and mentions no replication. A fresh internet search (2026-07-27) confirms the earlier finding: no subsequent paper self-identifies as reproducing this specific Iraqi trial. Two thematically similar studies by a different, Ethiopian research team were located - "The Effect of Reflection-Supported Learning of Writing on Students' Writing Performance and Writing Self-Efficacy" (Deti, T. W., Ferede, T., & Tiruneh, D., 2022, SSRN Electronic Journal) and "The Effect of Reflection Supported Learning of Writing on Students' Writing Attitude and Writing Achievement Goal Orientations" (Deti, T., Ferede, T., & Tiruneh, D., 2023, Research Square preprint). These studies examine "reflection-supported learning of writing" with a different student population and partly different dependent variables (writing self-efficacy, goal orientations, rather than the writing-attitude/performance pair used here), but full text access was blocked (403/paywall) during this check, so it could not be verified whether they replicate Abbas's specific Gibbs-model reflection-sheet plus Tompkins process-writing design, sample equalization procedure, and measures, or cite Abbas (2016) as their basis; no verbatim quotes from those papers can therefore be responsibly provided. Absent confirmed methodological correspondence and independent verification, they cannot be counted as a reproduction of this particular study. Criterion R is not met because no independent peer-reviewed replication of this specific study was confirmed; two thematically related but unverified studies by a different team were identified and flagged for reference.
    • A

      All-subject Exams

      • Only EFL writing performance and writing attitude were measured, criterion E is not met, and no other core subjects were assessed.
      • "At the end of the experiment, the two instruments of the study, i.e., a writing performance test and attitude toward writing scale are administered on both groups." (Abstract, p. 42)
      • Relevant Quotes: 1) "At the end of the experiment, the two instruments of the study, i.e., a writing performance test and attitude toward writing scale are administered on both groups." (Abstract, p. 42) 2) "To evaluate participants' writing performance, a writing test is administered in which participants are asked to write a composition of at least three paragraphs" (p. 50) Detailed Analysis: Criterion A requires standardised exam-based assessment of all main subjects taught at the educational level, and per the instructions it automatically fails when criterion E fails. Criterion E is not met here, since both outcome instruments are researcher-made. Moreover, only a single domain, EFL writing (plus attitude toward writing), was assessed; no other subjects of the English department curriculum or wider programme were measured, and while the intervention is specialised (EFL writing in a university English department), the assessments used are not standardised exams, so the specialised-intervention exception cannot rescue the criterion. Criterion A is not met because criterion E fails and only custom writing measures in a single domain were used.
    • G

      Graduation Tracking

      • Measurement stopped at the end of the 15-week experiment with no tracking of students to graduation, criterion Y is not met, and no follow-up publication tracking the same cohort was found.
      • "The attitude scale and writing performance test are administered to the sample of the study in one session at the end of the experiment." (p. 53)
      • Relevant Quotes: 1) "The attitude scale and writing performance test are administered to the sample of the study in one session at the end of the experiment." (p. 53) 2) "The experiment is inaugurated in the second semester of the academic year 2015/2016 and lasted for 15 weeks." (p. 52) Detailed Analysis: Criterion G requires following participants until graduation from their educational stage, and per the instructions automatically fails when criterion Y fails, which it does here. Participants were second-year undergraduates, and all outcome data were collected in a single session at the end of the 15-week experiment; the paper reports no follow-up of the cohort toward degree completion. A fresh internet search (2026-07-27) for later publications by the same author, Salam Hamid Abbas, tracking this specific 2015-2016 cohort of 88 second-year English majors found no such paper; the author's other located publications - "The Effect of Pair Writing Technique on Iraqi EFL University Students' Writing Performance and Anxiety" (Abbas & Al-bakri, 2018, AWEJ 9(2)) and three 2025 co-authored papers with Dehham on metacognitive awareness and self-regulation (Arab World English Journal 16(1); Theory and Practice in Language Studies 15(7); Journal of Posthumanism 5(5)) - address different research questions, interventions, and student samples, and none describes graduation tracking of the 2016 cohort. Criterion G is not met because tracking ended with the post-test at the end of the semester, far short of graduation, criterion Y also fails, and no follow-up publication tracking this cohort to graduation was located.
    • P

      Pre-Registered

      • The paper contains no mention of any pre-registered protocol, registry, or registration date, and no external pre-registration record was found.
      • Relevant Quotes: 1) "The non-randomized pretest/posttest control group design is adopted in this study." (p. 49) 2) "The experiment is inaugurated in the second semester of the academic year 2015/2016 and lasted for 15 weeks." (p. 52) Detailed Analysis: Criterion P requires that the full study protocol, including hypotheses, methods, and planned analyses, be publicly pre-registered before data collection begins, with a verifiable registry reference and date. The paper describes its design, hypotheses, and instruments but never mentions any registry (e.g., ClinicalTrials.gov, OSF, AEA registry), a registration ID, or a registration date. A fresh search (2026-07-27) of the paper's own text and available external sources found no pre-registration record for this study; this is unsurprising given the study is a 2016 quasi-experimental EFL education thesis-style paper from a context where pre-registration of such studies is not standard practice. Criterion P is not met because no pre-registration statement or registry reference exists in or for the paper.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.