Abstract
Prior studies have reported inconsistent findings with regard to the effects of small-group student talk on developing individual students' English-as-a-foreign-language (EFL) writing ability. To further explore the question under discussion, we designed a quasi-experimental study that included a pretest, a posttest, and a delayed posttest, and implemented it in two English-major groups at a university in China. We randomly assigned the students to an intervention group and a comparison group to investigate whether employing structured small-group student talk as collaborative prewriting discussions would effectively facilitate individual students' EFL writing development and whether such effects could be retained. The immediate and sustained effects after the quasi-experimental study was completed were measured by the analytic scores on five components of the writing task (content, organization, vocabulary, language, and mechanics) and the holistic writing scores cumulated of all these components. Statistical analyses revealed that the two groups were significantly distinguished by their analytic and holistic scores, indicating that students in the intervention group outperformed their comparison group peers in writing performance. The effects of collaborative prewriting discussions in the form of structured small-group student talk were found statistically significant in facilitating students' writing improvement in the content, organization, vocabulary, and language use, but not mechanics. The effects on content, organization, and vocabulary were retained as seen from the delayed posttest, while those on language use were not. The comparison group showed little improvement in their writing performance across the three tests. We concluded this study with a discussion on the implications for English-as-a-second/foreign-language (L2) writing instruction.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Two intact parallel classes (not individual students within one class) were randomly assigned to intervention and comparison conditions, satisfying class-level randomisation.
- "The two parallel groups were randomly assigned to an intervention group (n = 24) and a comparison group (n = 24)." (p. 6)
Relevant Quotes:
1) "Altogether, 48 sophomore students majoring in English Language and Literature from two intact English Writing groups participated in this study." (p. 6)
2) "The two parallel groups were randomly assigned to an intervention group (n = 24) and a comparison group (n = 24)." (p. 6)
3) "We randomly assigned the students to an intervention group and a comparison group to investigate whether employing structured small-group student talk as collaborative prewriting discussions would effectively facilitate individual students' EFL writing development..." (Abstract, p. 1)
4) "A quasi-experimental study [56] was adopted as the research design for the current study." (p. 5)
Detailed Analysis:
The unit of randomisation was the intact class: the study recruited "two intact English Writing groups" and "the two parallel groups were randomly assigned" to the intervention and comparison conditions. Students were not individually randomised within a shared classroom, so contamination between conditions inside one class is avoided - the two conditions met as separate classes with the same instructor. Although the paper self-describes as quasi-experimental (because intact convenience-sampled classes were used rather than individually randomised participants), random assignment of whole classes to conditions is exactly what criterion C requires. A caveat is that only two classes were randomised (one per condition), which is the minimal possible cluster count, but the criterion concerns the unit of randomisation, which here is the class.
Criterion C is met because entire intact classes, not students within a single class, were randomly assigned to the intervention and comparison conditions.
-
E
Exam-based Assessment
- The writing tests were taken from the database of TEM-4, a nationally standardised Chinese exam for English majors with reported high validity and reliability, and were scored with the well-established Jacobs et al. rubric.
- "The writing task for these tests came from the database of China's National English as a Foreign Language Test - Test for English Majors - Band 4 (TEM-4), which has been reported to have high validity and reliability [83,86]." (p. 6)
Relevant Quotes:
1) "The writing task for these tests came from the database of China's National English as a Foreign Language Test - Test for English Majors - Band 4 (TEM-4), which has been reported to have high validity and reliability [83,86]." (p. 6)
2) "TEM-4 is a nationally standardized annual test that is taken by Chinese university English-major students at the second semester of their second year." (pp. 6-7)
3) "The other five writing topics for intervention were randomly selected from the past TEM-4 test battery." (p. 7)
4) "Although TEM-4 rubric is nationally acknowledged in China as an official tool to assess English-major sophomore students' language proficiency, the current study intended to evaluate the writing competence of English learners in China with a more internationally recognized rubric." (p. 8)
5) "...the well-established and widely used writing rubric that was originally developed by Jacobs, Zinkgraf, Wormuth, Hartfiel, & Hughey [88] and modified and updated by Hedgcock and Lefkowitz [89] was used to rate and determine the overall quality of students' written texts, both holistically and analytically." (p. 8)
6) "Two Chinese raters, who held their Ph.D. degrees in second language acquisition or applied linguistics from well-known universities overseas and had no direct involvement in any other aspects of the current study, rated all the written texts. A blind assessment was implemented..." (p. 8)
Detailed Analysis:
The outcome measure was not a test specially designed by the researchers for this study. The pretest, posttest and delayed posttest task was drawn from the database of TEM-4, an official nationally standardised annual examination for Chinese English majors that the paper describes as having high validity and reliability; the intervention practice topics likewise came from the past TEM-4 battery. Scoring did not use the official TEM-4 rubric but instead the Jacobs et al. ESL Composition Profile (as updated by Hedgcock and Lefkowitz), which is a well-established, widely used and internationally recognised scoring instrument rather than a study-specific measure, applied blind by two independent raters with high interrater reliability (.955 holistic). Because both the test task (from a national standardised exam) and the scoring rubric (a widely recognised published instrument) are standard rather than custom-made for this study, the intent of criterion E - avoiding researcher-made tests overly aligned with the intervention - is satisfied.
Criterion E is met because outcomes were measured with argumentative writing tasks taken from the nationally standardised TEM-4 examination and scored with a widely recognised established rubric, not a custom-designed test.
-
T
Term Duration
- Outcomes were tracked across a full 16-week semester, with the delayed posttest administered about 14 weeks (approximately 3.5 months, one academic term) after the intervention began in week 2.
- "Each group had two 45-minute sessions per week with a total of 32 sessions in a 16-week semester." (p. 7)
Relevant Quotes:
1) "Each group had two 45-minute sessions per week with a total of 32 sessions in a 16-week semester." (p. 7)
2) "All the participants took part in the pretest, the posttest and the delayed posttest with the same writing tasks as used before, at the end of four weeks after the intervention (see Table 1)." (p. 6)
3) Table 1 "Procedures of the study": Week 1 - Pretest; Week 2 - Practice session of small-group student talk for planning (20 min) + individual writing (40 min); Weeks 3, 5, 7, 9, 11 - Intervention sessions; Week 12 - Posttest; Week 16 - Delayed posttest. (p. 7)
4) "Altogether, five rounds of structured small-group student talk were administered as the intervention..." (p. 7)
Detailed Analysis:
The intervention began with the practice session in week 2 of a 16-week semester, with intervention sessions in weeks 3, 5, 7, 9 and 11. The immediate posttest was in week 12 (about 10 weeks after intervention start) and the delayed posttest - which was a primary outcome for the study's core question of whether effects were retained - was in week 16, approximately 14 weeks (about 3.3-3.5 months) after the intervention began. A term is defined by the standard as a semester or equivalent (approximately 3-4 months); here the study spanned an entire 16-week university semester from pretest (week 1) to delayed posttest (week 16), and the interval from intervention start to the final measurement covers essentially the whole semester/term. The intervention dose itself was modest (six 20-minute discussion sessions), but the standard explicitly allows short interventions provided follow-up tracking reaches at least one term from intervention start, which the week-16 delayed posttest achieves.
Criterion T is met because outcome measurement extended to the end of the 16-week semester, roughly one full academic term after the intervention began.
-
D
Documented Control Group
- The comparison group's size, gender, age, language background, baseline scores and exact activities (identical structured task with individual planning) are documented in detail.
- "Participants in the comparison group included 16 females and 8 males between the ages of 18 and 21 (M = 19.5, SD = .93)." (p. 6)
Relevant Quotes:
1) "Participants in the comparison group included 16 females and 8 males between the ages of 18 and 21 (M = 19.5, SD = .93)." (p. 6)
2) "The voluntary participants all grew up in China with Chinese as their mother tongue. They had studied English previously in primary and secondary schools for an average of 10.6 years (SD = 1.2)." (p. 6)
3) "Both groups followed the same teaching syllabus and plan as required by the English Department of the selected university." (p. 7)
4) "However, participants in the comparison group planned individually for 20 minutes following the same structured writing task as used in the intervention group and after that they proceeded to individual writing of the task for 40 minutes." (p. 8)
5) "The results (see Table 3) showed that no significant between-subject differences were found in different measures at the time of pretest (overall, p = .943; content, p = .858; organization, p = .604; vocabulary, p = .528; language use, p = .882; mechanics, p = .526)." (p. 9)
6) Table 2 reports means and SDs for the comparison group (CG) on holistic and all five analytic scores at pretest, posttest and delayed posttest (e.g., "Overall CG 72.92 3.73..."). (p. 9)
Detailed Analysis:
The comparison group is documented in considerable detail: its size (n = 24), gender composition, age (M, SD), language background and years of prior English study, the fact that both groups took the same course with the same instructor, syllabus, textbook, tasks, timing and test conditions, and exactly what the comparison group did instead of the intervention (20 minutes of individual planning on the same structured task followed by 40 minutes of individual writing). Baseline equivalence is demonstrated with pretest descriptive statistics (Table 2) and formal t-tests showing no significant pretest differences on any measure. This satisfies the requirement for demographic information, baseline performance and description of the conditions the control group received.
Criterion D is met because the comparison group's composition, baseline performance and treatment conditions are clearly and thoroughly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation involved only two intact classes within a single university; no schools or institutions were randomised.
- "The current study was conducted in a comprehensive university in Central China in a teacher-fronted and test-driven context [79]..." (p. 6)
Relevant Quotes:
1) "The current study was conducted in a comprehensive university in Central China in a teacher-fronted and test-driven context [79], where a compulsory English Writing course was offered to second-year English-major students." (p. 6)
2) "Altogether, 48 sophomore students majoring in English Language and Literature from two intact English Writing groups participated in this study." (p. 6)
3) "The two parallel groups were randomly assigned to an intervention group (n = 24) and a comparison group (n = 24)." (p. 6)
Detailed Analysis:
Criterion S requires randomisation among schools or equivalent implementing institutions. This study took place at a single university, and the randomised units were two intact class groups within that one institution. No multiple schools, campuses or sites were involved, and therefore no school-level random assignment occurred.
Criterion S is not met because randomisation was at the class level within one university, not at the school/institution level.
-
I
Independent Conduct
- The authors designed the intervention, administered all tests, and analysed the data themselves; independent blind raters scored the texts, but there was no independent conduct or third-party oversight of the study as a whole.
- "All the three tests were administrated by the first author of the current study." (p. 7)
Relevant Quotes:
1) "All the three tests were administrated by the first author of the current study." (p. 7)
2) "Conceptualization: Hui Helen Li, Lawrence Jun Zhang. Data curation: Hui Helen Li. Formal analysis: Hui Helen Li, Lawrence Jun Zhang. ... Investigation: Hui Helen Li, Lawrence Jun Zhang. Methodology: Hui Helen Li, Lawrence Jun Zhang. Project administration: Hui Helen Li." (Author Contributions, p. 15)
3) "Two Chinese raters, who held their Ph.D. degrees in second language acquisition or applied linguistics from well-known universities overseas and had no direct involvement in any other aspects of the current study, rated all the written texts. A blind assessment was implemented in which the raters did not know which group of students they were rating, nor did they know if they were rating a pre-, post-, or a delayed post-test written text." (p. 8)
4) "This study was supported by Hubei Provincial Department of Education of China in the form of a grant awarded to HHL (2018112)... However, the funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript for publication." (p. 2)
Detailed Analysis:
The same two authors conceptualised the study, designed the structured small-group talk intervention, collected the data (the first author personally administered all three tests) and performed the formal analysis. There is no external evaluation team, agency, or third-party oversight of design, implementation or analysis. The one element of independence is the outcome scoring: two external Ph.D.-qualified raters with no other involvement rated all texts blind to group and test occasion, which mitigates rater bias. However, under criterion I blind scoring alone does not amount to an independently conducted study - the intervention designers themselves ran the trial, administered the tests and analysed the data, and the funder's non-involvement statement concerns the grant provider, not an independent evaluator.
Criterion I is not met because the intervention designers themselves conducted, administered and analysed the study, with independence limited to blind external scoring of the written texts.
-
Y
Year Duration
- Tracking from intervention start to the final delayed posttest lasted about 14 weeks (one semester), far short of 75% of a full academic year.
- "Each group had two 45-minute sessions per week with a total of 32 sessions in a 16-week semester." (p. 7)
Relevant Quotes:
1) "Each group had two 45-minute sessions per week with a total of 32 sessions in a 16-week semester." (p. 7)
2) Table 1: Week 1 - Pretest; Week 2 - Practice session; Weeks 3, 5, 7, 9, 11 - Intervention sessions; Week 12 - Posttest; Week 16 - Delayed posttest. (p. 7)
3) "The course spanned two semesters in the selected university. ... This study was conducted in the second semester." (p. 6)
Detailed Analysis:
Criterion Y requires outcomes to be measured at least 75% of a full academic year (roughly 9-10 months) after the intervention begins. Here the entire study - pretest, intervention, posttest and delayed posttest - was completed within a single 16-week semester. The interval from intervention start (week 2) to the last measurement (week 16) is about 14 weeks, roughly 3.3 months, which is well under 75% of an academic year. No longer-term follow-up is reported.
Criterion Y is not met because the tracking period was one semester (about 14 weeks), far below 75% of an academic year.
-
B
Balanced Control Group
- Both groups received identical instruction, tasks, and time (20 minutes planning plus 40 minutes writing on the same structured task); the only difference was collaborative versus individual planning, so inputs were fully balanced.
- "However, participants in the comparison group planned individually for 20 minutes following the same structured writing task as used in the intervention group and after that they proceeded to individual writing of the task for 40 minutes." (p. 8)
Relevant Quotes:
1) "Both groups followed the same teaching syllabus and plan as required by the English Department of the selected university." (p. 7)
2) "Specifically, in each session, each small group in the intervention group first talked for 20 minutes based on the structured writing task and then separated to write the task individually for 40 minutes." (p. 8)
3) "However, participants in the comparison group planned individually for 20 minutes following the same structured writing task as used in the intervention group and after that they proceeded to individual writing of the task for 40 minutes." (p. 8)
4) "During these sessions, neither the intervention group nor the comparison group was allowed to use any external resources. Meanwhile, the course instructor mainly remained silent as an observer unless students particularly asked for her help." (p. 8)
5) "Apart from that, the writing prompt, test time, and procedures were kept constant in both groups regarding the pretest, the posttest and the delayed posttest." (p. 7)
Detailed Analysis:
Following the criterion B decision tree: did the intervention add extra time or budget? No. Both groups attended the same course with the same instructor, syllabus, textbook and session schedule (two 45-minute sessions per week). In every intervention round, both groups spent exactly 20 minutes on planning with the identical structured writing task and 40 minutes on individual writing; the sole difference was whether planning was done collaboratively in small groups or individually. No additional materials, technology, instructional time or adult support was given to the intervention group - the instructor remained a silent observer in both conditions, and test conditions were kept constant across groups. Because there were no extra resources at all, the balance requirement is trivially satisfied (an active, time-matched comparison condition using the very same structured task).
Criterion B is met because both conditions received identical time, tasks, materials and teacher input, with collaborative versus individual planning being the only difference.
-
Level 3 Criteria
-
R
Reproduced
- No independent published replication of this specific study was found after a fresh internet search; related collaborative-prewriting studies are either earlier work, by the same authors reusing the same dataset, or different designs rather than replications.
Relevant Quotes:
1) "To advance the current knowledge of such an issue in L2 writing, this study adopted a quasi-experimental design [56] with a pretest, a posttest and a delayed posttest to address whether structured small-group student talk ... exerts immediate and sustained effects on Chinese tertiary EFL students' individual writing performance." (pp. 2-3)
2) "In order to document a comprehensive and thorough understanding of such effects, additional studies that employ pre-, post-, and delayed post-test measures are needed." (p. 5)
3) "Given this limitation, future studies can enlarge the size of sampling to enhance the reliability of the results." (p. 15)
Detailed Analysis:
The paper itself frames the study as filling a gap and calls for future studies, and it cites related prior work (Neumann and McDonough; McDonough et al.; Shin; Shi; Pu; Liao; Jiang et al.), all of which predate this study and examine related but distinct designs - they are the literature the study builds on, not replications of it. A renewed internet search for independent replications of this specific 2021 study (title, authors, and DOI 10.1371/journal.pone.0251569 as search terms) found: (a) Li HH, Zhang LJ (2022), "Investigating Effects of Small-Group Student Talk on the Quality of Argument in Chinese Tertiary English as a Foreign Language Learners' Argumentative Writing," Frontiers in Psychology 13:868045 - the same two authors report on "48 undergraduate students in their second year of study... sampled from a School of Foreign Languages at a large public comprehensive university in Central China ... randomly assigned into a treatment class and a comparison class with 24 students in each," matching this study's sample and design exactly, so this is a re-analysis of the same dataset with different outcome measures (argument quality) rather than an independent replication; and (b) later, unrelated studies of small-group/scaffolded prewriting discussions in other contexts (e.g., a scaffolded small-group discussion and picture-series study in Indonesia published in TELL-US Journal), which investigate related concepts with different designs and populations and do not reference-and- reproduce this specific study's design, intervention, and measures as a replication. No independent, peer-reviewed replication of this particular trial by a different research team was identified.
Criterion R is not met because no independent published replication of this specific study by a different research team was found.
-
A
All-subject Exams
- Only English argumentative writing was assessed; no other main subjects were measured and no specialised-context justification per the standard's exception applies.
- "These written texts were used to determine the effects of small-group student talk on students' argumentative writing performance in terms of analytic (content, organization, vocabulary, language use, and mechanics) and holistic scores." (p. 8)
Relevant Quotes:
1) "These written texts were used to determine the effects of small-group student talk on students' argumentative writing performance in terms of analytic (content, organization, vocabulary, language use, and mechanics) and holistic scores." (p. 8)
2) "The immediate and sustained effects after the quasi-experimental study was completed were measured by the analytic scores on five components of the writing task (content, organization, vocabulary, language, and mechanics) and the holistic writing scores cumulated of all these components." (Abstract, p. 1)
Detailed Analysis:
Criterion A requires standardised exam-based assessment across all main subjects taught at the educational level. This study measured only one outcome domain: English argumentative writing quality (five analytic subscores plus a holistic score of the same task). Even within the English major curriculum, other core components (reading, listening, speaking, translation - all parts of TEM-4 as described in the paper) were not assessed, and no other university subjects were measured. The standard's exception applies to highly specialised interventions in upper secondary or vocational education with an explicit rationale; the paper offers no such justification, and the context is a general undergraduate degree programme. Criterion E is met, so A is not blocked by the prerequisite, but subject coverage is limited to a single skill in a single subject.
Criterion A is not met because outcomes were measured only for English writing, not across all main subjects.
-
G
Graduation Tracking
- Measurement ended with the delayed posttest in week 16 of the semester; participants were not tracked to graduation, no follow-up publication tracking this cohort further was found, and criterion Y is not met, which per the standard's dependency rule also rules out G.
- "All the participants took part in the pretest, the posttest and the delayed posttest with the same writing tasks as used before, at the end of four weeks after the intervention (see Table 1)." (p. 6)
Relevant Quotes:
1) "All the participants took part in the pretest, the posttest and the delayed posttest with the same writing tasks as used before, at the end of four weeks after the intervention (see Table 1)." (p. 6)
2) Table 1: "12 Posttest ... 16 Delayed posttest" - the final measurement was in week 16 of the semester. (p. 7)
3) "Besides, such a pedagogical practice should be implemented over a longer period with more rounds of sessions than those administered in this study, since enhancing argumentative writing skills needs time and practice [98]." (p. 15)
Detailed Analysis:
Participants were second-year (sophomore) English majors, so graduation from their degree programme would be roughly two years after the study. Measurement stopped at the delayed posttest four weeks after the posttest, within the same semester as the intervention. No tracking to the end of the degree (or even the end of the academic year) is reported, and the authors' own recommendations imply no longer follow-up occurred. A renewed internet search for subsequent publications by the same authors tracking this cohort found only Li HH, Zhang LJ (2022), Frontiers in Psychology 13:868045, which reports the identical 48-student, treatment/comparison-class sample re-analysed for a different set of argument-quality outcome measures at the same pretest/ posttest/delayed-posttest time points, not a longer-term or graduation follow-up; no later graduation-tracking study of this cohort was found. In addition, per the prompt's dependency rule, criterion Y is not met, which also rules out G.
Criterion G is not met because participants were only tracked for one semester, with no follow-up to graduation.
-
P
Pre-Registered
- The paper contains no mention of any pre-registration or registry entry, and a renewed search of the full text and external sources found no registration record for this study.
Relevant Quotes:
1) "This study was reviewed and approved by The University of Auckland Ethics Committee on Human Participants." (p. 5)
2) "Received: October 2, 2020; Accepted: April 29, 2021; Published: May 28, 2021" (p. 1)
3) "Data Availability Statement: All relevant data are within the manuscript and its Supporting information files in a Zip folder." (p. 1)
Detailed Analysis:
The paper reports ethics approval and data availability but nowhere mentions pre-registration of hypotheses, methods or analysis plans on any registry platform (e.g., ClinicalTrials.gov, OSF, AsPredicted, or a trial registry). A renewed full-text check of the article (via its PMC mirror, PMC8162705) confirmed there is no reference anywhere in the text, methods, or supplementary information to prospective registration of the study protocol. A further web search for a registration record associated with this study (DOI 10.1371/journal.pone.0251569) likewise found no pre-registration entry. Ethics approval is not a substitute for pre-registration of the study protocol before data collection.
Criterion P is not met because there is no evidence the study protocol was pre-registered before data collection.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.