Abstract
This mixed-methods study evaluates the impact of AI-assisted language learning on Chinese English as a Foreign Language (EFL) students' writing skills and writing motivation. As artificial intelligence (AI) becomes more prevalent in educational settings, understanding its effects on language learning outcomes is crucial. The study employs a comprehensive approach, combining quantitative and qualitative methods. The quantitative phase utilizes a pre-test and post-test design to assess writing skills. Fifty EFL students, matched for proficiency, are randomly assigned to experimental (AI-assisted instruction via ChatGPT) or control (traditional instruction) groups. Writing samples are evaluated using established scoring rubrics. Concurrently, semi-structured interviews are conducted with a subset of participants to explore writing motivation and experiences with AI-assisted learning. Quantitative analysis reveals significant improvements in both writing skills and motivation among students who received AI-assisted instruction compared to the control group.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was at the individual student level, not at the class or school level, and the intervention was classroom instruction rather than one-to-one tutoring.
- "Random assignment was used to assign participants to either the experimental group or the control group."
Relevant Quotes:
1) "Fifty EFL students, matched for proficiency, are randomly assigned to experimental (AI-assisted instruction via ChatGPT) or control (traditional instruction) groups." (p. 1, Abstract)
2) "Random assignment was used to assign participants to either the experimental group or the control group. This randomization process aimed to minimize bias and ensure comparability between the groups, giving each participant an equal chance of being assigned to either group." (p. 5)
3) "The study included 50 Chinese EFL students who were enrolled in a Bachelor's degree program at a national university in China." (p. 5)
4) "Furthermore, a potential limitation is the potential for contamination between the experimental and control groups. Despite random assignment, it is challenging to completely separate instructional methods or interactions in real-world educational settings. Some degree of crossover between the groups may have occurred, introducing the possibility of contamination effects and influencing the observed results." (p. 12)
Detailed Analysis:
Criterion C requires randomisation at the class level (or the stronger school level), so that whole classes rather than individual students within the same setting are assigned to conditions. The quotes show unambiguously that randomisation was performed at the individual student level: 50 students from one university were individually and randomly assigned to the experimental or control group. No class or school unit was randomised.
The tutoring exception was considered: the intervention offers individualised AI feedback, but it was delivered as group writing instruction in scheduled sessions in a shared computer laboratory ("The instruction sessions of both groups were held in a dedicated computer laboratory"), not as a one-to-one personal tutoring programme, so the exception for personal teaching does not apply. The authors themselves acknowledge the contamination risk that class-level randomisation is designed to prevent.
Criterion C is not met because randomisation was conducted at the individual student level within a single institution without a valid tutoring exception.
-
E
Exam-based Assessment
- Writing outcomes were measured with IELTS academic writing tasks scored using official IELTS band descriptors, a widely recognised standardised assessment.
- "The study utilized the International English Language Testing System (IELTS) academic writing tasks 1 and 2 as pre-tests and post-tests to evaluate participants' academic writing skills."
Relevant Quotes:
1) "The study utilized the International English Language Testing System (IELTS) academic writing tasks 1 and 2 as pre-tests and post-tests to evaluate participants' academic writing skills." (p. 6)
2) "The participants' writing proficiency was assessed using the IELTS writing band descriptors for task 1 and task 2 (see Appendix A), which encompassed criteria such as task achievement, coherence and cohesion, lexicon, and grammatical range and accuracy (University of Cambridge ESOL Examinations, 2011)." (p. 6)
3) "To enhance objectivity in the scoring process, two independent raters evaluated the writing samples for the pre-tests and post-tests... Inter-rater reliability between the two raters was assessed, demonstrating a high level of agreement with a correlation coefficient of 0.88." (p. 6)
4) "In line with our commitment to maintaining the comparability of assessments, the post-test featured the same writing tasks, Task 1 and Task 2, extracted from the IELTS writing assessment." (p. 7, Table 1 note)
Detailed Analysis:
Criterion E requires the use of a standardised, widely recognised exam-based assessment rather than a custom test designed for the study. The primary academic outcome (writing skill) was measured using IELTS academic writing tasks 1 and 2, scored with the official IELTS band descriptors. IELTS is an internationally recognised standardised English language examination, and the tasks and rubrics were taken from established IELTS materials rather than being author-created instruments. Scoring was performed by two raters, including a trained IELTS instructor, with good inter-rater reliability (0.88). The secondary outcome (motivation) used a questionnaire, but the educational achievement outcome itself rests on a standardised exam framework.
Criterion E is met because the writing outcome was assessed with tasks and band descriptors from IELTS, a widely recognised standardised examination.
-
T
Term Duration
- Outcomes were measured after a 12-week (three-month) intervention period, which covers approximately one academic term.
- "The AI-assisted writing instruction sessions using ChatGPT were conducted twice a week over a period of 12 weeks."
Relevant Quotes:
1) "Over a three-month period, both the experimental and control groups were taught by a male experienced teacher who used the same instructional materials for both groups." (p. 5)
2) "The AI-assisted writing instruction sessions using ChatGPT were conducted twice a week over a period of 12 weeks. Each session lasted for approximately 60 min." (p. 6)
3) "Following the completion of the 12-week intervention period, all participants, both in the control and experimental groups, underwent a post-test." (p. 7, Table 1 note)
4) "During the three-month intervention period, control group participants engaged in writing exercises and activities..." (p. 7)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months, e.g., a semester) after the intervention begins. The intervention ran for 12 weeks (described by the authors as a three-month period) during the regular academic year, and the post-test was administered at the completion of this 12-week period. The interval from intervention start to outcome measurement is therefore approximately three months, which corresponds to one academic term/semester as defined by the standard.
Criterion T is met because the outcome measurement occurred 12 weeks (about three months) after the intervention began, satisfying the one-term minimum.
-
D
Documented Control Group
- The control group's composition, baseline pre-test scores, and business-as-usual treatment are documented in detail.
- "The control group, on the other hand, received traditional writing instruction from an experienced male teacher who used the same instructional materials as the experimental group."
Relevant Quotes:
1) "The control group, on the other hand, received traditional writing instruction from an experienced male teacher who used the same instructional materials as the experimental group. Participants attended in-person writing classes led by the teacher." (p. 7)
2) "The participants came from diverse backgrounds, representing various regions of China. On average, they had been studying English as a foreign language for five years. Their ages ranged from 18 to 22 years old, with the majority being in their second or third year of undergraduate studies." (p. 5)
3) "To establish a comparable proficiency level, an English proficiency test was administered as part of the screening process. This comprehensive test assessed participants' reading, writing, listening, and speaking skills, confirming their similar level of English proficiency confirming their intermediate to upper-intermediate level of English proficiency." (p. 5)
4) "The mean pre-test score for overall writing was 39.26 (SD = 12.03) in the experimental group and 37.31 (SD = 17.26) in the control group." (p. 8, Table 2)
5) "Unlike the experimental group, the control group did not receive AI-assisted feedback from ChatGPT. Instead, their feedback was based on the teacher's expertise and teaching experience." (p. 7)
Detailed Analysis:
Criterion D requires detailed documentation of the control group, including its composition, baseline performance, and the treatment it received. The paper documents the control condition in a dedicated section (3.3.2) describing what the control group received (traditional teacher-led writing instruction, same materials, same teacher, same 12-week duration, teacher feedback), and Table 1 summarises the intervention details side by side. Baseline pre-test means and standard deviations for the control group are reported for all outcome measures in Table 2, and the sample's demographics and screening-verified comparable proficiency are described. This is sufficient to assess comparability of the control group.
Criterion D is met because the control group's conditions, baseline performance, and characteristics are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- All participants came from one university and were randomised individually, so there was no school-level randomisation.
- "The study included 50 Chinese EFL students who were enrolled in a Bachelor's degree program at a national university in China."
Relevant Quotes:
1) "Random assignment was used to assign participants to either the experimental group or the control group." (p. 5)
2) "The study included 50 Chinese EFL students who were enrolled in a Bachelor's degree program at a national university in China." (p. 5)
3) "The instruction sessions of both groups were held in a dedicated computer laboratory equipped with the necessary technology and internet access." (p. 7)
Detailed Analysis:
Criterion S requires randomisation among schools or equivalent implementing institutions. This study took place at a single university, and randomisation was performed at the level of individual students within that one institution. There was no assignment of multiple schools, sites, or institutional units to conditions.
Criterion S is not met because randomisation occurred at the individual student level within a single university, not at the school level.
-
I
Independent Conduct
- The authors themselves designed, implemented, rated, and analysed the study with no independent third-party evaluation.
- "The raters consisted of the researcher/instructor (second author) and an experienced IELTS instructor who had received official training at an IELTS training center."
Relevant Quotes:
1) "The raters consisted of the researcher/instructor (second author) and an experienced IELTS instructor who had received official training at an IELTS training center." (p. 6)
2) "CS: Conceptualization, Data curation, Investigation, Methodology, Project administration, Resources, Validation, Visualization, Writing - original draft, Writing - review & editing. YS: Data curation, Formal analysis, Investigation, Methodology, Project administration, Resources, Software, Visualization, Writing - original draft, Writing - review & editing." (p. 13, Author contributions)
3) "An experienced EFL researcher, proficient in coding and labeling, collaborated in the validation of the processes." (pp. 7-8)
Detailed Analysis:
Criterion I requires that the study be conducted independently of those who designed the intervention, or at least that data collection and analysis be handled by an external, third-party team. Here the two authors designed the study, administered the intervention, collected the data, and analysed it themselves. Notably, one of the two raters who scored the writing samples was the researcher/instructor (second author), meaning outcome assessment was partly performed by the study team itself. The involvement of one external IELTS-trained rater and an EFL researcher assisting with qualitative coding does not constitute independent conduct of the trial, and no external evaluation agency or third-party oversight is mentioned anywhere in the paper.
Criterion I is not met because the same authors designed, delivered, scored, and analysed the intervention without documented independent oversight.
-
Y
Year Duration
- The tracking interval was only 12 weeks, far below the required 75% of a full academic year.
- "The AI-assisted writing instruction sessions using ChatGPT were conducted twice a week over a period of 12 weeks."
Relevant Quotes:
1) "The AI-assisted writing instruction sessions using ChatGPT were conducted twice a week over a period of 12 weeks." (p. 6)
2) "Following the completion of the 12-week intervention period, all participants, both in the control and experimental groups, underwent a post-test." (p. 7, Table 1 note)
3) "Another limitation of the study is the relatively short duration of the intervention. The impact of AI-assisted language learning on writing skills and motivation was evaluated over a limited period, potentially limiting our understanding of the sustained effects of such instruction." (p. 12)
Detailed Analysis:
Criterion Y requires outcome measurement at least 75% of a full academic year (roughly 7+ months of a 9-10 month year) after the intervention begins. The interval from intervention start to post-test in this study was 12 weeks (about three months), well short of 75% of an academic year. The authors themselves flag the short duration as a limitation, and no later follow-up measurement was conducted ("the study did not include a long-term follow-up", p. 12).
Criterion Y is not met because outcomes were measured only about three months after the intervention began, far less than 75% of an academic year.
-
B
Balanced Control Group
- The active control received the same duration, teacher, materials, and verified-equivalent practice time, with only the feedback source (ChatGPT vs teacher) differing as the treatment variable.
- "It is also worth noting that the equivalence of time spent on out-of-class practice between the experimental and control groups was meticulously maintained through the implementation of a time log system."
Relevant Quotes:
1) "The control group, on the other hand, received traditional writing instruction from an experienced male teacher who used the same instructional materials as the experimental group." (p. 7)
2) "During the three-month intervention period, control group participants engaged in writing exercises and activities that included a combination of classroom exercises and topics of interest, similar to the experimental group." (p. 7)
3) "The amount of time devoted to outside-class practice was designed to be comparable to that of the experimental group." (p. 7)
4) "It is also worth noting that the equivalence of time spent on out-of-class practice between the experimental and control groups was meticulously maintained through the implementation of a time log system." (p. 7)
5) "The systematic scrutiny of these time logs served as a robust mechanism to affirm the comparability of out-of-class writing practice time between the two groups. This rigorous monitoring and verification process unequivocally ensured that both groups had equitable opportunities for practice and improvement." (p. 7)
6) "The instruction sessions of both groups were held in a dedicated computer laboratory equipped with the necessary technology and internet access. This setting ensured a controlled and consistent environment for both the experimental and control groups." (p. 7)
7) Table 1: "Duration: 12 weeks / 12 weeks; Writing tasks: Classroom exercises + topics of interest (both groups); Feedback: AI-powered (ChatGPT) / Teacher; Interactive learning: Yes (with ChatGPT) / Yes (teacher-led classes); Progress tracking: Yes (writing portfolio) / Yes (writing portfolio)." (p. 7)
Detailed Analysis:
Criterion B requires that the control group receive comparable time, budget, and educational inputs, unless the extra resource is itself the treatment variable. Applying the decision tree: the intervention (ChatGPT feedback) replaces rather than supplements teacher feedback. Both groups had the same 12-week duration, the same instructional materials, the same teacher, the same types of writing tasks, the same computer laboratory setting, interactive learning and progress tracking in both conditions, and out-of-class practice time explicitly matched and verified through a time log system. The control was an active condition receiving teacher-led instruction and individualized teacher feedback, i.e., a comparable substitute for the AI feedback. The difference in feedback source (AI vs teacher) is the treatment contrast itself, and the experimental group's flexible home access was monitored so practice time stayed comparable. No unmatched extra time or budget was given to the intervention group.
Criterion B is met because both groups received matched instructional time, materials, teacher, tasks, and monitored out-of-class practice, with the feedback source being the integral treatment variable.
-
Level 3 Criteria
-
R
Reproduced
- The authors explicitly call for future replication, and an internet citation-index search found no independent peer-reviewed replication of this specific study.
- "Replication studies with more diverse populations are needed to validate the effectiveness of AI-assisted language learning in different settings."
Relevant Quotes:
1) "Replication studies with more diverse populations are needed to validate the effectiveness of AI-assisted language learning in different settings." (p. 12)
2) "Similarly, Yan (2023) investigated the impact of ChatGPT, an AI-assisted language learning tool, on the writing skills of EFL learners and reported significant improvements in their writing performance as a result of AI-assisted language learning." (p. 2)
3) "Nazari et al. (2021) conducted a true experimental study investigating the effects of AI-assisted language learning on EFL learners' writing performance." (p. 4)
Detailed Analysis:
Criterion R requires that this specific study be independently replicated by a different team, in a different context, in a peer-reviewed journal. The paper itself calls for future replication, indicating none existed at publication. Related studies cited in the literature review (e.g., Yan 2023; Nazari et al. 2021) predate this study and investigate related but distinct AI writing tools, designs, and populations; they are background literature, not replications of this particular 12-week ChatGPT RCT with Chinese EFL undergraduates.
Internet Verification (2026-07-27): The citation record for this paper (DOI 10.3389/fpsyg.2023.1260843) was queried via Semantic Scholar's citing-papers index to identify any independent replication published after this study. Papers citing this article were reviewed, including studies on ChatGPT versus teacher feedback for EFL writing with other populations (e.g., a study of Turkish B2-level learners comparing ChatGPT and teacher feedback, and a study of Thai university students examining AI feedback on IELTS writing). These studies address related research questions but use different populations, contexts, and study designs, and none of them describe themselves as an independent replication of this specific Song and Song (2023) trial. No peer-reviewed publication reproducing this exact study's design, sample, and intervention by a different research team was found.
Criterion R is not met because no independent peer-reviewed replication of this specific study is documented in the paper or found via citation-index search.
-
A
All-subject Exams
- Only English writing outcomes were measured, not all main subjects studied by the participants.
- "Our dependent variables encompassed global writing performance and specific writing skills, including writing content, organization, and language use, as well as writing motivation."
Relevant Quotes:
1) "The study utilized the International English Language Testing System (IELTS) academic writing tasks 1 and 2 as pre-tests and post-tests to evaluate participants' academic writing skills." (p. 6)
2) "Our dependent variables encompassed global writing performance and specific writing skills, including writing content, organization, and language use, as well as writing motivation." (p. 7)
3) "Participants pursued majors in disciplines such as engineering, business, social sciences, and humanities." (p. 5)
Detailed Analysis:
Criterion A requires standardised exam-based assessment across all main subjects taught at the educational level, to detect possible negative spillovers on non-target subjects. This study measured only English academic writing (plus a writing motivation questionnaire). The participants were undergraduates in varied majors (engineering, business, social sciences, humanities), yet no outcomes in their major subjects or any other academic domain were assessed. The specialised-intervention exception applies to upper secondary or vocational education with an explicit rationale; the paper provides no such justification for restricting measurement to one skill area. Only a single subject domain was assessed.
Criterion A is not met because only English writing was assessed, with no measurement of other main subjects and no qualifying justification.
-
G
Graduation Tracking
- Measurement ended at the 12-week post-test, the authors explicitly state no long-term follow-up was conducted, and no graduation-tracking follow-up paper was found online.
- "Lastly, the study did not include a long-term follow-up to assess the sustainability of the observed improvements in writing skills and motivation."
Relevant Quotes:
1) "Lastly, the study did not include a long-term follow-up to assess the sustainability of the observed improvements in writing skills and motivation. Without a post-intervention evaluation, it remains unknown whether the benefits of AI-assisted instruction persist over an extended period of time." (p. 12)
2) "Following the completion of the 12-week intervention period, all participants, both in the control and experimental groups, underwent a post-test." (p. 7, Table 1 note)
Detailed Analysis:
Criterion G requires tracking participants until graduation from their educational stage. The participants were second- and third-year undergraduates, and measurement ended immediately after the 12-week intervention. The authors explicitly state that no long-term follow-up was conducted.
Internet Verification (2026-07-27): A citation-index search (Semantic Scholar citing-papers for DOI 10.3389/fpsyg.2023.1260843) was conducted to check for subsequent publications by Cuiping Song or Yanping Song that might track this same cohort toward graduation. No such follow-up publication by either author was found among the papers citing this study, and no later paper by these authors referencing this cohort could be located. Additionally, per the ranking instructions, criterion G cannot be met when criterion Y (Year Duration) is not met, and Y is not met here.
Criterion G is not met because tracking stopped at the 12-week post-test with no follow-up to graduation found in the paper or via internet search, and the prerequisite criterion Y is also unmet.
-
P
Pre-Registered
- The paper contains no mention of a pre-registered protocol or any trial registry entry, and none was found via internet search.
Relevant Quotes:
1) "The studies involving humans were approved by the School of Foreign Studies, North Minzu University, Yinchuan, Ningxia. The studies were conducted in accordance with the local legislation and institutional requirements." (p. 13, Ethics statement)
2) "Prior to their involvement, participants provided written informed consent after receiving detailed explanations of the study's purpose, procedures, potential risks, and benefits." (p. 5)
Detailed Analysis:
Criterion P requires pre-registration of the full study protocol (hypotheses, methods, planned analyses) on a public registry before data collection began. The paper contains no mention of any registry (e.g., ClinicalTrials.gov, OSF, AEA/AsPredicted), no registration ID, and no registration date. The ethics approval and informed consent statements concern institutional review, which is distinct from public pre-registration of a study protocol.
Internet Verification (2026-07-27): Since the paper contains no registry name or registration ID to check, no corresponding pre-registration record could be located or verified in any registry database.
Criterion P is not met because no pre-registration statement, registry link, or registration date appears anywhere in the paper, and no registry entry for this study was found.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.