Abstract
A major challenge in educational technology integration is to engage students with different affective characteristics. Also, how technology shapes attitude and learning behavior is still lacking. Findings from educational psychology and learning sciences have gained less traction in research. The present study was conducted to examine the efficacy of a group format of an Artificial Intelligence (AI) powered writing tool for English second postgraduate students in the English academic writing context. In the present study, (N = 120) students were randomly allocated to either the equipped AI (n = 60) or non-equipped AI (NEAI). The results of the parametric test of analyzing of covariance revealed that at post-intervention, students who participated in the AI intervention group demonstrated statistically significant improvement in the scores, of the behavioral engagement (Cohen's d = .75, 95% CI [0.38, 1.12]), of the emotional engagement Cohen's d = .82, 95% CI [0.45, 1.25], of the cognitive engagement, Cohen's d = .39, 95% CI [0.04, .76], of the self-efficacy for writing, Cohen's d = .54, 95% CI [0.18, 0.91], of the positive emotions Cohen's d = .44, 95% CI [0.08, 0.80], and of the negative emotions, Cohen's d = -.98, 95% CI [-1.36, -0.60], compared with NEAI. The results suggest that AI-powered writing tools could be an efficient tool to promote learning behavior and attitudinal technology acceptance through formative feedback and assessment for non-native postgraduate students in English academic writing.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was carried out at the individual student level using a stratified block design, not at the class or school level, and the intervention was a group course rather than one-to-one tutoring.
- "Finally, the allocation was carried out using a block size method, stratified by gender and field of the study."
Relevant Quotes:
1) "In the present study, (N = 120) students were randomly allocated to either the equipped AI (n = 60) or non-equipped AI (NEAI)." (p. 1, Abstract)
2) "Finally, the allocation was carried out using a block size method, stratified by gender and field of the study." (p. 4)
3) "Randomization was performed using the block size stratified method. The block size was 6, stratified by gender and field. An independent statistician carried out the randomization and informed the participants and research team members about the allocation." (p. 5)
Detailed Analysis:
The paper explicitly describes randomisation as being performed on individual students (using a block-size stratified method with strata of gender and field of study), not on entire classes or schools. Participants were recruited individually through public online forums and social media invitations rather than as pre-existing classes, and were then allocated one by one to the AI or NEAI condition. There is no description of classes or schools being the unit of randomisation. The ERCT exception for personal tutoring/one-to-one teaching interventions does not apply here because the AI intervention was delivered as a 12-week group course to cohorts of about 60 students in a virtual classroom, not as individual one-to-one tutoring.
Because randomisation occurred at the student level without a valid tutoring exception, criterion C is not met.
-
E
Exam-based Assessment
- Outcomes were measured with self-report psychological scales (engagement, self-efficacy, emotions), not with a standardised, widely recognised exam-based assessment.
- "Student Engagement scale. (Fredricks et al., 2005). The student engagement scale is a 19-item self-report measure that determines to measure the students' engagement in three subscales."
Relevant Quotes:
1) "Student Engagement scale. (Fredricks et al., 2005). The student engagement scale is a 19-item self-report measure that determines to measure the students' engagement in three subscales." (p. 4)
2) "Self-Efficacy for Writing Scale (SEWS; Bruning et al., 2013). The SEWS, a 16-item self-report measure, was used to assess self-efficacy for writing." (p. 4)
3) "Achievement emotion questionnaire (AEQ; Pekrun et al., 2011). The AEQ was used to measure negative and positive emotions." (p. 4)
Detailed Analysis:
All primary outcomes in this study (behavioral, emotional and cognitive engagement; self-efficacy for writing; positive and negative academic emotions) are derived from self-report Likert-scale questionnaires measuring psychological constructs. None of these instruments is a standardised academic exam (e.g. a national curriculum test or writing proficiency exam) designed to objectively assess academic knowledge or skill; they are self-report psychometric scales developed for research on motivation/engagement, not standardised achievement tests. No standardised exam-based assessment of writing performance or any other subject was used anywhere in the study.
Because no standardised exam-based assessment was employed, criterion E is not met.
-
T
Term Duration
- The intervention lasted only 12 weeks and outcomes were measured immediately at the end of the intervention, so there is no documented interval of at least one full academic term from intervention start to measurement.
- "AI group intervention consists of 12 weekly two-hour sessions."
Relevant Quotes:
1) "AI group intervention consists of 12 weekly two-hour sessions." (p. 5)
2) "The NEAI intervention consists of 12 weekly two-hour sessions." (p. 5)
3) "The outcomes were assessed at two time-points: Time 1: pre-intervention to pre-allocation includes baseline, Time 2: immediate after intervention: post-intervention assessment." (p. 4-5)
Detailed Analysis:
The intervention consisted of 12 weekly two-hour sessions, i.e. approximately 12 weeks (roughly 2.5-3 months) of total duration, and the outcome measure ("Time 2") was collected immediately after the intervention concluded rather than after any additional term-length follow-up period. No explicit calendar dates are given for intervention start and end, and no statement clarifies that the local academic term is shorter than the conventional 3-4-month definition used by the standard. Since the total elapsed time from intervention start to outcome measurement is at or below the lower bound of a typical academic term, and there is no explicit confirmation that this interval reaches a full term, the documentation does not clearly establish that the term-duration threshold was met.
Because the interval from intervention start to outcome measurement is not clearly documented as covering at least one full academic term, criterion T is not met.
-
D
Documented Control Group
- The control (NEAI) group's demographics, size, and the content it received are clearly documented, including baseline and post-test descriptive statistics.
- "The NEAI course was knowledge-based. The NEAI intervention consists of 12 weekly two-hour sessions. The course contents were the same as the AI group, except for Grammarly contents."
Relevant Quotes:
1) "The NEAI included 46.7% females (n = 28) and 53.3% males (n = 32). NEAI participants' ages ranged from 26 to 39 years (M = 32.23 years, SD = 3.16). Also, NEAI group consists of 20% (n = 12) in humanity sciences, 41.7% (n = 25) technology sciences, and 38.3% (n = 23) health sciences." (p. 5)
2) "The NEAI course was knowledge-based. The NEAI intervention consists of 12 weekly two-hour sessions. The course contents were the same as the AI group, except for Grammarly contents." (p. 5)
3) "Also, 30% (n = 18) of the NEAI group dropped out post-intervention." (p. 5)
4) Table 3 reports NEAI Time 1 and Time 2 means and standard deviations for self-efficacy, engagement, positive emotion, and negative emotion. (p. 6)
Detailed Analysis:
The paper provides a clear demographic breakdown of the control (NEAI) group (gender split, age range and mean, field of study distribution), states exactly what the control group received (a 12-session knowledge-based course identical in structure to the AI group except for the Grammarly component), reports attrition figures, and provides baseline and post-intervention descriptive statistics for the control group in Table 3. This level of documentation allows readers to assess the control group's comparability and treatment.
Because the control group's characteristics, size, and content received are documented in sufficient detail, criterion D is met.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred at the individual student level, not at the school (or institution) level.
- "Randomization was performed using the block size stratified method. The block size was 6, stratified by gender and field."
Relevant Quotes:
1) "Finally, the allocation was carried out using a block size method, stratified by gender and field of the study." (p. 4)
2) "Randomization was performed using the block size stratified method. The block size was 6, stratified by gender and field. An independent statistician carried out the randomization..." (p. 5)
Detailed Analysis:
The unit of randomisation described throughout the paper is the individual postgraduate student, recruited via online forums/social media across what appears to be a single national university system, not distinct schools or institutions randomised as whole units. There is no mention of multiple schools being randomly assigned to conditions.
Since randomisation was at the individual student level rather than the school level, criterion S is not met.
-
I
Independent Conduct
- The same author who conceived and designed the intervention also performed the experiment and analyzed the data, with no independent external evaluation team running or analyzing the trial.
- "Nabi Nazari: Conceived and designed the experiments; Performed the experiments; Analyzed and interpreted the data; Contributed reagents, materials, analysis tools or data; Wrote the paper."
Relevant Quotes:
1) "Nabi Nazari: Conceived and designed the experiments; Performed the experiments; Analyzed and interpreted the data; Contributed reagents, materials, analysis tools or data; Wrote the paper." (p. 8, Author contribution statement)
2) "An independent statistician carried out the randomization and informed the participants and research team members about the allocation." (p. 5)
3) "The instructors (two associate professors) were not informed about the groups and aim of the study." (p. 5)
Detailed Analysis:
While the randomisation sequence was generated by an independent statistician and the instructors/assessors were blinded to condition, the author contribution statement shows that the same author (Nazari) who conceived and designed the intervention/experiment also personally performed the experiment and analyzed and interpreted the data. There is no external, independent evaluation agency or third-party researcher who was separate from the intervention's designers and who conducted the overall data collection and analysis of the trial. Blinding of instructors and use of an independent randomiser are partial safeguards, but they do not amount to the independent conduct of the evaluation as a whole required by this criterion.
Because the same team that designed the intervention also performed and analyzed the study without an independent evaluating body, criterion I is not met.
-
Y
Year Duration
- Since criterion T (Term Duration) is not met and the total study duration (12 weeks) falls far short of an academic year, criterion Y is not met.
- "AI group intervention consists of 12 weekly two-hour sessions."
Relevant Quotes:
1) "AI group intervention consists of 12 weekly two-hour sessions." (p. 5)
2) "The outcomes were assessed at two time-points: Time 1: pre-intervention to pre-allocation includes baseline, Time 2: immediate after intervention: post-intervention assessment." (p. 4-5)
Detailed Analysis:
The study's entire intervention and follow-up period spans only 12 weeks, with outcomes measured immediately at the end of the intervention. This is far shorter than the required 75% of an academic year (roughly 9-10 months). Per the standard's dependency rule, since criterion T (Term Duration) is not met, criterion Y is automatically not met as well.
Criterion Y is not met.
-
B
Balanced Control Group
- Both the AI and NEAI groups received an identical number and length of sessions and the same course content, differing only in whether the Grammarly tool (the treatment variable) was included.
- "The NEAI intervention consists of 12 weekly two-hour sessions. The course contents were the same as the AI group, except for Grammarly contents."
Relevant Quotes:
1) "AI group intervention consists of 12 weekly two-hour sessions." (p. 5)
2) "The NEAI intervention consists of 12 weekly two-hour sessions. The course contents were the same as the AI group, except for Grammarly contents." (p. 5)
3) "To increase retention, the research team provides the AI course for NEAI participants after the study." (p. 5)
Detailed Analysis (re-applied against the updated criterion B decision tree):
EXTRA_RESOURCES_PRESENT = true: the AI group had access to the Grammarly writing tool, which the NEAI group did not. RESOURCES_ARE_TREATMENT = true: the study's explicit purpose is to test the efficacy of the AI-powered writing assistant against a non-equipped condition, so Grammarly access is precisely the treatment variable under investigation, not a supplementary add-on. Per the decision tree, when RESOURCES_ARE_TREATMENT is true the control may remain business-as-usual by design and the criterion is met without needing the control to be given the tool itself. Independently of that branch, both groups also received exactly the same number of sessions (12 weekly two-hour sessions), the same instructors, and the same course curriculum apart from the Grammarly component, so all other time and instructional resources are matched. The NEAI group was also given access to the AI course after the study concluded, indicating no permanent resource disadvantage.
Because the only differing resource is the specific technology under investigation (the treatment variable itself), and all other time and instructional inputs are balanced between arms, criterion B is met.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific trial by a different research team was found in the paper or through internet searches of citing literature.
Relevant Quotes:
No quote in the paper announces or references a replication of this study.
Detailed Analysis:
The paper cites several other studies on automated writing evaluation/Grammarly (e.g. Koltovskaia, 2020; Zhang, 2020; Parra and Calero, 2019; Cavaleri and Dianati, 2016), but these are prior studies on related but distinct designs and outcomes, not independent replications of this specific randomized controlled trial (same design, same outcome constructs, different research team, different context) published in a peer-reviewed journal.
Internet verification: this paper's OpenAlex record (W3160941578, DOI 10.1016/j.heliyon.2021.e07014) shows 384 citing works. Filtering the citing-works list for full-text mentions of "replicat" returned zero results, and a manual scan of citing authors did not surface any team reporting an independent replication of this specific trial (same AI writing-assistant vs. NEAI design in a comparable postgraduate ESL population). No subsequent peer-reviewed replication of this precise trial was identified.
Because no independent replication of this specific study was found either in the paper or via internet search, criterion R is not met.
-
A
All-subject Exams
- Since criterion E (Exam-based Assessment) is not met and only affective/engagement constructs were measured rather than academic performance across subjects, criterion A is not met.
- "Student Engagement scale. (Fredricks et al., 2005). The student engagement scale is a 19-item self-report measure..."
Relevant Quotes:
1) "Student Engagement scale. (Fredricks et al., 2005)... a 19-item self-report measure that determines to measure the students' engagement in three subscales." (p. 4)
2) "Self-Efficacy for Writing Scale (SEWS; Bruning et al., 2013)." (p. 4)
3) "Achievement emotion questionnaire (AEQ; Pekrun et al., 2011)." (p. 4)
Detailed Analysis:
Per the standard, criterion E is a prerequisite for criterion A, and E was not met because no standardised exam-based assessment was used. In addition, the study only measured writing-related engagement, self-efficacy, and academic emotions -- it did not assess performance across other core subjects (e.g. mathematics, science) using standardised exams.
Because the exam-based assessment prerequisite is not satisfied and no all-subject standardised exam coverage exists, criterion A is not met.
-
G
Graduation Tracking
- Criterion Y (Year Duration) is not met, and no follow-up publication tracking this cohort toward graduation was found; the study measured outcomes only immediately after the 12-week intervention ended.
- "The outcomes were assessed at two time-points: Time 1: pre-intervention to pre-allocation includes baseline, Time 2: immediate after intervention: post-intervention assessment."
Relevant Quotes:
1) "The outcomes were assessed at two time-points: Time 1: pre-intervention to pre-allocation includes baseline, Time 2: immediate after intervention: post-intervention assessment." (p. 4-5)
2) "The results from this trial should be interpreted in the context of several limitations." (p. 7, Limitation section, with no mention of long-term follow-up)
Detailed Analysis:
The paper explicitly describes only two measurement time-points: a pre-intervention baseline and an immediate post-intervention assessment. There is no mention of any further follow-up, let alone tracking participants until graduation, and the Limitations section does not reference any planned or completed long-term follow-up study.
Per the standard's dependency rule, since criterion Y (Year Duration) is not met, criterion G is automatically not met as well.
Internet verification: an OpenAlex author search confirmed the correct Nabi Nazari record (OpenAlex ID A5075725391, ORCID 0000-0003-3771-3699, Lorestan University). Reviewing this author's listed publications (clinical/health psychology topics and two AEA RCT registry entries on unrelated transdiagnostic-therapy trials) found no follow-up paper tracking the AI-writing-assistant cohort from this study, and no other paper by any author reporting graduation-level outcomes for these participants was identified.
Because tracking stopped immediately after the intervention, criterion Y (a prerequisite) is not met, and no graduation-level follow-up publication was found, criterion G is not met.
-
P
Pre-Registered
- The paper mentions IRB ethical review but provides no registry name, ID, or date confirming pre-registration before data collection began, and no internet search surfaced a matching trial registration.
- "The IRB reviewed the research protocol to ensure ethical considerations, participant confidentiality, sampling, and obtaining informed consent."
Relevant Quotes:
1) "The IRB reviewed the research protocol to ensure ethical considerations, participant confidentiality, sampling, and obtaining informed consent." (p. 4)
2) "Additional information" / "No additional information is available for this paper." (p. 8)
Detailed Analysis:
The only procedural/ethical safeguard mentioned is institutional review board (IRB) approval, which concerns ethics oversight, not a public pre-registration of the study's hypotheses, methods, and planned analyses on a trial registry (e.g. ClinicalTrials.gov, ISRCTN, OSF Registries, AEA RCT Registry). No registry name, registration ID, or registration date is given anywhere in the paper, and the "Additional information" declaration explicitly states none is available.
Internet verification: searches of OpenAlex and the AEA RCT Registry for this specific Grammarly/AI writing-assistant trial (by title and by author Nabi Nazari) returned no matching pre-registration record; the AEA registry entries found under this author's name concern unrelated transdiagnostic-therapy trials, not this study.
Because there is no evidence of a pre-registered protocol for this study in the paper or in trial registries, criterion P is not met.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.