AI meets education: How ChatGPT transforms reading skills in Omani EFL learners

Behnam Behforouz, Ali Al Ghaithi

Published:
ERCT Check Date:
DOI: 10.22363/2521-442X-2025-9-3-10-21
  • reading
  • L2 languages
  • higher education
  • Asia
  • blended learning
  • EdTech app
0
  • C

    Randomisation was carried out at the individual student level within a single institution's cohort, not at the class or school level, and the study is classroom-based rather than one-to-one tutoring, so the tutoring exception does not apply.

    "50 Omani EFL learners from a higher education institution in Oman were randomly assigned to an experimental group and a control group, with 25 students in each group..." (Section 3.1)

  • E

    The reading tests used to measure outcomes were researcher-developed and modified rather than a pre-existing widely recognised standardised exam.

    "To compare and measure the performance of students in both groups on reading skills, three sets of tests, including pretest, posttest, and delayed posttest, were designed." (Section 3.2.1, p. 13)

  • T

    Outcomes were measured within about two months of intervention start (one month of treatment plus a delayed posttest three weeks after treatment ended), well short of a full academic term.

    "The treatment period lasted for a month... A delayed posttest was conducted to measure students' knowledge retention... three weeks after the treatment period." (Section 3.4, p. 14)

  • D

    The control group's size, demographic composition, and the instruction/activities it received are clearly documented in the methods section.

    "The control group received training and instruction on reading skills, finding answers to various questions, and some reading techniques, such as skimming and scanning, with extra practices within the classroom and through traditional face-to-face teaching techniques." (Section 3.4, p. 14)

  • S

    Randomisation occurred among individual students within a single institution, not among schools or institutions.

    "50 Omani EFL learners from a higher education institution in Oman were randomly assigned to an experimental group and a control group..." (Section 3.1, p. 13)

  • I

    The same two authors designed the intervention, delivered the instruction/oversight, and conducted the analysis, with no independent evaluator described.

    "CREdIT AUTHOR STATEMENT: Behnam Behforouz: Formal Analysis, Writing – Review & Editing, Project Administration. Ali Al Ghaithi: Methodology, Writing – Original Draft, Resources." (p. 10)

  • Y

    Since criterion T (Term Duration) is not met, this stronger Year Duration criterion is automatically not met; in any case the total tracked period was about two months, far short of a year.

    "The treatment period lasted for a month... A delayed posttest was conducted... three weeks after the treatment period." (Section 3.4, p. 14)

  • B

    The additional ChatGPT-based practice is the explicit treatment variable the study set out to test, with the control group receiving a comparable "business as usual" package of face-to-face instruction and extra quizzes.

    "...to what extent does the use of ChatGPT as a facilitator improve the reading comprehension abilities of Omani EFL learners compared to traditional face-to-face instruction?" (Section 1, p. 11-12)

  • R

    No independent replication of this specific study by a different research team was found; the closest related work is a prior study by the same authors.

    "Behforouz, B., & Al Ghaithi, A. (2024). Investigating the effect of an interactive educational chatbot on reading comprehension skills." (References, p. 18)

  • A

    Since criterion E (Exam-based Assessment) is not met, this stronger All-subject Exams criterion is automatically not met; in any case only reading was assessed.

    "This study depended primarily on reading tests for assessment; hence, other aspects of language learning had not been taken into consideration, like writing or speaking." (Section 6, p. 18)

  • G

    Since criterion Y (Year Duration) is not met, this stronger Graduation Tracking criterion is automatically not met, and no follow-up beyond the three-week delayed posttest is reported.

    "A delayed posttest was conducted to measure students' knowledge retention and the efficacy of using ChatGPT three weeks after the treatment period." (Section 3.4, p. 14)

  • P

    No pre-registration of the study protocol, registry link, or registration date is mentioned anywhere in the paper, and no registry entry was found via internet search.

Abstract

The main objective of this study was to measure the impact of ChatGPT on the reading skills of language learners. Therefore, a total of fifty Omani students with intermediate English proficiency were selected and randomly assigned into two groups, one control group and an experimental group, with an equal number of students in each group. Both groups received the traditional face-to-face training to engage with the reading comprehension skills, techniques, and strategies in understanding the texts and finding answers for various types of questions arising from the test, but the experimental group received extra explanation and practice from ChatGPT. To compare the results in both groups, the researcher developed and modified some reading tests. Their reliability and validity were measured and monitored by the experts. After a month of treatment, the findings revealed that both groups initially received higher scores in the reading posttests compared to their pretests, but the experimental group performed significantly better than the control group. Additionally, further analysis of the delayed posttests of reading showed that the control group had no increment in their scores while the experimental group continued its progress, and performance was significantly higher, suggesting improved retention of reading abilities. The results of this study are useful for teachers, students, and educational institutions.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was carried out at the individual student level within a single institution's cohort, not at the class or school level, and the study is classroom-based rather than one-to-one tutoring, so the tutoring exception does not apply.
      • "50 Omani EFL learners from a higher education institution in Oman were randomly assigned to an experimental group and a control group, with 25 students in each group..." (Section 3.1)
      • Relevant Quotes: 1) "To conduct this quasi-experimental research study, 50 Omani EFL learners from a higher education institution in Oman were randomly assigned to an experimental group and a control group, with 25 students in each group, comprising both males and females." (Section 3.1, p. 13) 2) "These students were studying in the Foundation Programme, in which they had to study some modules on English, Math, and IT..." (Section 3.1, p. 13) 3) "The control group received training and instruction on reading skills... within the classroom and through traditional face-to-face teaching techniques... the experimental group received in-class instructions on reading skills and techniques... the students were instructed to use ChatGPT to practice extra reading comprehension activities..." (Section 3.4, p. 14) Detailed Analysis: The paper explicitly states that individual students, not entire classes, were randomly assigned to the experimental and control conditions ("50 Omani EFL learners... were randomly assigned"). There is no mention of randomising intact classes or the school/ institution as the unit of assignment; instead, individual students from a single cohort in a Foundation Programme were split into two groups of 25. This is a student-level design, which under the ERCT Standard is insufficient unless the tutoring exception applies. The tutoring exception is designed for one-to-one personal instruction interventions. Here, both groups received standard classroom instruction in reading (a group/classroom activity, "two sessions, each lasting one hour and forty minutes"), and the experimental group additionally used ChatGPT individually outside class as a supplementary practice tool. Since the core instructional delivery remained a classroom-based group activity rather than personal one-to-one tutoring, the exception for personal/tutoring interventions does not clearly apply here. Final: Criterion C is not met because randomisation occurred at the individual student level within one cohort, without qualifying for the tutoring exception.
    • E

      Exam-based Assessment

      • The reading tests used to measure outcomes were researcher-developed and modified rather than a pre-existing widely recognised standardised exam.
      • "To compare and measure the performance of students in both groups on reading skills, three sets of tests, including pretest, posttest, and delayed posttest, were designed." (Section 3.2.1, p. 13)
      • Relevant Quotes: 1) "To compare the results in both groups, the researcher developed and modified some reading tests. Their reliability and validity were measured and monitored by the experts." (Abstract, p. 10) 2) "To compare and measure the performance of students in both groups on reading skills, three sets of tests, including pretest, posttest, and delayed posttest, were designed. The tests were aligned with three types of questions, including five multiple-choice questions, five true and false questions, five matching questions, three fill-in-the-blank questions, and two short-answer questions." (Section 3.2.1, p. 13) 3) "To ensure the reliability of the tests, a pilot study was conducted before the main round of the study with 25 random Omani EFL learners... Following the reliability, the questions were reviewed by two Omani PhD holders in applied Linguistics to validate the questions." (Section 3.2.1, p. 13) Detailed Analysis: The paper is explicit that the pretest, posttest, and delayed posttest were "designed" and "developed and modified" by the researchers themselves, drawing on reading passages from a commercial textbook (NorthStar3). Although the authors report Cronbach's Alpha reliability values (0.856-0.870) and had the items reviewed for content validity by two PhD holders, this is a custom-built assessment created specifically for this study, not a widely recognised standardised exam (e.g., a national or international standardised reading test). The ERCT Standard requires a standard, widely recognised test rather than a locally validated but custom-made instrument. Final: Criterion E is not met because the outcome measure is a custom-developed test rather than a standardised exam.
    • T

      Term Duration

      • Outcomes were measured within about two months of intervention start (one month of treatment plus a delayed posttest three weeks after treatment ended), well short of a full academic term.
      • "The treatment period lasted for a month... A delayed posttest was conducted to measure students' knowledge retention... three weeks after the treatment period." (Section 3.4, p. 14)
      • Relevant Quotes: 1) "The present investigation was conducted during the autumn semester of 2024-2025... The treatment period lasted for a month, and according to the curriculum and delivery plan, two sessions, each lasting one hour and forty minutes, focused on reading comprehension activities..." (Section 3.4, p. 14) 2) "The following week after the end of the treatment, a posttest was conducted to compare the performance of both groups before and after the treatment. A delayed posttest was conducted to measure students' knowledge retention and the efficacy of using ChatGPT three weeks after the treatment period." (Section 3.4, p. 14) 3) "Another limitation was that the intervention period was short, lasting only one month; this might not be enough to gauge the long-term effects or sustainability of improvements in reading comprehension." (Section 6, p. 18) Detailed Analysis: The intervention itself ran for one month, the posttest followed about a week after the treatment ended, and the final (delayed) measurement occurred three weeks after that. In total, the interval from intervention start to final measurement is roughly two months, which is shorter than a typical academic term (~3-4 months). The authors themselves flag the short one-month intervention period as a limitation of the study, and no year-long tracking is reported that would otherwise satisfy this criterion by the stronger Y route. Final: Criterion T is not met because the total follow-up from intervention start to final measurement is approximately two months, short of a full academic term.
    • D

      Documented Control Group

      • The control group's size, demographic composition, and the instruction/activities it received are clearly documented in the methods section.
      • "The control group received training and instruction on reading skills, finding answers to various questions, and some reading techniques, such as skimming and scanning, with extra practices within the classroom and through traditional face-to-face teaching techniques." (Section 3.4, p. 14)
      • Relevant Quotes: 1) "50 Omani EFL learners from a higher education institution in Oman were randomly assigned to an experimental group and a control group, with 25 students in each group, comprising both males and females. The English proficiency level of these learners was determined to be at the intermediate level based on the university's placement test. These students were native Arabic speakers, and their ages ranged from 18 to 20 years old." (Section 3.1, p. 13) 2) "The control group received training and instruction on reading skills, finding answers to various questions, and some reading techniques, such as skimming and scanning, with extra practices within the classroom and through traditional face-to-face teaching techniques. The teacher regularly monitored the students in the class through observation and question-and-answer sessions, and provided extra quizzes." (Section 3.4, p. 14) 3) "Table 2 shows that for the pretest, the control had a statistic of 0.229 at 0.002 significance, while for the experimental group, it was 0.222 at 0.002." (Section 4, p. 14) — indicating baseline pretest data are reported separately for the control group. Detailed Analysis: The paper describes the control group's size (n=25), demographic profile (age range, native language, proficiency level, shared with the experimental group via random assignment from the same cohort), and the specific instructional activities it received (traditional face-to-face reading instruction, teacher monitoring, and extra quizzes). Baseline (pretest) performance is also reported and statistically compared between groups (Mann-Whitney U test, Table 7, showing no significant baseline difference, p=.478), confirming comparability at the start of the study. This level of detail satisfies the requirement for a well-documented control group. Final: Criterion D is met because the control group's size, demographics, baseline performance, and classroom treatment are clearly described.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred among individual students within a single institution, not among schools or institutions.
      • "50 Omani EFL learners from a higher education institution in Oman were randomly assigned to an experimental group and a control group..." (Section 3.1, p. 13)
      • Relevant Quotes: 1) "To conduct this quasi-experimental research study, 50 Omani EFL learners from a higher education institution in Oman were randomly assigned to an experimental group and a control group, with 25 students in each group..." (Section 3.1, p. 13) Detailed Analysis: Only one higher education institution is involved, and the randomisation unit is the individual student, not the institution/school. There is no description of multiple schools being randomised to conditions. Since the weaker class-level criterion (C) is also not met, the stronger school-level criterion cannot be met either. Final: Criterion S is not met because the study was conducted at a single institution with student-level randomisation, not school-level randomisation.
    • I

      Independent Conduct

      • The same two authors designed the intervention, delivered the instruction/oversight, and conducted the analysis, with no independent evaluator described.
      • "CREdIT AUTHOR STATEMENT: Behnam Behforouz: Formal Analysis, Writing – Review & Editing, Project Administration. Ali Al Ghaithi: Methodology, Writing – Original Draft, Resources." (p. 10)
      • Relevant Quotes: 1) "CREdIT AUTHOR STATEMENT: Behnam Behforouz: Formal Analysis, Writing – Review & Editing, Project Administration. Ali Al Ghaithi: Methodology, Writing – Original Draft, Resources." (p. 10) 2) "All participants in the experimental group were allocated a ChatGPT account established by the researchers to enable the teacher to oversee the students' progress in ChatGPT, verify adherence to instructions, and ensure the completion of assignments as stipulated by the researchers. Subsequently, the investigator facilitated a one-hour workshop to instruct the treatment groups on utilising ChatGPT..." (Section 3.4, p. 14) 3) There is no methods, acknowledgments, or disclosure statement mentioning an external or third-party evaluator, agency, or independent data collection team. Detailed Analysis: The CRediT statement shows both named authors handled the methodology, formal analysis, and writing, with no mention of an independent evaluation team. The narrative further indicates that "the researchers" and "the investigator" directly set up the ChatGPT accounts, ran the training workshop, and monitored student compliance themselves. This indicates the same team that designed the intervention also implemented and analysed it, without independent third-party oversight of data collection or analysis. Final: Criterion I is not met because the study was designed, delivered, and analysed by the same authors without independent conduct.
    • Y

      Year Duration

      • Since criterion T (Term Duration) is not met, this stronger Year Duration criterion is automatically not met; in any case the total tracked period was about two months, far short of a year.
      • "The treatment period lasted for a month... A delayed posttest was conducted... three weeks after the treatment period." (Section 3.4, p. 14)
      • Relevant Quotes: 1) "The treatment period lasted for a month, and according to the curriculum and delivery plan, two sessions, each lasting one hour and forty minutes, focused on reading comprehension activities..." (Section 3.4, p. 14) 2) "A delayed posttest was conducted to measure students' knowledge retention and the efficacy of using ChatGPT three weeks after the treatment period." (Section 3.4, p. 14) Detailed Analysis: Per the ERCT specification, if criterion T is not met, criterion Y is automatically not met. In this study, the entire tracked period from intervention start to final (delayed) measurement is approximately two months, which is far short of the required 75% of an academic year (~9-10 months). Final: Criterion Y is not met, both because T is not met and because the actual duration is far below a year.
    • B

      Balanced Control Group

      • The additional ChatGPT-based practice is the explicit treatment variable the study set out to test, with the control group receiving a comparable "business as usual" package of face-to-face instruction and extra quizzes.
      • "...to what extent does the use of ChatGPT as a facilitator improve the reading comprehension abilities of Omani EFL learners compared to traditional face-to-face instruction?" (Section 1, p. 11-12)
      • Relevant Quotes: 1) "...the following question will be covered in this study: to what extent does the use of ChatGPT as a facilitator improve the reading comprehension abilities of Omani EFL learners compared to traditional face-to-face instruction?" (Section 1, p. 11-12) 2) "The control group received training and instruction on reading skills, finding answers to various questions, and some reading techniques, such as skimming and scanning, with extra practices within the classroom and through traditional face-to-face teaching techniques. The teacher regularly monitored the students in the class through observation and question-and-answer sessions, and provided extra quizzes." (Section 3.4, p. 14) 3) "...although the experimental group received in-class instructions on reading skills and techniques to cope with different types of questions, the students were instructed to use ChatGPT to practice extra reading comprehension activities, such as creating questions, understanding the general idea, and the main idea." (Section 3.4, p. 14) Detailed Analysis: Both groups received the same core in-class instruction on reading skills and techniques (skimming, scanning, etc.), delivered face-to-face. The experimental group additionally received a ChatGPT account and extra out-of-class practice (following the pseudocode decision tree: extra resources are present, and the difference is not negligible). However, the study's explicit research question is precisely whether adding ChatGPT- facilitated practice improves outcomes relative to traditional face-to-face instruction alone — i.e., the additional resource (ChatGPT practice) is the primary treatment variable being tested (integral to the design), not an incidental add-on to an otherwise identical curriculum. The control group's business- as-usual package (traditional instruction plus teacher-provided extra quizzes) represents a reasonable comparator against which the added ChatGPT-based practice can be isolated. Applying the Criterion B decision tree: extra resources are present and not negligible, but they are integral to the treatment being tested, so the control group may remain at the business-as-usual level without violating the criterion. Final: Criterion B is met because the additional ChatGPT-based resources given to the experimental group are the explicit treatment variable under study, and the control group received a standard, business-as-usual face-to-face condition.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study by a different research team was found; the closest related work is a prior study by the same authors.
      • "Behforouz, B., & Al Ghaithi, A. (2024). Investigating the effect of an interactive educational chatbot on reading comprehension skills." (References, p. 18)
      • Relevant Quotes: 1) "Behforouz and Al Ghaithi (2024) investigated the role of an interactive chatbot as a facilitator in reading, illustrating that chatbots can serve as advantageous tools in education, especially in improving reading proficiency." (Section 5, p. 17) 2) "Wang and Feng (2024) examined the effect of ChatGPT assistance on reading skills over 4 weeks involving 83 Chinese undergraduate students..."; "Zhang et al. (2025) conducted a study on a novel reading platform powered by ChatGPT... Sixty-four undergraduate students..." (Section 2.2, pp. 11-12) Detailed Analysis: The most closely related prior work cited is the authors' own 2024 study on an interactive chatbot and reading comprehension among Omani EFL students — this is not an independent replication since it shares authorship with the present paper. Other cited studies (Wang & Feng, 2024; Zhang et al., 2025; Muman, 2025; Amimi & Saragih, 2025) investigate ChatGPT's effect on reading in different populations (Chinese undergraduates, vocational students, business education students) with different designs, sample sizes, and instruments; they are conceptually related but do not constitute an independent reproduction of this specific study's design, population, and measures. An internet search for independent replications of this specific study (published September 2025) was conducted, covering academic search engines and the authors' subsequent publications (e.g., their 2025 work on vocabulary/self-regulation and AI-assisted writing evaluation with Omani EFL learners). No peer-reviewed study by a different research team replicating this exact design, population, and measures was found; given how recently the paper was published, independent replication has likely not had time to appear. Final: Criterion R is not met because no independent replication of this specific study was found; the most similar prior work shares the same authors.
    • A

      All-subject Exams

      • Since criterion E (Exam-based Assessment) is not met, this stronger All-subject Exams criterion is automatically not met; in any case only reading was assessed.
      • "This study depended primarily on reading tests for assessment; hence, other aspects of language learning had not been taken into consideration, like writing or speaking." (Section 6, p. 18)
      • Relevant Quotes: 1) "To compare and measure the performance of students in both groups on reading skills, three sets of tests, including pretest, posttest, and delayed posttest, were designed." (Section 3.2.1, p. 13) 2) "This study depended primarily on reading tests for assessment; hence, other aspects of language learning had not been taken into consideration, like writing or speaking." (Section 6, p. 18) Detailed Analysis: Per the ERCT specification, criterion A cannot be met if criterion E is not met, which is the case here since the assessment instrument was custom-built rather than a standardised exam. Independently, the study only measured reading comprehension, and the authors explicitly acknowledge that other language subskills such as writing and speaking were not assessed, so even setting aside the E dependency, not all main subjects/skills were covered. Final: Criterion A is not met, both due to the dependency on criterion E and because only reading was assessed.
    • G

      Graduation Tracking

      • Since criterion Y (Year Duration) is not met, this stronger Graduation Tracking criterion is automatically not met, and no follow-up beyond the three-week delayed posttest is reported.
      • "A delayed posttest was conducted to measure students' knowledge retention and the efficacy of using ChatGPT three weeks after the treatment period." (Section 3.4, p. 14)
      • Relevant Quotes: 1) "A delayed posttest was conducted to measure students' knowledge retention and the efficacy of using ChatGPT three weeks after the treatment period." (Section 3.4, p. 14) 2) No mention anywhere in the paper of tracking participants beyond the delayed posttest, nor any reference to a planned or published follow-up study tracking the same cohort to graduation. Detailed Analysis: Per the ERCT specification, criterion G cannot be met if criterion Y is not met, which is the case here. Independently, the study's follow-up ends at the delayed posttest, three weeks after the one-month treatment period, with no tracking through course completion or graduation, and no evidence of a follow-up publication tracking the same cohort. An internet search was conducted for subsequent papers by Behforouz and/or Al Ghaithi that might track this same cohort of 50 Foundation Programme students toward graduation. Their other identified publications (2024-2025) involve different chatbot tools, vocabulary, self-regulation, and AI-assisted writing evaluation topics, and none describe longitudinal tracking of this cohort to graduation. No such follow-up paper was found. Final: Criterion G is not met, both due to the dependency on criterion Y and because tracking stopped shortly after the intervention.
    • P

      Pre-Registered

      • No pre-registration of the study protocol, registry link, or registration date is mentioned anywhere in the paper, and no registry entry was found via internet search.
      • Relevant Quotes: 1) The Methods, Ethical Considerations, and Data Availability Statement sections make no mention of a pre-registration platform, registry ID, or registration date. "DATA AVAILABILITY STATEMENT: The data supporting the findings of this study are included within the paper, or it could be shared by the corresponding author upon request." (p. 10) 2) "Before the study began, ethical approvals were received from the Research Department and the related authorities within the institution." (Section 3.3, p. 14) — this refers only to ethics approval, not protocol pre-registration. Detailed Analysis: The paper describes obtaining institutional ethics approval and informed consent, but at no point references a public pre-registration of the study's hypotheses, methods, or analysis plan (e.g., via a registry such as OSF, ISRCTN, or AsPredicted). No registration date or registry ID is provided. An internet search for a pre-registration record of this study (by title, authors, and topic) on common registries did not surface any matching entry, consistent with the paper's own silence on pre-registration. Final: Criterion P is not met because there is no evidence of pre-registration in the paper or in public registries.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.