Starfall as a Catalyst for Kuwaiti EFL Young Learners' Reading Comprehension: A Teacher's Reflections

Ruba Fahmi Bataineh, May Bader Alghareeb

Published:
ERCT Check Date:
DOI: 10.29333/ejecs/2338
  • reading
  • L2 languages
  • K12
  • Asia
  • gamification
  • EdTech website
0
  • C

    Two intact second-grade classes were randomly assigned to experimental and control conditions, so randomisation was at the class level (albeit with only two clusters).

    "The sections were randomly assigned as control and experimental groups; the former (n=23) taught reading comprehension using the conventional method outlined in the teacher's book, and the latter (n=21) used the instructional program incorporating Starfall." (p. 144)

  • E

    Outcomes were measured with a researcher-made observation card and a teacher's journal, not with any standardised exam.

    "The data were collected using an observation card and a teacher's journal." (p. 145)

  • T

    Outcomes were observed only during the ten-week treatment, which is shorter than the required full academic term of roughly 3-4 months.

    "The treatment spanned ten weeks in the first semester of the academic year 2023/2024 in the form of two weekly 45-minute sessions." (pp. 144-145)

  • D

    The control group is described only by size and its conventional instruction, with no demographic or baseline performance documentation.

    "The sections were randomly assigned as control and experimental groups; the former (n=23) taught reading comprehension using the conventional method outlined in the teacher's book..." (p. 144)

  • S

    The study randomised two classes within a single Kuwaiti school, so no school-level randomisation took place.

    "The sample comprised two intact second-grade classes from Othman Abdullatif Othman School, a public school in the city of Kuwait where the second researcher is a teacher." (p. 144)

  • I

    The researchers designed, delivered, observed, and analysed the intervention themselves, with no independent evaluators.

    "Over the course of the treatment, the teacher reflected on the use of Starfall in the experimental group's reading class in a personal journal after each session." (p. 145)

  • Y

    The ten-week study falls far short of the required 75% of an academic year, and the prerequisite term-duration criterion is also unmet.

    "The treatment spanned ten weeks in the first semester of the academic year 2023/2024 in the form of two weekly 45-minute sessions." (pp. 144-145)

  • B

    Both groups had the same regular reading class time, and the Starfall-based instruction itself was the integral treatment variable tested against conventional teaching.

    "The sections were randomly assigned as control and experimental groups; the former (n=23) taught reading comprehension using the conventional method outlined in the teacher's book, and the latter (n=21) used the instructional program incorporating Starfall." (p. 144)

  • R

    The study is presented as among the first of its kind and no independent published replication of it was found.

    "To the best of the researchers' knowledge, this study may be one of the first empirical studies into the use of Starfall as a tool to improve reading comprehension and engagement among young Kuwaiti EFL learners." (p. 142)

  • A

    No standardised exams were used at all (criterion E fails) and only English reading was observed, not all main subjects.

    "The data were collected using an observation card and a teacher's journal." (p. 145)

  • G

    Data collection ended with the ten-week treatment and no graduation tracking or follow-up of the cohort exists.

    "Future research could explore the effectiveness of other interactive platforms or examine the long-term effects of such interventions." (p. 149)

  • P

    The paper contains no reference to any pre-registration, registry ID, or published protocol.

Abstract

This study examines the effect of a Starfall-based instructional program on second-grade pupils' reading comprehension in a Kuwaiti public school in the first semester of the academic year 2023/2024. The participants were divided into a control group taught conventionally per the guidelines of the Ministry of Education and an experimental group taught using Starfall-based instruction. Using an observation card and a teacher's journal, the research highlights improvements in the experimental group's engagement, motivation and comprehension. Emerging themes include increased student participation, improved classroom dynamics, and technology integration as a catalyst for literacy. The findings emphasize the potential of interactive applications, of which Starfall is one, to support EFL instruction. The researchers put forth practical recommendations for incorporating technology in the early-grade EFL classroom, as the findings provide valuable insights for educators aiming to innovate EFL teaching practices.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Two intact second-grade classes were randomly assigned to experimental and control conditions, so randomisation was at the class level (albeit with only two clusters).
      • "The sections were randomly assigned as control and experimental groups; the former (n=23) taught reading comprehension using the conventional method outlined in the teacher's book, and the latter (n=21) used the instructional program incorporating Starfall." (p. 144)
      • Relevant Quotes: 1) "The sample comprised two intact second-grade classes from Othman Abdullatif Othman School, a public school in the city of Kuwait where the second researcher is a teacher." (p. 144) 2) "The sections were randomly assigned as control and experimental groups; the former (n=23) taught reading comprehension using the conventional method outlined in the teacher's book, and the latter (n=21) used the instructional program incorporating Starfall." (p. 144) Detailed Analysis: The unit of randomisation is the intact class (section), not individual students within one classroom. The paper explicitly states that the two second-grade sections "were randomly assigned as control and experimental groups," which satisfies the requirement that entire classes, rather than students within a class, be allocated to conditions. The design is very weak in absolute terms (only two clusters, so randomisation cannot balance confounders, and no description of the randomisation mechanism is provided), but under the ERCT procedure the criterion asks whether allocation was at class level or stronger, which is clearly stated here. The intervention is classroom instruction, not one-to-one tutoring, so the tutoring exception is not needed. This same quote also supports rct: true at the top level, since it documents a genuine random allocation of intact classes to a control and an experimental condition; the paper's self-description as "qualitative" (see criterion E) reflects its outcome instruments, not the presence or absence of randomisation. Criterion C is met because the two intact classes (sections) were randomly assigned to the experimental and control conditions, i.e., randomisation occurred at the class level.
    • E

      Exam-based Assessment

      • Outcomes were measured with a researcher-made observation card and a teacher's journal, not with any standardised exam.
      • "The data were collected using an observation card and a teacher's journal." (p. 145)
      • Relevant Quotes: 1) "The data were collected using an observation card and a teacher's journal. The 17-item observation card, based on a 3-point Likert Scale (i.e., often (3), sometimes (2), and never (1)), was used by the second researcher to gauge the pupils' engagement and understanding during the reading sessions, with specific focus on active participation, use of strategies, and interaction with the material." (p. 145) 2) "Conducted at Othman Abdullatif Othman School, a public school in the city of Kuwait in the first semester of the academic year 2023/2024, this qualitative study investigated the effect of using Starfall on Kuwaiti second-grade pupils' reading comprehension in English through class observation and the teacher's daily journal." (p. 142) 3) "Over the course of the treatment, the teacher reflected on the use of Starfall in the experimental group's reading class in a personal journal after each session." (p. 145) Detailed Analysis: Criterion E requires that outcomes be measured with a widely recognised, standardised exam-based assessment. This study used no exam at all: outcomes were gauged with a researcher-made 17-item observation card scored on a 3-point Likert scale by the teacher-researcher herself, plus her personal reflective journal. Both instruments were custom-designed for this study and are subjective observational measures, not standardised tests. No pre-test or post-test of reading achievement, let alone a national or state standardised exam, is mentioned anywhere in the paper. The authors themselves frame this as "a qualitative study," reinforcing that no standardised exam instrument was used. Criterion E is not met because outcomes were measured with a custom observation card and a teacher's journal rather than any standardised exam-based assessment.
    • T

      Term Duration

      • Outcomes were observed only during the ten-week treatment, which is shorter than the required full academic term of roughly 3-4 months.
      • "The treatment spanned ten weeks in the first semester of the academic year 2023/2024 in the form of two weekly 45-minute sessions." (pp. 144-145)
      • Relevant Quotes: 1) "The treatment spanned ten weeks in the first semester of the academic year 2023/2024 in the form of two weekly 45-minute sessions." (pp. 144-145) 2) "The card was used once a week for both the control and the experimental groups during the treatment (n=16) to observe their improvement." (p. 145) 3) "Future research could explore the effectiveness of other interactive platforms or examine the long-term effects of such interventions." (p. 149) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. Here the treatment lasted ten weeks (roughly 2.5 months), and the outcome data (weekly observation cards and journal entries) were collected during the treatment itself, ending when the ten-week programme ended. There is no delayed follow-up measurement; indeed the authors themselves recommend that future research "examine the long-term effects of such interventions." Ten weeks falls short of a full academic term of 3-4 months, and no quote establishes tracking beyond the treatment window. Criterion T is not met because measurement ended with the ten-week treatment, which is shorter than one full academic term from intervention start.
    • D

      Documented Control Group

      • The control group is described only by size and its conventional instruction, with no demographic or baseline performance documentation.
      • "The sections were randomly assigned as control and experimental groups; the former (n=23) taught reading comprehension using the conventional method outlined in the teacher's book..." (p. 144)
      • Relevant Quotes: 1) "The sections were randomly assigned as control and experimental groups; the former (n=23) taught reading comprehension using the conventional method outlined in the teacher's book, and the latter (n=21) used the instructional program incorporating Starfall." (p. 144) 2) "The card was used once a week for both the control and the experimental groups during the treatment (n=16) to observe their improvement." (p. 145) 3) "Table 2 Means and Standard Deviations of the Items on the Observation Card on the Utility of Starfall in Teaching Reading Comprehension to the Control Group" (p. 146) Detailed Analysis: Criterion D requires detailed documentation of the control group, including demographic information, baseline performance, and treatments received. The paper documents the control group's size (n=23), its setting (a second-grade section of the same Kuwaiti public school), and its treatment (conventional reading instruction per the Ministry of Education teacher's book), and Table 2 reports the control group's observation-card scores. However, no demographic characteristics (gender, socioeconomic background, age detail) are given for either group, and, critically, no baseline (pre-treatment) performance data are reported, so comparability of the two intact classes before the intervention cannot be assessed. The documentation is limited to group size and instructional condition, which falls short of the "detailed description of control group characteristics" the criterion demands. Criterion D is not met because the control group is described only by size and instructional condition, with no demographic information or baseline performance data.
  • Level 2 Criteria

    • S

      School-level RCT

      • The study randomised two classes within a single Kuwaiti school, so no school-level randomisation took place.
      • "The sample comprised two intact second-grade classes from Othman Abdullatif Othman School, a public school in the city of Kuwait where the second researcher is a teacher." (p. 144)
      • Relevant Quotes: 1) "The sample comprised two intact second-grade classes from Othman Abdullatif Othman School, a public school in the city of Kuwait where the second researcher is a teacher." (p. 144) 2) "The sections were randomly assigned as control and experimental groups..." (p. 144) Detailed Analysis: Criterion S requires randomisation among schools (or equivalent implementing units). This study took place within a single public school, with the two second-grade sections of that school assigned to conditions. No schools were randomised; the unit of randomisation was the class within one school. Criterion S is not met because randomisation occurred between two classes inside a single school, not across schools.
    • I

      Independent Conduct

      • The researchers designed, delivered, observed, and analysed the intervention themselves, with no independent evaluators.
      • "Over the course of the treatment, the teacher reflected on the use of Starfall in the experimental group's reading class in a personal journal after each session." (p. 145)
      • Relevant Quotes: 1) "Both experienced EFL practitioners, the researchers are often disheartened by learners' weakness and lack of interest in learning English. More specifically, the second researcher, an English second-grade teacher in Kuwait, has noticed her pupils' weakness in reading comprehension..." (p. 141) 2) "These activities were then used to design an instructional program targeting literal comprehension skills appropriate for EFL second-grade learners." (p. 144) 3) "The 17-item observation card ... was used by the second researcher to gauge the pupils' engagement and understanding during the reading sessions..." (p. 145) 4) "Over the course of the treatment, the teacher reflected on the use of Starfall in the experimental group's reading class in a personal journal after each session." (p. 145) 5) "Another teacher observed one session per group and filled the observation card to ensure inter-rater reliability." (p. 145) Detailed Analysis: Criterion I requires the study to be conducted independently of the intervention's designers. Here the two researchers designed the Starfall-based instructional program themselves, and the second researcher is the classroom teacher who delivered the intervention, scored the observation card, and wrote the reflective journal that constitutes the main outcome data. Data collection, implementation, and analysis were therefore all carried out by the intervention's designers, with an evident conflict of roles (the teacher rated her own pupils and her own programme). The only outside involvement was a second teacher observing one session per group for inter-rater reliability, which is far from independent third-party conduct or oversight of the evaluation. Criterion I is not met because the intervention designers themselves delivered the programme and collected and analysed all outcome data, with no independent evaluation team.
    • Y

      Year Duration

      • The ten-week study falls far short of the required 75% of an academic year, and the prerequisite term-duration criterion is also unmet.
      • "The treatment spanned ten weeks in the first semester of the academic year 2023/2024 in the form of two weekly 45-minute sessions." (pp. 144-145)
      • Relevant Quotes: 1) "The treatment spanned ten weeks in the first semester of the academic year 2023/2024 in the form of two weekly 45-minute sessions." (pp. 144-145) 2) "Future research could explore the effectiveness of other interactive platforms or examine the long-term effects of such interventions." (p. 149) Detailed Analysis: Criterion Y requires that outcomes be tracked for at least 75% of a full academic year (roughly 9-10 months) after the intervention begins. This study covered only ten weeks within one semester, with all data collected during that period and no follow-up afterwards. Since the weaker Term Duration criterion (T) is already not met, this stronger Year Duration criterion cannot be met either, per the standard's instructions. Criterion Y is not met because tracking lasted only ten weeks, far short of 75% of an academic year, and criterion T is not met.
    • B

      Balanced Control Group

      • Both groups had the same regular reading class time, and the Starfall-based instruction itself was the integral treatment variable tested against conventional teaching.
      • "The sections were randomly assigned as control and experimental groups; the former (n=23) taught reading comprehension using the conventional method outlined in the teacher's book, and the latter (n=21) used the instructional program incorporating Starfall." (p. 144)
      • Relevant Quotes: 1) "The sections were randomly assigned as control and experimental groups; the former (n=23) taught reading comprehension using the conventional method outlined in the teacher's book, and the latter (n=21) used the instructional program incorporating Starfall." (p. 144) 2) "The treatment spanned ten weeks in the first semester of the academic year 2023/2024 in the form of two weekly 45-minute sessions." (pp. 144-145) 3) "The card was used once a week for both the control and the experimental groups during the treatment (n=16) to observe their improvement." (p. 145) 4) "Seeking a potential solution, Starfall was incorporated into the EFL teaching/learning process as a teaching technique recommended for improving pupils' reading comprehension, motivation, and interest..." (pp. 141-142) 5) "The teacher also noted increased parental involvement and feedback on their children's use of Starfall at home." (p. 148) Detailed Analysis: Criterion B asks whether the control condition received comparable time and resources, unless the extra resource is itself the treatment variable. Both groups received their regular second-grade EFL reading instruction during the same semester: the control group was taught reading comprehension conventionally per the Ministry teacher's book, while the experimental group covered the same textbook-aligned reading content (Fun with English 2A) through Starfall-based sessions. The intervention thus replaced the method of instruction within normal class time rather than adding extra instructional hours or budget; Starfall is a free website and its use is the very treatment variable being tested (technology-based instruction versus conventional instruction). Applying the decision tree: no extra time or budget beyond ordinary class time is documented for the experimental group, so the criterion is trivially satisfied at step 1; even if the switch to Starfall were treated as an added resource, it is explicitly the integral treatment variable under test, so business-as-usual instruction for the control group remains acceptable by design. A minor caveat is that experimental pupils also used Starfall voluntarily at home, adding informal exposure, but this out-of-class engagement is an inherent product of the motivational intervention itself rather than a separable resource the researchers gave only to one group. Criterion B is met because both groups received their normal reading class time, and the Starfall technology is the integral treatment variable tested against business-as-usual instruction.
  • Level 3 Criteria

    • R

      Reproduced

      • The study is presented as among the first of its kind and no independent published replication of it was found.
      • "To the best of the researchers' knowledge, this study may be one of the first empirical studies into the use of Starfall as a tool to improve reading comprehension and engagement among young Kuwaiti EFL learners." (p. 142)
      • Relevant Quotes: 1) "To the best of the researchers' knowledge, this study may be one of the first empirical studies into the use of Starfall as a tool to improve reading comprehension and engagement among young Kuwaiti EFL learners." (p. 142) 2) "While the findings are promising, their generalizability is limited by the small sample size and the use of a single technological tool (viz., Starfall)." (p. 149) Detailed Analysis: Criterion R requires independent replication of this specific study by a different team, published in a peer-reviewed journal. The authors themselves frame the study as possibly "one of the first" of its kind in Kuwait, and cite earlier Starfall-related work (e.g., Halsey, 2009; Metis Associates, 2014; Zamora & Pittman, 2018; Cando Sanchez, 2021), but those are prior studies of the Starfall tool in other contexts and with other outcomes, not replications of this 2025 Kuwaiti trial. A fresh internet search (July 2026) via Google Scholar/Semantic Scholar found only the original article (Journal of Ethnic and Cultural Studies, 12(1), 2025) and a single incidental bibliographic citation to it in an unrelated 2025 paper by Al-Karasneh and Kanaan; that citation is merely a reference-list entry, not a replication attempt, and does not reproduce this study's design, sample, or findings. No independent peer-reviewed replication of this specific study could be located. Criterion R is not met because no independent replication of this specific study has been published; earlier Starfall studies predate it and do not reproduce it.
    • A

      All-subject Exams

      • No standardised exams were used at all (criterion E fails) and only English reading was observed, not all main subjects.
      • "The data were collected using an observation card and a teacher's journal." (p. 145)
      • Relevant Quotes: 1) "Due to the proficiency level of the pupils under study, comprehension was limited to the literal level, as neither inferential nor critical comprehension was gauged." (p. 148) 2) "The data were collected using an observation card and a teacher's journal." (p. 145) Detailed Analysis: Criterion A requires standardised exam-based assessment of all main school subjects, and explicitly presupposes criterion E. Since criterion E is not met (no standardised exam of any kind was used), criterion A automatically fails. Moreover, only one narrow outcome domain (English reading comprehension at the literal level) was observed; no other subjects such as mathematics, science, or Arabic were assessed in any form. Criterion A is not met because criterion E fails and only English reading comprehension was measured, not all main subjects.
    • G

      Graduation Tracking

      • Data collection ended with the ten-week treatment and no graduation tracking or follow-up of the cohort exists.
      • "Future research could explore the effectiveness of other interactive platforms or examine the long-term effects of such interventions." (p. 149)
      • Relevant Quotes: 1) "The treatment spanned ten weeks in the first semester of the academic year 2023/2024 in the form of two weekly 45-minute sessions." (pp. 144-145) 2) "Future research could explore the effectiveness of other interactive platforms or examine the long-term effects of such interventions." (p. 149) Detailed Analysis: Criterion G requires tracking participants until graduation from their educational stage, and presupposes criterion Y, which is not met. All data collection ended with the ten-week treatment; the second-grade pupils were not followed to the end of primary school, and the authors explicitly defer long-term effects to future research. A fresh internet search (July 2026) for subsequent publications by Bataineh and/or Alghareeb tracking this same Kuwaiti cohort found none; given the article was only published in early 2025 and the underlying data were collected in the 2023/2024 semester, no graduation-tracking follow-up would be expected to exist yet in any case. Criterion G is not met because measurement stopped at the end of the ten-week treatment, with no tracking toward graduation, and the prerequisite criterion Y also fails.
    • P

      Pre-Registered

      • The paper contains no reference to any pre-registration, registry ID, or published protocol.
      • Relevant Quotes: No quotes mentioning pre-registration, a trial registry, a registration ID, or a published protocol appear anywhere in the paper. Detailed Analysis: Criterion P requires that the full study protocol be registered on a public registry before data collection began. The paper contains no mention of any registry (e.g., ClinicalTrials.gov, ISRCTN, OSF, AEA registry), no registration number, and no pre-specified analysis plan. The methods describe only the instructional program design and its validation by a jury of experts, which is not pre-registration. A fresh internet check (July 2026) of the journal's article page and general registry search likewise found no pre-registration record for this study. Criterion P is not met because there is no evidence of any pre-registered protocol for this study.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.