Investigating the impact of Kahoot! on the ESP vocabulary knowledge development of Chinese EFL learners: a quasi-experimental study

Yiou Sun and Debbita Ai Lin Tan

Published:
ERCT Check Date:
DOI: 10.1080/2331186X.2026.2614815
  • L2 languages
  • higher education
  • China
  • gamification
  • EdTech app
  • EdTech platform
0
  • C

    The study is explicitly a quasi-experimental design using two pre-existing intact classes, not a true randomised controlled trial, so the class-level RCT criterion is not satisfied.

    "The designation of each intact class as either Experimental or Control was established using the coin-toss method (Neuman, 2020)."

  • E

    The vocabulary test used was a modified version of the VKS built around a custom, study-specific list of 40 target words, not a widely recognised standardised exam.

    "Finally, the researchers selected 40 target words (listed in Supplementary Appendix A)."

  • T

    Outcomes were tracked only to Week 12 of the research timeline (about 11-12 weeks from intervention start), which falls short of a full academic term (approximately 3-4 months).

    "In Week 12, delayed post-testing was implemented using the same VKS test as the pre- and post-tests."

  • D

    The control group's size, standard-instruction condition, and baseline demographic/academic characteristics are documented and statistically compared with the experimental group.

    "In contrast, participants in the Control group followed a standard approach as outlined in the ESP syllabus and did not receive any additional treatment."

  • S

    Both intact classes were drawn from a single technical college, so randomisation (such as it was) occurred at the class level, not across multiple schools.

    "A purposive sampling technique was used to select 61 students from a technical college for the research."

  • I

    The researchers themselves (including a lecturer at the participating institution) designed and directly supervised the intervention and testing procedures, with no independent third-party evaluator involved.

    "In partnership with authorities in a technical college in China, the researchers supervised the entire procedure to ensure adherence to protocols."

  • Y

    Since criterion T (term duration) is not met, criterion Y is automatically not met; in any case the roughly 12-week tracking period is far short of 75% of an academic year.

    "Phase 4 (Week 8, Week 12) Immediate post- testing, Delayed post-testing, Post-intervention questionnaire."

  • B

    The intervention did not add extra class time or budget; both groups followed the same standard ESP syllabus during the same period, differing only in teaching method (Kahoot! versus traditional instruction).

    "In contrast, participants in the Control group followed a standard approach as outlined in the ESP syllabus and did not receive any additional treatment."

  • R

    No independent replication of this newly published study exists; a citation-database check confirms zero citing works to date.

  • A

    Since criterion E (Exam-based Assessment) is not met, criterion A is automatically not met; in addition, only ESP vocabulary was assessed, not all main subjects.

    "This study investigates the efficacy of Kahoot! in improving learners' recall and retention of ESP vocabulary within the Chinese EFL context."

  • G

    Since criterion Y (Year Duration) is not met, criterion G is automatically not met; no follow-up publication tracking this cohort was found either.

    "In Week 12, delayed post-testing was implemented using the same VKS test as the pre- and post-tests."

  • P

    No pre-registration of the study protocol is mentioned anywhere in the paper, and none was found via external registry/metadata checks.

Abstract

Research over the past decade has identified gamification as an emerging approach in language teaching, particularly in updating vocabulary instruction through game-based methods. This study investigates the efficacy of Kahoot! in improving learners' recall and retention of ESP vocabulary within the Chinese EFL context. A total of 61 participants were recruited and randomly assigned to two intact cohorts: an experimental group (N = 31) receiving instruction through Kahoot!, and a control group (N = 30) following traditional teaching practice. The Vocabulary Knowledge Scale (VKS) was employed to measure lexical recall and retention across three testing phases: pre-, post-, and delayed post-testing. Additionally, a questionnaire was administered to explore learners' perceptions of Kahoot!. The findings indicate that integrating Kahoot! into vocabulary instruction enhances both the recall and retention of ESP vocabulary items and has an overall positive impact on affective factors related to vocabulary learning. Further research is required to establish the generalizability of these results across various contexts, diverse EFL populations, and a wider range of ESP genres.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • The study is explicitly a quasi-experimental design using two pre-existing intact classes, not a true randomised controlled trial, so the class-level RCT criterion is not satisfied.
      • "The designation of each intact class as either Experimental or Control was established using the coin-toss method (Neuman, 2020)."
      • Relevant Quotes: 1) "A purposive sampling technique was used to select 61 students from a technical college for the research, which employed a quasi- experimental design with two intact cohorts: an experimental group (N = 31) and a control group (N = 30)." (p. 6) 2) "The designation of each intact class as either Experimental or Control was established using the coin-toss method (Neuman, 2020)." (p. 5) 3) "As a quasi-experimental study, several limitations warrant consideration." (p. 19) 4) "Additionally, using convenience sampling and intact class groups introduces potential selection bias, despite efforts to establish baseline equivalence. The absence of random assignment increases the likelihood of confounding variables influencing the results." (p. 19) Detailed Analysis: Criterion C requires a genuine RCT in which randomisation is properly implemented at the class (or school) level. This paper repeatedly and explicitly self-identifies as a "quasi- experimental" study (in its title, abstract, methods section, and limitations), using two pre-existing "intact cohorts" rather than randomly constituted groups. The only random element was a coin toss used to decide which of the two already-formed classes would be labelled "Experimental" and which "Control" - this is not random assignment of participants to conditions, and with only one class per condition there is no true randomisation of a participant pool. The authors themselves state in the Limitations section that "the absence of random assignment increases the likelihood of confounding variables," confirming that this is not treated by the authors as a genuine RCT. Because there is no bona fide random allocation process (only a coin-toss labelling of two convenience cohorts), the class-level RCT criterion is not met. Criterion C is not met because the study is an explicitly self-labelled quasi-experimental design using intact convenience cohorts, not a randomised controlled trial.
    • E

      Exam-based Assessment

      • The vocabulary test used was a modified version of the VKS built around a custom, study-specific list of 40 target words, not a widely recognised standardised exam.
      • "Finally, the researchers selected 40 target words (listed in Supplementary Appendix A)."
      • Relevant Quotes: 1) "The Vocabulary Knowledge Scale (VKS) was employed to measure lexical recall and retention across three testing phases: pre-, post-, and delayed post-testing." (p. 1) 2) "The eight modules of the textbook and their glossary were scanned and checked in Microsoft Word for errors. The resulting files were then converted into .txt format and run through EAP Vocabulary Profiler, generating a frequency- based list of 234 vocabularies from the glossary... Finally, the researchers selected 40 target words (listed in Supplementary Appendix A)." (p. 6) 3) "Rosszell (2007) modified VKS by adding: 'I have seen this word before, and I think it is related to the following word/idea: ______', to better reflect partial understanding... In the present study, the modified version of the VKS was adapted to evaluate ESP lexical knowledge development in the Chinese EFL context." (p. 6-7) Detailed Analysis: Criterion E requires a standard, widely recognised exam-based assessment rather than an instrument specially designed for the study. The VKS methodology itself has prior scholarly precedent, but the actual test administered here was built from a bespoke list of 40 target words drawn from this specific college's own textbook glossary, filtered through a vocabulary profiler, and rated for usefulness by the two involved TEFL teachers and the course instructor. This is, in effect, a custom-made assessment tailored precisely to the intervention's content and the researchers' target items, rather than a nationally or internationally standardised exam such as a curriculum-wide or state-administered test. The further modification of the VKS format itself (adding an L1 translation step) reinforces that this is a researcher-customised instrument. Criterion E is not met because the assessment used a custom-built list of 40 target words and a modified self-report scale specific to this study, rather than a widely recognised standardised exam.
    • T

      Term Duration

      • Outcomes were tracked only to Week 12 of the research timeline (about 11-12 weeks from intervention start), which falls short of a full academic term (approximately 3-4 months).
      • "In Week 12, delayed post-testing was implemented using the same VKS test as the pre- and post-tests."
      • Relevant Quotes: 1) "Phase 2 (Week 1) Pretesting; Phase 3 (Week 1 – Week 8) Intervention period; Phase 4 (Week 8, Week 12) Immediate post-testing, Delayed post- testing, Post-intervention questionnaire." (Table 2, p. 8) 2) "During Phase 3 (Week 1 to Week 8), all participants of the Experimental group engaged in Kahoot!-based teaching and learning based on a standard ESP teaching syllabus issued for a technical college." (p. 8) 3) "In Week 8, immediate post-testing was conducted using an identical Vocabulary Knowledge Scale (VKS) test of the same target words. In Week 12, delayed post-testing was implemented using the same VKS test as the pre- and post-tests." (p. 8) Detailed Analysis: The intervention began in Week 1 and ran through Week 8 (an eight-week teaching period), with the final, delayed outcome measurement occurring in Week 12. This means the interval from intervention start to the last measurement point spans roughly 11-12 weeks (about 2.5-2.8 months). The ERCT standard requires outcomes to be tracked for at least one full academic term, generally defined as approximately 3-4 months (or a full semester, which in Chinese higher- education institutions typically runs about 16- 18 weeks). The paper provides no statement that this 12-week window constitutes a full term in its local academic calendar, and the tracked interval is noticeably shorter than the standard 3-4 month benchmark. Criterion T is not met because the longest tracked interval (about 12 weeks from intervention start) falls short of the one-term minimum duration required by the standard.
    • D

      Documented Control Group

      • The control group's size, standard-instruction condition, and baseline demographic/academic characteristics are documented and statistically compared with the experimental group.
      • "In contrast, participants in the Control group followed a standard approach as outlined in the ESP syllabus and did not receive any additional treatment."
      • Relevant Quotes: 1) "A total of 61 participants were recruited and randomly assigned to two intact cohorts: an experimental group (N = 31) receiving instruction through Kahoot!, and a control group (N = 30) following traditional teaching practice." (p. 1) 2) "In contrast, participants in the Control group followed a standard approach as outlined in the ESP syllabus and did not receive any additional treatment." (p. 8) 3) "Table 21. Descriptive statistics of personal information... Age Experimental 31 18.87 0.846 0.891 / Control 30 18.90 0.803 ... Year Experimental 31 11.71 1.575 0.589 / Control 30 11.50 1.432 ... NEMT ... CET..." (Table 21, p. 14) 4) "Table 5. Pre-test results: ESP receptive vocabulary knowledge... Experimental 31 15.69 14.73 / Control 30 15.43 18.46." (Table 5, p. 9) Detailed Analysis: Criterion D requires clear documentation of the control group's size, demographic/baseline characteristics, and the conditions they experienced. The paper specifies the control group's size (N = 30), clearly states it received standard, "business as usual" ESP instruction with "no additional treatment," and separately reports its age, academic-year, National Matriculation English Test (NMET), and College English Test (CET) statistics compared against the experimental group (Table 21), plus its pre-test receptive and productive VKS baseline scores (Tables 5 and 7), confirming baseline comparability. This level of detail satisfies the documentation requirement even though the brief demographic Table 1 itself left the control row largely blank. Criterion D is met because the control group's size, treatment condition, and baseline characteristics are documented and statistically compared to the experimental group.
  • Level 2 Criteria

    • S

      School-level RCT

      • Both intact classes were drawn from a single technical college, so randomisation (such as it was) occurred at the class level, not across multiple schools.
      • "A purposive sampling technique was used to select 61 students from a technical college for the research."
      • Relevant Quotes: 1) "A purposive sampling technique was used to select 61 students from a technical college for the research, which employed a quasi- experimental design with two intact cohorts." (p. 6) 2) "In partnership with authorities in a technical college in China, the researchers supervised the entire procedure to ensure adherence to protocols." (p. 9) Detailed Analysis: Criterion S requires randomisation across multiple schools or institutions. This study involved a single technical college with only two intact classes (cohorts) drawn from that one institution. There is no indication that multiple schools were involved or that whole schools were randomised to conditions. Since even the weaker class-level RCT criterion (C) is not met, the stronger school-level criterion cannot be met either. Criterion S is not met because the study was conducted within a single institution using two intact classes, not across multiple randomised schools.
    • I

      Independent Conduct

      • The researchers themselves (including a lecturer at the participating institution) designed and directly supervised the intervention and testing procedures, with no independent third-party evaluator involved.
      • "In partnership with authorities in a technical college in China, the researchers supervised the entire procedure to ensure adherence to protocols."
      • Relevant Quotes: 1) "In partnership with authorities in a technical college in China, the researchers supervised the entire procedure to ensure adherence to protocols." (p. 9) 2) "The list of words was rated by two involved TEFL teachers, as well as the engaged bilingual instructor, for their usefulness and learnability." (p. 6) 3) "Yiou Sun is a lecturer at the Language Centre of Jiangsu University of Science & Technology (mainland China)." (About the authors, p. 20) Detailed Analysis: Criterion I requires that the evaluation be conducted independently of the people who designed and delivered the intervention, to avoid bias in implementation and analysis. Here, the researchers, together with in-house TEFL teachers and "the engaged bilingual instructor," directly selected the target words, designed the lesson plan, and personally "supervised the entire procedure" within the partner college. There is no mention of an external, independent organisation or blinded third-party evaluators handling data collection or analysis; interrater reliability (Cohen's kappa) pertains only to scoring consistency between two markers, not to independence of the overall study conduct from its designers. Criterion I is not met because the same research team that designed the intervention also directly supervised its delivery and evaluation, with no independent third-party conduct documented.
    • Y

      Year Duration

      • Since criterion T (term duration) is not met, criterion Y is automatically not met; in any case the roughly 12-week tracking period is far short of 75% of an academic year.
      • "Phase 4 (Week 8, Week 12) Immediate post- testing, Delayed post-testing, Post-intervention questionnaire."
      • Relevant Quotes: 1) "Phase 3 (Week 1 – Week 8) Intervention period; Phase 4 (Week 8, Week 12) Immediate post-testing, Delayed post-testing." (Table 2, p. 8) Detailed Analysis: Per the ERCT standard's specific instruction, criterion Y cannot be met if criterion T is not met. The total study duration here, from intervention start (Week 1) to the final delayed post-test (Week 12), is only about 12 weeks - far short of 75% of a typical 9-10 month academic year (which would require roughly 27-30 weeks of tracking). Criterion Y is not met, both because criterion T was not met and because the approximately 12-week tracking period is far shorter than 75% of a full academic year.
    • B

      Balanced Control Group

      • The intervention did not add extra class time or budget; both groups followed the same standard ESP syllabus during the same period, differing only in teaching method (Kahoot! versus traditional instruction).
      • "In contrast, participants in the Control group followed a standard approach as outlined in the ESP syllabus and did not receive any additional treatment."
      • Relevant Quotes: 1) "During Phase 3 (Week 1 to Week 8), all participants of the Experimental group engaged in Kahoot!-based teaching and learning based on a standard ESP teaching syllabus issued for a technical college. In contrast, participants in the Control group followed a standard approach as outlined in the ESP syllabus and did not receive any additional treatment." (p. 8) 2) "Table 22. Descriptive statistics of exposure time to English... all calculated Sig. (p) values for each item exceeded the threshold of 0.05 (p > 0.05), suggesting that there were no statistically significant differences between the two cohorts [regarding contact time]." (p. 14-15) Detailed Analysis: Following the criterion B decision procedure: the first question is whether the intervention added extra time or budget beyond business-as-usual instruction. Both groups were taught under the "same standard ESP teaching syllabus" during the identical Week 1-8 period; the only difference described is the instructional method (Kahoot!- based delivery versus traditional teaching), not additional class time, materials budget, or out-of-class resources. Table 22 further confirms that both cohorts reported statistically equivalent levels of general English exposure/ contact time in and out of class. Because no extra resources were introduced for the Experimental group beyond swapping the delivery method within the same allotted syllabus time, the "no extra resources present" branch of the decision tree applies, and the balance requirement is trivially satisfied. Criterion B is met because the intervention involved no additional time or budget beyond the shared standard syllabus; only the teaching method differed between the groups.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this newly published study exists; a citation-database check confirms zero citing works to date.
      • Relevant Quotes: (No quotes found; the paper does not reference any independent replication of its own findings, and given its recent publication (accepted 15 September 2025, published online 16 January 2026), no external replication could plausibly exist yet.) Detailed Analysis: Criterion R requires independent replication of this specific study by a different research team in a peer-reviewed outlet. The paper cites related but distinct prior Kahoot!/gamification vocabulary studies (e.g. Reynolds & Taylor, 2020; Ahmed et al., 2022) as supporting context, but none of these are replications of this particular Chinese-EFL, ESP-vocabulary, VKS-based design. An internet check of this article's citation record (OpenAlex/CrossRef, checked July 2026) confirms it currently has zero citing works (cited_by_count = 0), consistent with its very recent publication date, so no subsequent independent replication could exist at the time of this review. Criterion R is not met because no independent replication of this specific study was found, and citation-database checks confirm zero citing works to date.
    • A

      All-subject Exams

      • Since criterion E (Exam-based Assessment) is not met, criterion A is automatically not met; in addition, only ESP vocabulary was assessed, not all main subjects.
      • "This study investigates the efficacy of Kahoot! in improving learners' recall and retention of ESP vocabulary within the Chinese EFL context."
      • Relevant Quotes: 1) "This study investigates the efficacy of Kahoot! in improving learners' recall and retention of ESP vocabulary within the Chinese EFL context." (p. 1) 2) "The Vocabulary Knowledge Scale (VKS) was employed to measure lexical recall and retention across three testing phases." (p. 1) Detailed Analysis: Per the ERCT standard's specific instruction, criterion A cannot be met if criterion E is not met, which is the case here. Independently, the study's outcome measures are confined entirely to ESP vocabulary knowledge (via the VKS); no other main subjects (e.g. general English skills beyond vocabulary, or other academic disciplines) are assessed, and no exception rationale for a narrowly focused specialised assessment is articulated. Criterion A is not met both because criterion E is not met and because only ESP vocabulary outcomes were assessed.
    • G

      Graduation Tracking

      • Since criterion Y (Year Duration) is not met, criterion G is automatically not met; no follow-up publication tracking this cohort was found either.
      • "In Week 12, delayed post-testing was implemented using the same VKS test as the pre- and post-tests."
      • Relevant Quotes: (No quotes found describing graduation tracking, either in this paper or in any located follow-up publication by the same authors.) 1) "Phase 4 (Week 8, Week 12) Immediate post- testing, Delayed post-testing, Post-intervention questionnaire." (Table 2, p. 8) 2) "The current findings open promising avenues for future research... Future research should investigate whether gamification leads to long- term vocabulary retention." (p. 19) Detailed Analysis: Per the ERCT standard's specific instruction, criterion G cannot be met if criterion Y is not met, which is the case here. The paper's own conclusion explicitly frames long-term retention tracking as a topic for "future research," and there is no evidence of any follow-up beyond the Week 12 delayed post-test, let alone tracking through graduation. An internet check of this article's citation record (OpenAlex/CrossRef, checked July 2026) confirms zero citing works to date, so no subsequent follow-up paper by these authors tracking this cohort toward graduation could yet exist. Criterion G is not met because criterion Y was not met, no tracking beyond the 12-week delayed post-test is reported, and no follow-up publication tracking this cohort was located.
    • P

      Pre-Registered

      • No pre-registration of the study protocol is mentioned anywhere in the paper, and none was found via external registry/metadata checks.
      • Relevant Quotes: (No quotes found; the Methods, Data Collection Procedures, and Disclosure statement sections contain no mention of a trial registry, protocol registration platform, or registration date.) Detailed Analysis: Criterion P requires a quoted reference to a registry platform and a registration date prior to data collection. This paper's disclosure statement reads only "No potential conflict of interest was reported by the authors" (p. 18), and no section of the paper references pre- registration of hypotheses, methods, or analysis plans on any registry. A CrossRef metadata check for this DOI (performed July 2026) likewise returns no protocol-registration or clinical- trial-registration information for this work, and no pre-registration record for this study was located under its title or authors. Criterion P is not met because no evidence of pre-registration is present anywhere in the paper, and none was found via external checks.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.