Effects of a mobile game-based English vocabulary learning app on learners' perceptions and learning performance: A case study of Taiwanese EFL learners

Chih-Ming Chen, Huimei Liu, Hong-Bin Huang

Published:
ERCT Check Date:
DOI: 10.1017/S0958344018000228
  • L2 languages
  • higher education
  • Asia
  • gamification
  • EdTech app
  • mobile learning
  • digital assessment
0
  • C

    Randomisation was carried out at the individual student level within a single cohort at one university, not at the class or school level, and no personal-tutoring exception is invoked.

    "the study recruited 20 sophomores (second-year university students) and randomly assigned them to the experimental group with MEVLA-GF support or the control group with MEVLA-NGF support for English vocabulary learning." (p. 172)

  • E

    The vocabulary tests were drawn from the New TOEIC Official Test-Preparation Guide III, a widely recognised, standardised, ETS-produced test, not a custom instrument.

    "this study adopted questions from the New TOEIC Official Test-Preparation Guide III (ETS, 2011), an official test-preparation guide, to evaluate participants' vocabulary." (p. 176)

  • T

    Outcomes were measured only 4 weeks after the intervention began, with retention tested 2 weeks later, so the total interval (6 weeks) is far shorter than one academic term.

    "During a four-week experiment, 20 sophomore students were randomly assigned..." combined with "a vocabulary delayed post-test two weeks later to evaluate vocabulary retention." (p. 170, 173)

  • D

    The control group's size, gender composition, baseline-matching procedure, and detailed app-usage statistics are clearly documented in the text and Table 8.

    "the remaining 10 students (five males, five females) were assigned to the control group (i.e. MELVA-NGF)." (p. 173)

  • S

    Randomisation occurred among 20 individual students within a single university cohort, not among schools.

    "the study recruited 20 sophomores (second-year university students) and randomly assigned them to the experimental group with MEVLA-GF support or the control group with MEVLA-NGF support..." (p. 172)

  • I

    The mobile app (the intervention) was developed by a separate commercial studio, and the academic research team conducting the trial had no described role in designing the app, supporting independent conduct of the evaluation.

    "This study adopted the PHONE Words app, developed by the Alice English Education Studio in Taiwan, as the vocabulary learning tool..." (p. 173)

  • Y

    As the Term Duration criterion (T) is not met, the stronger Year Duration criterion cannot be met either; the total tracked interval was only about six weeks.

    "During a four-week experiment, 20 sophomore students were randomly assigned..." combined with the two-week delayed post-test. (p. 170, 173)

  • B

    The only substantive difference between conditions is the gamified assessment/ranking feature itself, the explicit treatment variable, while both groups used the same core app functions, word list, and minimum weekly usage expectation.

    "The main difference in the functions provided between the MEVLA-GF and MEVLA-NGF is that the MEVLA-NGF does not have the gamified assessment with ranking among learning peers, but the other functions provided by the two apps are the same." (p. 173-174)

  • R

    Neither the paper itself nor an internet search for citing or follow-up literature indicates that this specific study has been independently replicated by a different research team.

  • A

    Only English vocabulary outcomes were assessed; no other core school subjects were measured, and the paper does not provide the kind of upper- secondary/vocational justification the standard requires for a narrower exception.

    "Learning performance based on vocabulary acquisition and vocabulary retention was assessed using three vocabulary tests..." (p. 172)

  • G

    Tracking ended two weeks after the four-week intervention with no further follow-up; since the Year Duration criterion (Y) is not met, this criterion cannot be met either, and no subsequent graduation-tracking papers by these authors were found.

    "a vocabulary delayed post-test two weeks later to evaluate vocabulary retention." (p. 173)

  • P

    The paper contains no statement of pre-registration on any registry platform, nor any registration date or protocol reference, and no internet search of registry platforms located a matching entry.

Abstract

Many studies have demonstrated that vocabulary size plays a key role in learning English as a foreign language (EFL). In recent years, mobile game-based learning (MGBL) has been considered a promising scheme for successful acquisition and retention of knowledge. Thus, this study applies a mixed methodology that combines quantitative and qualitative approaches to assess the effects of PHONE Words, a novel mobile English vocabulary learning app (application) designed with game-related functions (MEVLA-GF) and without game-related functions (MEVLA-NGF), on learners' perceptions and learning performance. During a four-week experiment, 20 sophomore students were randomly assigned to the experimental group with MEVLA-GF support or the control group with MEVLA-NGF support for English vocabulary learning. Analytical results show that performance in vocabulary acquisition and retention by the experimental group was significantly higher than that of the control group. Moreover, questionnaire results confirm that MEVLA-GF is more effective and satisfying for English vocabulary learning than MEVLA-NGF. Spearman rank correlation results show that involvement and dependence on gamified functions were positively correlated with vocabulary learning performance.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was carried out at the individual student level within a single cohort at one university, not at the class or school level, and no personal-tutoring exception is invoked.
      • "the study recruited 20 sophomores (second-year university students) and randomly assigned them to the experimental group with MEVLA-GF support or the control group with MEVLA-NGF support for English vocabulary learning." (p. 172)
      • Relevant Quotes: 1) "the study recruited 20 sophomores (second-year university students) and randomly assigned them to the experimental group with MEVLA-GF support or the control group with MEVLA-NGF support for English vocabulary learning." (p. 172) 2) "The research participants were 20 EFL learners who were Taiwanese sophomores studying at the College of Liberal Arts at National Chengchi University (NCCU)." (p. 173) 3) "Based on these considerations, 10 students (five males, five females) were assigned to the experimental group (i.e. MEVLA-GF), and the remaining 10 students (five males, five females) were assigned to the control group (i.e. MELVA-NGF)." (p. 173) 4) "The experimental treatment did not use any formal classroom instruction, as learners can learn vocabulary anytime and anywhere via MEVLA-GF or MEVLA-NGF in the mobile context." (p. 172-173) Detailed Analysis: The paper explicitly describes randomisation of 20 individual sophomore volunteers, drawn from a single college cohort at one university, into an experimental (n=10) and control (n=10) group. This is student-level randomisation, not class-level or school-level. The ERCT standard allows an exception when the intervention is specifically designed as personal tutoring or one-to-one teaching; here the app is self-paced, autonomous, individual study material used "anytime and anywhere," but the paper never frames this as tutoring or one-to-one teaching, nor does it invoke any such exception. It is a self-directed EdTech tool used outside any classroom setting, which is a different situation from one-to-one tutoring by an instructor. Because there is no class-level (or stronger school-level) randomisation and no explicit tutoring exception quote, this criterion is not satisfied. Criterion C is not met because randomisation occurred at the individual student level within one cohort, without a stated tutoring exception.
    • E

      Exam-based Assessment

      • The vocabulary tests were drawn from the New TOEIC Official Test-Preparation Guide III, a widely recognised, standardised, ETS-produced test, not a custom instrument.
      • "this study adopted questions from the New TOEIC Official Test-Preparation Guide III (ETS, 2011), an official test-preparation guide, to evaluate participants' vocabulary." (p. 176)
      • Relevant Quotes: 1) "Learning performance based on vocabulary acquisition and vocabulary retention was assessed using three vocabulary tests, a pre-test, a post-test, and a delayed post-test obtained randomly from the New TOEIC Official Test-Preparation Guide III published by the Educational Testing Service (ETS; 2011), an organization devoted to educational measurements and research in educational policy." (p. 172) 2) "this study adopted questions from the New TOEIC Official Test-Preparation Guide III (ETS, 2011), an official test-preparation guide, to evaluate participants' vocabulary. The TOEIC test was developed in 1979." (p. 176) 3) "The adopted questions consisted of incomplete sentences that the student must complete." (p. 176) Detailed Analysis: The primary vocabulary outcome measures (pre-test, immediate post-test, delayed post-test) were all drawn from an official ETS preparation guide for the TOEIC test, a long-established (since 1979), widely recognised, internationally standardised English-proficiency exam, rather than a bespoke instrument created solely for this study. Although the items were selected/rearranged by the researchers for pre/post/delayed administration, the underlying item bank and construct come from an official, standardised test-preparation source. This satisfies the requirement for a standardised exam-based assessment. Criterion E is met because the vocabulary assessment was based on the standardised, ETS-produced TOEIC test-preparation materials rather than a fully custom-made instrument.
    • T

      Term Duration

      • Outcomes were measured only 4 weeks after the intervention began, with retention tested 2 weeks later, so the total interval (6 weeks) is far shorter than one academic term.
      • "During a four-week experiment, 20 sophomore students were randomly assigned..." combined with "a vocabulary delayed post-test two weeks later to evaluate vocabulary retention." (p. 170, 173)
      • Relevant Quotes: 1) "During a four-week experiment, 20 sophomore students were randomly assigned to the experimental group with MEVLA-GF support or the control group with MEVLA-NGF support for English vocabulary learning." (p. 170, Abstract) 2) "learners in both groups utilized the MEVLA-GF and MEVLA-NGF as portable learning tools in an autonomous learning context during the four-week experimental period." (p. 172-173) 3) "a vocabulary post-test was performed immediately at the end of the four-week experimental period, and was followed by a vocabulary delayed post-test two weeks later to evaluate vocabulary retention." (p. 173) Detailed Analysis: The intervention itself ran for four weeks, and the final (delayed) outcome measurement occurred two weeks after that, for a total tracked interval of approximately six weeks from intervention start to final measurement. An academic term is typically defined as roughly 3-4 months (12-16 weeks), so six weeks falls well short of this threshold, and the paper gives no indication of a longer follow-up period. Criterion T is not met because the total interval from intervention start to final measurement (about six weeks) is substantially shorter than one academic term.
    • D

      Documented Control Group

      • The control group's size, gender composition, baseline-matching procedure, and detailed app-usage statistics are clearly documented in the text and Table 8.
      • "the remaining 10 students (five males, five females) were assigned to the control group (i.e. MELVA-NGF)." (p. 173)
      • Relevant Quotes: 1) "The pre-test results were further adopted as the group-assigning criterion, which was anticipated to be the prerequisite for establishing two evenly distributed groups with the same original vocabulary level. In addition, gender balance in both groups was also considered." (p. 173) 2) "the remaining 10 students (five males, five females) were assigned to the control group (i.e. MELVA-NGF)." (p. 173) 3) "Table 8 shows the descriptive statistics of MEVLA-NGF usage behaviors in the control group. Among the three functions, the traditional assessment had the highest mean for the number of clicks (M = 222.40, SD = 28.19)... The mean of total clicks of all the functions provided by the MEVLA-NGF during the four weeks was 531.20 times, and the mean of total usage time was as high as 21.33 hours." (p. 181-182) 4) "the mean pre-test score for the experimental group was not significantly different from that of the control group (p value = .9993 > 0.05), indicating that learners in the two groups had similar vocabulary skills before performing the experiment." (p. 178) Detailed Analysis: The paper documents the control group's size (n=10), gender split (5 male/5 female), the criterion used to allocate participants to balance baseline vocabulary level and gender, statistically confirms baseline equivalence between groups (pre-test p=.9993), and provides detailed descriptive statistics (Table 8) of the control group's app usage behaviour (clicks per function, total usage time) over the four weeks. This level of detail allows a reader to assess comparability of the control group at baseline and during the study. Criterion D is met because the control group's demographics, baseline characteristics, and usage behaviour are clearly and quantitatively documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred among 20 individual students within a single university cohort, not among schools.
      • "the study recruited 20 sophomores (second-year university students) and randomly assigned them to the experimental group with MEVLA-GF support or the control group with MEVLA-NGF support..." (p. 172)
      • Relevant Quotes: 1) "the study recruited 20 sophomores (second-year university students) and randomly assigned them to the experimental group with MEVLA-GF support or the control group with MEVLA-NGF support for English vocabulary learning." (p. 172) 2) "The research participants were 20 EFL learners who were Taiwanese sophomores studying at the College of Liberal Arts at National Chengchi University (NCCU)." (p. 173) Detailed Analysis: All 20 participants were drawn from a single institution (NCCU), and randomisation was performed at the individual-student level, not across multiple schools or institutions. There is no school-level cluster randomisation described anywhere in the paper. Criterion S is not met because the study involved a single-site, student-level randomisation rather than randomisation across schools.
    • I

      Independent Conduct

      • The mobile app (the intervention) was developed by a separate commercial studio, and the academic research team conducting the trial had no described role in designing the app, supporting independent conduct of the evaluation.
      • "This study adopted the PHONE Words app, developed by the Alice English Education Studio in Taiwan, as the vocabulary learning tool..." (p. 173)
      • Relevant Quotes: 1) "This study adopted the PHONE Words app, developed by the Alice English Education Studio in Taiwan, as the vocabulary learning tool because it is a novel mobile English vocabulary learning app designed for learning English vocabulary on such formal vocabulary tests as TOEIC and TOEFL." (p. 173) 2) "Since passing the TOEIC test is a graduation requirement at several universities in Taiwan -- including NCCU -- this study thus chose the PHONE Words app as the research instrument to conduct the experiment." (p. 173) 3) Author affiliations: "Graduate Institute of Library, Information and Archival Studies, National Chengchi University" and "Department of Statistics, National Chengchi University" (p. 170), with no author affiliated with Alice English Education Studio. Detailed Analysis: The intervention (the PHONE Words / MEVLA-GF and MEVLA-NGF app) was designed and built by a distinct commercial entity, the Alice English Education Studio, while the study was designed, conducted, and analysed entirely by academic researchers at National Chengchi University who are not affiliated with that studio and did not design the app themselves. This structurally separates the intervention designer from the evaluators, analogous to cases where a government or academic team independently evaluates a third-party developer's product. The paper does not contain an explicit "independence" or conflict-of- interest statement, which is a limitation, but there is no evidence that the authors were involved in designing or building the app, and the app is presented as a pre-existing external commercial product chosen as a research instrument. Criterion I is met because the intervention (app) was designed by an independent commercial developer, and the evaluating researchers were not involved in its design.
    • Y

      Year Duration

      • As the Term Duration criterion (T) is not met, the stronger Year Duration criterion cannot be met either; the total tracked interval was only about six weeks.
      • "During a four-week experiment, 20 sophomore students were randomly assigned..." combined with the two-week delayed post-test. (p. 170, 173)
      • Relevant Quotes: 1) "During a four-week experiment, 20 sophomore students were randomly assigned to the experimental group with MEVLA-GF support or the control group with MEVLA-NGF support for English vocabulary learning." (p. 170) 2) "a vocabulary post-test was performed immediately at the end of the four-week experimental period, and was followed by a vocabulary delayed post-test two weeks later to evaluate vocabulary retention." (p. 173) Detailed Analysis: The total time from intervention start to final (delayed) outcome measurement is approximately six weeks, far short of the 75% of an academic year (roughly 7 months) required by criterion Y. Per the ERCT standard, since the weaker Term Duration criterion (T) is already not met, the stronger Year Duration criterion (Y) is automatically not met as well. Criterion Y is not met because the study duration (about six weeks) falls far short of a full academic year, and the prerequisite Term Duration criterion is also not met.
    • B

      Balanced Control Group

      • The only substantive difference between conditions is the gamified assessment/ranking feature itself, the explicit treatment variable, while both groups used the same core app functions, word list, and minimum weekly usage expectation.
      • "The main difference in the functions provided between the MEVLA-GF and MEVLA-NGF is that the MEVLA-NGF does not have the gamified assessment with ranking among learning peers, but the other functions provided by the two apps are the same." (p. 173-174)
      • Relevant Quotes: 1) "The six main functions provided by the PHONE Words app are the word list, customized word list, pre-established learning path, traditional assessment, gamified assessment, and ranking among learning peers. The main difference in the functions provided between the MEVLA-GF and MEVLA-NGF is that the MEVLA-NGF does not have the gamified assessment with ranking among learning peers, but the other functions provided by the two apps are the same." (p. 173-174) 2) "learners in both groups could make their own learning schedules; however, they were expected to use MEVLA-GF or MEVLA-NGF for at least five hours per week." (p. 172-173) 3) "Both the MEVLA-GF and MEVLA-NGF had the same ETS TOEIC word list." (p. 173) 4) Measured usage: experimental group "mean of total usage time was as high as 23.10 hours" (p. 181) versus control group "the mean of total usage time was as high as 21.33 hours" (p. 182). Detailed Analysis: Applying the criterion B decision tree: extra resources are present, since MEVLA-GF adds a gamified assessment and peer-ranking mechanism not present in MEVLA-NGF. However, these added functions are explicitly the treatment variable being tested, as the study's central research question, and its title, is about the effect of adding "game-related functions" to an otherwise identical vocabulary app. Both groups received the identical underlying app platform, the same TOEIC-based word list, the same four shared core functions (word list, customized word list, pre-established learning path, traditional assessment), and the same minimum expected weekly usage (five hours). The only feature withheld from the control group is the gamified assessment and peer ranking mechanism, matching the ERCT exception for Criterion B: when the additional resource is itself the explicit treatment being tested, the control group may reasonably use the non-gamified, business-as-usual version of the same tool. Measured total usage time was also similar between groups (23.10 vs 21.33 hours), indicating no gross time/resource imbalance beyond the gamification feature itself. Criterion B is met because the sole substantive difference between conditions is the gamification feature that is the explicit object of study (the treatment variable), with all other resources, materials, and time expectations held equal.
  • Level 3 Criteria

    • R

      Reproduced

      • Neither the paper itself nor an internet search for citing or follow-up literature indicates that this specific study has been independently replicated by a different research team.
      • Relevant Quotes: No quotes in the paper reference a prior or subsequent independent replication of this specific MEVLA-GF/MEVLA-NGF vocabulary app study. The discussion section (p. 182-184) cites related prior studies on mobile/game-based vocabulary learning (e.g., Dolati & Mikaili, 2011; Sukstrienwong & Vongsumedh, 2013; Uzun et al., 2013) as supporting context for the general MGBL literature, but none of these are replications of this particular intervention, app, or study design. Detailed Analysis: Criterion R requires that this specific study (or its central experimental claim, in the same design and context) be independently replicated by a different research team and published in a peer-reviewed outlet. An internet search (general web search, Semantic Scholar) for the PHONE Words app, the MEVLA-GF/MEVLA-NGF labels, and this specific Chen, Liu and Huang (2019) study found later papers on mobile/game-based vocabulary learning that cite this study as supporting literature, but no paper that reproduces this specific app, design, or claim in a new context by an independent team was located. No such replication paper (title, authors, year) could be identified. Criterion R is not met because no independent replication of this specific study was found either in the paper's own content or through internet search.
    • A

      All-subject Exams

      • Only English vocabulary outcomes were assessed; no other core school subjects were measured, and the paper does not provide the kind of upper- secondary/vocational justification the standard requires for a narrower exception.
      • "Learning performance based on vocabulary acquisition and vocabulary retention was assessed using three vocabulary tests..." (p. 172)
      • Relevant Quotes: 1) "Learning performance based on vocabulary acquisition and vocabulary retention was assessed using three vocabulary tests, a pre-test, a post-test, and a delayed post-test obtained randomly from the New TOEIC Official Test-Preparation Guide III..." (p. 172) 2) "this study adopted questions from the New TOEIC Official Test-Preparation Guide III (ETS, 2011)... to evaluate participants' vocabulary." (p. 176) 3) The study's stated goal: "assess the effects of PHONE Words, a novel mobile English vocabulary learning app... on learners' perceptions and learning performance." (p. 170, Abstract) Detailed Analysis: The study measured only English vocabulary acquisition and retention; no other core subjects (mathematics, science, social studies, etc.) were assessed. While the ERCT standard allows a narrower subject-scope exception for "highly specialised interventions in upper secondary or vocational education" with clear rationale, this study involves general university sophomores in a liberal-arts college using a general-purpose vocabulary app for a graduation English requirement, not an upper- secondary or vocational specialisation with an articulated rationale for restricting outcome scope. Criterion A is not met because only a single subject (English vocabulary) was assessed, without a qualifying vocational/specialisation exception.
    • G

      Graduation Tracking

      • Tracking ended two weeks after the four-week intervention with no further follow-up; since the Year Duration criterion (Y) is not met, this criterion cannot be met either, and no subsequent graduation-tracking papers by these authors were found.
      • "a vocabulary delayed post-test two weeks later to evaluate vocabulary retention." (p. 173)
      • Relevant Quotes: 1) "a vocabulary post-test was performed immediately at the end of the four-week experimental period, and was followed by a vocabulary delayed post-test two weeks later to evaluate vocabulary retention." (p. 173) 2) The paper's future-work section only proposes studying learner characteristics and rural-urban divides in later research (p. 185); it does not mention tracking these participants further, let alone through graduation. Detailed Analysis: Data collection concluded with a delayed post-test administered two weeks after the four-week intervention ended (about six weeks total). There is no mention of any further follow-up, and certainly no tracking of participants until graduation. An internet search for later publications by Chen, Liu, or Huang that might track this same cohort of NCCU sophomores through graduation did not surface any such follow-up study; subsequent papers found citing this work are unrelated third-party studies, not cohort follow-ups by these authors. Per the ERCT standard, since the prerequisite Year Duration criterion (Y) is not met, Graduation Tracking (G) cannot be met either. Criterion G is not met because tracking stopped shortly after the intervention with no continued follow-up found in the paper or in later publications, and the prerequisite Year Duration criterion is also not met.
    • P

      Pre-Registered

      • The paper contains no statement of pre-registration on any registry platform, nor any registration date or protocol reference, and no internet search of registry platforms located a matching entry.
      • Relevant Quotes: No quotes anywhere in the Methods, Research Instruments, Results, Discussion, or Ethical Statement sections mention a study registry, a pre-registration platform (e.g., OSF, AsPredicted, ClinicalTrials.gov), or a pre-specified analysis plan published before data collection. The only ethics-related text found is: "written informed consent was obtained from the participants after the experiment was explained in full." (p. 184, Ethical statement) Detailed Analysis: Criterion P requires explicit evidence that the study protocol, hypotheses, and planned analyses were registered on a public platform before data collection began. The paper's Ethical statement addresses informed consent and data privacy only, with no mention of pre-registration. An internet search for a pre-registration record for this study (by title, authors, and topic) found no matching entry on common registry platforms, and no secondary source referring to a pre-registration for this study was identified. Criterion P is not met because no pre-registration statement, registry link, or registration date is provided anywhere in the paper or found via internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.