Learning academic vocabulary with digital flashcards: Comparing the outcomes from computers and smartphones

Zahra Zarrati, Mohammad Zohrabi, Hakimeh Abedini, Ismail Xodabande

Published:
ERCT Check Date:
DOI: 10.1016/j.ssaho.2024.100900
  • L2 languages
  • higher education
  • Asia
  • EdTech app
  • mobile learning
0
  • C

    Randomisation occurred at the individual student level within one TEFL cohort, not at the class or school level, and no tutoring exception applies.

    "The participants were randomly divided into three distinct groups: two experimental groups and one control group." (p. 4)

  • E

    The study used a custom-selected, self-report vocabulary scale rather than a standardised exam, so this criterion is not met.

    "VKS was employed as a self-assessment tool to evaluate participants' vocabulary knowledge both before and after the intervention." (p. 4-5)

  • T

    The five-week intervention plus three-week delayed follow-up totals roughly eight weeks, short of the minimum one-term duration required.

    "Over the following five weeks, participants dedicated time to studying the target vocabulary items... a delayed post-test, which was conducted three weeks after the treatment." (p. 5)

  • D

    The control group's size, baseline scores, and treatment are documented and confirmed comparable to the experimental groups at baseline.

    "The Control Group, with 42 students, was allocated to use traditional paper-based flashcards." (p. 4)

  • S

    Randomisation was at the individual student level within one institution, not at the school level.

    "The participants were randomly divided into three distinct groups: two experimental groups and one control group." (p. 4)

  • I

    The same author team designed, delivered, and analysed the study with no independent evaluators involved.

    "Zahra Zarrati: Writing – original draft, Validation, Investigation, Formal analysis, Data curation, Conceptualization." (CRediT statement, p. 8)

  • Y

    Since Term Duration (T) is not met and the total tracked period is only about eight weeks, Year Duration is not met.

    "Over the following five weeks, participants dedicated time to studying the target vocabulary items... a delayed post-test, which was conducted three weeks after the treatment." (p. 5)

  • B

    All groups received identical content and schedule; the only extra element (a short app tutorial) is negligible and tied to the medium being tested.

    "Each week, they received a new set of flashcards, each comprising 10 words accompanied by relevant definitions and two illustrative sentences to demonstrate correct usage." (p. 5)

  • R

    No independent replication of this specific study is reported or evident; related prior studies and 2024-2025 citing papers are not replications of this particular experiment.

  • A

    Since criterion E is not met and only vocabulary knowledge (a single subject) was assessed, this criterion is not met.

    "The focus of this study was on 50 academic words, carefully selected from a corpus..." (p. 4)

  • G

    Tracking ended shortly after the intervention with no follow-up toward graduation, Y is not met, and no later graduation-tracking publication was found.

    "...a delayed post-test, which was conducted three weeks after the treatment." (p. 5)

  • P

    No pre-registration of the study protocol is mentioned anywhere in the paper, and none was found in external registries.

Abstract

This study systematically evaluates the comparative efficacy of digital flashcards on mobile and computer platforms versus traditional paper-based flashcards in augmenting academic vocabulary knowledge development among Iranian undergraduate university students. A randomized controlled trial was conducted with 112 subjects, allocated into three distinct groups: one utilizing digital flashcards via smartphones, another via laptops, and a control group employing paper-based flashcards. Over a five-week period, these participants were exposed to 50 frequently used academic vocabulary items in a self-directed learning mode. Academic vocabulary knowledge of the study participants was assessed pre- and post-intervention, employing an Analysis of Variance (ANOVA) for statistical evaluation of the test scores. The results demonstrated statistically significant learning gains in vocabulary knowledge across all groups, with the smartphone group showing the most pronounced improvements. The performance of this group notably surpassed that of the laptop group and the control group, underscoring the superior efficacy of mobile devices in facilitating academic vocabulary learning. These findings illuminate the potential of mobile-assisted language learning tools in academic settings and suggest a differential impact of device type on vocabulary acquisition. The broader implications of these outcomes for the design and implementation of technology-enhanced language learning strategies are discussed.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation occurred at the individual student level within one TEFL cohort, not at the class or school level, and no tutoring exception applies.
      • "The participants were randomly divided into three distinct groups: two experimental groups and one control group." (p. 4)
      • Relevant Quotes: 1) "This study engaged 112 undergraduate students enrolled in a Teaching English as a Foreign Language (TEFL) program at a university in Tehran, Iran." (p. 4) 2) "To investigate the efficacy of technology- assisted vocabulary learning, the participants were randomly divided into three distinct groups: two experimental groups and one control group." (p. 4) 3) "Experimental Group 1, comprising 35 students, was provided with a 1-h tutorial on the installation and usage of the Anki digital flashcard (AnkiDroid, 2020) application on their smartphones. Similarly, Experimental Group 2, also consisting of 35 students, received identical instruction for utilizing Anki on laptop computers. The Control Group, with 42 students, was allocated to use traditional paper-based flashcards." (p. 4) Detailed Analysis: The paper describes randomisation explicitly at the level of individual participants drawn from a single TEFL program cohort, not at the level of classes or schools. There is no quote describing intact classes being assigned as whole units to a condition; rather, 112 individual students were allocated across three groups. This is a self-directed, out-of-class vocabulary-study intervention (using personal smartphones/laptops/paper cards) rather than a classroom-delivered instructional program, but the ERCT exception is specifically for one-to-one personal tutoring interventions, and this study does not describe a tutoring relationship; it is an individually-randomized comparison of study media among students who may share classes, creating some risk of exposure/discussion between conditions that class-level randomisation is designed to prevent. No quote confirms class- or school-level allocation. Criterion C is not met because randomisation was conducted at the individual student level within a single program, without a description of class-level assignment or a clearly stated personal-tutoring exception.
    • E

      Exam-based Assessment

      • The study used a custom-selected, self-report vocabulary scale rather than a standardised exam, so this criterion is not met.
      • "VKS was employed as a self-assessment tool to evaluate participants' vocabulary knowledge both before and after the intervention." (p. 4-5)
      • Relevant Quotes: 1) "The focus of this study was on 50 academic words, carefully selected from a corpus comprising research articles published in 20 prominent journals within the field of applied linguistics." (p. 4) 2) "The Vocabulary Knowledge Scale (VKS) (Wesche & Paribakht, 1996) was then utilized to determine the words that were mostly unknown or less known by the participants, leading to the selection of the target words as presented in Table 1." (p. 4) 3) "Vocabulary Knowledge Scale (VKS): VKS was employed as a self-assessment tool to evaluate participants' vocabulary knowledge both before and after the intervention." (p. 4-5) 4) "Considering the number of target words (Table 1), the test had 50 items, and the score range was from 0 (not knowing any of the words) to 200 which implies knowing all the words." (p. 5) Detailed Analysis: The outcome measure is the Vocabulary Knowledge Scale (VKS), a self-report scale on which participants rate their own familiarity with words (see Fig. 1: "I don't remember having seen this word before" ... "I can use this word in a sentence"). Although VKS as a research instrument is widely used in second-language-acquisition studies, it is not a standardised, externally validated exam comparable to a national curriculum test; furthermore, the 50 specific vocabulary items tested were custom- selected by the researchers themselves for this study via corpus analysis, meaning the assessment content itself was purpose-built rather than drawn from an existing, widely recognised standardised test. This mirrors the "custom-made assessment" failure pattern described in the ERCT standard. Criterion E is not met because the outcome measure is a researcher-selected, self-assessment vocabulary scale rather than a standardised, widely recognised exam-based assessment.
    • T

      Term Duration

      • The five-week intervention plus three-week delayed follow-up totals roughly eight weeks, short of the minimum one-term duration required.
      • "Over the following five weeks, participants dedicated time to studying the target vocabulary items... a delayed post-test, which was conducted three weeks after the treatment." (p. 5)
      • Relevant Quotes: 1) "Over the following five weeks, participants dedicated time to studying the target vocabulary items through different types of flashcards. Each week, they received a new set of flashcards, each comprising 10 words..." (p. 5) 2) "Following this intensive study period, participants' vocabulary knowledge was reassessed using the VKS in both an immediate post-test and a delayed post-test, which was conducted three weeks after the treatment." (p. 5) Detailed Analysis: The intervention itself lasted five weeks, and the final (delayed) outcome measurement occurred three weeks after the treatment ended, i.e. roughly eight weeks after the intervention began. An academic term is typically 3-4 months (about 12-16 weeks), so the total tracked interval from intervention start to final measurement falls well short of one full term. Criterion T is not met because the interval from intervention start to the final (delayed) measurement is approximately eight weeks, considerably shorter than one academic term.
    • D

      Documented Control Group

      • The control group's size, baseline scores, and treatment are documented and confirmed comparable to the experimental groups at baseline.
      • "The Control Group, with 42 students, was allocated to use traditional paper-based flashcards." (p. 4)
      • Relevant Quotes: 1) "The Control Group, with 42 students, was allocated to use traditional paper-based flashcards." (p. 4) 2) "Control Group (paper flashcards) Pre-test 42 77.24 15.10 / Post-test 42 127.93 25.33 / Delayed Post-test 42 128.08 27.37" (Table 2, p. 5) 3) "In the initial VKS pre-test, the mean scores were closely aligned across the groups... and the control group (paper flashcards) 77.24." (p. 5) 4) "Initially, in the VKS pre-test, no significant differences were observed among the groups, suggesting a uniform starting level in vocabulary knowledge." (p. 6) Detailed Analysis: The paper documents the control group's size (n=42), its treatment (traditional paper-based flashcards covering the same target words on the same weekly schedule as the experimental groups), and its baseline (pre-test) performance with mean and SD in Table 2, along with a formal statistical confirmation (Table 3) that the control group did not differ significantly from the experimental groups at baseline. While a separate demographic breakdown (age, gender) per group is not provided, the size, baseline scores, and nature of the control condition are clearly and quantitatively documented. Criterion D is met because the control group's size, baseline performance, and the business-as-usual nature of its treatment (paper flashcards) are clearly documented and statistically compared to the experimental groups at baseline.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was at the individual student level within one institution, not at the school level.
      • "The participants were randomly divided into three distinct groups: two experimental groups and one control group." (p. 4)
      • Relevant Quotes: 1) "This study engaged 112 undergraduate students enrolled in a Teaching English as a Foreign Language (TEFL) program at a university in Tehran, Iran." (p. 4) 2) "To investigate the efficacy of technology- assisted vocabulary learning, the participants were randomly divided into three distinct groups: two experimental groups and one control group." (p. 4) Detailed Analysis: The study was conducted at a single university with a single cohort of 112 students from one TEFL program, randomised individually into three groups. There is no mention of multiple schools/institutions being randomised as units; the entire sample comes from one institution and randomisation was at the individual level. Criterion S is not met because there is no school-level (or higher unit) randomisation; the study involves individual-level assignment within a single institution.
    • I

      Independent Conduct

      • The same author team designed, delivered, and analysed the study with no independent evaluators involved.
      • "Zahra Zarrati: Writing – original draft, Validation, Investigation, Formal analysis, Data curation, Conceptualization." (CRediT statement, p. 8)
      • Relevant Quotes: 1) "Zahra Zarrati: Writing – original draft, Validation, Investigation, Formal analysis, Data curation, Conceptualization. Mohammad Zohrabi: Writing – original draft, Software, Methodology, Investigation, Conceptualization. Hakimeh Abedini: Writing – original draft, Software, Investigation, Formal analysis. Ismail Xodabande: Writing – review & editing, Supervision, Project administration, Methodology, Conceptualization." (CRediT statement, p. 8) 2) "The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper." (p. 8) Detailed Analysis: The CRediT statement shows that the same four authors handled conceptualization, methodology, investigation, formal analysis, and writing; there is no mention of an external/independent team conducting data collection or analysis, nor of blinded administrators or a third-party evaluation agency. The intervention (selecting words, designing materials, delivering the tutorial, and analysing results) was designed and executed entirely by the same research team. Criterion I is not met because the same authors designed, delivered, and analysed the study, with no evidence of independent or third-party conduct.
    • Y

      Year Duration

      • Since Term Duration (T) is not met and the total tracked period is only about eight weeks, Year Duration is not met.
      • "Over the following five weeks, participants dedicated time to studying the target vocabulary items... a delayed post-test, which was conducted three weeks after the treatment." (p. 5)
      • Relevant Quotes: 1) "Over the following five weeks, participants dedicated time to studying the target vocabulary items through different types of flashcards." (p. 5) 2) "...a delayed post-test, which was conducted three weeks after the treatment." (p. 5) Detailed Analysis: Per the criterion-specific rule, since criterion T (Term Duration) is not met, criterion Y (Year Duration) is automatically not met. Independently, the total tracked period (five-week intervention plus a three-week delayed post-test, roughly eight weeks) is far shorter than 75% of an academic year (~9-10 months). Criterion Y is not met both because criterion T failed and because the total duration is far short of a full academic year.
    • B

      Balanced Control Group

      • All groups received identical content and schedule; the only extra element (a short app tutorial) is negligible and tied to the medium being tested.
      • "Each week, they received a new set of flashcards, each comprising 10 words accompanied by relevant definitions and two illustrative sentences to demonstrate correct usage." (p. 5)
      • Relevant Quotes: 1) "Experimental Group 1, comprising 35 students, was provided with a 1-h tutorial on the installation and usage of the Anki digital flashcard (AnkiDroid, 2020) application on their smartphones. Similarly, Experimental Group 2... received identical instruction for utilizing Anki on laptop computers." (p. 4) 2) "Over the following five weeks, participants dedicated time to studying the target vocabulary items through different types of flashcards. Each week, they received a new set of flashcards, each comprising 10 words accompanied by relevant definitions and two illustrative sentences to demonstrate correct usage." (p. 5) 3) "The participants used their own devices to study the target words." (p. 5) Detailed Analysis: The additional resource provided to the experimental groups beyond the control group is a one-hour setup tutorial for installing/using the Anki app; no extra class time, extra study time, or budget was allocated to the experimental groups beyond this brief onboarding session. All three groups (mobile, laptop, paper) studied the same 50 target words, delivered in identical weekly sets of 10 words over the same five-week period, and used their own personal devices/materials. Since the very variable under investigation is the study medium (smartphone vs. laptop vs. paper), the differing medium is integral to (indeed, is) the treatment being tested, and the one-hour app-installation tutorial is a negligible, necessary onboarding step rather than a substantive extra educational resource. Content, dosage (10 words/week), and duration were identical across groups. Criterion B is met because all three groups received equivalent content and study time, and the only difference (a brief app-installation tutorial) is a negligible, necessary onboarding step tied directly to the study medium being tested, not an unbalanced extra resource.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study is reported or evident; related prior studies and 2024-2025 citing papers are not replications of this particular experiment.
      • Relevant Quotes: 1) "Among the few studies that explored this issue, Sage et al. (2019) undertook a comparative analysis of college students' vocabulary learning using paper, computer, and tablet flashcards... Extending this inquiry, Sage et al. (2020) further investigated the impacts of paper, laptop, and smartphone flashcards on college students' vocabulary learning outcomes." (p. 4) 2) "These findings partially align with Sage et al. (2019), who reported superior outcomes with mobile device flashcards compared to computer-based ones... our results are not in accord with Sage et al. (2020), who observed equivalent learning gains from both mobile and computer platforms." (p. 6-7) Internet Research (2026-07-29): A citation search (Semantic Scholar, keyed to DOI 10.1016/j.ssaho.2024.100900) found 26 papers citing this study, none of which independently replicate its specific design (Iranian TEFL undergraduates, 112 participants, smartphone vs. laptop vs. paper Anki flashcards for 50 corpus-derived academic words). The closest candidates were examined and ruled out: Najafi Karimi and Kheradmandi Amiri (2025), "The effects of using digital flashcards versus paper flashcards on vocabulary learning and retention of Iranian intermediate EFL learners," compares digital flashcards, paper flashcards, and word lists among 90 Iranian EFL learners aged 13-15 (not undergraduates, and only two flashcard modalities are compared, not the three-way smartphone/laptop/paper design of the target study). Xodabande, Atai, and Hashemi (2024), "Exploring the effectiveness of mobile assisted learning with digital flashcards in enhancing long-term retention of technical vocabulary among university students," shares an author (Xodabande) with the target paper and is therefore not an independent replication by a different team; it also targets technical rather than academic vocabulary and does not compare smartphone vs. laptop vs. paper conditions. Detailed Analysis: The paper itself is not a replication of a prior study but discusses related, independently conducted prior studies (Sage et al., 2019, 2020) that compared similar device-based flashcard conditions in a different population (US college students) with partially conflicting results. However, these are related but distinct studies with different populations, word sets, and designs, rather than independent replications of this specific study (Iranian TEFL students, 50 corpus-derived academic words, Anki app). No quote in the paper, nor any indication found via citation search, shows that this specific study (or an equivalent close replication of it) has itself been independently reproduced by a separate research team. Criterion R is not met because no independent replication of this specific study was found in the paper or via internet-based citation search.
    • A

      All-subject Exams

      • Since criterion E is not met and only vocabulary knowledge (a single subject) was assessed, this criterion is not met.
      • "The focus of this study was on 50 academic words, carefully selected from a corpus..." (p. 4)
      • Relevant Quotes: 1) "The present study seeks to fill this research void by comparing the learning outcomes of academic vocabulary acquisition among university students using digital flashcards on mobile phones and computers." (p. 4) 2) "The focus of this study was on 50 academic words, carefully selected from a corpus..." (p. 4) Detailed Analysis: Per the criterion-specific rule, since criterion E (Exam-based Assessment) is not met, criterion A (All-subject Exams) cannot be met either. Independently, the study measured only academic vocabulary knowledge (English/L2 vocabulary) via the VKS; no other core subjects (mathematics, science, etc.) were assessed, and no exception rationale for a specialised, single-subject focus is provided. Criterion A is not met both because criterion E failed and because only vocabulary knowledge was assessed, with no other subjects measured.
    • G

      Graduation Tracking

      • Tracking ended shortly after the intervention with no follow-up toward graduation, Y is not met, and no later graduation-tracking publication was found.
      • "...a delayed post-test, which was conducted three weeks after the treatment." (p. 5)
      • Relevant Quotes: 1) "...participants' vocabulary knowledge was reassessed using the VKS in both an immediate post-test and a delayed post-test, which was conducted three weeks after the treatment." (p. 5) 2) "Future research should aim to address these limitations, possibly by including larger and more diverse samples, extending the duration of treatments, and exploring the application of vocabulary in authentic academic contexts." (p. 7) Internet Research (2026-07-29): A search for subsequent publications by the same author team (Zarrati, Zohrabi, Abedini, Xodabande) that might report longer-term or graduation-level follow-up on this same 112-student cohort did not find any such paper. The only related 2024 paper sharing an author (Xodabande, Atai, and Hashemi, "Exploring the effectiveness of mobile assisted learning with digital flashcards in enhancing long-term retention of technical vocabulary among university students") concerns a different sample and different (technical, not academic) vocabulary, and does not extend tracking of this cohort toward graduation. No graduation-tracking follow-up study for this specific cohort could be identified. Detailed Analysis: Per the criterion-specific rule, since criterion Y (Year Duration) is not met, criterion G (Graduation Tracking) is automatically not met. Independently, the paper's own limitations section calls for "extending the duration of treatments" in future work, confirming that no long-term or graduation- level follow-up was conducted; tracking ended three weeks after the treatment concluded, and no subsequent follow-up publication tracking this cohort toward graduation was found. Criterion G is not met because tracking stopped three weeks after treatment, with no follow-up toward graduation, criterion Y also failed, and no follow-up publication was identified.
    • P

      Pre-Registered

      • No pre-registration of the study protocol is mentioned anywhere in the paper, and none was found in external registries.
      • Relevant Quotes: 1) "All data collection procedures were in agreement with the ethical standards of the institutional research committee and the 1964 Helsinki Declaration. Ethical review and approval was not required for this study in accordance with the local legislation and institutional requirements in Iran issued by National Committee for Ethics in Research (NAOEC, 2020)." (p. 8) 2) No statement referencing a study registry platform (e.g., ClinicalTrials.gov, ISRCTN, OSF Registries, AsPredicted) or a pre-registration date appears anywhere in the manuscript. Internet Research (2026-07-29): A search of the OSF Registries platform for a pre-registration matching this study (title, authors, or topic) did not return any matching entry, consistent with the absence of any pre-registration statement in the manuscript itself. Detailed Analysis: The paper includes an ethics approval statement but explicitly states that formal ethical review was not even required for this study, and there is no mention anywhere of a pre-registered protocol, hypotheses, or analysis plan being published before data collection began. No pre-registration record was found in external registries either. Criterion P is not met because no pre-registration statement, registry link, or registration date is provided anywhere in the paper or found externally.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.