The facilitative effect of the keyword mnemonic on L2 vocabulary retrieval practice

Kejia Qu, Tianzhi Liu, Yihuan Qiao, Pengcheng Wang

Published:
ERCT Check Date:
DOI: 10.1016/j.heliyon.2024.e25212
  • L2 languages
  • higher education
  • China
0
  • C

    Randomization occurred at the individual student level in a lab setting, not at the class or school level, and no tutoring exception applies.

    "Participants were randomly assigned to one of four conditions: SRR (study-restudy-restudy), STT (study-test-test), KimposedTT (imposed keyword-test-test) or KinducedTT (induced keyword-test-test)." (p. 4)

  • E

    The outcome was measured using a custom-made cued-recall test of researcher-selected vocabulary, not a standardized exam.

    "Participants returned to Laboratory 1D at a later time for the final cue recall test, in which they were provided with 20 English words as cues on one A4 sheet of paper... and asked to recollect their Chinese meanings within 3 min without feedback." (p. 5)

  • T

    The intervention and follow-up test spanned only about one day, far short of the required one-term interval.

    Fig. 1 depicts a "1 day" gap between the initial session (study, retrieval, distracter) and the "FINAL TEST"; "The Day 1 procedure took approximately 13 min to complete." (p. 5)

  • D

    The comparison/baseline conditions' sample sizes and exact procedures are clearly documented, even though demographic data are only reported in aggregate.

    "Sample sizes were as follows: SRR (n = 26), STT (n = 28), KimposedTT (n = 28), and Kinduced TT (n = 28)." (p. 4)

  • S

    Randomization was at the individual participant level within a single university lab, not among schools.

    "A total of 112 Liaoning Normal University undergraduate students participated... Participants were randomly assigned to one of four conditions." (p. 4)

  • I

    The same research team designed, conducted, and analyzed the study; no independent or external evaluation is documented.

    CRediT authorship statement shows all four authors performed conceptualization, investigation, data curation, formal analysis, and writing, with no independent/external evaluator mentioned. (p. 8)

  • Y

    Because criterion T is not met, and the study tracked outcomes only one day after the intervention, criterion Y is not met.

    Fig. 1 depicts only a "1 day" gap between the initial session and the final test. (p. 5)

  • B

    Session time/structure was matched across all four conditions, and the keyword-mnemonic instructions distinguishing two groups are integral to the treatment variable being tested.

    "The KimposedTT group and KinducedTT group differed in that, in addition to the Chinese and English word pairs, the subjects were initially given general instructions on the keyword mnemonic..." (p. 5)

  • R

    No independent replication of this specific study by a different research team was found.

  • A

    Because criterion E is not met, and only a single custom vocabulary-recall measure was used, criterion A is not met.

  • G

    Because Y is not met, and the study only tracked participants for one day with no further follow-up and no subsequent graduation-tracking publications were found, criterion G is not met.

  • P

    No pre-registration of the study protocol is mentioned; only ethics/IRB approval and a post-hoc data repository link are provided.

    "This study was reviewed and approved by the Ethics Committee of Liaoning Normal University, with the approval number: No. LL2022024." (p. 8)

Abstract

Keyword mnemonics and retrieval practice are two learning strategies that facilitate foreign language vocabulary learning. This study examined the combination of these strategies for learning English L2 vocabulary with a limited retrieval time. We recruited 110 Chinese college students studying English as a foreign language to investigate the effects of four learning strategies on the retention of English-Chinese word pairs: restudy, retrieval practice, imposed keyword mnemonic combined with retrieval practice, and induced keyword mnemonic combined with retrieval practice. The results revealed that when retrieval practice was constrained to two times, the final performance of the retrieval practice group did not exceed that of the restudy group; however, the combined keyword-retrieval group outperformed the restudy group, regardless of whether the keyword was imposed or induced. Furthermore, there was no significant difference in memory retention performance between the induced and imposed keyword-retrieval combinations. The findings suggest that when retrieval practice is constrained to two times, the keyword-retrieval strategy combination significantly enhances English L2 vocabulary learning compared to restudy or retrieval practice alone, and both the imposed and induced keyword mnemonics can strengthen its efficiency.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomization occurred at the individual student level in a lab setting, not at the class or school level, and no tutoring exception applies.
      • "Participants were randomly assigned to one of four conditions: SRR (study-restudy-restudy), STT (study-test-test), KimposedTT (imposed keyword-test-test) or KinducedTT (induced keyword-test-test)." (p. 4)
      • Relevant Quotes: 1) "The experiment was based on a factorial between-subject design." (p. 4) 2) "Participants were randomly assigned to one of four conditions: SRR (study-restudy-restudy), STT (study-test-test), KimposedTT (imposed keyword-test-test) or KinducedTT (induced keyword-test-test)." (p. 4) 3) "A total of 112 Liaoning Normal University undergraduate students participated in exchange for course credit for psychology courses or ¥ 15." (p. 4) Detailed Analysis: The paper explicitly states that individual participants (undergraduate students at a single university) were randomly assigned to one of four experimental conditions in a laboratory setting. There is no classroom or school unit that was randomized; randomization occurred at the individual-student level. The exception for personal tutoring/one-to-one teaching does not apply here, since this is a laboratory memory strategy experiment administered via computer to individuals, not a personalized tutoring intervention. Criterion C is not met because randomization was performed at the individual student level within a single laboratory sample, with no valid exception applying.
    • E

      Exam-based Assessment

      • The outcome was measured using a custom-made cued-recall test of researcher-selected vocabulary, not a standardized exam.
      • "Participants returned to Laboratory 1D at a later time for the final cue recall test, in which they were provided with 20 English words as cues on one A4 sheet of paper... and asked to recollect their Chinese meanings within 3 min without feedback." (p. 5)
      • Relevant Quotes: 1) "Thirty university students who did not engage in the main experiment were asked to rate the familiarity of the words on a 7-point Likert scale... Following the evaluation, 20 pairs of English–Chinese word pairs were chosen as the experimental material." (p. 4) 2) "Students were all non-English majors and had never taken the International English Language Testing System (IELTS) test." (p. 4) 3) "Participants returned to Laboratory 1D at a later time for the final cue recall test, in which they were provided with 20 English words as cues on one A4 sheet of paper from the same Chinese–English word pairs as in the initial study and asked to recollect their Chinese meanings within 3 min without feedback." (p. 5) Detailed Analysis: The outcome measure is a bespoke cued-recall test built from 20 word pairs that the researchers themselves selected and rated for familiarity, not a widely recognized, standardized exam. The paper explicitly notes participants had never taken a standardized language exam (IELTS), and no standardized instrument was used to assess vocabulary learning; instead, a researcher-designed recall task specific to this study's word list was used. Criterion E is not met because the assessment is a custom-made recall test rather than a standardized, widely recognized exam.
    • T

      Term Duration

      • The intervention and follow-up test spanned only about one day, far short of the required one-term interval.
      • Fig. 1 depicts a "1 day" gap between the initial session (study, retrieval, distracter) and the "FINAL TEST"; "The Day 1 procedure took approximately 13 min to complete." (p. 5)
      • Relevant Quotes: 1) "The Day 1 procedure took approximately 13 min to complete." (p. 5) 2) Fig. 1 ("The experiment procedure") shows the sequence Initial Study -> Initial Retrieval -> Distracter -> "1 day" -> Final Test. (p. 4) 3) "Participants returned to Laboratory 1D at a later time for the final cue recall test..." (p. 5) Detailed Analysis: The entire study, from initial exposure to the word pairs through the distractor task, took only about 13 minutes on a single day, and the final outcome (retention) test was administered after just one day, as shown in Fig. 1. This is drastically shorter than the minimum interval of one academic term required by the ERCT standard. Criterion T is not met because the interval from intervention start to outcome measurement was only about one day, far short of a full academic term.
    • D

      Documented Control Group

      • The comparison/baseline conditions' sample sizes and exact procedures are clearly documented, even though demographic data are only reported in aggregate.
      • "Sample sizes were as follows: SRR (n = 26), STT (n = 28), KimposedTT (n = 28), and Kinduced TT (n = 28)." (p. 4)
      • Relevant Quotes: 1) "Participants were randomly assigned to one of four conditions: SRR (study-restudy-restudy), STT (study-test-test), KimposedTT (imposed keyword-test-test) or KinducedTT (induced keyword-test-test)." (p. 4) 2) "Sample sizes were as follows: SRR (n = 26), STT (n = 28), KimposedTT (n = 28), and Kinduced TT (n = 28)." (p. 4) 3) "Participants in the SRR group repeatedly learned 20 English words and their Chinese equivalents twice for 10 s each. The instructions of SRR group remained the same as those in the initial learning." (p. 5) 4) "A total of 112 Liaoning Normal University undergraduate students participated in exchange for course credit for psychology courses or ¥ 15 (M Age = 20.26; SD = 2.15; 85 % female)." (p. 4) Detailed Analysis: This four-arm design compares learning strategies rather than a single treatment vs. an untreated control; the SRR (restudy) and STT (test-test, no keyword) conditions function as the comparison/baseline arms against the keyword-combination conditions. The paper documents the exact sample size (n = 26) and the precise procedure (twice-repeated restudy, 10 s per pair) of the SRR baseline group, and reports aggregate sample demographics (age, gender) and per-group outcome statistics (Tables 1-2). Group-level demographic breakdowns are not reported separately per condition, but the comparison group's size, exact procedure, and "treatment received" (restudy only, no keyword or test) are clearly and specifically documented. Criterion D is met because the comparison (restudy) group's size and exact procedure/conditions are clearly documented, even though demographics are only reported in aggregate across the whole sample.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomization was at the individual participant level within a single university lab, not among schools.
      • "A total of 112 Liaoning Normal University undergraduate students participated... Participants were randomly assigned to one of four conditions." (p. 4)
      • Relevant Quotes: 1) "A total of 112 Liaoning Normal University undergraduate students participated..." (p. 4) 2) "Participants were randomly assigned to one of four conditions..." (p. 4) Detailed Analysis: Randomization occurred among individual undergraduate participants recruited from a single university's psychology courses, not among schools. There is no mention of multiple schools or institutions being randomized as units. Criterion S is not met because the study was conducted at a single institution with individual-level randomization, not school-level randomization.
    • I

      Independent Conduct

      • The same research team designed, conducted, and analyzed the study; no independent or external evaluation is documented.
      • CRediT authorship statement shows all four authors performed conceptualization, investigation, data curation, formal analysis, and writing, with no independent/external evaluator mentioned. (p. 8)
      • Relevant Quotes: 1) "School of Psychology, Liaoning Normal University, Dalian, Liaoning, 116029, China" (author affiliation, p. 1) 2) "Kejia Qu: Writing - review & editing, Writing - original draft, Supervision, Project administration, Funding acquisition, Conceptualization. Tianzhi Liu: Writing - original draft, Visualization, Investigation, Data curation. Yihuan Qiao: Writing - original draft, Visualization, Methodology, Investigation, Formal analysis, Data curation. Pengcheng Wang: Writing - review & editing, Methodology." (p. 8) 3) "This research in this article was funded by the 14th Five-years Plan of National Science of Education the Key Research Topics of the Ministry of Education (DBA230365)..." (p. 9) Detailed Analysis: All four authors are affiliated with the same institution (School of Psychology, Liaoning Normal University) and, per the CRediT statement, the same author team handled conceptualization, investigation, data curation, formal analysis, and writing. There is no mention of an external evaluation team, independent data collectors, or blinded administrators; the acknowledgment section only credits funding, not independent conduct. Criterion I is not met because the same research team designed, conducted, and analyzed the study with no independent third-party evaluation.
    • Y

      Year Duration

      • Because criterion T is not met, and the study tracked outcomes only one day after the intervention, criterion Y is not met.
      • Fig. 1 depicts only a "1 day" gap between the initial session and the final test. (p. 5)
      • Relevant Quotes: 1) "The Day 1 procedure took approximately 13 min to complete." (p. 5) 2) Fig. 1 shows only a "1 day" interval between the initial session and the "FINAL TEST". (p. 4) Detailed Analysis: Per the criterion-specific instruction, if criterion T (Term Duration) is not met, Y is automatically not met. Independently, the data confirm this: the study tracked participants for only about one day between intervention and outcome measurement, far short of 75% of an academic year. Criterion Y is not met both because T is not met and because the actual follow-up interval was only one day.
    • B

      Balanced Control Group

      • Session time/structure was matched across all four conditions, and the keyword-mnemonic instructions distinguishing two groups are integral to the treatment variable being tested.
      • "The KimposedTT group and KinducedTT group differed in that, in addition to the Chinese and English word pairs, the subjects were initially given general instructions on the keyword mnemonic..." (p. 5)
      • Relevant Quotes: 1) "During the initial study phase, all participants were randomly shown 20 English words and their Chinese equivalents using E-prime 2.0... Each pair of word will be displayed for 10 s" (p. 4-5) 2) "The KimposedTT group and KinducedTT group differed in that, in addition to the Chinese and English word pairs, the subjects were initially given general instructions on the keyword mnemonic (i.e., what the keyword mnemonic is and how to create an image that incorporates the keyword and its Chinese meaning)." (p. 5) 3) "In the initial retrieval phase, unlike the SRR group, participants in the STT, KimposedTT, and KinducedTT groups took two rounds of cue recall tests following the initial study... Participants in the SRR group repeatedly learned 20 English words and their Chinese equivalents twice for 10 s each." (p. 5) 4) "After the initial retrieval phase, participants engaged in a distractor task of a 3-min simple arithmetic computation. The Day 1 procedure took approximately 13 min to complete." (p. 5) Detailed Analysis: Determining study intent: the explicit purpose of the study is to compare the presence/absence and generation method of keyword mnemonics against retrieval practice alone, i.e., the keyword instruction itself is the treatment variable under investigation (research question a/b, Section 1.5). Identifying resources: the only extra element given to the KimposedTT/KinducedTT groups beyond the other groups is the keyword-generation instruction and brief explanation of the technique; total session structure (initial study, two further 10-s rounds, distractor task, single-day gap to final test) is otherwise identical and time-matched across all four conditions. Resource integration assessment: this additional instructional content is explicitly the "treatment" being tested (keyword mnemonic vs. no keyword; imposed vs. induced), not a supplementary, separable add-on, so per the ERCT decision framework this qualifies as integral to the design (RESOURCES_ARE_TREATMENT branch of the criterion B decision tree). Criterion B is met because session time and structure are matched across all four conditions, and the additional keyword-mnemonic instructions given only to two groups are integral to the treatment variable being tested rather than an extraneous, unbalanced resource.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study by a different research team was found.
      • Relevant Quotes: 1) "Received 5 September 2023; Received in revised form 16 January 2024; Accepted 23 January 2024. Available online 24 January 2024." (p. 1) 2) No statement in the paper references any prior or subsequent independent replication of this specific Chinese-English keyword-retrieval experiment by a different research team. Detailed Analysis: This paper reports a single experiment published in early 2024. It cites related but non-replicating prior studies (e.g., Miyatsu and McDaniel's Lithuanian-English study, Karpicke and Smith's study), which are different designs/materials, not independent replications of this specific study's Chinese-English, two-time retrieval design. Internet searches (PubMed Central full-text index and Google Scholar citation search for the paper title) identified only one indexed work citing this paper: Nisya & Asy'ari (2026), "Instructional Mnemonic Strategies in Fiqh Learning for Non-Residential Students: Enhancing Memory Retention and Active Learning," which applies mnemonic techniques to Islamic religious education and is not a replication of this specific keyword-retrieval design or population. No independent replication of this specific study by a different research team was found in any available source. Criterion R is not met because no independent replication of this specific study was identified.
    • A

      All-subject Exams

      • Because criterion E is not met, and only a single custom vocabulary-recall measure was used, criterion A is not met.
      • Relevant Quotes: 1) "20 pairs of English-Chinese word pairs were chosen as the experimental material." (p. 4) 2) "Participants returned to Laboratory 1D at a later time for the final cue recall test, in which they were provided with 20 English words as cues..." (p. 5) Detailed Analysis: Per the criterion-specific instruction, since criterion E (Exam-based Assessment) is not met, criterion A is automatically not met. Independently, the study only assessed retention of a single, researcher-selected list of 20 vocabulary word pairs; no other subject areas or broader curriculum outcomes were measured. Criterion A is not met both because E is not met and because only a single narrow vocabulary-recall outcome was assessed.
    • G

      Graduation Tracking

      • Because Y is not met, and the study only tracked participants for one day with no further follow-up and no subsequent graduation-tracking publications were found, criterion G is not met.
      • Relevant Quotes: 1) Fig. 1 shows only a "1 day" gap between the initial session and the "FINAL TEST", with no further follow-up depicted. (p. 4) 2) No statement anywhere in the paper mentions any tracking of participants beyond the single one-day follow-up test, let alone until graduation. Detailed Analysis: Per the criterion-specific instruction, since criterion Y (Year Duration) is not met, criterion G is automatically not met. Independently, internet searches (PubMed Central full-text index and Google Scholar) for follow-up or longitudinal publications by the same author team (Kejia Qu, Tianzhi Liu, Yihuan Qiao, Pengcheng Wang) tracking this cohort found no such papers; only the original 2024 Heliyon article and one unrelated 2026 citing paper on mnemonic strategies in Fiqh learning were located. The study only followed participants for one day after the initial session, with no mention of any longer-term or graduation tracking, nor any indication of planned or published follow-up studies on this cohort. Criterion G is not met both because Y is not met and because no follow-up beyond one day is reported, and no subsequent graduation-tracking publications by the same authors were found.
    • P

      Pre-Registered

      • No pre-registration of the study protocol is mentioned; only ethics/IRB approval and a post-hoc data repository link are provided.
      • "This study was reviewed and approved by the Ethics Committee of Liaoning Normal University, with the approval number: No. LL2022024." (p. 8)
      • Relevant Quotes: 1) "This study was reviewed and approved by the Ethics Committee of Liaoning Normal University, with the approval number: No. LL2022024. All participants provided informed consent to participate in the study." (Ethics statement, p. 8) 2) "The data can be downloaded at: https://osf.io/tr5n9/." (Data availability statement, p. 8) Detailed Analysis: The paper documents institutional ethics/IRB approval and a post-hoc OSF link for the data, but nowhere states that the study protocol (hypotheses, methods, planned analyses) was pre-registered on a registry platform before data collection began. An internet check of the linked OSF project (https://osf.io/tr5n9/) did not surface any pre-registration record with a timestamp preceding data collection; the link functions as a post-hoc data-sharing repository rather than a pre-registration. Ethics approval is not equivalent to pre-registration of the study protocol. Criterion P is not met because no pre-registration statement, registry link, or registration date is provided anywhere in the paper, and no external pre-registration record was found.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.