Abstract
Healthcare students often struggle with learning medical terminology due to its complexity and abstract nature. This randomized controlled trial assessed the effectiveness of MedQuiz, a digital serious game, in enhancing immediate terminology acquisition and user satisfaction among 60 undergraduate students in health-related programs. Participants were assigned to either a control group with traditional instruction or an intervention group using MedQuiz alongside lectures. Post-test scores were significantly higher in the intervention group (P < .001), and user experience predicted performance. Usability metrics (SUS = 90.36%) and playability ratings indicated strong engagement. The MEEGA+ (Metrics for Educational Game Assessment + ) framework showed positive perceptions of usability and learning, while entertainment value was moderate. Findings support MedQuiz as a scalable and engaging tool for short-term medical terminology learning. Long-term retention was not assessed. Further studies should examine delayed learning outcomes and explore integration with artificial intelligence-based spaced repetition systems for sustained knowledge acquisition.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomization was at the individual student level within a single course, not at the class or school level, and the intervention is not one-to-one tutoring.
- "Eligible students were randomly assigned to either the intervention group (MedQuiz) or the control group (traditional instruction) using a 1:1 ratio." (p. 9)
Relevant Quotes:
1) "All eligible participants consented to participate in the study and were randomly allocated into two equal groups: the intervention group (30 participants) and the control group (30 participants)." (p. 2)
2) "To ensure balanced group sizes and reduce selection bias, block randomization with a fixed block size of 4 was employed. The randomization sequence was generated prior to participant recruitment using SealedEnvelope.com, a secure online platform for computer-generated random allocation." (p. 9)
3) "Eligible students were randomly assigned to either the intervention group (MedQuiz) or the control group (traditional instruction) using a 1:1 ratio." (p. 9)
4) "The study population comprised undergraduate students enrolled in the Medical Terminology course at the School of Paramedical Sciences, MUMS." (p. 9)
Detailed Analysis:
The ERCT 'C' criterion requires randomization at the class level or stronger (school level), unless the intervention is one-to-one tutoring. In this study, randomization was performed at the individual student level using block randomization with a 1:1 ratio, assigning 30 students to the intervention and 30 to the control group. All students were drawn from the same Medical Terminology course at a single institution (MUMS), meaning treatment and control students were mixed within the same classes/cohort. This is a student-level RCT, not a class-level or school-level design. The intervention (a competitive multiplayer digital game used alongside lectures) is not a one-to-one personal tutoring intervention, so the tutoring exception does not apply. Because treatment and control students were in the same course, there is a clear risk of contamination, which the criterion is designed to prevent (the authors themselves acknowledge contamination risk and used one-time activation codes to mitigate it).
Criterion C is not met because randomization was conducted at the individual student level within a single course rather than at the class or school level, and the intervention is not a one-to-one tutoring exception.
-
E
Exam-based Assessment
- The outcome was measured with an author-developed, study-specific terminology test rather than a widely recognised standardised exam.
- "This assessment, originally developed as a 20-item test in the published protocol, was expanded to 40 expert-validated items prior to implementation to ensure broader coverage of medical terminology domains." (p. 9)
Relevant Quotes:
1) "The primary outcome of this study was the improvement in medical terminology knowledge, measured by the change in scores on a standardized multiple-choice Medical Terminology Assessment administered pre- and post-intervention. This assessment, originally developed as a 20-item test in the published protocol, was expanded to 40 expert-validated items prior to implementation to ensure broader coverage of medical terminology domains." (p. 9)
2) "Knowledge acquisition was assessed pre- and post-intervention using a standardized Medical Terminology Assessment, consisting of 40 multiple-choice questions validated by a panel of subject matter experts." (p. 13)
3) "Although the assessment was expanded to 40 items prior to implementation—following curriculum mapping and expert panel recommendations to ensure comprehensive content coverage across all major body systems—this post-hoc analysis aimed to verify whether the observed effects were dependent on the additional items." (p. 5)
Detailed Analysis:
The ERCT 'E' criterion requires a standardised, widely recognised exam-based assessment (e.g., a national curriculum exam or state-wide standardised test), not a test specially designed by the researchers for the study. Although the authors label their instrument a "standardized multiple-choice Medical Terminology Assessment," the description makes clear it was developed by the authors themselves: it originated as a 20-item test in their own published protocol and was then expanded to 40 items and "validated by a panel of subject matter experts" for this study. This is an author-developed, study-specific instrument aligned to the intervention content, not an external, widely recognised standardised exam. The word "standardized" here refers to internal validation/consistency, not to a nationally or internationally recognised standardised examination. This is precisely the type of custom test the criterion warns against.
Criterion E is not met because the outcome was measured with an author-developed, study-specific Medical Terminology Assessment rather than a widely recognised standardised exam.
-
T
Term Duration
- The intervention-to-measurement interval was only about two months with no delayed follow-up, shorter than one full academic term (~3-4 months).
- "This two-month period allowed for adequate exposure to the MedQuiz and timely post-intervention assessment." (p. 9)
Relevant Quotes:
1) "A two-month RCT was conducted to evaluate the effectiveness of the MedQuiz digital serious game in enhancing medical terminology learning." (p. 13)
2) "The duration of the intervention was determined based on the academic structure of the Medical Terminology course, which spans approximately four months in a standard semester. To ensure curricular alignment and educational equity, the intervention was implemented during the first half of the semester. This two-month period allowed for adequate exposure to the MedQuiz and timely post-intervention assessment." (p. 9)
3) "While previous studies have explored both short-term and long-term learning effects of digital serious games, the current study focuses solely on immediate learning outcomes. Long-term retention was not assessed in our design." (p. 2)
4) "Second, the current study specifically assessed immediate post-intervention learning outcomes. No delayed post-tests were conducted; therefore, conclusions regarding long-term knowledge retention cannot be drawn." (p. 8)
Detailed Analysis:
The ERCT 'T' criterion requires that outcomes be measured at least one full academic term (typically ~3-4 months) after the intervention begins. The standard allows naturally short interventions but insists on term-long follow-up tracking from intervention start to measurement. Here the intervention spanned roughly two months ("first half of the semester"), with the post-test administered immediately after this two-month period. The interval from intervention start to outcome measurement is therefore about two months, which is shorter than one full academic term (the paper itself describes the semester / standard term as approximately four months). The authors explicitly state only immediate post-intervention outcomes were measured and that no delayed post-tests or longer follow-up were conducted.
Criterion T is not met because the interval from intervention start to outcome measurement was only about two months, shorter than one full academic term, and no longer-term follow-up was conducted.
-
D
Documented Control Group
- The control group's size, demographics, baseline scores, and business-as-usual conditions are clearly documented in Table 1 and Fig. 2.
- "The demographic characteristics of participants in the intervention and control groups are presented in Table 1. No significant differences were observed between the groups in terms of gender, age, academic grade point averages (GPA), or field of study (P > 0.05)..." (p. 3)
Relevant Quotes:
1) "All eligible participants consented to participate in the study and were randomly allocated into two equal groups: the intervention group (30 participants) and the control group (30 participants). Notably, there were no dropouts throughout the study, and all participants completed the entire study protocol." (p. 2)
2) "The demographic characteristics of participants in the intervention and control groups are presented in Table 1. No significant differences were observed between the groups in terms of gender, age, academic grade point averages (GPA), or field of study (P > 0.05), indicating baseline homogeneity prior to the intervention." (p. 3)
3) Table 1 reports, for the Control group: Gender Male 13 (43.3) / Female 17 (56.7); GPA 17.05 ± 1.72; Age 21.10 ± 1.47; Field of Study HIT 18 (60) / Speech therapy 12 (40); Total n = 30. (Table 1, p. 3)
4) "the control group, which received traditional lecture-based instruction only." (p. 9)
5) Figure 2: "the control group had a smaller gain (16.60 ± 5.51 to 20.80 ± 7.79)." (Fig. 2, p. 4)
Detailed Analysis:
The ERCT 'D' criterion requires the control group to be well-documented, including demographics, baseline performance, size, and the conditions/treatment it received. The paper provides this: Table 1 gives the control group's size (n = 30), gender distribution, mean GPA, mean age, and field-of-study composition, with statistical comparison confirming baseline homogeneity. The control group's baseline and post-test scores are reported (pre 16.60 ± 5.51; post 20.80 ± 7.79 in Fig. 2). The conditions are clearly described: the control group received traditional lecture-based instruction only and did not receive the MedQuiz game during the study (gaining access only after completion). This is detailed documentation sufficient to confirm comparability at baseline and the treatment received.
Criterion D is met because the control group's size, demographics, baseline scores, and conditions (traditional instruction only) are clearly documented in Table 1 and Fig. 2.
-
Level 2 Criteria
-
S
School-level RCT
- The study randomized individual students at a single institution, with no school-level randomization.
- "since the study was conducted at a single academic institution, there remains a possibility of information exchange between the intervention and control groups (contamination bias)." (p. 8)
Relevant Quotes:
1) "Eligible students were randomly assigned to either the intervention group (MedQuiz) or the control group (traditional instruction) using a 1:1 ratio." (p. 9)
2) "To ensure balanced group sizes and reduce selection bias, block randomization with a fixed block size of 4 was employed." (p. 9)
3) "This RCT involved 60 undergraduate students from MUMS, enrolled in the Health Information Technology (HIT) and Speech Therapy programs." (p. 2)
4) "Third, since the study was conducted at a single academic institution, there remains a possibility of information exchange between the intervention and control groups (contamination bias)." (p. 8)
Detailed Analysis:
The ERCT 'S' criterion requires randomization at the school level (the educational institution or unit implementing the intervention), with multiple schools randomly assigned. This study was conducted at a single institution (MUMS) and randomized individual students, not schools. There is only one institution and no randomization of schools, centers, or sites. The authors explicitly note the single-institution design as a limitation and recommend future multi-site trials.
Criterion S is not met because randomization was at the individual student level within a single institution, with no school-level (or higher-unit) randomization.
-
I
Independent Conduct
- The same team designed, administered, analyzed, and reported the trial, with no independent third-party evaluation of the intervention they created.
- "we acknowledge that the serious game was introduced by an instructor affiliated with the study. This instructional role may have introduced social desirability bias in participants' evaluations." (p. 8)
Relevant Quotes:
1) "To bridge this research gap, we developed MedQuiz, an innovative digital serious game tool designed to enhance medical terminology acquisition..." (p. 2)
2) "K.K. supervised the project, acquired funding, and coordinated the technical development of the MedQuiz game... M.S.F. designed and implemented the MedQuiz platform... K.K., S.F.M.B. and M.S. also authored the RCT protocol." (Author contributions, p. 15)
3) "we acknowledge that the serious game was introduced by an instructor affiliated with the study. This instructional role may have introduced social desirability bias in participants' evaluations. Future studies should consider involving third-party facilitators to administer the intervention and its evaluation, thereby minimizing potential bias related to instructor influence." (p. 8)
4) "The statistician performing the analyses was unaware of group assignments. Moreover, all assessments were coded, preventing evaluators from associating test results with specific groups." (p. 9)
Detailed Analysis:
The ERCT 'I' criterion requires that the study be conducted independently from the authors who designed the intervention, to reduce bias in implementation, measurement, analysis, and reporting. Here, the same research team designed and developed MedQuiz, authored the RCT protocol, administered the study, and analyzed and reported the results. The authors explicitly acknowledge that the game "was introduced by an instructor affiliated with the study," which "may have introduced social desirability bias," and recommend that future studies use third-party facilitators — confirming the absence of independent conduct in this trial. Although the statistician was blinded to group assignment and assessments were coded, these are blinding measures within the same team and do not amount to an external, third-party evaluator independent of the intervention designers, as the criterion requires.
Criterion I is not met because the same team that designed and developed MedQuiz also administered, analyzed, and reported the trial, with the intervention introduced by an affiliated instructor and no independent third-party evaluation.
-
Y
Year Duration
- The study covered only about two months from start to measurement, far short of a year, and criterion T was not met.
- "A two-month RCT was conducted to evaluate the effectiveness of the MedQuiz digital serious game in enhancing medical terminology learning." (p. 13)
Relevant Quotes:
1) "A two-month RCT was conducted to evaluate the effectiveness of the MedQuiz digital serious game in enhancing medical terminology learning." (p. 13)
2) "the intervention was implemented during the first half of the semester. This two-month period allowed for adequate exposure to the MedQuiz and timely post-intervention assessment." (p. 9)
3) "No delayed post-tests were conducted; therefore, conclusions regarding long-term knowledge retention cannot be drawn." (p. 8)
Detailed Analysis:
The ERCT 'Y' criterion requires outcomes to be measured at least 75% of one full academic year (~9-10 months) after the intervention begins. This study spanned only about two months from intervention start to outcome measurement, with no delayed follow-up. This is far short of a full academic year. Additionally, per the prompt's criterion-specific rule, if criterion T (Term Duration) is not met, then criterion Y cannot be met; criterion T is not met here.
Criterion Y is not met because the intervention-to-measurement interval was only about two months, far below 75% of an academic year, and because criterion T was not met.
-
B
Balanced Control Group
- The intervention group received substantial extra study/gameplay time as a supplementary add-on that the business-as-usual control group did not get in equivalent matched form.
- "the platform served purely as an optional enhancement to the existing curriculum." (p. 5)
Relevant Quotes:
1) "Participants were randomly assigned to one of two groups: the intervention group, which engaged with the MedQuiz, digital serious game alongside traditional instruction, and the control group, which received traditional lecture-based instruction only." (p. 9)
2) "All educational content embedded within the MedQuiz platform, including instructional materials and quiz questions, was identical to the resources available to the control group in the form of lecture notes and handouts. This ensured that the only difference between the groups was the digital serious game component, serving as an instructional supplement rather than a content replacement." (p. 13)
3) "There was no prescribed or mandatory duration for game use. Students in the intervention group were free to engage with the platform at their discretion, and the platform served purely as an optional enhancement to the existing curriculum." (p. 5)
4) "Students in the intervention group used MedQuiz for an average of 3.53 h per week and participated in approximately 7.43 competitive sessions per week." (p. 5)
5) "the intervention group was additionally provided access to the MedQuiz game for the duration of the study, integrating digital serious games into their routine educational activities." (p. 13)
Detailed Analysis:
Criterion B compares the time, budget, and materials provided to intervention versus control conditions. The intervention group received substantial additional resources: access to the MedQuiz digital game, on which they spent an average of 3.53 hours per week of extra study/gameplay over the 8-week period, plus the game's multiplayer competition, analytics, and feedback features. The control group received only traditional lecture-based instruction with the identical content in the form of lecture notes and handouts, and received no comparable supplementary activity or matched extra time. Thus extra time/resources were given only to the intervention group (EXTRA_RESOURCES_PRESENT = true), and the difference (~3.5 h/week of additional engaged learning time for 8 weeks) is not negligible.
Applying the decision tree: the key question is whether these additional resources (the digital game and the extra time-on-task it generates) are the explicit treatment variable being tested, or a separable, confounding add-on. The paper frames MedQuiz as an "instructional supplement rather than a content replacement" and "an optional enhancement to the existing curriculum" used "alongside traditional instruction." The study's own analyses show that engagement intensity (hours played, number of games) strongly correlated with post-test gains (r = 0.715 and r = 0.726), indicating that the additional time-on-task itself — not merely the game format — drove outcomes. Because the game is positioned as a supplementary add-on that increased study time on top of business-as-usual, rather than the resource increase itself being the core treatment variable with a matched-time control, the extra time and engagement constitute an unbalanced, potentially confounding resource. The control group was not given equivalent additional supervised practice time or an active alternative activity, so the effect of the game cannot be cleanly isolated from the effect of extra time-on-task.
Criterion B is not met because the intervention group received substantial additional study/gameplay time (~3.53 h/week over 8 weeks) as a supplementary add-on that was not matched by any equivalent time or activity in the business-as-usual control group, and this extra time was a confounding resource rather than the explicitly framed treatment variable.
-
Level 3 Criteria
-
R
Reproduced
- The MedQuiz trial is novel and has not been independently replicated by another team; the authors themselves call for future replication.
- "represents one of the first RCTs specifically evaluating a digital serious game developed for medical terminology acquisition among healthcare students." (p. 8)
Relevant Quotes:
1) "This study provides valuable insights into the role of digital serious games in medical education and represents one of the first RCTs specifically evaluating a digital serious game developed for medical terminology acquisition among healthcare students." (p. 8)
2) "Replication in larger cohorts is needed to confirm the stability and generalizability of these associations." (p. 5)
3) "Nonetheless, replication with larger samples is recommended to further enhance model generalizability." (p. 8)
Detailed Analysis:
The ERCT 'R' criterion requires the specific study to be independently replicated by a different research team in a different context, published in a peer-reviewed journal. MedQuiz is a newly developed, proprietary tool, and the authors describe this as "one of the first RCTs" evaluating a digital serious game for medical terminology — explicitly novel. The paper cites other digital-serious-game studies (e.g., TERMInator, HistoRM, Tan et al.) for context, but these use different tools and are not replications of the MedQuiz trial. An internet search across Nature, PubMed, and general web sources for an independent replication of this specific MedQuiz trial returned none; the study was received 8 October 2025, accepted 27 February 2026, and published online 6 April 2026, and the source code is proprietary and not publicly available. The authors themselves call for future replication, confirming none currently exists.
Criterion R is not met because there is no independent replication of this specific MedQuiz trial by another research team; the study is described as one of the first of its kind and the authors call for future replication.
-
A
All-subject Exams
- Only a single subject (medical terminology) was assessed, and criterion E was not met, so all-subject exams cannot be satisfied.
- "The primary outcome of this study was the improvement in medical terminology knowledge, measured by the change in scores on a standardized multiple-choice Medical Terminology Assessment administered pre- and post-intervention." (p. 9)
Relevant Quotes:
1) "The primary outcome of this study was the improvement in medical terminology knowledge, measured by the change in scores on a standardized multiple-choice Medical Terminology Assessment administered pre- and post-intervention." (p. 9)
2) "Although the assessment was expanded to 40 items prior to implementation—following curriculum mapping and expert panel recommendations to ensure comprehensive content coverage across all major body systems..." (p. 5)
3) "The secondary outcomes included: User engagement... Learner experience, including usability, motivation, and perceived learning, was measured using the MEEGA+ questionnaire." (p. 9)
Detailed Analysis:
The ERCT 'A' criterion requires that the study measure impact on all main subjects taught, using standardised exam-based assessments, and explicitly depends on criterion E being met. This study measured only a single subject — medical terminology — via an author-developed test; it did not assess any other core subjects of the curriculum. The "all major body systems" coverage refers to subtopics within the single medical-terminology domain, not multiple school subjects. Furthermore, per the prompt's criterion-specific rule, if criterion E (Exam-based Assessment) is not met, then criterion A cannot be met; criterion E is not met here.
Criterion A is not met because outcomes were assessed in only one subject (medical terminology) and because criterion E (standardised exam-based assessment) was not met.
-
G
Graduation Tracking
- The study measured only immediate outcomes with no follow-up to graduation, no follow-up papers were found, and criterion Y was not met.
- "Long-term retention was not assessed in our design." (p. 2)
Relevant Quotes:
1) "While previous studies have explored both short-term and long-term learning effects of digital serious games, the current study focuses solely on immediate learning outcomes. Long-term retention was not assessed in our design." (p. 2)
2) "No delayed post-tests were conducted; therefore, conclusions regarding long-term knowledge retention cannot be drawn. Although students maintained weekly engagement during the 8-week intervention, future trials should include follow-up assessments at 4–8 weeks post-intervention to evaluate the durability of learning gains over time." (p. 8)
Detailed Analysis:
The ERCT 'G' criterion requires tracking participants until their graduation to assess long-term impact. This study measured only immediate post-intervention outcomes, with no delayed post-tests and no follow-up tracking, let alone tracking to graduation. The authors explicitly state long-term retention was not assessed. An internet search for subsequent or follow-up publications by the same authors tracking this cohort to graduation returned none; the most recent related work (the protocol paper, 2022, and the MARS-based and AI-perception studies, 2024–2025) does not report graduation tracking of this trial cohort. Additionally, per the prompt's criterion-specific rule, if criterion Y (Year Duration) is not met, then criterion G cannot be met; criterion Y is not met here.
Criterion G is not met because the study conducted no follow-up beyond the immediate post-test and did not track students to graduation, no follow-up publications doing so were found, and criterion Y was not met.
-
P
Pre-Registered
- The authors explicitly state the study was not registered in any clinical trial registry, and no registry record was found online, so registry pre-registration is absent.
- "The study was not registered in a clinical trial registry." (p. 14)
Relevant Quotes:
1) "The complete RCT protocol, outlining the full methodological framework of this study, is available for reference." (p. 9) [reference 24: Mousavi Baigi et al., Stud. Health Technol. Inform. 295, 51-54 (2022), a published protocol]
2) "The study protocol was published (DOI: 10.3233/SHTI220658). The study was approved by the Ethics Committee of Mashhad University of Medical Sciences (Approval No.: IR.MUMS.REC.1400.336). The study was not registered in a clinical trial registry." (pp. 13-14)
3) "This study was approved by the Ethics Committee of Mashhad University of Medical Sciences (MUMS), under Approval ID: IR.MUMS.REC.1400.336, dated 29 January 2022." (p. 13)
Detailed Analysis:
The ERCT 'P' criterion requires the full study protocol to be pre-registered on a registry platform before data collection begins, with a verifiable registration date. The authors published a protocol paper (DOI: 10.3233/SHTI220658, 2022) describing the planned trial, but they explicitly and unambiguously state that "The study was not registered in a clinical trial registry." An internet search of trial registries (including the Iranian Registry of Clinical Trials, IRCT, commonly used by MUMS studies, as well as ClinicalTrials.gov and ISRCTN) found no registry entry for this MedQuiz trial, consistent with the authors' statement. Publication of a protocol paper in a journal is not the same as pre-registration of the protocol in a trial registry with a time-stamped registration entry that the criterion requires. Because the authors directly confirm the absence of registry pre-registration and no registry record was found, the criterion's core requirement is not satisfied.
Criterion P is not met because the study was explicitly not registered in any clinical trial registry (and no registry record was found online); a published protocol paper does not substitute for time-stamped registry pre-registration before data collection.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.