Abstract
Background: Active learning methods like flipped classrooms, case-based learning (CBL), and game-based learning (GBL) are increasingly important in medical and pharmacy education. While studies suggest integrating these methods may improve outcomes, direct comparisons of CBL and GBL within flipped classrooms are limited, often focusing on small sample sizes and different student populations. This study compares the effectiveness of GBL and CBL in a flipped classroom for pharmacy education which are held completely virtual, aiming to assess learning outcomes and student satisfaction. Methods: Participants were randomly assigned to the GBL or CBL group. In-class activities for both groups followed virtual flipped instruction classrooms on pharmacotherapy topics. Knowledge-based tests were used to assess learning outcomes, and a reliable and validated questionnaire was employed to measure student satisfaction. Results: The 56 fourth-year PharmD students completed the study. Whereas both groups showed significant learning outcome gains, the scores for all tests were invariably higher in GBL; however, none of them are significant. Satisfaction with the learning method was also greater in the GBL group (3.82) than in the CBL group (3.61). No statistically significant differences were observed for either the test scores or the satisfaction levels of either group. Conclusions: Both CBL and GBL prove to be equally effective in pharmacy education in a virtual environment.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomization was performed at the individual student level within a single university cohort, not at the class or school level, and no tutoring exception applies.
- "Participants were allocated using a computer-generated constrained randomization algorithm (Python software) designed to minimize differences between groups in sex distribution and cumulative GPA: the CBL group (29 students) and the GBL group (27 students)."
Relevant Quotes:
1) "Participants were randomly assigned to the GBL or CBL group." (Abstract, Methods)
2) "Participants were allocated using a computer-generated constrained randomization algorithm (Python software) designed to minimize differences between groups in sex distribution and cumulative GPA: the CBL group (29 students) and the GBL group (27 students)." (Methods)
3) "The participants were then randomly assigned to either the case-based learning (CBL) group or the game-based learning (GBL) group, ensuring a balanced distribution of cumulative GPA and gender." (Intervention)
4) "The selected population comprises fourth-year students pursuing their PharmD at TUMS who have already completed Pharmacology 1 and Pharmacotherapy 1." (Participants and Sample Size)
Detailed Analysis:
The ERCT 'C' criterion requires randomization by entire class (or the stronger school level), unless the intervention is one-to-one tutoring. Here the unit of randomization is the individual student: a constrained algorithm assigned individual PharmD students within a single fourth-year cohort at one institution (TUMS) to the CBL or GBL arm. This is student-level randomization, not class-level or school-level. The intervention is not personal tutoring (CBL was group-based case discussion and GBL was an individually-played board game within shared virtual classes), so the tutoring exception does not apply. Because both arms were drawn from the same cohort and subgroups were led by facilitators in shared online classes, contamination between conditions cannot be excluded.
Criterion C is not met because randomization was carried out at the individual student level within a single class/cohort rather than at the class or school level.
-
E
Exam-based Assessment
- Learning outcomes were measured with a researcher-built 26-item knowledge test, not a widely recognized standardized exam.
- "The knowledge-based test questions were prepared in compliance with Millan's checklist (19) and comprehensively covered the educational content."
Relevant Quotes:
1) "A knowledge-based test and a satisfaction questionnaire were used to measure learning outcomes and satisfaction. The knowledge-based test questions were prepared in compliance with Millan's checklist (19) and comprehensively covered the educational content." (Data collection)
2) "The test was checked by fifteen experts in the field of clinical pharmacy for the content validity ratio using the Lawshe method, where all the items had values more than 0.49." (Psychometric Properties)
3) "The finalized test includes 26 multiple-choice questions with 0.7 acceptable reliability." (Psychometric Properties)
Detailed Analysis:
The 'E' criterion requires a standardized, widely recognized exam that was not created specifically for the study. In this paper the outcome measure was a 26-item multiple-choice knowledge test developed by the research team specifically for this hypertension pharmacotherapy course. Although the authors report content-validity, face-validity, and KR-20 reliability procedures, this is a locally constructed instrument rather than a national or otherwise standard exam. No state-wide, national, or externally recognized standardized test is named.
Criterion E is not met because outcomes were assessed with a custom-made knowledge test rather than a recognized standardized exam.
-
T
Term Duration
- The intervention and outcome measurement spanned only about one to two weeks, far shorter than a full academic term.
- "Test 1 was then sent to the participants, and once all participants completed it, the electronic educational content was provided as the self-directed portion of the flipped classroom. The participants were given one week to study the content."
Relevant Quotes:
1) "Test 1 was then sent to the participants, and once all participants completed it, the electronic educational content was provided as the self-directed portion of the flipped classroom. The participants were given one week to study the content." (Intervention)
2) "After the study period, Test 2 was administered. The participants were then randomly assigned to either the case-based learning (CBL) group or the game-based learning (GBL) group..." (Intervention)
3) "Finally, the participants were asked to complete Test 3 and the Satisfaction Questionnaire." (Intervention)
4) "Firstly, our sample was relatively small (n=56)... suggesting both instructional approaches... were similarly effective in improving short-term learning outcomes in this cohort." (Discussion)
Detailed Analysis:
The 'T' criterion requires that outcomes be measured at least one full academic term (~3-4 months) after the intervention begins. The described sequence is Test 1 (baseline), one week of self-directed study, Test 2, then a single in-class CBL/GBL session, immediately followed by Test 3. The entire intervention-to-final- measurement window is on the order of one to two weeks, and the authors themselves repeatedly describe the effects as "short-term." Although the overall project ran across 2021-2023, the per-participant intervention and follow-up interval is only a couple of weeks.
Criterion T is not met because the interval from intervention start to outcome measurement was roughly one to two weeks, well short of one academic term.
-
D
Documented Control Group
- Both study arms, including the comparator group, are documented with size, demographics, and baseline characteristics.
- "There were no statistically significant differences in the number, sex distribution (χ²(1) = 1.03, p = 0.31), or distribution of cumulative GPA between the CBL and GBL groups (Table 1)."
Relevant Quotes:
1) "There were no statistically significant differences in the number, sex distribution (χ²(1) = 1.03, p = 0.31), or distribution of cumulative GPA between the CBL and GBL groups (Table 1)." (Results)
2) "Table 1: Descriptive characteristics of the study groups... Participants number 29 (51.8%) CBL / 27 (48.2%) GBL; Female 19 (66%) / 21 (78%); Mean of age (years) 21.48 ±0.57 / 21.56 ±0.89; Cumulated GPA 16.53 ±1.17 / 16.20 ±1.24." (Table 1)
3) "In the CBL group, each subgroup was tasked with answering two scenarios, each consisting of four parts... a clinical pharmacy professor in the main class reviewed and discussed the scenarios." (Intervention)
4) Baseline Test 1 scores are reported for both arms (CBL 14.31 ± 2.98; GBL 14.7 ± 3.208). (Table 3)
Detailed Analysis:
This is a two-arm comparative-effectiveness RCT in which the CBL arm serves as the active comparator (control) for GBL and vice versa. The 'D' criterion requires a clear description of the control/comparator group's characteristics, size, and conditions. The paper provides exactly this: Table 1 details each arm's sample size, sex distribution, mean age, and cumulative GPA, and Table 3 reports baseline (Test 1) scores, confirming baseline comparability. The conditions each arm received are also described. This level of documentation allows a proper comparison.
Criterion D is met because the comparator group's size, demographics, baseline performance, and conditions are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomization occurred at the individual student level in a single institution, not across schools.
- "Participants were allocated using a computer-generated constrained randomization algorithm (Python software)... the CBL group (29 students) and the GBL group (27 students)."
Relevant Quotes:
1) "Participants were allocated using a computer-generated constrained randomization algorithm (Python software) designed to minimize differences between groups in sex distribution and cumulative GPA: the CBL group (29 students) and the GBL group (27 students)." (Methods)
2) "The selected population comprises fourth-year students pursuing their PharmD at TUMS." (Participants and Sample Size)
Detailed Analysis:
The 'S' criterion requires randomization of entire schools or equivalent implementing units. In this study the randomized unit is the individual student, and all participants belong to a single institution (Tehran University of Medical Sciences). No multiple schools, centers, or sites were randomized.
Criterion S is not met because randomization was at the student level within one institution, not at the school level.
-
I
Independent Conduct
- The same authors designed the intervention, ran the study, and analyzed the data, with no independent or third-party evaluator.
- "Mahtab Amini contributed to the conception and design of the work, data collection, and writing of the manuscript."
Relevant Quotes:
1) "For this study, we have designed and used both a CBL-based flipped classroom and a GBL-based flipped classroom that was fully virtual." (Methods)
2) "Maryam Alizadeh contributed to the conception and design of the work and substantively revised the manuscript. Mahtab Amini contributed to the conception and design of the work, data collection, and writing of the manuscript. Shahideh Amini contributed to the conception and design of the work and writing the manuscript." (Authors' contributions)
3) "This work was supported by the Health Professions Research Center, Tehran, Iran, under Grant Code 1401-3-255-62769." (Funding)
Detailed Analysis:
The 'I' criterion requires that the evaluation be conducted independently of those who designed the intervention, to reduce bias in implementation, measurement, and analysis. Here the same author team designed the CBL/GBL flipped classrooms and the custom board game, collected the data, and performed the analysis. There is no statement of an external evaluation team, blinded independent assessors, or third-party oversight of data collection or analysis.
Criterion I is not met because the intervention designers also conducted and analyzed the study with no independent oversight.
-
Y
Year Duration
- The study covered only about one to two weeks, far short of a full academic year, and criterion T is not met.
- "The participants were given one week to study the content."
Relevant Quotes:
1) "The participants were given one week to study the content." (Intervention)
2) "Finally, the participants were asked to complete Test 3 and the Satisfaction Questionnaire." (Intervention)
3) "...were similarly effective in improving short-term learning outcomes in this cohort." (Discussion)
Detailed Analysis:
The 'Y' criterion requires outcomes to be measured at least 75% of an academic year (~9-10 months) after the intervention begins, and per the specification it cannot be met if 'T' is not met. The intervention-to- measurement window here is roughly one to two weeks, and the authors describe the outcomes as short-term. Because T is not met and the duration is only weeks, Y cannot be satisfied.
Criterion Y is not met because the study duration was on the order of weeks rather than most of an academic year.
-
B
Balanced Control Group
- Both arms received the same pre-class content and comparable in-class time, with the pedagogical modality itself being the treatment variable compared.
- "Both groups had the same educational material concerning the treatment of hypertension, consisting of audio and video presentations, along with additional material, which was provided online before the classes."
Relevant Quotes:
1) "Both groups had the same educational material concerning the treatment of hypertension, consisting of audio and video presentations, along with additional material, which was provided online before the classes." (Content, cases, and game development)
2) "Both classes were conducted online. The first fifteen minutes were dedicated to a review session and an explanation of the class process. Next, the participants were randomly divided into subgroups of three to four members, each led by a facilitator." (Intervention)
3) "For the GBL group, the same educational domains and learning objectives were incorporated into a board game format... Pedagogically, the CBL scenarios were designed to ensure that all these levels were covered in a systematic manner so that students from both groups underwent similar cognitive demands." (Content, cases, and game development)
4) "In the GBL group... After an hour of gameplay, any questions the participants had regarding the educational content were addressed by the clinical pharmacy professor." (Intervention)
Detailed Analysis:
Applying the criterion B decision tree: this is a head-to-head comparison of two active learning modalities rather than an intervention-versus-nothing design. Both arms received identical pre-class materials (the same audio/video and additional content), the same flipped-classroom structure, the same review session, and comparable in-class time (a single online session, with roughly an hour of core activity), and were explicitly designed so that "students from both groups underwent similar cognitive demands." No arm was given extra instructional time or budget over the other; the only systematic difference is the pedagogical format (case discussion vs. board game), which is precisely the treatment variable under investigation. Thus there is no unmatched extra time/budget imbalance. (The authors note a design difference - CBL was collaborative while GBL was individual - but this concerns social interaction, not time or budget balance.)
Criterion B is met because both arms received equivalent educational content and comparable time/resources, with the instructional modality itself being the variable compared.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study by a different team is reported or identified.
Relevant Quotes:
1) "There is a distinct lack of research evaluating how these two strategies compare when delivered in a completely virtual, online-only environment (28, 29)." (Background)
2) "We expand upon the current body of evidence by conducting a direct comparison of CBL and GBL within the context of flipped classroom instruction in pharmacy." (Discussion)
Detailed Analysis:
The 'R' criterion requires that this specific study be independently replicated by a different research team, in a different context, and published in a peer-reviewed journal. The authors present the study as filling a gap and providing a novel direct comparison, and cite related but distinct studies (e.g., Telner et al. 2010 on game-based vs. case-based stroke CME) that are not replications of this particular virtual pharmacy trial. An external check via the DOI landing page and the paper's own reference list found no independent reproduction of this specific study; as an in-press 2026 trial with n=56, no replication would yet be expected.
Criterion R is not met because there is no independent replication of this specific study.
-
A
All-subject Exams
- Only a single domain (hypertension pharmacotherapy) was assessed, and criterion E is not met, so A cannot be met.
- "Both groups had the same educational material concerning the treatment of hypertension..."
Relevant Quotes:
1) "Both groups had the same educational material concerning the treatment of hypertension, consisting of audio and video presentations..." (Content, cases, and game development)
2) "all questions and scenarios were designed to assess the same hypertension pharmacotherapy learning objectives." (Content, cases, and game development)
3) "The finalized test includes 26 multiple-choice questions." (Psychometric Properties)
Detailed Analysis:
The 'A' criterion requires measurement of impact across all main subjects using standardized exams, and per the specification it cannot be met if 'E' is not met. Here the outcome was confined to a single topic - hypertension pharmacotherapy - measured with a custom knowledge test. No other subjects were assessed, and the assessment instrument is not a standardized exam.
Criterion A is not met because only one narrow subject area was assessed and criterion E (standardized exam) is not satisfied.
-
G
Graduation Tracking
- Measurement stopped immediately after the in-class session with no follow-up to graduation, and criterion Y is not met.
- "Finally, the participants were asked to complete Test 3 and the Satisfaction Questionnaire."
Relevant Quotes:
1) "Finally, the participants were asked to complete Test 3 and the Satisfaction Questionnaire." (Intervention)
2) "In future studies, it is recommended that more participants are involved, in-person implementation is performed, and long-term knowledge retention is investigated." (Discussion)
Detailed Analysis:
The 'G' criterion requires tracking participants through graduation, and per the specification it cannot be met if 'Y' is not met. Outcome measurement ended immediately after the single in-class session (Test 3), and the authors explicitly note that long-term retention was not investigated and recommend it for future work. An external check for a follow-up/graduation-tracking paper by the same authors found none; the study measures no graduation outcome and Y is not met.
Criterion G is not met because there was no long-term or graduation follow-up and criterion Y is not met.
-
P
Pre-Registered
- Only an institutional research-ethics approval is reported; no pre-registered protocol on a trial registry is provided.
- "This study is based on a Pharm D thesis, approved with research ethics code IR.TUMS.TIPS.REC.1400.152 at TUMS."
Relevant Quotes:
1) "This study is based on a Pharm D thesis, approved with research ethics code IR.TUMS.TIPS.REC.1400.152 at TUMS." (Declaration: Ethics approval)
2) "The datasets used and/or analyzed during the current study are available from the corresponding author upon reasonable request." (Availability of data and material)
Detailed Analysis:
The 'P' criterion requires public pre-registration of the full study protocol (hypotheses, methods, planned analyses) before data collection, typically on a recognized trial registry, with a verifiable registration date. The paper reports only an institutional ethics-committee approval code (IR.TUMS.TIPS.REC.1400.152), which is an ethics approval, not a pre-registration of the analysis plan on a registry such as IRCT or ClinicalTrials.gov. No registry identifier or registration date is provided anywhere in the paper, so no registry search was possible.
Criterion P is not met because no pre-registered protocol on a trial registry is reported.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.