Abstract
This pre-test post-test quasi-experimental study was grounded in a mixed method embedded design to delve into the quality and efficiency of flipped classroom model in enhancing university prep students' overall academic performance in EFL and that in its sub-skills in addition to the durability of that performance. The study has also pioneered to reveal the impact of gender on flipped classroom EFL learners' post-test scores. Quantitative data was gathered through the administration of EFL Achievement Test to 41 EFL students enrolled at Foreign Language School, Gebze Technical University in two different classrooms randomly assigned as experimental (N= 21) and control group (N=20). The intervention lasted during the whole 2016-2017 fall term. On the other hand, qualitative data was collected through follow-up semi-controlled interviews with 9 experiment group students from different achievement groups. All the quantitative data was analyzed using Statistical Packages for Social Sciences (SPSS) 21 for Windows and ITEMAN4 while qualitative data was analyzed manually by employing content analysis procedures. The results of the study revealed flipped classroom model as a significant facilitator of EFL performance and long-term retention of this performance at universities in Turkey. More specifically, students in flipped classroom significantly outperformed those in the traditional lecture based classroom in all skill areas except for listening. Furthermore, qualitative results supported this impact of flipped classroom model on EFL performance. As a unique aspect of the study, EFL students' performance in the flipped classroom was explored to be independent of their gender. To conclude, the present study has promised a bulk of valuable results that set flipping EFL classrooms as an efficient way of dealing with failure in EFL in Turkey.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- The paper self-identifies as a quasi-experimental study using convenience-sampled intact classes taught by the researcher himself, with no described randomisation procedure for assigning the two classes to conditions.
- "The researcher employed convenience sampling method to draw the sample of the study... Among the available classes, two ones where the researcher teaches himself were adopted..." (Section 2.3)
Relevant Quotes:
1) "This pre-test post-test quasi-experimental study was grounded in a mixed method embedded design..." (p. 120, Abstract)
2) "Quantitative data was gathered through the administration of EFL Achievement Test to 41 EFL students enrolled at Foreign Language School, Gebze Technical University in two different classrooms randomly assigned as experimental (N= 21) and control group (N=20)." (p. 120, Abstract)
3) "The researcher employed convenience sampling method to draw the sample of the study." (p. 127, Section 2.3 Participants)
4) "As a result of this technique and such constraints as time, fund and nature of the data collection instruments, the researcher included prep EFL students from two different classes at Gebze Technical University, whose populations range from 20 to 34. Among the available classes, two ones where the researcher teaches himself were adopted and they were exposed to pre-test administrations of EFL Achievement Test." (p. 127)
5) "To see if these classes are equated in terms of the variables within the interest of the study, the researcher computed Independent Samples T-Tests." (p. 127)
Detailed Analysis:
Criterion C requires a genuine RCT with a clearly described randomisation process, ideally at the class level. The abstract claims the classrooms were "randomly assigned," but the Methodology section (2.1 and 2.3) explicitly labels the design "quasi-experimental" and describes convenience sampling of two pre-existing intact classes that the researcher already taught. No quote anywhere in the paper describes an actual randomisation procedure (e.g., a lottery, random number generator, or coin flip) used to decide which of the two classes became the experimental group and which became control. The use of a post-hoc Independent Samples T-Test to check that the two classes were "equated" at baseline is a hallmark of a nonequivalent-groups quasi-experimental design, not of a randomised trial, since true randomisation would not require statistically verifying comparability after the fact in this way (though it is good practice regardless). Given the explicit self-classification as quasi-experimental and the convenience-based selection of pre-existing classes with no described random-assignment mechanism, this criterion is not met.
Criterion C is not met because the study is explicitly quasi-experimental, using convenience sampling of two pre-existing intact classes with no documented randomisation procedure for assigning experimental/control condition.
-
E
Exam-based Assessment
- The outcome measure was the EFL Achievement Test, an instrument developed by the researcher himself for this study, not a widely recognised standardised exam.
- "In order to elicit the participant EFL learners' scores of EFL performance, the EFL Achievement Test developed by the researcher was applied." (Section 2.4.1)
Relevant Quotes:
1) "In order to elicit the participant EFL learners' scores of EFL performance, the EFL Achievement Test developed by the researcher was applied. This test aimed to explore the progress the participants make during the English course in a prep EFL class at Gebze Technical University." (p. 128, Section 2.4.1)
2) "Despite the lack of precise steps to follow during the process of constructing an achievement test, to ensure the quality of the test to be developed, the researcher of the present thesis followed the stages in the form of the composition of Turgut, Baykul (2012) and Ivanova (2011)." (p. 128)
3) "...the researcher developed and ensured EFL Achievement Test as a valid and reliable assessor of EFL prep learners' EFL performance at Gebze University (KR20= 0,91)." (p. 129)
Detailed Analysis:
Criterion E requires a widely recognised, standardised exam rather than a test custom-built for the study. The paper is explicit that the EFL Achievement Test was designed, piloted, and validated entirely by the researcher specifically for this study population, following a test-construction methodology (item writing, piloting, item analysis via ITEMAN). While the authors took reasonable psychometric steps (content/face validity panels, KR20 reliability), this is fundamentally a bespoke, study-specific instrument rather than a nationally or internationally recognised standardised achievement test. Therefore this criterion is not met.
Criterion E is not met because the assessment was a custom instrument built by the researcher for this specific study, not a standardised exam.
-
T
Term Duration
- The intervention spanned the full 15-week fall term, and the primary outcome (post-test) was measured at the conclusion of that term, which constitutes at least one full academic term.
- "The instruction lasted 15 weeks from September to December." (Section 2.5)
Relevant Quotes:
1) "The intervention lasted during the whole 2016-2017 fall term." (p. 120, Abstract)
2) "Based on the permission from Institutional Review of Board (IBR), the data for the study was collected in during the fall term of 2016-2017 academic year. During this period, the same syllabus was followed in both classrooms. The instruction lasted 15 weeks from September to December." (p. 130, Section 2.5)
3) Tables 5 and 6 report both an immediate post-test comparison and a separate "durability" test administered later in the same term. (pp. 134-135)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Here, the intervention itself ran for 15 weeks (September to December), which is itself approximately one full term, and the post-test/durability measurements were taken at and after the conclusion of this 15-week period. The interval from intervention start to primary outcome measurement therefore meets or exceeds the one-term threshold required by the standard.
Criterion T is met because the 15-week intervention and its associated post-tests span a full academic term from the intervention's start date.
-
D
Documented Control Group
- The control group's size, gender composition, baseline scores, and instructional conditions (traditional lecture, same syllabus) are clearly documented with supporting tables.
- "The results indicated that the students did not differ significantly in terms of the mean pre-tests scores of concerned variables of the study (t = 0765; p > .05)." (Section 2.3, Table 1)
Relevant Quotes:
1) "Table 1: Independent Samples T-Tests for Equating Groups... EFL Achievement Pre-test Control 20 50,65 17,03... Exper. 21 46,52 17,50..." (p. 127, Table 1)
2) "The results indicated that the students did not differ significantly in terms of the mean pre-tests scores of concerned variables of the study (t = 0765; p > .05). In other words, it was explored that two groups were assumed to be similar based on their pretest results." (p. 127)
3) "Table 2: Descriptive Statistics for the Participants... Experiment Male 12 29,3%... Female 9 21,9%... Control Male 11 26,9%... Female 9 21,9%..." (pp. 127-128, Table 2)
4) "During this period, the same syllabus was followed in both classrooms... he adopted traditional lecture-based instruction in the control group." (p. 130, Section 2.5)
Detailed Analysis:
Criterion D requires clear documentation of the control group's demographics, baseline performance, and instructional conditions. The paper provides a dedicated control-group sample size (N=20), a gender breakdown (Table 2: 11 male, 9 female in the control group), a baseline pre-test comparison confirming equivalence with the experimental group (Table 1), and an explicit description of what the control group received (traditional lecture-based instruction under the same syllabus and timeframe as the experimental group). This level of detail satisfies the documentation requirement.
Criterion D is met because the control group's size, demographics, baseline scores, and instructional condition are all clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation (such as it was) occurred at the level of two intact classrooms within a single university, not across multiple schools.
- "...the researcher included prep EFL students from two different classes at Gebze Technical University..." (Section 2.3)
Relevant Quotes:
1) "...the researcher included prep EFL students from two different classes at Gebze Technical University, whose populations range from 20 to 34." (p. 127, Section 2.3)
2) "Among the available classes, two ones where the researcher teaches himself were adopted..." (p. 127)
Detailed Analysis:
Criterion S requires assignment at the level of whole schools/institutions, not individual classes within a single institution. This study involves exactly two classes at a single university (Gebze Technical University), both taught by the same instructor. There is no school-level comparison or multi-site design of any kind.
Criterion S is not met because the study involved only two classrooms at a single university, not multiple schools.
-
I
Independent Conduct
- The same researcher designed the intervention and the assessment instrument, personally taught both the experimental and control classes, and conducted the interviews and analysis, with no independent or external party involved.
- "...the researcher followed Flipped Classroom Model in the experimental class while he adopted traditional lecture-based instruction in the control group." (Section 2.5)
Relevant Quotes:
1) "In order to elicit the participant EFL learners' scores of EFL performance, the EFL Achievement Test developed by the researcher was applied." (p. 128)
2) "...the researcher followed Flipped Classroom Model in the experimental class while he adopted traditional lecture-based instruction in the control group." (p. 130, Section 2.5)
3) "Taking the commonly cited advantages of the flipped classroom model into account, the researcher, also as the teacher in the class, aimed to conduct a study informing the concerned bodies about a better EFL practice." (p. 129, Section 2.4.2)
4) "...the researcher interviewed with nine experiment group students developing differently from pre- to post- academic achievement test." (p. 130)
Detailed Analysis:
Criterion I requires that the evaluation be conducted independently of the person(s) who designed and delivered the intervention. Here, the same individual (the paper's first author, an instructor at the institution) designed the flipped-classroom intervention, personally taught both the experimental and control sections, developed and validated the outcome test, conducted the qualitative interviews, and performed the data analysis. No external evaluator, blinded administrator, or third-party agency is mentioned anywhere in the paper.
Criterion I is not met because the researcher who designed and delivered the intervention also taught both groups, collected data, and analysed the results without any independent oversight.
-
Y
Year Duration
- The intervention and outcome tracking lasted only one 15-week academic term, far short of the 75% of an academic year required by this criterion.
- "The instruction lasted 15 weeks from September to December." (Section 2.5)
Relevant Quotes:
1) "The intervention lasted during the whole 2016-2017 fall term." (p. 120, Abstract)
2) "The instruction lasted 15 weeks from September to December." (p. 130, Section 2.5)
Detailed Analysis:
Criterion Y requires the study to track outcomes over at least 75% of a full academic year (roughly 9-10 months). Here, the entire intervention and its associated durability follow-up occurred within a single 15-week fall term (roughly 3.5 months), which is well under the required threshold. There is no indication that tracking extended into the spring term or beyond.
Criterion Y is not met because the study's duration (15 weeks, one fall term) falls far short of 75% of an academic year.
-
B
Balanced Control Group
- Both groups followed the same syllabus over the same 15-week period with no described extra time or budget given to the experimental group beyond the change in instructional method.
- "During this period, the same syllabus was followed in both classrooms." (Section 2.5)
Relevant Quotes:
1) "During this period, the same syllabus was followed in both classrooms. The instruction lasted 15 weeks from September to December." (p. 130, Section 2.5)
2) "...the researcher followed Flipped Classroom Model in the experimental class while he adopted traditional lecture-based instruction in the control group." (p. 130)
3) Qualitative student quotes describe the flipped-classroom group's out-of-class video viewing and quizzes (e.g., "Thanks to the presentations I watched and quizzes I took, I always went to school having studied the subject"), but no comparable description of the control group's out-of-class workload is given. (pp. 137-139)
Detailed Analysis:
Applying the decision procedure for Criterion B: the paper explicitly states both classes followed the "same syllabus" over the identical 15-week period, and there is no quote indicating the experimental group received additional class contact hours, budget, or materials beyond substituting take-home videos/quizzes for traditional homework within the same overall course structure. The flipped model reallocates when content delivery versus practice occurs rather than adding net new instructional time or resources (EXTRA_RESOURCES_PRESENT is not clearly established, and any home-viewing time appears to substitute for, rather than add to, ordinary homework). Since the paper does not document any extra time or budget given to the experimental group relative to the control group, the "no extra resources present" branch of the decision tree applies and the criterion is met by default. Note, however, that the paper does not explicitly quantify or compare out-of-class time spent by each group, which introduces some uncertainty into this determination.
Criterion B is met because both groups followed the same syllabus over the same period, and no quotes indicate the experimental group received additional educational time or budget beyond the change in teaching method.
-
Level 3 Criteria
-
R
Reproduced
- No evidence of an independent replication of this specific study by a different research team is present in the paper or discoverable externally.
Relevant Quotes:
1) The Discussion section cites related but distinct studies on flipped classrooms in EFL (e.g., Kang, 2015; Al-Harbi, Alshumaimeri, 2016; Hung, 2015; Boyraz, 2014; Ekmekci, 2014), but these are separate, independently designed studies with different populations and instruments, not replications of this specific study's design and test. (pp. 140-143)
2) "...the present study produced results supported by insufficient number of studies conducted in international context..." (p. 141)
Detailed Analysis:
Criterion R requires that this specific study be independently replicated by a different research team in a peer-reviewed outlet. The paper situates itself among a small body of related flipped classroom/EFL research, but none of the cited works are replications of this particular study (same intervention, same EFL Achievement Test, same context); they are independent studies on the broader topic. An internet search (Semantic Scholar, Google Scholar, DuckDuckGo, ERIC, and citing records such as a Dergipark/RumeliDE literature-review article that cites this paper in its bibliography) found only citations of this paper within broader literature reviews on blended/flipped EFL instruction, not an independent replication attempt of this specific study's design (same EFL Achievement Test, same Gebze Technical University context) by a different research team. No such replication study could be found.
Criterion R is not met because no independent replication of this specific study was found either in the paper or through internet search.
-
A
All-subject Exams
- Only EFL performance and its sub-skills were assessed; no other core subjects were measured, and criterion E (a prerequisite for A) is not met.
- "...administration of EFL Achievement Test to 41 EFL students..." (Abstract)
Relevant Quotes:
1) "Quantitative data was gathered through the administration of EFL Achievement Test to 41 EFL students enrolled at Foreign Language School, Gebze Technical University..." (p. 120, Abstract)
2) Tables 3-6 report only EFL sub-skills (Listening, Grammar, Reading, Vocabulary, Writing) and Total EFL score; no other subjects (math, science, etc.) are assessed. (pp. 132-135)
Detailed Analysis:
Per the criterion-specific instructions, if Criterion E is not met, Criterion A cannot be met either. Since the EFL Achievement Test used here is a non-standardised, researcher-made instrument (Criterion E: not met), Criterion A automatically fails. Independently, the study also only measures one subject (EFL and its sub-skills), with no assessment of other core subjects, which would fail the "all-subject" requirement on its own merits as well.
Criterion A is not met because Criterion E is not met, and additionally only a single subject (EFL) was assessed.
-
G
Graduation Tracking
- Tracking stopped at a within-term durability test; there was no follow-up until graduation, and criterion Y (a prerequisite for G) is not met.
Relevant Quotes:
1) "The intervention lasted during the whole 2016-2017 fall term." (p. 120, Abstract)
2) Table 6 ("Independent Samples T-Tests for Durability of EFL Achievement") reports a delayed post-test still administered within the same study/term, with no mention of any further follow-up. (p. 135)
3) No statement anywhere in the paper describes tracking students beyond the 2016-2017 fall term or through to graduation.
Detailed Analysis:
Per the criterion-specific instructions, if Criterion Y is not met, Criterion G cannot be met either; since Y (Year Duration) was not met here (the entire study spanned one 15-week term), Criterion G automatically fails. Independently, the paper also provides no evidence of any follow-up beyond a "durability" test taken shortly after the intervention within the same term, let alone tracking through to the students' graduation. An internet search for subsequent publications by the same authors (Orhan Iyitoğlu, Yavuz Erişen) tracking this same cohort found only a separate, later study by the same first author on flipped-classroom EFL outcomes ("The Impact of Flipped Classroom Model on EFL Learners' Academic Achievement, Attitudes and Self-Efficacy Beliefs", 2018); this could not be confirmed to track the same 41 students, and it does not describe graduation tracking, so it does not establish that this criterion is met.
Criterion G is not met because Criterion Y is not met, and neither the paper nor a follow-up search found tracking beyond a single-term durability test.
-
P
Pre-Registered
- The paper mentions only Institutional Review Board approval, with no reference to a public pre-registration of the study protocol before data collection began.
- "Based on the permission from Institutional Review of Board (IBR), the data for the study was collected..." (Section 2.5)
Relevant Quotes:
1) "Based on the permission from Institutional Review of Board (IBR), the data for the study was collected in during the fall term of 2016-2017 academic year." (p. 130, Section 2.5)
2) No other statement in the paper mentions a registry platform (e.g., ClinicalTrials.gov, ISRCTN, OSF Registries) or a pre-registration date.
Detailed Analysis:
Criterion P requires a publicly registered protocol (hypotheses, methods, planned analyses) filed before data collection began. The paper only mentions ethics/IRB approval, which is a distinct requirement from public pre-registration of the study design and analysis plan. No registry reference or pre-registration date is present anywhere in the text. An internet search for a pre-registration record of this study (e.g., on OSF Registries or similar platforms) found no matching entry.
Criterion P is not met because no pre-registration of the study protocol is mentioned anywhere in the paper, and none could be found through internet search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.