Level 1 Criteria
-
C Class-level RCT
- Randomisation was carried out on individual students rather than on intact classes or schools, and the tutoring exception does not apply.
- Fifty participants were randomly assigned to experimental or control groups using an online randomization tool (https://www.randomizer.org/) with a 1:1 allocation ratio.
- Relevant Quotes: 1) "A randomized controlled trial was used to examine the effects of the flipped classroom on creativity. Fifty participants were randomly assigned to experimental or control groups using an online randomization tool (https://www.randomizer.org/) with a 1:1 allocation ratio." (p. 2) 2) "An initial pool of 283 students enrolled at the School of Art and Design was considered... After applying these criteria, 50 participants were retained." (p. 2) 3) "Intervention group (n = 25) - Received allocated intervention (n = 25) ... Control group (n = 25) - Received allocated intervention (n = 25)" (Figure 1, CONSORT flow diagram, p. 3) 4) "Additionally, as a single instructor was involved, teaching style and interaction patterns may have influenced the results. Even with controlled conditions, some degree of informal student interaction between groups cannot be entirely excluded." (p. 7) Detailed Analysis: Criterion C requires randomisation at the class level or stronger (school level), with an exception only for interventions that are inherently personal, such as one-to-one tutoring. Here the unit of randomisation was explicitly the individual student: fifty individually recruited volunteers were allocated 1:1 to the flipped or traditional condition using an online randomiser. No intact classes or schools were assigned. The intervention is a classroom pedagogy (flipped classroom with pre-class materials, in-class brainstorming, collaborative design projects), not personal tutoring, so the tutoring exception does not apply. The authors themselves acknowledge the contamination risk that class-level randomisation is meant to prevent, noting that "some degree of informal student interaction between groups cannot be entirely excluded" and that a single instructor taught both conditions. Criterion C is not met because randomisation was performed at the individual student level within a single institution rather than at the class or school level, and no tutoring exception applies.
-
E Exam-based Assessment
- Outcomes were measured with psychometric creativity tasks (GAU, RAT) and a custom survey rather than with a recognised standardised exam.
- Creativity was assessed before and after the intervention using the Guilford Alternate Uses Test and the Remote Associates Test.
- Relevant Quotes: 1) "Creativity was assessed before and after the intervention using the Guilford Alternate Uses Test and the Remote Associates Test." (p. 1) 2) "Convergent thinking was measured using the remote associates test (RAT), first devised by Mednick (1962)." (p. 2) 3) "The Guilford alternate uses test (GAU) (Guilford, 1967) was used to evaluate divergent thinking." (p. 3) 4) "Participants were asked to consider two common objects, such as a toothbrush or a brick, and generate as many new and useful ways to use them as possible within 8 min (4 min for each object)... To avoid learning effects, the objects were changed to be a newspaper, a plastic bottle and an electric wire in the post-test." (p. 3) 5) "Responses were assessed by two trained independent raters who were blinded to group assignment. The raters were calibrated using a pilot dataset to ensure scoring consistency." (p. 3) 6) "Although the GAU and RAT are widely used and validated measures of divergent and convergent thinking, they primarily assess specific cognitive components rather than creativity in its entirety." (p. 7) 7) "The survey was aimed at measuring students' beliefs around their experience in the flipped classroom, with regard to SRL... The survey consisted of Likert scale items..." (p. 3) Detailed Analysis: Criterion E requires standardised, widely recognised exam-based assessment of educational outcomes - the kind of state-wide, national or curriculum-linked achievement exam that yields objective and comparable measures of school learning. The outcome instruments used here are the Guilford Alternate Uses Test and the Remote Associates Test. These are established psychometric laboratory tasks with published reliability evidence, but they are psychological measures of divergent and convergent thinking, not exam-based academic assessments of subject attainment. No course exam, national exam or standardised achievement test was administered. Moreover, the administration was locally adapted by the research team: the specific objects were chosen and changed between pre- and post-test, and scoring rules for fluency, flexibility, originality and elaboration were defined and applied by the study's own trained raters rather than by an external standardised scoring protocol. The third outcome, the self-reflection survey, is a custom Likert instrument. The authors also concede these measures "may not fully reflect real-world creative problem solving." Criterion E is not met because the outcomes were measured with psychometric creativity tasks and a custom survey, locally administered and scored, rather than with a recognised standardised exam.
-
T Term Duration
- The post-test was administered 16 weeks after the intervention began, which equals roughly one full academic semester.
- The flipped classroom intervention was implemented over a 16-week period.
- Relevant Quotes: 1) "This study employed a 16-week self-regulated flipped classroom intervention to examine creativity outcomes, focusing on divergent and convergent thinking." (p. 1) 2) "The flipped classroom intervention was implemented over a 16-week period. Students were required to complete preparatory materials before class, while in-class time was devoted to interactive and creative activities designed to promote engagement and higher-order thinking." (p. 3) 3) "Creative performance (divergent and convergent thinking) was assessed at both pre-test and post-test, while the SRL survey was administered at the post-test to capture students' perceptions of their learning experience following the intervention." (p. 3) 4) "Further screening criteria included: (1) participation in at least one art or design course during the semester..." (p. 2) 5) "As the course progressed, students engaged in more complex tasks requiring them to apply creative thinking and technical skills, culminating in final project presentations and reflective discussions." (p. 4) Detailed Analysis: Criterion T requires that the primary outcome be measured at least one full academic term (approximately 3-4 months) after the intervention begins. The paper states repeatedly that the intervention ran for 16 weeks and that the post-test was administered after the intervention. Sixteen weeks is approximately four months and corresponds to a full semester; the paper frames the study as running across a "semester" course. The interval from intervention start to primary outcome measurement is therefore at least 16 weeks, which satisfies the one-term minimum. The measurement occurred immediately at the end of the intervention rather than at a delayed follow-up, but the standard permits this because the interval from intervention start to measurement is what matters. Criterion T is met because the intervention began and the post-test was taken 16 weeks (roughly one full semester) later, meeting the one-term minimum.
-
D Documented Control Group
- The control group's size, attrition, baseline scores on all outcomes and business-as-usual condition are documented, though demographic detail is thin.
- The experimental group received the flipped instruction, while the control group was taught using traditional methods.
- Relevant Quotes: 1) "The experimental group received the flipped instruction, while the control group was taught using traditional methods." (p. 2) 2) "A randomized controlled trial was conducted with art and design students assigned to either an intervention group (n = 24) or a control group (n = 24)." (p. 1) 3) "Control group (n = 25) - Received allocated intervention (n = 25) - Did not receive allocated intervention (n = 0) ... Lost to follow-up (quit training) (n = 0) Discontinued intervention (n = 1) ... Analysed (n = 24)" (Figure 1, CONSORT flow diagram, p. 3) 4) "Fluency Pretest 3.864 +/- 0.089 [Intervention] 3.864 +/- 0.147 [Control] ... Originality Pretest 4.893 +/- 0.119 / 4.881 +/- 0.090 ... Flexibility Pretest 3.130 +/- 0.635 / 3.241 +/- 0.689 ... Elaboration Pretest 4.626 +/- 0.224 / 4.629 +/- 0.234 ... Convergent thinking Pretest 3.645 +/- 0.087 / 3.633 +/- 0.092" (Table 1, p. 5) 5) "An initial pool of 283 students enrolled at the School of Art and Design was considered. To ensure sufficient language proficiency, students who had not passed the College English Test Band 4 (CET-4) were excluded... Further screening criteria included: (1) participation in at least one art or design course during the semester, (2) no prior experience with flipped classroom instruction, and (3) willingness to participate in all course activities and assessments." (p. 2) 6) "In traditional lecture-based classrooms, learning can be sufficiently passive to provide fewer opportunities for creativity." (p. 4) Detailed Analysis: Criterion D requires the control group to be documented in terms of size, composition, baseline performance and the conditions it experienced. The paper reports the control group size at every stage of the CONSORT flow (25 allocated, 1 discontinued, 24 analysed), and Table 1 gives the control group's baseline (pre-test) means and standard deviations on all five outcome measures, so baseline comparability with the intervention arm can be directly assessed and is evidently close. The eligibility and screening criteria that define the population from which both arms were drawn are also stated (CET-4 pass, enrolled in an art or design course, no prior flipped classroom experience). The condition the control group received is stated: they "were taught using traditional methods", i.e. the lecture-based business-as-usual instruction described in the discussion, and no special treatment beyond normal teaching is indicated. The documentation is not exhaustive - no demographic breakdown (age, gender distribution, year of study) is given for either arm, and the traditional condition is not described session by session - but the size, baseline performance and treatment received by the control arm are all clearly quotable. Criterion D is met because the control group's size, attrition, baseline scores on every outcome and business-as-usual instructional condition are explicitly documented, despite the absence of a demographic table.