Level 1 Criteria
-
C Class-level RCT
- Randomisation was at the individual student level within intact sections rather than at the class or school level, and the intervention is not personal tutoring.
- "Within each section, students were individually randomized (1:1) to either the STEAM-RB curriculum or the traditional Roliball syllabus, stratified by section to preserve existing social and timetable structures."
- Relevant Quotes: 1) "This single-site, section-stratified, individual-level randomized trial evaluated a STEAM-integrated Roliball (STEAM-RB) curriculum in university physical education." (p. 2, Section 2.1) 2) "Within each section, students were individually randomized (1:1) to either the STEAM-RB curriculum or the traditional Roliball syllabus, stratified by section to preserve existing social and timetable structures." (p. 3, Section 2.3) 3) "This produced two mixed-composition classes in each section (STEAM-RB and traditional), with instruction delivered in parallel." (p. 3, Section 2.3) 4) "Two intact first-year Roliball PE sections at a public university in central China were recruited at the beginning of the semester (N = 70)." (p. 2, Section 2.1) 5) "Time-on-task was matched across sections; contamination was minimized by tool-use restrictions in the traditional condition." (Figure 2 note, p. 5) Detailed Analysis: Criterion C requires that the unit of randomisation be the whole class (or a stronger unit such as the school), unless the intervention is personal tutoring or one-to-one teaching, in which case student-level randomisation is acceptable. The paper is explicit and repeated that randomisation was at the individual student level: it self-describes as an "individual-level randomized trial" and states that "students were individually randomized (1:1)". Section was used only as a stratification variable, not as the randomised unit. Students drawn from the same intact section were split between the two arms, and the resulting instructional groups were "mixed-composition classes" formed after randomisation from within the same sections. This is precisely the design pattern that criterion C is intended to guard against: students in the same social and timetable structure are allocated to different conditions, creating contamination risk. The authors themselves recognise this exposure, noting that "contamination was minimized by tool-use restrictions in the traditional condition" - an administrative mitigation, not a design-level isolation of the groups. The tutoring exception does not apply. The STEAM-RB curriculum is a group-delivered university PE syllabus taught in standard indoor PE facilities in two 90-min classes per week; it is not one-to-one or personal tutoring. Criterion C is not met because randomisation was performed at the individual student level within intact sections rather than at the class or school level, and no tutoring exception applies.
-
E Exam-based Assessment
- All outcomes used bespoke, curriculum-aligned instruments (custom STEAM test, study-specific rubric and kinematic proxy) rather than any recognised standardised exam.
- "Task-specific STEAM knowledge was assessed through a 15-item criterion-referenced test aligned with the core STEAM concepts embedded in the curriculum..."
- Relevant Quotes: 1) "Kinematic performance was indexed by peak racket-head angular velocity (deg/s) during a standardized Roliball forehand drive task at pre- and post-test." (p. 4, Section 2.5.1) 2) "Technical quality of the forehand drive was also assessed using blinded expert ratings from standardized video clips ... Two experienced Roliball instructors not involved in day-to-day teaching independently rated each clip on a 10-point holistic scale, using a brief rubric emphasizing swing path continuity, segment coordination, timing, and overall control." (p. 4, Section 2.5.2) 3) "Task-specific STEAM knowledge was assessed through a 15-item criterion-referenced test aligned with the core STEAM concepts embedded in the curriculum ... Items included multiple-choice, short-answer, and diagram-based questions, each scored on a 0-3 scale (0 = incorrect, 3 = fully correct), yielding total scores from 0 to 45." (p. 4, Section 2.5.3) 4) "Content validity was evaluated by an expert panel of eight PE and STEAM educators using content validity index (CVI) procedures. Item-level CVIs ranged from 0.88 to 1.00, and the scale-level CVI was 0.94 ... In the present sample, internal consistency was acceptable (Cronbach's alpha = 0.85)." (p. 4, Section 2.5.3) 5) "The STEAM knowledge test showed how curricular tasks can be mapped onto a focused assessment instrument with acceptable content validity, rather than relying on generic science or mathematics tests." (p. 8, Section 4.4) 6) "Finally, the STEAM knowledge test was tightly aligned to RB-specific content, so scores should not be interpreted as reflecting general STEAM competence..." (p. 9, Section 4.6) 7) "Motivation for physical activity in the PE context was assessed with the Chinese University Students' Physical Activity Motivation Scale (CUSPAMS), developed and validated under contemporary PE policy constraints." (p. 4, Section 2.6.1) Detailed Analysis: Criterion E requires outcomes to be measured with standardised, widely recognised exam-based assessments rather than instruments designed for the study. None of the outcome measures in this trial qualifies. The STEAM knowledge test is a bespoke 15-item criterion-referenced instrument built specifically for this curriculum; the author states plainly that it was designed "rather than relying on generic science or mathematics tests" and that it was "tightly aligned to RB-specific content". A locally constructed test whose content mirrors the intervention's own inquiry tasks is exactly the alignment bias criterion E exists to exclude, and the reported CVI and Cronbach's alpha figures speak to internal quality, not to external standardisation. The kinematic performance measure (peak racket-head angular velocity from smartphone video) is a study-specific instrumented performance proxy, not an exam. The blinded expert ratings use "a brief rubric" devised for the trial and rated on a 10-point holistic scale by two instructors; the excellent inter-rater ICC (0.94) establishes rater agreement but not exam standardisation. CUSPAMS is an externally developed and validated psychometric scale, but it is a self-report motivation questionnaire, not a standardised exam-based assessment of educational attainment. The same applies to the revised Creative Personality Scale. There is no national, state-wide, or otherwise widely recognised standardised exam anywhere in the outcome battery, and no university course grade or official examination result is reported. Criterion E is not met because all attainment outcomes relied on custom-built, curriculum-aligned instruments rather than recognised standardised exams.
-
T Term Duration
- Outcomes were measured at the end of an 8-week intervention with no later follow-up, which is shorter than one full academic term.
- "Both groups completed two 90-min Roliball classes per week for 8 weeks (16 sessions) in standard indoor PE facilities."
- Relevant Quotes: 1) "Participants were allocated to either an 8-week STEAM-RB intervention or a traditional syllabus." (p. 1, Abstract) 2) "Both groups completed two 90-min Roliball classes per week for 8 weeks (16 sessions) in standard indoor PE facilities." (p. 3, Section 2.4) 3) "The following outcomes were assessed at pretest and posttest: kinematic performance, blinded expert ratings, STEAM knowledge, physical activity motivation, and creative disposition." (p. 4, Section 2.5) 4) "Thus, there was no attrition over the 8-week intervention, and all randomized participants were included in the primary analyses." (p. 5, Section 3.1) 5) "The 8-week duration precluded examination of long-term maintenance and may have been too short to affect more stable constructs such as creative disposition." (p. 9, Section 4.6) 6) "Future research could extend this work through multi-site trials across different types of universities and regions, longer interventions spanning a semester or year..." (p. 9, Section 4.6) Detailed Analysis: Criterion T requires at least one full academic term (approximately 3-4 months) to elapse between the start of the intervention and the measurement of primary outcomes. Short interventions are permitted, but term-long follow-up tracking is not optional. Here the intervention ran for 8 weeks (16 sessions across weeks 1-8, as also depicted in Figure 2), and the posttest was administered at the end of that 8-week block. There is no delayed or maintenance follow-up: the paper reports only pretest and posttest, and explicitly acknowledges that "the 8-week duration precluded examination of long-term maintenance". The author's own future-directions statement contrasts this trial with "longer interventions spanning a semester or year", confirming that 8 weeks falls short of a semester in this institutional context. Eight weeks is roughly two months, well under the 3-4 month minimum, and since no measurement occurred after the intervention ended, the interval from intervention start to primary outcome measurement is itself only 8 weeks. Criterion T is not met because the interval from intervention start to outcome measurement was only 8 weeks, shorter than one full academic term, with no later follow-up.
-
D Documented Control Group
- The control group's size, baseline demographics and pretest scores (Table 1), and instructional conditions are documented in full, with fidelity checks confirming no STEAM elements leaked in.
- "The traditional curriculum followed the existing university syllabus, emphasizing teacher-directed demonstration, part-whole practice of core techniques (e.g., basic swings, directional control), and summative skill testing."
- Relevant Quotes: 1) "All 70 students who consented were randomized to the STEAM-RB (n = 35) or traditional Roliball group (n = 35) and completed both pretest and posttest assessments (Figure 2)." (p. 5, Section 3.1) 2) "The traditional curriculum followed the existing university syllabus, emphasizing teacher-directed demonstration, part-whole practice of core techniques (e.g., basic swings, directional control), and summative skill testing." (p. 3, Section 2.4) 3) "TABLE 1 Baseline characteristics with ASMD, t/chi-square, and p (N = 70)." (p. 6) 4) "The two groups were similar in age (STEAM-RB: 18.5 +/- 0.6 years; traditional: 18.4 +/- 0.7 years; ASMD = 0.15; t approx. 0.64, p = 0.523) and body mass index (21.6 +/- 2.8 vs. 21.2 +/- 2.5 kg/m2; ASMD = 0.15; p = 0.531)." (p. 5, Section 3.2) 5) "Groups were also balanced on prior sport exposure (5.8 +/- 2.1 vs. 5.5 +/- 2.3 years; ASMD = 0.13; p = 0.571), weekly physical activity (3.4 +/- 1.5 vs. 3.6 +/- 1.6 h; ASMD = 0.13; p = 0.591), and autonomous motivation at pretest (3.30 +/- 0.58 vs. 3.38 +/- 0.54; ASMD = 0.14; p = 0.550)." (p. 5, Section 3.2) 6) "Pretest STEAM knowledge scores were similarly low in both groups (15.2 +/- 4.8 vs. 14.7 +/- 5.2; ASMD = 0.10; p = 0.677). Sex distribution was comparable (21/35 [60.0%] vs. 19/35 [54.3%] male; ASMD = 0.11; chi-square(1) = 0.23, p = 0.629)." (p. 5, Section 3.2) 7) "In week 4, the first author observed one session per group using a structured observation form to verify that STEAM-specific elements were implemented only in the STEAM-RB group. No major protocol deviations were identified." (p. 4, Section 2.6) Detailed Analysis: Criterion D requires that the control group be well documented: size, demographic composition, baseline performance, and the conditions or treatment it received. All four elements are present. The control arm size is stated exactly (n = 35 of 70, with zero attrition). Table 1 provides a full column of baseline data for the traditional group covering age, sex, BMI, prior sport exposure, weekly physical activity, pretest peak racket-head angular velocity, pretest blinded expert rating, pretest autonomous motivation, and pretest STEAM knowledge, each with means, SDs, ASMDs, test statistics, and p values. The condition the control group received is described substantively rather than merely labelled "business as usual": it followed the existing university Roliball syllabus with teacher-directed demonstration, part-whole technique practice, and summative skill testing, delivered in the same facilities on the same two 90-min-per-week schedule. Fidelity monitoring further documents that adherence to the traditional syllabus was checked after each class and that STEAM-specific elements were verified as absent from the control condition, confirming no unintended interventions leaked in. Criterion D is met because the control group's size, baseline demographic and performance characteristics, and instructional conditions are all documented in detail and verified by fidelity checks.