Level 1 Criteria
-
C Class-level RCT
- Eight intact classes were randomly allocated as clusters using a computer-generated, sealed allocation sequence, satisfying class-level randomisation.
- "Randomization was conducted at the class level. ... a research team member who was not involved in teaching or assessment generated the random allocation sequence using computer-generated random numbers, assigned the eight intact classes to four experimental clusters and four control clusters, and sealed the allocation results."
- Relevant Quotes: 1) "This study employed a cluster-randomized controlled trial design, with intact classes serving as clusters. A total of eight intact classes were included and randomly assigned to either the experimental group (four clusters) or the control group (four clusters)." (p. 2) 2) "Randomization was conducted at the class level. Before the study began, a research team member who was not involved in teaching or assessment generated the random allocation sequence using computer-generated random numbers, assigned the eight intact classes to four experimental clusters and four control clusters, and sealed the allocation results." (p. 3) 3) "They were drawn from eight intact classes, with approximately 24-28 students in each class." (p. 3) 4) "Because classes were used as the unit of randomization, observations from students within the same class could not be treated as fully independent; therefore, the sample size estimation was adjusted for class-level clustering." (p. 3) Detailed Analysis: Criterion C requires that randomisation be performed at the class level (or stronger), with the process clearly described, so as to prevent contamination between treatment and control students sharing the same classroom. The paper is explicit and unambiguous on this point. Intact classes were the unit of randomisation: eight whole classes were allocated, four to the autonomy-supportive teaching condition and four to conventional teaching. The allocation sequence was computer-generated by a team member not involved in teaching or assessment, and the results were sealed, which is a properly described allocation procedure. The sample size (216 analysed participants, 228 enrolled) and the cluster structure are reported, and the analysis accounted for clustering with class-level random intercepts and reported class-level ICCs. No individual student within a class was assigned to a different condition from classmates, so contamination risk of the type this criterion targets is avoided; the paper further notes that instructor training covered "requirements for preventing contamination between teaching conditions" (p. 4). Criterion C is met because randomisation was clearly described and performed at the level of entire intact classes rather than individual students within a class.
-
E Exam-based Assessment
- Outcomes were measured with a study-specific ski-technique rating rubric and self-report psychological scales rather than any widely recognised standardised exam.
- "Skill acquisition was evaluated using a pre-established standardized technical scoring rubric comprising five core dimensions: posture control (20%), center-of-mass transfer (25%), turning stability (25%), speed control (15%), and trajectory control (15%)."
- Relevant Quotes: 1) "Skill acquisition was evaluated using a pre-established standardized technical scoring rubric comprising five core dimensions: posture control (20%), center-of-mass transfer (25%), turning stability (25%), speed control (15%), and trajectory control (15%). Each dimension was rated on a 0-20 scale." (p. 4) 2) "Three raters with national ski instructor certification and university-level teaching experience independently scored all performances. The final score for each participant was the mean of the three ratings." (p. 4) 3) "Perceived teacher autonomy support was assessed using the 15-item scale developed by Tilga et al. (2017). ... Several items were minimally adapted to the skiing context by replacing references to 'physical education class' with 'ski class'..." (p. 4) 4) "Self-efficacy was measured using the General Self-Efficacy Scale (Schwarzer and Jerusalem, 1995)." (p. 4) 5) "Mental toughness was assessed using the Sports Mental Toughness Questionnaire (SMTQ; Sheard et al., 2009)." (p. 5) 6) "Inter-rater reliability for skill scores was high (ICC = 0.89, 95% CI [0.84, 0.93])." (p. 6) Detailed Analysis: Criterion E requires a standardised, exam-based assessment that is a widely recognised test rather than an instrument created for the study. The primary outcome, skill acquisition, was measured with a rater-scored technical rubric constructed by the research team for this trial. The paper calls it a "pre-established standardized technical scoring rubric," but "standardized" here refers only to consistency of application across raters within this study - it names no external test, cites no source, and reports no external validation, norms, or recognised body behind the rubric. The weighting scheme (20/25/25/15/15 percent across five ski-technique dimensions) is bespoke to this beginner ski course. The high inter-rater ICC (0.89) demonstrates reliable scoring, not that the instrument is a recognised standardised exam. The remaining outcomes are validated psychometric questionnaires (Tilga et al. autonomy support scale, the General Self-Efficacy Scale, the SMTQ). These are published, previously validated self-report scales, but they are psychological questionnaires, not exam-based assessments of academic attainment; the autonomy support scale was additionally reworded for the ski context and serves only as a manipulation check. Consequently no widely recognised standardised exam (e.g. a national curriculum test or state-wide achievement test) was used to measure the educational outcome. Criterion E is not met because the primary outcome was scored with a study-specific ski-technique rubric, and the other measures are self-report psychological questionnaires, not standardised exams.
-
T Term Duration
- The whole programme lasted three weeks with final measurement taken immediately at its end, far short of one academic term.
- "The instructional program lasted 3 weeks and comprised three 200-min sessions per week, for a total of nine sessions."
- Relevant Quotes: 1) "The instructional program lasted 3 weeks and comprised three 200-min sessions per week, for a total of nine sessions." (p. 2) 2) "The program consisted of a 1-week standardized teaching phase followed by a 2-week differential intervention phase." (p. 2) 3) "The differential intervention was implemented during Weeks 2 and 3. ... T0 was completed at the end of Week 1, after the standardized teaching phase and before the differential intervention began. T1 and T2 were completed after the instructional sessions in Weeks 2 and 3, respectively." (p. 3) 4) "Third, the intervention period was short, and the findings mainly reflect short-term changes during the beginner stage." (p. 12) 5) "Third, future work should extend follow-up periods and examine retention, transfer, and safety-related outcomes." (p. 13) Detailed Analysis: Criterion T requires that the primary outcome be measured at least one full academic term (roughly 3-4 months) after the intervention begins; short interventions are acceptable, but term-length follow-up tracking is not optional. Here the entire programme spanned three weeks, and the differential intervention that constitutes the treatment contrast ran for only two weeks (Weeks 2 and 3). The final measurement, T2, was taken immediately after the last instructional session of Week 3. The interval from the start of the differential intervention to the primary outcome measurement is therefore approximately two weeks, and even counting from the start of the standardised phase it is only three weeks. The authors themselves characterise the design as an "intensive short-term schedule" and concede in the limitations that "the intervention period was short" and that longer follow-up is required; no delayed or retention-phase assessment was conducted. Two to three weeks falls far short of one academic term. Criterion T is not met because outcomes were measured immediately at the end of a two-week differential intervention, roughly three weeks from programme start, with no term-length follow-up.
-
D Documented Control Group
- The control group's size, participant characteristics, baseline scores on all outcomes, and the conventional instruction it received are documented in detail and verified by fidelity observation.
- "The control group maintained the conventional 'demonstration-explanation-practice' instructional routine. ... Students in the control group completed the same practice tasks according to the instructor's arrangement."
- Relevant Quotes: 1) "The control group maintained the conventional 'demonstration-explanation-practice' instructional routine. Specifically, instructors followed the unified instructional progression by providing movement demonstrations, explaining technical points, organizing centralized practice, giving safety reminders, and offering corrective feedback. Students in the control group completed the same practice tasks according to the instructor's arrangement." (p. 4) 2) "A total of 228 participants were enrolled in the study (experimental group: n = 114; control group: n = 114). ... Twelve participants (experimental group: n = 4; control group: n = 8) did not meet the prespecified attendance requirements and were excluded from the analyses. The final analytic sample comprised 216 participants (experimental group: n = 110; control group: n = 106)." (p. 5) 3) "Table 1 presents descriptive statistics (means and standard deviations) for perceived teacher autonomy support, skill acquisition, self-efficacy, and mental toughness in the experimental and control groups at T0, T1, and T2. At T0, the two groups showed similar mean values across all four measures." (p. 5) 4) Table 1 control-group baseline (T0) values: Skill Acquisition 64.62 +/- 4.86; Self-Efficacy 1.91 +/- 0.19; Mental Toughness 2.05 +/- 0.16; Perceived Teacher Autonomy Support 1.98 +/- 0.25. (p. 6) 5) "Participants were first-year undergraduate students from Shenyang Sport University who were not majoring in skiing. ... Before the course began, none of the participants had received systematic ski training..." (p. 3) 6) Table 2 reports classroom observation fidelity scores for the control group on all seven dimensions, with a control-group total score of 2.25 +/- 1.04 against 12.25 +/- 1.16 in the experimental group. (p. 6) 7) "No significant group difference was detected at T0 (b = 0.41, p = 0.839)." (p. 6) Detailed Analysis: Criterion D requires the control group to be well documented: size, composition, baseline performance, and the conditions or treatment it received. The paper documents all of these. The control arm's size is given at enrolment (n = 114) and after exclusions (n = 106), with reasons for exclusion stated. Baseline (T0) means and standard deviations for the control group are reported in Table 1 for every outcome, and formal tests confirm no significant baseline group differences on skill acquisition, self-efficacy, mental toughness, or perceived autonomy support. Participant characteristics (first-year undergraduates at Shenyang Sport University, non-ski majors, no prior systematic ski training, health inclusion and exclusion criteria) apply to and describe the control arm. What the control group actually received is described in detail: the conventional "demonstration-explanation-practice" routine with unified progression, demonstrations, technical explanation, centralised practice, safety reminders, and corrective feedback, plus the same Week 1 standardised safety and foundational instruction. Table 2's fidelity observations quantitatively confirm that the control classes did not systematically implement the autonomy-supportive behaviours, verifying the absence of unintended contamination. Criterion D is met because the control group's size, demographic profile, baseline scores on every outcome, and the instruction it actually received are all explicitly documented and verified by fidelity observation.