Effects of autonomy-supportive ski instruction on skill acquisition, mental toughness, and self-efficacy in novice skiers: a cluster-randomized controlled trial

Huanyu Gao, Dawei Cao, Yangyang Ren, Hongbo Shi, Yuxiang Liang, Yuanguo Liu

Published:
ERCT Check Date:
DOI: 10.3389/fpsyg.2026.1828839
  • higher education
  • China
0
  • C

    Eight intact classes were randomly allocated as clusters using a computer-generated, sealed allocation sequence, satisfying class-level randomisation.

    "Randomization was conducted at the class level. ... a research team member who was not involved in teaching or assessment generated the random allocation sequence using computer-generated random numbers, assigned the eight intact classes to four experimental clusters and four control clusters, and sealed the allocation results."

  • E

    Outcomes were measured with a study-specific ski-technique rating rubric and self-report psychological scales rather than any widely recognised standardised exam.

    "Skill acquisition was evaluated using a pre-established standardized technical scoring rubric comprising five core dimensions: posture control (20%), center-of-mass transfer (25%), turning stability (25%), speed control (15%), and trajectory control (15%)."

  • T

    The whole programme lasted three weeks with final measurement taken immediately at its end, far short of one academic term.

    "The instructional program lasted 3 weeks and comprised three 200-min sessions per week, for a total of nine sessions."

  • D

    The control group's size, participant characteristics, baseline scores on all outcomes, and the conventional instruction it received are documented in detail and verified by fidelity observation.

    "The control group maintained the conventional 'demonstration-explanation-practice' instructional routine. ... Students in the control group completed the same practice tasks according to the instructor's arrangement."

  • S

    Randomisation was among eight intact classes within one single university, not among schools or institutions.

    "Randomization was conducted at the class level."

  • I

    The same research team adapted the protocol, trained the instructors, monitored fidelity, and analysed the data, with no external or third-party evaluator involved.

    "The autonomy-supportive teaching protocol used in the experimental group was developed based on the seven-category framework of autonomy-supportive teacher behaviors proposed by Cheon et al. (2023), with contextual adaptation for skiing instruction."

  • Y

    The study tracked outcomes for only about three weeks, nowhere near 75% of an academic year, and criterion T was also not met.

    "Although the differential intervention lasted only 2 weeks, students received three 200-min sessions per week, resulting in a concentrated learning schedule."

  • B

    Time, technical content, practice volume, and instructor qualifications were explicitly matched across arms so that interaction style was the only substantive difference, with the brief autonomy-support instructor training being integral to the intervention tested.

    "During the differential intervention phase, the two groups were matched on class duration, technical content, instructional progression, core practice tasks, and total practice volume. The only between-group difference concerned instructional interaction style and the organization of practice activities."

  • R

    Internet searching found no independent replication of this specific beginner-ski cluster RCT; the authors themselves list replication as future work.

    "First, future studies should increase the number of clusters and conduct replication across multiple contexts."

  • A

    Only ski technique was assessed, using a non-standardised study-specific rubric, and criterion E was not met, so the all-subject requirement fails.

    "This study aimed to compare the effects of autonomy-supportive teaching and conventional teaching on skill acquisition, self-efficacy, and mental toughness in beginner skiing instruction."

  • G

    Measurement stopped immediately after the three-week programme with no tracking towards graduation, no follow-up publication was found, and criterion Y was also not met.

    "Third, future work should extend follow-up periods and examine retention, transfer, and safety-related outcomes."

  • P

    Neither the paper nor any searched trial registry records a pre-registration for this study; only institutional ethics approval and CONSORT reporting are reported.

Abstract

Background: Instructional interactions in beginner skiing classes may influence both objective skill performance and psychological outcomes. However, longitudinal intervention evidence in high-risk skill-learning contexts remains limited, and objective skill performance and psychological outcomes have often been examined separately. Objective: This study aimed to compare the effects of autonomy-supportive teaching and conventional teaching on skill acquisition, self-efficacy, and mental toughness in beginner skiing instruction. Methods: A cluster-randomized controlled trial was conducted with 216 novice skiers from eight intact classes. The instructional program lasted 3 weeks and followed an intensive short-term schedule, with three 200-min sessions per week. Skill assessments and questionnaires were administered at three time points: after the standardized teaching phase in Week 1, after the Week 2 instructional sessions, and after the Week 3 instructional sessions. Linear mixed-effects models were used to examine Group x Time interaction effects. Results: Learners receiving autonomy-supportive teaching reported significantly greater increases in perceived teacher autonomy support than those receiving conventional teaching. Skill acquisition, self-efficacy, and mental toughness all showed significant Group x Time interaction effects, with the autonomy-supportive teaching group demonstrating greater incremental improvements over time. Conclusion: Autonomy-supportive teaching can simultaneously improve skill acquisition, self-efficacy, and mental toughness among novice skiers without increasing instructional time or practice volume. This study provides longitudinal empirical evidence for evaluating pedagogical intervention effects in high-risk skill-learning contexts.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Eight intact classes were randomly allocated as clusters using a computer-generated, sealed allocation sequence, satisfying class-level randomisation.
      • "Randomization was conducted at the class level. ... a research team member who was not involved in teaching or assessment generated the random allocation sequence using computer-generated random numbers, assigned the eight intact classes to four experimental clusters and four control clusters, and sealed the allocation results."
      • Relevant Quotes: 1) "This study employed a cluster-randomized controlled trial design, with intact classes serving as clusters. A total of eight intact classes were included and randomly assigned to either the experimental group (four clusters) or the control group (four clusters)." (p. 2) 2) "Randomization was conducted at the class level. Before the study began, a research team member who was not involved in teaching or assessment generated the random allocation sequence using computer-generated random numbers, assigned the eight intact classes to four experimental clusters and four control clusters, and sealed the allocation results." (p. 3) 3) "They were drawn from eight intact classes, with approximately 24-28 students in each class." (p. 3) 4) "Because classes were used as the unit of randomization, observations from students within the same class could not be treated as fully independent; therefore, the sample size estimation was adjusted for class-level clustering." (p. 3) Detailed Analysis: Criterion C requires that randomisation be performed at the class level (or stronger), with the process clearly described, so as to prevent contamination between treatment and control students sharing the same classroom. The paper is explicit and unambiguous on this point. Intact classes were the unit of randomisation: eight whole classes were allocated, four to the autonomy-supportive teaching condition and four to conventional teaching. The allocation sequence was computer-generated by a team member not involved in teaching or assessment, and the results were sealed, which is a properly described allocation procedure. The sample size (216 analysed participants, 228 enrolled) and the cluster structure are reported, and the analysis accounted for clustering with class-level random intercepts and reported class-level ICCs. No individual student within a class was assigned to a different condition from classmates, so contamination risk of the type this criterion targets is avoided; the paper further notes that instructor training covered "requirements for preventing contamination between teaching conditions" (p. 4). Criterion C is met because randomisation was clearly described and performed at the level of entire intact classes rather than individual students within a class.
    • E

      Exam-based Assessment

      • Outcomes were measured with a study-specific ski-technique rating rubric and self-report psychological scales rather than any widely recognised standardised exam.
      • "Skill acquisition was evaluated using a pre-established standardized technical scoring rubric comprising five core dimensions: posture control (20%), center-of-mass transfer (25%), turning stability (25%), speed control (15%), and trajectory control (15%)."
      • Relevant Quotes: 1) "Skill acquisition was evaluated using a pre-established standardized technical scoring rubric comprising five core dimensions: posture control (20%), center-of-mass transfer (25%), turning stability (25%), speed control (15%), and trajectory control (15%). Each dimension was rated on a 0-20 scale." (p. 4) 2) "Three raters with national ski instructor certification and university-level teaching experience independently scored all performances. The final score for each participant was the mean of the three ratings." (p. 4) 3) "Perceived teacher autonomy support was assessed using the 15-item scale developed by Tilga et al. (2017). ... Several items were minimally adapted to the skiing context by replacing references to 'physical education class' with 'ski class'..." (p. 4) 4) "Self-efficacy was measured using the General Self-Efficacy Scale (Schwarzer and Jerusalem, 1995)." (p. 4) 5) "Mental toughness was assessed using the Sports Mental Toughness Questionnaire (SMTQ; Sheard et al., 2009)." (p. 5) 6) "Inter-rater reliability for skill scores was high (ICC = 0.89, 95% CI [0.84, 0.93])." (p. 6) Detailed Analysis: Criterion E requires a standardised, exam-based assessment that is a widely recognised test rather than an instrument created for the study. The primary outcome, skill acquisition, was measured with a rater-scored technical rubric constructed by the research team for this trial. The paper calls it a "pre-established standardized technical scoring rubric," but "standardized" here refers only to consistency of application across raters within this study - it names no external test, cites no source, and reports no external validation, norms, or recognised body behind the rubric. The weighting scheme (20/25/25/15/15 percent across five ski-technique dimensions) is bespoke to this beginner ski course. The high inter-rater ICC (0.89) demonstrates reliable scoring, not that the instrument is a recognised standardised exam. The remaining outcomes are validated psychometric questionnaires (Tilga et al. autonomy support scale, the General Self-Efficacy Scale, the SMTQ). These are published, previously validated self-report scales, but they are psychological questionnaires, not exam-based assessments of academic attainment; the autonomy support scale was additionally reworded for the ski context and serves only as a manipulation check. Consequently no widely recognised standardised exam (e.g. a national curriculum test or state-wide achievement test) was used to measure the educational outcome. Criterion E is not met because the primary outcome was scored with a study-specific ski-technique rubric, and the other measures are self-report psychological questionnaires, not standardised exams.
    • T

      Term Duration

      • The whole programme lasted three weeks with final measurement taken immediately at its end, far short of one academic term.
      • "The instructional program lasted 3 weeks and comprised three 200-min sessions per week, for a total of nine sessions."
      • Relevant Quotes: 1) "The instructional program lasted 3 weeks and comprised three 200-min sessions per week, for a total of nine sessions." (p. 2) 2) "The program consisted of a 1-week standardized teaching phase followed by a 2-week differential intervention phase." (p. 2) 3) "The differential intervention was implemented during Weeks 2 and 3. ... T0 was completed at the end of Week 1, after the standardized teaching phase and before the differential intervention began. T1 and T2 were completed after the instructional sessions in Weeks 2 and 3, respectively." (p. 3) 4) "Third, the intervention period was short, and the findings mainly reflect short-term changes during the beginner stage." (p. 12) 5) "Third, future work should extend follow-up periods and examine retention, transfer, and safety-related outcomes." (p. 13) Detailed Analysis: Criterion T requires that the primary outcome be measured at least one full academic term (roughly 3-4 months) after the intervention begins; short interventions are acceptable, but term-length follow-up tracking is not optional. Here the entire programme spanned three weeks, and the differential intervention that constitutes the treatment contrast ran for only two weeks (Weeks 2 and 3). The final measurement, T2, was taken immediately after the last instructional session of Week 3. The interval from the start of the differential intervention to the primary outcome measurement is therefore approximately two weeks, and even counting from the start of the standardised phase it is only three weeks. The authors themselves characterise the design as an "intensive short-term schedule" and concede in the limitations that "the intervention period was short" and that longer follow-up is required; no delayed or retention-phase assessment was conducted. Two to three weeks falls far short of one academic term. Criterion T is not met because outcomes were measured immediately at the end of a two-week differential intervention, roughly three weeks from programme start, with no term-length follow-up.
    • D

      Documented Control Group

      • The control group's size, participant characteristics, baseline scores on all outcomes, and the conventional instruction it received are documented in detail and verified by fidelity observation.
      • "The control group maintained the conventional 'demonstration-explanation-practice' instructional routine. ... Students in the control group completed the same practice tasks according to the instructor's arrangement."
      • Relevant Quotes: 1) "The control group maintained the conventional 'demonstration-explanation-practice' instructional routine. Specifically, instructors followed the unified instructional progression by providing movement demonstrations, explaining technical points, organizing centralized practice, giving safety reminders, and offering corrective feedback. Students in the control group completed the same practice tasks according to the instructor's arrangement." (p. 4) 2) "A total of 228 participants were enrolled in the study (experimental group: n = 114; control group: n = 114). ... Twelve participants (experimental group: n = 4; control group: n = 8) did not meet the prespecified attendance requirements and were excluded from the analyses. The final analytic sample comprised 216 participants (experimental group: n = 110; control group: n = 106)." (p. 5) 3) "Table 1 presents descriptive statistics (means and standard deviations) for perceived teacher autonomy support, skill acquisition, self-efficacy, and mental toughness in the experimental and control groups at T0, T1, and T2. At T0, the two groups showed similar mean values across all four measures." (p. 5) 4) Table 1 control-group baseline (T0) values: Skill Acquisition 64.62 +/- 4.86; Self-Efficacy 1.91 +/- 0.19; Mental Toughness 2.05 +/- 0.16; Perceived Teacher Autonomy Support 1.98 +/- 0.25. (p. 6) 5) "Participants were first-year undergraduate students from Shenyang Sport University who were not majoring in skiing. ... Before the course began, none of the participants had received systematic ski training..." (p. 3) 6) Table 2 reports classroom observation fidelity scores for the control group on all seven dimensions, with a control-group total score of 2.25 +/- 1.04 against 12.25 +/- 1.16 in the experimental group. (p. 6) 7) "No significant group difference was detected at T0 (b = 0.41, p = 0.839)." (p. 6) Detailed Analysis: Criterion D requires the control group to be well documented: size, composition, baseline performance, and the conditions or treatment it received. The paper documents all of these. The control arm's size is given at enrolment (n = 114) and after exclusions (n = 106), with reasons for exclusion stated. Baseline (T0) means and standard deviations for the control group are reported in Table 1 for every outcome, and formal tests confirm no significant baseline group differences on skill acquisition, self-efficacy, mental toughness, or perceived autonomy support. Participant characteristics (first-year undergraduates at Shenyang Sport University, non-ski majors, no prior systematic ski training, health inclusion and exclusion criteria) apply to and describe the control arm. What the control group actually received is described in detail: the conventional "demonstration-explanation-practice" routine with unified progression, demonstrations, technical explanation, centralised practice, safety reminders, and corrective feedback, plus the same Week 1 standardised safety and foundational instruction. Table 2's fidelity observations quantitatively confirm that the control classes did not systematically implement the autonomy-supportive behaviours, verifying the absence of unintended contamination. Criterion D is met because the control group's size, demographic profile, baseline scores on every outcome, and the instruction it actually received are all explicitly documented and verified by fidelity observation.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was among eight intact classes within one single university, not among schools or institutions.
      • "Randomization was conducted at the class level."
      • Relevant Quotes: 1) "Randomization was conducted at the class level." (p. 3) 2) "This study employed a cluster-randomized controlled trial design, with intact classes serving as clusters. A total of eight intact classes were included and randomly assigned to either the experimental group (four clusters) or the control group (four clusters)." (p. 2) 3) "Participants were first-year undergraduate students from Shenyang Sport University who were not majoring in skiing. They were drawn from eight intact classes..." (p. 3) 4) "To minimize the influence of instructor variability on between-group comparisons, the intervention was delivered by four instructors ... A paired allocation strategy was used such that each instructor taught one experimental class and one control class." (p. 4) 5) "First, future studies should increase the number of clusters and conduct replication across multiple contexts. Larger cluster-randomized trials conducted across different schools, ski-field conditions, and instructional teams would improve the precision of class-level inference..." (p. 13) Detailed Analysis: Criterion S requires randomisation among schools, i.e. among the educational institutions or delivery units implementing the intervention, rather than among classes within a single institution. All participants in this trial came from one institution, Shenyang Sport University, and the randomised units were eight intact classes within that single university. There was no second institution, campus, or ski-site to randomise; the same four instructors crossed both conditions, each teaching one experimental and one control class, which confirms that the delivery unit was the class and not an institution. The authors explicitly signal this gap by proposing that future work should conduct "larger cluster-randomized trials ... across different schools, ski-field conditions, and instructional teams," acknowledging that the present design did not randomise at that level. Criterion S is not met because randomisation occurred among eight classes inside a single university rather than among schools or institutional delivery units.
    • I

      Independent Conduct

      • The same research team adapted the protocol, trained the instructors, monitored fidelity, and analysed the data, with no external or third-party evaluator involved.
      • "The autonomy-supportive teaching protocol used in the experimental group was developed based on the seven-category framework of autonomy-supportive teacher behaviors proposed by Cheon et al. (2023), with contextual adaptation for skiing instruction."
      • Relevant Quotes: 1) "The autonomy-supportive teaching protocol used in the experimental group was developed based on the seven-category framework of autonomy-supportive teacher behaviors proposed by Cheon et al. (2023), with contextual adaptation for skiing instruction." (p. 3) 2) "After completion of the T0 assessments, the research team disclosed group allocation to the instructors and provided standardized training before the start of Week 2. ... Formal intervention began only after the research team confirmed adherence to the assigned teaching protocol." (p. 4) 3) "To assess intervention fidelity, the research team conducted classroom observation records during the differential intervention phase. ... The observation records were completed by one researcher who was not involved in teaching..." (p. 4) 4) "To reduce rating bias, skill performance was video-recorded and anonymized using coded identifiers, and three raters independently scored all performances." (p. 3) 5) "Personnel responsible for data management and statistical analysis had no access to group information until data cleaning and model specification had been completed, thereby ensuring blinding at the statistical analysis stage." (p. 3) 6) "HG: Writing - original draft, Writing - review & editing. DC: Writing - review & editing. YR: Writing - original draft. HS: Writing - original draft. YuxL: Writing - original draft. YuaL: Writing - review & editing." (p. 13) 7) "The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest." (p. 13) Detailed Analysis: Criterion I requires that the evaluation be conducted independently of the people who designed the intervention, and looks for an explicit statement of third-party evaluation or oversight. Here the same research team performed every function. The team adapted the autonomy-supportive protocol for skiing, trained the instructors, confirmed protocol adherence before formal intervention, conducted the fidelity observations, collected the data, and analysed the results. All authors are from the sponsoring sport universities; there is no external evaluation agency, government body, or independent evaluator named anywhere in the paper. The study does include several internal bias-reduction safeguards: skill videos were anonymised and scored by three certified raters, questionnaires were anonymously coded, and analysts were kept blind to group until model specification was complete. These are genuine methodological strengths, but they are internal procedures administered by the same team; they are not third-party conduct or oversight. The single fidelity observer was a researcher on the team, merely "not involved in teaching," and the authors themselves list single-observer fidelity as a limitation. The conflict-of-interest and author-contribution statements likewise contain no assertion of evaluator independence from the intervention designers. Criterion I is not met because the authors designed, adapted, trained, delivered, monitored, and analysed the intervention themselves, with no external or third-party evaluator.
    • Y

      Year Duration

      • The study tracked outcomes for only about three weeks, nowhere near 75% of an academic year, and criterion T was also not met.
      • "Although the differential intervention lasted only 2 weeks, students received three 200-min sessions per week, resulting in a concentrated learning schedule."
      • Relevant Quotes: 1) "The instructional program lasted 3 weeks and comprised three 200-min sessions per week, for a total of nine sessions." (p. 2) 2) "The program consisted of a 1-week standardized teaching phase followed by a 2-week differential intervention phase." (p. 2) 3) "T1 and T2 were completed after the instructional sessions in Weeks 2 and 3, respectively." (p. 3) 4) "Although the differential intervention lasted only 2 weeks, students received three 200-min sessions per week, resulting in a concentrated learning schedule." (p. 11) 5) "Third, the intervention period was short, and the findings mainly reflect short-term changes during the beginner stage." (p. 12) Detailed Analysis: Criterion Y requires outcomes to be measured at least 75% of a full academic year (roughly 7 or more months) after the intervention begins. In addition, the ranking instructions state that if criterion T is not met then criterion Y is automatically not met. Criterion T was not met: the tracking interval was about two to three weeks. The paper confirms this repeatedly - the whole programme was three weeks, the differential intervention two weeks, and the last measurement (T2) was taken immediately after the final Week 3 session. Even the most generous reading yields roughly 3 weeks against a requirement of about 7 to 10 months, i.e. under 10% of the required duration. There is no follow-up wave in a subsequent term or ski season; the authors list extending follow-up as future work. Criterion Y is not met because tracking lasted about three weeks rather than at least 75% of an academic year, and criterion T was also not met.
    • B

      Balanced Control Group

      • Time, technical content, practice volume, and instructor qualifications were explicitly matched across arms so that interaction style was the only substantive difference, with the brief autonomy-support instructor training being integral to the intervention tested.
      • "During the differential intervention phase, the two groups were matched on class duration, technical content, instructional progression, core practice tasks, and total practice volume. The only between-group difference concerned instructional interaction style and the organization of practice activities."
      • Relevant Quotes: 1) "During the differential intervention phase, the two groups were matched on class duration, technical content, instructional progression, core practice tasks, and total practice volume. The only between-group difference concerned instructional interaction style and the organization of practice activities." (p. 3) 2) "The control and experimental groups were matched in class duration, technical content, core practice tasks, and total practice volume, but the control group did not systematically implement the autonomy-supportive strategies specified in the experimental protocol, such as bounded choice, emotional acknowledgment, negotiative language, and limited autonomous arrangements." (p. 4) 3) "During the standardized phase, both groups received identical safety instruction, including ski slope safety rules and fall-protection techniques, as well as foundational technical instruction, including flat-ground exercises and basic stance training." (p. 2) 4) "To minimize the influence of instructor variability on between-group comparisons, the intervention was delivered by four instructors, all of whom held national ski instructor certification and had experience teaching university skiing courses. A paired allocation strategy was used such that each instructor taught one experimental class and one control class." (p. 4) 5) "Autonomy-supportive teaching can simultaneously improve skill acquisition, self-efficacy, and mental toughness among novice skiers without increasing instructional time or practice volume." (p. 1) 6) "The practical significance of this study does not lie in increasing instructional time, expanding practice volume, or altering technical content per se. Its central contribution is to show that, within a standardized instructional structure and fixed safety boundaries, optimizing classroom interaction style can simultaneously improve skill performance and key psychological resources." (p. 12) 7) "The training lasted 3 h and covered the principles of autonomy-supportive teaching, implementation of the seven categories of autonomy-supportive behaviors, boundary conditions for conventional teaching in the control group, classroom interaction language norms, and requirements for preventing contamination between teaching conditions." (p. 4) Detailed Analysis: Applying the criterion B decision tree, the first question is whether the intervention added time or budget relative to the control condition. On instructional time and practice resources the answer is clearly no. The paper states three separate times that the arms were matched on class duration, technical content, instructional progression, core practice tasks, and total practice volume, and the abstract and discussion both make the absence of extra time or practice volume a headline claim. Both arms shared the identical Week 1 standardised safety and foundational instruction. Instructor quality was also held constant by design: four equally certified instructors each taught one experimental and one control class, so no arm received better-qualified teaching staff. The sole manipulated variable was interaction style - bounded choice, task-value rationales, emotional acknowledgment, non-controlling language, and limited autonomy over repetitions and rest within a fixed class duration. One resource asymmetry deserves explicit mention. The 3-hour instructor training before Week 2 was oriented to the autonomy-supportive protocol, so the experimental condition carried a small teacher-preparation cost the control condition did not. This must be assessed carefully rather than waved through. Two considerations make it non-disqualifying. First, that training is precisely the operationalisation of the intervention being tested - autonomy-supportive teaching cannot be delivered without teaching instructors what it is, so it is integral to the treatment package rather than a separable add-on. Second, the same 3-hour session also covered "boundary conditions for conventional teaching in the control group," meaning instructors were prepared for both roles in one block; the amount is negligible relative to the 1,200 minutes of differential-phase instruction each student received (and 1,800 minutes across the whole programme), and it reached students only through interaction style, adding no student-facing time, materials, or budget. Because no meaningful extra time, budget, or materials reached the intervention students, and the only preparation differential is integral to the treatment and negligible in magnitude, the decision tree returns "met" at the no-extra-resources / negligible-difference branch without needing the treatment-variable exception. Criterion B is met because class duration, technical content, core practice tasks, total practice volume, and instructor qualifications were deliberately matched across arms, leaving instructional interaction style as the only substantive difference, with the modest autonomy-support instructor training being integral to the intervention itself.
  • Level 3 Criteria

    • R

      Reproduced

      • Internet searching found no independent replication of this specific beginner-ski cluster RCT; the authors themselves list replication as future work.
      • "First, future studies should increase the number of clusters and conduct replication across multiple contexts."
      • Relevant Quotes: 1) "However, few studies have integrated randomized controlled designs, skill acquisition assessment, and psychological resource measures, making it difficult to determine the overall intervention effect of autonomy-supportive teaching in beginner skiing instruction." (p. 2) 2) "For this reason, beginner skiing provides a suitable context for examining the applicability of autonomy-supportive teaching in high-risk skill learning." (p. 2) 3) "First, future studies should increase the number of clusters and conduct replication across multiple contexts. Larger cluster-randomized trials conducted across different schools, ski-field conditions, and instructional teams would improve the precision of class-level inference..." (p. 13) 4) "The present study extends this evidence from general classroom and conventional physical education contexts to a high-risk, short-cycle instructional context, providing longitudinal evidence with greater contextual relevance for pedagogical intervention research in skiing education." (p. 11) Detailed Analysis: Criterion R requires that this specific study, or its central experimental claim in the same context and design, have been independently replicated by a different research team and published in a peer-reviewed journal. The paper reports no replication of itself. On the contrary, it positions itself as filling a gap - it states that few studies have combined an RCT design with skill acquisition and psychological measures in beginner skiing, and describes itself as extending existing autonomy-support evidence into a novel high-risk instructional context. The authors explicitly call for replication as future work, which is an acknowledgement that none exists. Internet verification was carried out for this verification pass, searching for the article title, its DOI, the author names, and topic-level queries combining "autonomy-supportive teaching", "beginner skiing", "cluster-randomized" and "replication". No study by any other research team replicating this beginner-ski cluster RCT was found in any indexed source, and no citing literature of any kind was located. No verbatim quote from a replication study can therefore be provided, because no such study exists to quote. The paper does cite other autonomy-supportive teaching trials by other teams (for example Cheon et al., 2022, 2023, 2024; Reeve and Cheon, 2024; Wang et al., 2025) and meta-analyses such as Vasconcellos et al. (2020) and Patzak and Zhang (2025). These establish that autonomy-supportive teaching as a general pedagogical approach has been tested widely, but they are conventional classroom and physical education studies, not replications of this particular beginner-ski cluster RCT with its specific rubric-based skill outcome and three-week intensive design. Under the standard's own worked example on this point, prior trials of the same broad programme in different populations do not constitute independent reproduction of the specific study under review. The article was published on 29 May 2026, under two months before this check, so any independent replication would in any case be extremely unlikely to exist yet. Criterion R is not met because internet searching found no independent replication of this specific ski-instruction cluster RCT by any other research team.
    • A

      All-subject Exams

      • Only ski technique was assessed, using a non-standardised study-specific rubric, and criterion E was not met, so the all-subject requirement fails.
      • "This study aimed to compare the effects of autonomy-supportive teaching and conventional teaching on skill acquisition, self-efficacy, and mental toughness in beginner skiing instruction."
      • Relevant Quotes: 1) "Skill acquisition was evaluated using a pre-established standardized technical scoring rubric comprising five core dimensions: posture control (20%), center-of-mass transfer (25%), turning stability (25%), speed control (15%), and trajectory control (15%)." (p. 4) 2) "This study aimed to compare the effects of autonomy-supportive teaching and conventional teaching on skill acquisition, self-efficacy, and mental toughness in beginner skiing instruction." (p. 1) 3) "Participants were first-year undergraduate students from Shenyang Sport University who were not majoring in skiing." (p. 3) 4) "Self-efficacy was measured using the General Self-Efficacy Scale (Schwarzer and Jerusalem, 1995)." (p. 4) 5) "Mental toughness was assessed using the Sports Mental Toughness Questionnaire (SMTQ; Sheard et al., 2009)." (p. 5) Detailed Analysis: Criterion A requires that all main subjects taught at that educational level be assessed, using standardised exam-based assessments, and the ranking instructions state that if criterion E is not met then criterion A cannot be met either. Criterion E was not met, so criterion A fails on that prerequisite alone. Independently, the substantive requirement also fails. The participants were first-year undergraduates taking a university skiing course, and the only performance outcome measured was ski technique on a five-dimension rubric. No other subject from these students' degree programmes was assessed, and the two remaining outcomes (general self-efficacy and sport mental toughness) are psychological constructs rather than academic subjects. There is no evidence about whether the intervention helped or harmed attainment in any other course the students were taking, which is exactly the spillover risk criterion A exists to detect. The paper offers no rationale for restricting measurement to the single intervention domain. Criterion A is not met because criterion E was not met and because only ski technique was measured, with no assessment of any other subject the students were studying.
    • G

      Graduation Tracking

      • Measurement stopped immediately after the three-week programme with no tracking towards graduation, no follow-up publication was found, and criterion Y was also not met.
      • "Third, future work should extend follow-up periods and examine retention, transfer, and safety-related outcomes."
      • Relevant Quotes: 1) "Data were collected at three measurement time points (T0-T2), each scheduled at the end of the instructional sessions for the corresponding week." (p. 5) 2) "T1 and T2 were completed after the instructional sessions in Weeks 2 and 3, respectively." (p. 3) 3) "Third, the intervention period was short, and the findings mainly reflect short-term changes during the beginner stage." (p. 12) 4) "Third, future work should extend follow-up periods and examine retention, transfer, and safety-related outcomes. Additional follow-up assessments would make it possible to evaluate skill retention, cross-task transfer, and stability across ski seasons." (p. 13) Detailed Analysis: Criterion G requires participants to be tracked through to graduation from their educational stage, and the ranking instructions state that if criterion Y is not met then criterion G cannot be met. Criterion Y was not met, so criterion G fails on that prerequisite. The substantive evidence points the same way. Measurement stopped at T2, immediately after the final instructional session of Week 3. The participants were first-year undergraduates who would remain enrolled for roughly three further years, yet no attempt was made to follow them to the end of their degree, or even to the end of the academic year or the ski season. As required by the verification procedure, an internet search was conducted for subsequent or follow-up publications by the same author team (Gao, Cao, Ren, Shi, Liang, Liu; Shenyang Sport University, Huaibei Normal University, Harbin Sport University) that might track this cohort further. No follow-up paper, cohort-tracking study, or graduation-outcome publication for this sample could be located in any indexed source, so no verbatim quote from such a paper can be supplied. Given publication on 29 May 2026, under two months before this check, this is expected. The authors themselves state only that extending the follow-up period is a task for future research. Criterion G is not met because data collection ended immediately after the three-week programme with no graduation tracking, no follow-up publication tracking this cohort could be found, and criterion Y was also not met.
    • P

      Pre-Registered

      • Neither the paper nor any searched trial registry records a pre-registration for this study; only institutional ethics approval and CONSORT reporting are reported.
      • Relevant Quotes: 1) "The reporting of this study was guided by the CONSORT extension for cluster randomized trials, with attention to the unit of randomization, number of clusters, allocation procedure, participant flow, handling of clustering effects, and multilevel statistical analysis." (p. 3) 2) "The studies involving humans were approved by Shenyang Sport University Academic Ethics Committee [Approval No. 202669]. The studies were conducted in accordance with the local legislation and institutional requirements." (p. 13) 3) "Sample size estimation was performed using G*Power 3.1, with skill acquisition specified as the primary outcome. A medium effect size of d = 0.50, statistical power of 1 - beta = 0.80, and a significance level of alpha = 0.05 were specified." (p. 3) 4) "During the study, any course interruption, injury-related withdrawal, absence, or incomplete assessment was documented with respect to group assignment, time point, and reason, and then handled according to the prespecified rules." (p. 5) 5) "The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation." (p. 13) Detailed Analysis: Criterion P requires a pre-registration of the full protocol on a registry platform before data collection began, with a verifiable registration reference and date. The paper contains no registration statement of any kind. There is no trial registry name, no registration identifier, and no registration date anywhere in the methods, the data availability statement, or the end matter. What the paper does provide is ethics committee approval (Approval No. 202669), which is institutional research governance and not public pre-registration; an a priori power calculation; a statement that certain handling rules were prespecified; and a claim that reporting followed the CONSORT extension for cluster trials. CONSORT is a reporting guideline applied at write-up, not a registration mechanism, and none of these items substitutes for a public, dated, pre-data-collection protocol registration. For this verification pass the published Frontiers article page was re-checked online and confirmed to contain no trial registration number, and registry-oriented searches (ChiCTR, ClinicalTrials.gov, OSF) for this trial returned no matching record. With no registry reference at all, timing relative to data collection cannot be verified, and the transparency purpose of the criterion - preventing selective outcome reporting and analysis flexibility - is unaddressed. Criterion P is not met because neither the paper nor any searched registry provides a trial registration, identifier, or registration date; only institutional ethics approval is reported.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.