Class-level RCT
Tests interventions at the classroom level to prevent cross-group contamination
The Educational Randomized Controlled Trial (ERCT) Standard is a rigorous framework addressing key challenges in educational studies. With 12 criteria across 3 progressive levels, it resolves issues like bias, limited scope, and short-term focus – enabling researchers to produce actionable results that improve education systems worldwide.
The ERCT Standard has 3 levels, each containing 4 criteria
Tests interventions at the classroom level to prevent cross-group contamination
Uses standardized exams for objective and comparable results
Ensures studies last at least one academic term to measure meaningful impacts
Requires detailed control group data for proper comparisons
Expands testing to whole schools for real-world relevance
Removes bias by using third-party evaluators.
Ensures studies last at least one academic year to measure meaningful impacts
Ensures equal time and resources for both groups to isolate the intervention's impact.
Independently replicated study
Assesses effects across all core subjects, avoiding imbalances
Tracks students until graduation to evaluate long-term impacts.
Increases transparency by publishing study plans before data collection
The study is randomized at the school level, which satisfies class-level randomization.
Outcomes were measured using Texas STAAR, a standardized state assessment.
The study ran across two school years, exceeding the minimum one-term requirement.
The control condition is clearly described and supported by detailed baseline and implementation documentation.
Random assignment was implemented at the school level (64 schools).
RAND conducted the evaluation independently of Zearn, supported by IES funding.
Outcomes were tracked over two full academic years (2022-2023 and 2023-2024).
Although Zearn provided additional implementation supports, teachers reported similar total math instructional time across groups and the supports appear integral to the tested intervention package.
An independent, peer-reviewed randomized study by a different author compared Zearn Math to another program, serving as an external replication effort on the intervention.
The paper reports standardized outcomes only for mathematics, not for all core subjects.
The study reports outcomes through the end of the second study year and does not track students to graduation.
The paper states a REES registry ID, but the public registry entry date could not be verified to be before study start.
Zearn Math is a popular software platform for K-8 mathematics learning, designed to enable all students to successfully access grade-level content. RAND researchers collaborated with Zearn, the product's developer, to design this evaluation. Then RAND conducted the study independently, randomly...
The trial randomised whole schools rather than classes or students, which is stronger than and automatically satisfies class-level randomisation.
The primary outcome used genuine past national KS2 Writing test papers externally marked, and the secondary outcomes (Reading, GPS) were drawn from compulsory national KS2 testing, both standardised measures.
Outcomes were measured around 20 months after the intervention began, far exceeding the minimum one-term requirement (and this is superseded by the Year Duration criterion also being met).
The control group is documented in detail, with a full comparison table of baseline demographic and attainment characteristics for both arms and a description of control conditions.
Randomisation was conducted at the level of whole schools using a formal minimisation procedure, directly satisfying the school-level RCT requirement.
The trial was independently evaluated by Sheffield Hallam University with blind, externally-contracted test markers, separate from the Enfield LA team that designed and delivered the intervention.
Pupils were tracked from the start of the intervention in Y5 through to final outcome testing roughly 20 months later at the end of Y6, comfortably exceeding 75% of an academic year.
The additional resource in this trial was teacher professional development delivered outside pupil instructional time; pupils in both arms received the same amount of normal classroom time, with the control condition appropriately remaining business-as-usual as this is precisely the treatment being tested.
This is described as the first large-scale evaluation of this programme, with no independent replication reported or evident in the paper, and none was located through internet research.
Only English-domain outcomes (writing, reading, grammar/ punctuation/spelling) were assessed; no other core subjects such as mathematics or science were measured.
Pupils were followed from Y5 through to the completion of Y6 and the KS2 assessment, which marks the end of primary schooling in England before transition to secondary school; no follow-up publication tracking this cohort further was found.
The trial has a public ISRCTN registry entry, but neither the paper nor further internet research could confirm that registration occurred before data collection began.
The Integrating English programme aimed to improve the writing ability of Year 5 and 6 pupils through a structured CPD programme (based on the Australian LiLAC course) that provided teachers with linguistic and pedagogical strategies, with the greatest expected impact...
Randomisation occurred at the school level, which meets or exceeds the class‑level RCT requirement.
The study uses the nationally standardized ENLACE exam for objective, comparable assessment.
Outcomes were collected approximately three years after the intervention began, exceeding the minimum of one academic term.
Baseline demographics and pre‑intervention scores for the control group are provided in detail.
Randomisation at the school level fulfills the School‑level RCT requirement.
The evaluation was conducted by researchers unaffiliated with SEP, ensuring independence of the study’s implementation and analysis.
Participants were followed for approximately three years, which is longer than one academic year from intervention start to final outcomes.
The difference in time or resources between groups is trivial and integral to the intervention, unlikely to bias the results.
The paper does not reference any independent replication studies, and none were found in external literature.
Only math and language outcomes are measured, omitting other main subjects.
The RCT measured on-time high school completion, meaning participants were followed through the completion of that educational level.
No pre-registration statement or registry ID is provided in the paper or its references.
We use data from the randomized control trial of the Percepciones pilot to study whether providing 10th grade students with information about the average earnings associated with different educational attainments, life expectancy, and obtaining funding for higher education can contribute...
Randomization occurred at the class (school) level, avoiding within-class mixing and meeting the class-level RCT standard.
Student learning was measured with the TerraNova standardized test, a well‑established, nationally normed exam.
The intervention spanned a full academic year of seventh-grade instruction, exceeding the one-term minimum.
The control group’s practices, demographics, and baseline scores are clearly documented for proper comparison.
Entire schools (not just classes) were randomized, satisfying the school-level RCT requirement.
An independent research team evaluated the intervention, and no authors had a financial interest in ASSISTments.
The trial ran for a full academic year (with data from the second year cohort), meeting the one‑year duration requirement.
Both groups followed identical homework policies and content; the ASSISTments tool and training were the core treatment.
An independent replication by another research team found similar positive results, confirming the original findings.
Only mathematics achievement was assessed; other subjects were not tested.
Students were originally tracked only through 7th grade; a separate follow-up measured outcomes at the end of 8th grade, but tracking did not extend to high school graduation.
The study was not pre-registered; no registry or preregistration statement is provided.
In a randomized field trial with 2,850 seventh-grade mathematics students, we evaluated whether an educational technology intervention increased mathematics learning. Assigning homework is common yet sometimes controversial. Building on prior research on formative assessment and adaptive teaching, we predicted that...
Randomisation was at the school (cluster) level, which meets or exceeds the class-level requirement.
The study used IDELA, a widely used and validated standardized early learning assessment tool.
Outcomes were tracked from October 2022 to October 2023, far exceeding a one-term minimum.
The control group is described as business-as-usual and baseline characteristics and scores are reported for comparison.
Schools were randomly assigned to treatment or control, satisfying the school-level RCT requirement.
The paper states the DPL provider did not participate in data collection, analysis, or conclusions, supporting independent conduct.
Outcomes were measured over about 13 months and four school terms, meeting the year-duration threshold.
The intervention added devices and implementation support, but these additional resources are integral to the DPL treatment being tested against business-as-usual.
No independent replication of this specific trial by a different research team could be identified from the paper or from web searching.
The study assessed standardized outcomes in literacy and numeracy only, not all main subject areas for the educational setting.
No evidence was found that the study tracked learners until graduation, and web searching did not identify follow-up publications tracking this cohort through graduation.
The protocol was registered/published after data collection began, so it does not meet the ERCT requirement for pre-registration before the study started.
Research on digital personalised learning (DPL) alongside classroom teaching is limited in low- and middle-income countries. This study investigates, for the first time, a DPL programme aligned with national curricula and teaching practices. A randomised trial evaluated the impact of...
Randomization was at the school level, which satisfies the class-level RCT requirement.
Outcomes were measured using EGRA and EGMA, which are standardized assessment systems.
Midline outcomes were measured about 1.5 years after implementation began, exceeding one academic term.
The control group is described with sample sizes and baseline characteristics, including a baseline balance table.
Treatment was assigned at the school level, meeting the school-level RCT requirement.
Data collection involved an external firm and the authors analyze secondary data rather than directly implementing the intervention.
Midline outcomes were measured about 1.5 years after implementation began, exceeding one academic year.
The intervention adds facilitated discussions, and those added resources are the treatment being tested rather than an uncontrolled add-on.
No independent reproduction of this specific study was identified in the paper, and none could be verified from accessible sources.
The study measures mathematics and reading only, not all core subjects.
Participants were followed through endline in 2016, which is not tracking the cohort to graduation.
The study reports AEA RCT Registry registration, but the registry entry timing could not be verified from accessible sources for this check.
This article evaluates the impact that facilitated discussions about girls’ education have on education outcomes for students in rural Zimbabwe. The staggered implementation of components of a randomized education project allowed for the causal analysis of a dialogue-based engagement campaign....
Centers (preschool sites) were randomly assigned, which is class-level or stronger randomization.
The outcomes use established, widely used assessments (IGDI, FastBridge, HTKS) rather than a custom test created for this study.
Outcomes are measured from fall to spring within a school year, which is at least one academic term (and in practice close to a full year).
The paper clearly defines a comparison control group (including sample sizes) and reports balance checks and baseline equivalence statements.
Randomization occurred at the preschool center level, which meets the school-level RCT requirement in this context.
The evaluation is conducted by NORC at the University of Chicago, while the program is implemented in collaboration with Kidango as the implementation partner.
The study measures outcomes from fall to spring within a school year, and the intervention is described as occurring during the 2017-2018 school year.
Any additional resources (PD and coaching) are the intervention being tested, and the comparison group is business-as-usual with delayed treatment.
No peer-reviewed, independent replication of this specific SEEDS PD RCT was found in the available literature.
The study focuses on language/literacy and executive function outcomes and does not assess all core subjects (for example, mathematics).
The study reports outcomes within preschool years (fall-to-spring) and does not track students through graduation from the educational stage.
The paper does not cite a pre-registration record or registry identifier, and no verified pre-registration entry was found.
NORC at the University of Chicago designed and implemented an impact evaluation of the SEEDS of Learning (SEEDS) professional development (PD) program on behalf of the Kenneth Rainin Foundation, in collaboration with Kidango. SEEDS of Learning is an evidence-based PD...
Whole schools were randomly assigned to intervention and control, exceeding the class-level randomisation requirement.
Outcomes were measured using the standardised NWEA MAP Growth assessment rather than a custom test.
The study followed students from Autumn 2023 through the July 2024 post-assessment, exceeding one full term.
The report documents the control group definition, baseline school characteristics, and control schools business as usual practices.
Twenty schools were randomised at the school level, satisfying the school-level RCT requirement.
The report separates the developer from the evaluator, with WhatWorked conducting randomisation and evaluation activities.
Outcomes were tracked from September 2023 to July 2024, covering an academic year.
The intervention did not add unbalanced time or resources: homework frequency and duration were similar and controls used other platforms.
No independent peer-reviewed replication of this specific school-level RCT was found.
Only mathematics attainment was assessed, not standardised outcomes in all core subjects.
Follow-up stops in July 2024 for the Year 7 cohort, with no graduation tracking reported.
The report provides no trial registration or protocol ID, and no public pre-registration could be verified.
This report evaluates the impact of Eedi, a digital mathematics platform, on raising maths attainment amongst Key Stage 3 students (Year 7). The study used a randomised controlled trial design with 20 schools, where schools were randomly assigned to either...
Randomization was conducted at the school level, satisfying the class-level requirement.
The study uses the ECE, Peru's national standardized assessment for primary schools.
Outcomes were measured approximately 9 months after the intervention began, exceeding the one-term requirement.
The control group is clearly defined as non-participating schools and their baseline characteristics are extensively documented in Table 2.
Randomization occurred at the school level across 6,218 schools.
The study evaluation was conducted by independent academics, distinct from the Ministry that designed the intervention.
The study measured outcomes after 1 and 3 years of program implementation.
The intervention explicitly tests the impact of adding significant resources (coaching), making the resource imbalance integral to the study design.
The study has not been independently reproduced.
Assessments were limited to mathematics and reading, omitting other core subjects like science.
Tracking ended at Grade 4, prior to primary school graduation.
No pre-registration of the study protocol is mentioned.
We evaluate the impact of a large-scale teacher coaching program in Peru, a context with high teacher turnover, on teachers' pedagogical skills and student learning. Previous studies find that small-scale coaching programs can improve teaching of reading and science in...
The study randomised at the school (cluster) level, which is class-level or stronger and reduces contamination risks.
The outcomes were measured using IDELA, a widely used and validated assessment tool rather than a study-created exam.
Outcomes were tracked from baseline (October 2022) to endline (October 2023), which exceeds one academic term.
The control group is clearly described as business-as-usual and baseline characteristics and scores are reported for both groups.
Schools were the cluster unit and were randomly allocated to treatment or control, satisfying school-level randomisation.
The paper explicitly states independence from the DPL provider for data collection, analysis, and conclusions, with independent enumerators collecting assessment data.
The study assessed outcomes from October 2022 to October 2023 (about 13 months), exceeding 75% of an academic year.
The treatment received devices and implementation support, but these additional resources are integral to the DPL intervention being tested against business-as-usual schooling.
No independent peer-reviewed replication of this specific RCT was found in the paper or through internet searching.
Although IDELA is used, the study assesses literacy and numeracy only and does not measure impacts across all main subject areas.
No evidence was found that the study tracked the same learners through a graduation milestone beyond the endline in October 2023, and no follow-up paper reporting graduation tracking was identified.
The study mentions protocol registration prior to analysis, but external checking indicates the publicly available protocol was published in October 2023 (after baseline began in October 2022), so it does not meet ERCT pre-registration timing requirements.
Research on digital personalised learning (DPL) alongside classroom teaching is limited in low- and middle-income countries. This study investigates, for the first time, a DPL programme aligned with national curricula and teaching practices. A randomised trial evaluated the impact of...
The study randomized entire villages (clusters), satisfying the requirement for class-level or stronger randomization.
The study used EGRA and EGMA, which are widely recognized standardized assessments.
The study duration spanned multiple years, significantly exceeding the one-term requirement.
The control group is well-documented, including demographics and confirmation that they received no educational intervention.
Randomization occurred at the village level, which serves as the implementation unit, satisfying the school-level RCT criterion.
Data collection was conducted by an independent organization (GHTC) with blinded administrators, ensuring independent conduct.
The study assessed outcomes approximately 30 months after the intervention began, satisfying the year duration criterion.
The study explicitly tested the impact of the additional resources (para-instructor intervention) as the treatment variable.
The intervention methodology has been replicated by independent teams (e.g., J-PAL/Banerjee et al.) as cited in the paper.
The study assessed outcomes in Reading and Mathematics, covering the main subjects for the target grade levels.
The study explicitly states that no further follow-up was conducted after the intervention ended, failing the graduation tracking requirement.
The paper cites a protocol published after the trial start but does not provide a specific pre-registration registry link or date in the text.
In common with many other low- and middle-income countries (LMICs), India has witnessed a massive expansion in school enrolment over the last 20 years, and yet many students finish primary education without the foundational literacy and numeracy skills that would...
Randomisation occurred at the TAC tutor zone level (clusters of 3-6 schools), which is stronger than class-level randomisation and therefore satisfies this criterion.
The literacy battery was built on the internationally recognised EGRA and PALS assessment frameworks, widely used across sub-Saharan Africa.
Outcomes were measured at 9 and 24 months after the intervention began, far exceeding one academic term.
Table 3 provides detailed baseline demographic, school, and household characteristics for the 50 control schools compared with intervention schools.
Schools (organised in randomly-allocated clusters of TAC tutor zones) were the unit of random allocation to literacy intervention or control.
The paper explicitly states, and lists as a limitation, that the same team that designed and implemented the intervention also conducted the evaluation.
The study tracked outcomes for 24 months, well beyond the required 75% of an academic year.
The additional teacher training, materials, and SMS support are themselves the explicit treatment variable being tested against a business-as-usual control, which satisfies the balance requirement by design.
No independent peer-reviewed replication of this specific cluster-randomized HALI study was found; later related programs are similar in concept but distinct interventions/designs, and a review of the citing literature confirms no independent reproduction exists.
In addition to literacy, the study assessed numeracy outcomes using an EGMA-based battery, covering the main subjects taught at this early-grade level.
Tracking stopped at the end of Grade 2 (24 months), with no evidence of follow-up through graduation from primary school, and a search of the authors' subsequent publications found no such extended tracking.
The trial was registered on ClinicalTrials.gov (NCT00878007) in April 2009, about nine months before data collection began in January 2010, as confirmed by independent verification of the registry record.
We evaluated a program to improve literacy instruction on the Kenyan coast using training workshops, semiscripted lesson plans, and weekly text-message support for teachers to understand its impact on students' literacy outcomes and on the classroom practices leading to those...
Randomisation occurred at the school level (193 schools), which is a stronger unit than class-level and therefore automatically satisfies the class-level RCT requirement.
The study used well-established standardized language assessments (CELF Preschool, Renfrew APT), not custom-built instruments.
Outcomes were measured roughly 7-8 months after the intervention began, far exceeding the one-term minimum.
The control group's size, demographics, baseline scores and the provision it received are documented in detail in the text, CONSORT diagram and Table 1.
The trial randomly allocated entire schools (193 schools) to intervention or control, directly satisfying the school-level criterion.
Despite independent randomisation and blinded testers, several co-authors are the intervention's original designers or hold direct financial stakes in commercial products central to the trial.
Tracking from intervention start to the final outcome measurement spans roughly 8 months, exceeding 75% of the UK academic year.
The additional teaching-assistant time and training is the treatment variable being tested; the control group's business-as-usual provision is the appropriate comparison, not an imbalance.
No independent replication of this specific large-scale cluster RCT by a different research team was found in the paper or via further internet search.
Only oral language and early word reading were assessed; no other core subject (e.g. mathematics) was measured, with no stated rationale for this narrow scope.
Same-team follow-ups extend tracking to only about two years post-intervention (around ages 6-7); no evidence of tracking through graduation was found.
The trial was preregistered on the ISRCTN registry on 2018-06-05, before screening (t0) began in September 2018, and the paper states its analyses followed this preregistered plan.
Background: It is well established that oral language skills provide a critical foundation for formal education. This study evaluated the effectiveness of the Nuffield Early Language Intervention (NELI) programme in ameliorating language difficulties in the first year of school when...
Randomization was conducted at the school level (30 schools), which is stronger than the class-level requirement, so the class-level RCT criterion is met.
The study used the MAP reading assessment (NWEA), a widely recognized standardized computer-adaptive test, as an outcome measure for domain-general reading comprehension.
The intervention began in February 2019 and outcomes were measured through fall of Grade 2 (September 2019) and the end of the Grade 2 science unit, an interval far exceeding one academic term.
The control group is documented in detail, including sample sizes, demographics, baseline test scores, the business-as- usual curriculum it received, and its instructional time.
Thirty elementary schools were the unit of randomization in a blocked cluster design, satisfying the school-level RCT criterion.
The same Harvard-led research team that developed the MORE intervention also implemented the trial and analyzed the outcomes, with no independent third-party evaluator reported.
The study was an explicitly 12-month intervention (February 2019 through the Grade 2 unit ending late January 2020) with final outcomes measured about a year after the start, exceeding 75% of an academic year of tracking.
The treatment package (thematic lessons, teacher PD, and free summer books) was itself the treatment variable tested against business-as-usual instruction, and in Grade 2 instructional time was empirically equivalent between conditions.
No independent replication of this study by a different research team is reported or found; prior and subsequent MORE studies located online were conducted by the same author team.
Standardized outcome testing covered only reading (MAP reading); mathematics and other core subjects were not assessed as outcomes, and the science measure was a custom researcher-made test.
Tracking ended in Grade 2 after COVID-19 school closures cancelled planned Grade 3 assessments; a possible same-author Grade 3 follow-up was located online but does not reach graduation and could not be verified with quotes.
The study design was pre-registered in the AEA RCT Registry (trial AEARCTR-0003489) on October 24, 2018, confirmed via the registry itself to predate the Grade 1 intervention that began in February 2019.
We developed a sustained content literacy intervention that emphasized building domain and topic knowledge from Grade 1 to Grade 2 and evaluated transfer effects on students' reading comprehension outcomes. The Model of Reading Engagement (MORE) intervention emphasizes thematic lessons that...
Randomization was conducted at the school (cluster) level, satisfying the class‑level RCT criterion.
The study measured outcomes using national examination scores, a standardized assessment.
Intervention and follow‑up lasted ten months, exceeding a full academic term.
The control group’s size and baseline characteristics are clearly documented in the methods and tables.
Entire schools were randomized, satisfying the school‑level RCT criterion.
The same team that designed the intervention conducted and analyzed the trial without independent oversight.
Participants were tracked from February through December—one full academic year.
HIIT was integrated into standard PE time without additional class time or resources, keeping groups balanced.
No independent replication of this intervention has been reported.
Outcomes were measured only in mathematics and Mongolian language, not all main subjects.
Participants were not followed until graduation; follow-up ended at study completion.
Trial registration was completed before the study began (registered 1st February 2018).
OBJECTIVES: Physical inactivity is an important health concern worldwide. We examined the effects of an exercise intervention on children’s academic achievement, cognitive function, physical fitness, and other health-related outcomes. METHODS: We conducted a population-based cluster RCT among 2301 fourth‑grade students...
The study randomised treatment at the grade‐by‐school level, satisfying the class‐level RCT requirement.
They used the Smarter Balanced standardized exams for Math and ELA as outcome measures.
The intervention lasted from late October through May, satisfying the full academic term requirement.
Control group characteristics and communications are clearly documented in the methods and Table 1.
Randomisation occurred at the grade‐by‐school level rather than entire schools.
The study was designed, implemented, and analyzed by the same team without external evaluation.
The study intervention covered a full academic year.
No additional instructional time or budget was provided, only low‑cost informational text messages.
A separate research team reproduced the intervention in another context and reported similar positive results.
Only Math and ELA were assessed via standardized exams, failing to cover all core subjects.
Participants were tracked only through the end of the school year, not until graduation.
The study’s analysis plan and outcomes were publicly pre-registered before data collection began.
While leveraging parents has the potential to increase student performance, programs that do so are often costly to implement or they target younger children. We partner text‐messaging technology with school information systems to automate the gathering and provision of information...
Randomization was conducted at the teacher/classroom level (whole focal classes assigned to treatment or control), which satisfies the class-level RCT requirement.
The study used a recognized state standardized exam (SBAC ELA) as an outcome measure alongside the study-created writing assessment, satisfying the exam-based assessment requirement.
Outcomes were measured in spring, roughly seven to eight months after the fall start of the year-long intervention, exceeding the one-term requirement.
The control group's size, demographics, baseline equivalence, attrition, and business-as-usual treatment are documented in detail.
Randomization was at the teacher/class level within school-by-grade blocks, so whole schools were not the unit of assignment.
SRI International, an independent nonprofit evaluator, performed randomization, data collection, and analysis with blinded scorers, ensuring independent conduct.
The year-long intervention was tracked from fall to a late April/early May posttest, covering at least 75% of the academic year.
The extra teacher PD is the integral treatment variable tested against business-as-usual, and student instructional time was equal across conditions.
All replications of the Pathway intervention, including a newer multi-state expansion trial, were run by the same developer/evaluator team (UCIWP/SRI) and none is published as an independent peer-reviewed replication.
Only ELA writing outcomes were assessed; no other main secondary school subjects were measured.
Measurement stopped at the end of Year 1; students were not tracked through graduation, and no follow-up publication with graduation data was found.
No public pre-registration of the study protocol (registry link, ID, or date) is mentioned in the paper or found via internet search; a later, separate Pathway study was registered, underscoring the absence of registration for this one.
This study reports findings from a multisite cluster randomized controlled trial designed to validate and scale up an existing successful professional development program that uses a cognitive strategies approach to text-based analytical writing. The Pathway to Academic Success Project worked...
Randomisation was carried out at the school/teacher level (a stratified school-level cluster trial), which meets or exceeds the class-level requirement.
The primary outcome was the standardised Progress in English (PiE) writing test from GL Assessment, a widely recognised standardised assessment.
The intervention ran as ten weekly one-hour lessons with outcome measurement roughly three months after baseline, spanning approximately one academic term.
The business-as-usual control group is documented with demographic and baseline data in Tables 1 and 2 and a description of its conditions.
The trial was an explicitly described school-level cluster RCT with 70 schools randomly allocated to intervention or control.
The Englicious approach was designed by the linguists on the research team, who also delivered the training, ran the study and analysed the data, so it was not independently conducted.
Outcomes were measured only about three months after baseline, far short of 75% of an academic year.
The additional writing practice and grammar training are integral to the Englicious approach that is the explicit treatment variable, tested against a business-as-usual control.
This is the first study of its kind with this age group and has not been independently replicated.
Only writing/grammar outcomes were assessed; no other core school subjects were measured.
Pupils were tested only one to two weeks after the intervention, with no tracking through to graduation.
A full evaluation protocol (Anders et al., 2019) and statistical analysis plan were pre-published before data collection, and the trial was prospectively registered on the SREE Registry (ID #1920.1v1).
Very few research studies of the teaching of grammar and writing had been carried out with children younger than eight-years-old prior to the research reported in this paper. The research evaluated a new approach to teaching grammar and writing called...
Schools were randomized (cluster RCT), which meets and exceeds the class-level randomization requirement.
The outcomes use standardized, widely recognized assessments (New PiRA and WIAT-III UK-T) administered under standardized procedures.
Outcomes were measured from Summer 2022 (baseline) to Summer 2023 (endline), which is well beyond one academic term after intervention began.
The control condition is explicitly described as business-as-usual, with clear sample sizes and baseline balance information reported.
Schools (not classes or students) were randomized, satisfying the school-level RCT criterion.
The paper indicates intervention development by study authors, but does not clearly document that evaluation and analysis were led independently from the intervention designers.
The intervention and outcome tracking span from October 2022 through endline in Summer 2023, which is consistent with at least 75% of an academic year after intervention start.
The intervention resources (structured peer-reading sessions plus training/materials) are integral to the treatment being tested and are not framed as a separable, confounding add-on.
No independent peer-reviewed replication by a different research team of this specific 2022/23 PALS-UK cluster RCT was found.
The study measures standardized reading outcomes only and does not assess all main subjects.
The paper reports outcomes through the end of Year 5 only and does not report tracking to the end of primary school graduation or later milestones.
The protocol was preregistered before randomization, but the OSF preregistration date appears to be after baseline data collection had already occurred.
Purpose: This study evaluates the impact of Peer Assisted Learning Strategies (PALS-UK) in developing pupils’ reading attainment, reading skills (comprehension and fluency) and affective factors (reading self-efficacy and motivation). Method: All Year 5 pupils (9–10 years old, N = 4840,...
Criterion C is met because random assignment occurred at the school level, which is class-level or stronger.
Criterion E is met because outcomes include standardized exam-based measures (EOG and MAP).
Criterion T is met because outcomes were measured well beyond one term after the intervention began.
Criterion D is met because the control condition and baseline comparability information are explicitly documented.
Criterion S is met because randomization occurred at the school level.
Criterion I is not met because the paper does not provide explicit evidence that the evaluation was conducted independently from the intervention designers.
Criterion Y is met because outcomes were tracked across multiple school years from the start of implementation through spring of Grade 4.
Criterion B is met because the additional resources (PD, lessons, books, and a digital app) are integral to the MORE intervention package being tested against business-as-usual.
Criterion R is not met because no independent replication by a different research team is identified or documented.
Criterion A is not met because the standardized exam outcomes reported are limited to reading and math, not all main subjects.
Criterion G is not met because follow-up is reported through spring of Grade 4, not through graduation.
Criterion P is not met because no pre-registration ID/link and pre-intervention registration date are provided or verifiable from accessible sources.
Far transfer—the application of learning across distant domains—remains elusive in intervention research, and even when it is found, its mechanisms remain unclear or unexplored. This study analyzes data from the Model of Reading Engagement (MORE), a sustained content literacy intervention...
Randomisation occurred at the ECD-centre (site) level, which is stronger than class-level randomisation and satisfies criterion C.
The study uses ELOM, described as a standardised tool with established validity and reliability, meeting criterion E.
Outcomes were measured at Time 2 after 26 weeks, which exceeds a typical academic term duration.
The paper documents the wait-list business-as-usual control group and reports group sizes and baseline characteristics, meeting D.
ECD centres (sites) were randomly assigned to conditions, meeting the school/site-level RCT requirement.
The paper does not clearly document evaluation independence from the intervention organisation/designers, and a competing interest is declared, so criterion I is not met.
Outcomes were assessed after 26 weeks, which is below 75% of the paper’s stated 36-week full school-year programme, so Y is not met.
Although the intervention adds resources (materials and teacher training/support), these are integral to the package being tested against business-as-usual, so criterion B is met.
No independent, peer-reviewed replication by a different research team was identified, so criterion R is not met.
The study assesses broad learning domains (including emergent literacy/language and emergent numeracy) using the standardised ELOM, satisfying criterion A for broad coverage in this context.
The study does not track participants to graduation, and because criterion Y is not met, criterion G is not met as well.
The paper links to an OSF pre-registration but the registration date (and its precedence to data collection) is not documented in the paper, so criterion P is not met.
Literacy rates in South Africa are low and many children start school without the requisite levels of emergent language and literacy skills needed to succeed. We report two RCTs of a story-based intervention delivered by preschool teachers to two language...
Although randomization was within classrooms, the intervention was delivered as supplemental small-group instruction (tutoring-like), so student-level randomization is acceptable under the ERCT exception.
The outcomes include widely used standardized, norm-referenced writing assessments (e.g., WIAT-III and TOWL-4).
Outcomes were measured after a 12-week instructional period (about one academic term) and again at an approximately 6-month follow-up.
The paper documents the control group’s size, baseline characteristics, and what activities it received during sessions.
Randomization occurred within classrooms, not at the school level.
The authors evaluate an established SRSD model developed by other researchers and report no conflicts of interest, so independence from intervention designers is reasonably documented.
From intervention start (12-week program) to follow-up assessment approximately 6 months later is roughly 9 months, meeting the 75%-of-academic-year threshold.
Both groups received the same scheduled instructional time, and the additional individualized supports (e.g., goal setting) appear to be integral to the SRSD package rather than a separable confound.
No peer-reviewed independent replication of this specific 2026 RCT was found, though an IES replication project is underway.
The standardized outcome assessments are focused on writing rather than covering all core school subjects.
Neither the paper nor identifiable follow-up publications track the cohort to graduation; follow-up is only to approximately 7th grade.
No protocol pre-registration record (registry/ID and pre-study date) was identified in the paper or via registry searching.
Purpose: Primary purpose of this study was to extend the research on the SRSD model by examining its efficacy in a randomized controlled trial within a Tier-2 framework for small groups of students at-risk for writing difficulties. Methods: 243 sixth...
Randomization occurred at the department-grade (cohort) level, meeting the class-level RCT requirement.
The study measured outcomes using official final grades from institutional examination records.
The study ran for the full Spring 2024 semester (February-May), meeting the term-duration requirement.
The control condition is clearly described as business-as-usual with unrestricted phone access and documented baseline balance.
Randomization was within institutions (departments/cohorts), not between institutions, so it is not a school-level RCT.
The authors designed, implemented, and analyzed the study themselves, with no independent evaluator described.
The study covers one semester rather than a full academic year, so the year-duration requirement is not met.
Phone boxes were installed in all classrooms and the intervention was a policy enforcement rather than added instructional resources, so inputs are balanced.
No peer-reviewed independent replication of this specific RCT was found.
The outcome reflects overall grades across (nearly) all courses, providing an all-subject style assessment rather than a single subject test.
The paper reports only semester-end outcomes and does not track students to graduation.
The paper links to a pre-registration, and the registry date is before the Spring 2024 data collection period.
Widespread smartphone bans are being implemented in classrooms worldwide, yet their causal effects on student outcomes remain unclear. In a randomized controlled trial involving nearly 17,000 students, we find that mandatory in-class phone collection led to higher grades - particularly...
Randomisation was conducted at the school level, which is stronger than and automatically satisfies the class-level requirement.
The study used the TELL, a standardised, externally validated Pearson assessment aligned to multiple states' language standards.
Outcomes were measured about 8 months after the intervention began, far exceeding one academic term.
The control group's curriculum, size, and demographic characteristics are clearly documented in the text and Table 1.
Randomisation was performed at the level of whole schools (four vs. four), satisfying the school-level RCT requirement.
Rosetta Stone employees (the intervention's designer and funder) are co-authors and remained involved throughout the study, and the "external" consultants who ran the analysis were hired directly by Rosetta Stone rather than by an independent body.
Outcomes were tracked from late August 2017 to early May 2018, covering the substantial majority of the full 2017-2018 academic year.
Both groups received the same total daily ESL instructional time; the software replaced existing curriculum content rather than adding extra resources to the treatment group.
No independent, peer-reviewed replication of this specific study was found in a citation search; the authors describe replication as a direction for future research.
Only English language proficiency (TELL) was assessed; no other main subjects (e.g., math, science) were measured, and no vocational/upper secondary exception applies.
Data collection ended with the end-of-year posttest; no tracking toward graduation was reported in the paper, and no follow-up publication by the same authors on this cohort was found via internet search.
The paper contains no reference to a pre-registered protocol, registry ID, or registration date, and none was found via internet search.
English learners (ELs) in K-12 schools must acquire English while simultaneously mastering content knowledge. Educational technology may support students' learning through the affordance of individualized language practice. The current randomized controlled trial intervention study examined the effects of Rosetta Stone...
Randomisation occurred at the school level, which is stronger than and therefore satisfies the class-level RCT requirement.
The study used the WMLS-R, a widely recognised standardised language and literacy assessment.
Outcomes were measured at the end of a full school year, far exceeding the one-term minimum.
The control group's demographics and baseline characteristics are extensively documented in Tables 2 and 3.
Randomisation was conducted at the school level, satisfying the school-level RCT requirement.
The same research team that designed the DCCS program also delivered coaching, conducted observations, and analysed the data, with no independent third-party evaluator involved.
The intervention and outcome measurement spanned a complete academic school year.
The additional professional development time is the explicit treatment variable under study, so the business-as-usual control group does not violate this criterion.
No independent replication of the DCCS professional development program by a different research team was found; a 2025 follow-up trial exists but was conducted by the same author team.
Only language and literacy outcomes were measured; no other core subjects were assessed, and no exception applies.
Tracking ended after one school year, with no follow-up of this cohort through graduation reported or found.
No pre-registration of the study protocol is referenced in the paper or found elsewhere.
Using a randomized controlled trial, we tested a new teacher professional development program for increasing the language and literacy skills of young Latino English learners with 45 teachers and 105 students in 12 elementary schools. School-based teams randomly assigned to...
Randomisation was via an individual lottery rather than by class or school, but winners and losers attend different schools, so classroom-level contamination (the concern behind this criterion) does not arise.
Outcomes were measured using OAKS, Oregon's official state-mandated standardized test, not a custom instrument.
Outcomes were measured multiple years after the intervention began (kindergarten through Grade 8/9), well beyond the one-term minimum.
Table 2 gives detailed demographic, baseline, and attrition statistics for the control group, with balance tests against the treatment group.
Randomisation was of individual student applicants for admission slots, not of whole schools to treatment or control status.
A co-author is the Portland Public Schools official who currently develops and supports the dual-language immersion programs under evaluation, compromising full independence of conduct.
Students are tracked for up to nine years after kindergarten entry, far exceeding the required 75% of one academic year.
The intervention changes the language of instruction within the standard school day rather than adding extra time or budget, so no resource imbalance arises.
An independently authored study (Morales, 2024, Educational Evaluation and Policy Analysis) uses a comparable lottery-based causal design in a different set of dual-language immersion programs and finds directionally and magnitude-consistent reading/math gains.
Social studies, a core subject heavily delivered in the partner language, was not assessed by any standardized exam, and the omission is not justified in the paper.
Tracking stops around ninth grade at most, far short of graduation, and no follow-up study extending tracking to graduation was found through internet search.
No pre-registration registry, protocol, or date is referenced anywhere in the paper, and none was found via registry search (AEA Social Science Registry, OSF).
Using data from seven cohorts of language immersion lottery applicants in a large, urban school district, we estimate the causal effects of immersion programs on students' test scores in reading, mathematics, and science and on English learners' (EL) reclassification. We...
Randomisation occurred at the school level with classes equally divided into intervention and control groups, satisfying the Class‑level RCT criterion.
The study used GL Assessment Progress Tests—standardized instruments in English, mathematics, and science—scored blind by the test publisher, fulfilling the ERCT Standard’s Exam‑based Assessment criterion.
Primary outcomes were measured 20 weeks after intervention start, exceeding a single academic term.
Control cohorts are described alongside intervention cohorts, with baseline comparisons implied.
The study’s RCT was conducted at the whole-school level: entire schools were randomly assigned to either the dialogic teaching intervention or the control condition, fulfilling the School‑level RCT criterion.
The RCT was conducted by an independent evaluation team, separate from the intervention’s designers.
The study lasted 20 weeks, well short of a full academic year.
Control group teachers did not receive the intervention’s teacher induction, training, or mentoring support.
No independent replication of this dialogic teaching trial is reported in the paper or elsewhere.
Student performance was assessed in English, mathematics and science, covering the core primary curriculum.
No follow‑up tracking to graduation is described in the study or in any subsequent publications.
No pre‑registration or protocol identifier is provided in the paper (the trial was only registered retrospectively on ISRCTN, after completion).
This paper considers the development and randomised control trial (RCT) of a dialogic teaching intervention designed to maximise the power of classroom talk to enhance students’ engagement and learning. Building on the author’s earlier work, the intervention’s pedagogical strand instantiates...
Although randomization was at the student level, the intervention is tutoring delivered to individuals/small groups, which the ERCT standard treats as an allowed exception for Criterion C.
The outcomes are measured using widely recognized standardized assessments (TPRI, STAAR, DIBELS, LEAP), satisfying the exam-based assessment requirement.
The evaluation reports outcomes over the full 2024-25 school year, which exceeds the minimum one-term follow-up requirement.
The control condition is defined as business-as-usual literacy instruction and the report provides baseline characteristics and equivalence checks for treatment vs. control.
Randomization occurred at the student level within schools, so the study does not meet the school-level RCT requirement.
The report does not explicitly document evaluator independence from the intervention provider across key evaluation steps, and it notes provider involvement in preparing the de-identified analytic dataset.
Outcomes are measured over the 2024-25 school year with a full-year tutoring implementation, meeting the year-duration threshold.
The treatment adds substantial tutoring time and resources, and this added tutoring is the intervention being tested versus business-as-usual, so the resource difference is integral to the intended treatment contrast.
No independent replication by a different research team in a peer-reviewed journal was found; the report only references a prior study by the same authors.
The study uses standardized reading assessments but does not assess all core school subjects, focusing on literacy outcomes only.
The study measures outcomes within a single school year and explicitly states that longer-term persistence requires future research; no graduation-tracking follow-up papers by the same authors were found.
The report contains no public pre-registration link/ID or registration date demonstrating protocol registration before data collection.
Air Reading partnered with the Center for Research and Reform in Education (CRRE) to conduct an evaluation of Air Reading in a project funded by Accelerate and Arnold Ventures. The trial was conducted across the 2024-25 school year in a...
The study randomizes within classrooms, but it evaluates tutoring (small-group personal instruction), which meets the tutoring exception under Criterion C.
The study measures outcomes using the NWEA MAP Math assessment, which the paper describes as a widely used standardized test.
Outcomes are measured from January 2024 (baseline) to May 2024 (endline), which spans roughly one academic term.
The control group is clearly defined and its baseline characteristics and sample size are documented in Table 1.
The study is conducted in one school and does not randomize treatment assignment at the school level.
The paper does not explicitly state that the evaluation was conducted by an independent third-party team separate from the author and implementation partners.
The evaluation runs from January 2024 to May 2024, which is substantially less than 75% of an academic year.
The intervention provides additional tutoring time and tutor labor relative to business-as-usual, but these added resources are the treatment being tested, so an unbalanced business-as- usual control is acceptable under Criterion B’s intent check.
No independent, peer-reviewed replication of this specific experiment was found during the ERCT check.
The study uses a standardized assessment (MAP) but reports outcomes only for math rather than for all main subjects.
The study does not track students through graduation, and per the ERCT dependency rule, Criterion G cannot be met because Criterion Y is not met.
The paper states the study was pre-registered (AEARCTR-0012858), but the registry page could not be accessed to verify the exact registration date relative to study start.
High-dosage tutoring has the potential to substantially raise adolescent academic achievement. However, at scale, schools may not have the financial ability to deliver small-group tutoring frequently. In this paper, I test the relative importance of group size (quality) versus tutoring...
Student-level randomization is acceptable here because the intervention is tutoring.
Outcomes were measured with standardized assessments (DIBELS-8 and i-Ready).
Outcomes were measured at end of year, well more than one term after start.
The control group is described as business-as-usual and baseline characteristics are reported.
Randomization occurred within classrooms rather than at the school level.
The paper evaluates externally developed, district-implemented programs rather than a researcher-designed intervention.
Implementation began in November, so the study does not span a full academic year from start of year.
Extra tutoring time is the treatment being tested, so a business-as-usual control is acceptable.
No independent replication of this paper's para-tutoring implementations and findings was found.
Only literacy and math outcomes are measured, not all core subjects.
The study reports end-of-year outcomes only and does not track students through graduation.
The paper claims preregistration but provides no registry link, ID, or date that can be verified.
Using embedded paraprofessionals to provide personalized instruction is a promising model for differentiating instruction within the classroom. This study examines two randomized controlled trials of paraprofessional-led tutoring in early-grade math and literacy. However, intent-to-treat (ITT) analyses revealed no overall achievement...
The unit of randomization was the classroom (teacher), meeting the class-level RCT requirement.
The primary outcomes were measured using established, standardized assessments (for example PAT-2, Woodcock-Johnson III, and PPVT-IV).
Outcomes were measured across the school year, exceeding a single academic term.
The control condition and baseline characteristics are documented, with both demographics (Table 1) and descriptions of control instruction.
Randomization was at the classroom (teacher) level rather than at the school level.
The intervention was developed and evaluated by the same research team, so the evaluation was not independent of the intervention designers.
The intervention and measurement spanned the full academic school year.
The treatment replaced part of the normal literacy block rather than adding extra student instruction time, and the study evaluates the full implementation package (curriculum plus required teacher training and coaching) against business-as-usual.
No independent replication by a different research team was identified.
Outcomes were limited to language and literacy, with no standardized measures reported for other core subjects.
The study did not track participants through graduation or an equivalent endpoint.
The paper does not report a pre-registration record or registry identifier for the trial.
The goal of the present study was to assess the effectiveness of Foundations for Literacy for deaf and hard-of-hearing (DHH) children. Forty-eight teachers in 14 states were randomly assigned to intervention or control groups. Teachers in the intervention group used...
Randomisation was student-level, but the intervention is targeted small- group instruction outside normal English lessons, fitting the tutoring exception.
Outcomes were measured using recognised standardised digital reading tests (NGRT and ART).
The outcome measurement occurred after approximately two terms (about six months), exceeding the one-term minimum.
The control group and its business-as-usual condition were described, and the paper reports baseline equivalence plus detailed counterfactual information about interventions used with control students.
Randomisation was at the individual student level, not the school level.
The intervention is attributed to Fischer Family Trust Literacy (FFTL), while the trial is authored by university researchers and funded by an independent foundation.
Outcomes were measured after approximately two terms (about six months), which is shorter than a full academic year.
The intervention adds time, training, and materials, but these resources are integral to the intervention being tested against business as usual.
I found no peer-reviewed independent replication of this specific high- school FFTL Reciprocal Reading evaluation by a different research team.
Only reading outcomes were assessed; impacts on all main subjects were not measured.
The study reports post-test at the end of the intervention and provides no evidence of tracking participants to graduation; also, criterion Y is not met so G cannot be met under the ERCT rules.
The trial was registered after enrolment had begun, so it does not meet the requirement for a prospectively pre-registered protocol.
Targeted reciprocal reading instruction can lead to improved reading attainment. Though tested in elementary schools, the technique is less studied with older students. This paper reports results from a Phase 3 definitive trial designed to detect attainment gains previously identified...
Randomisation occurred at the individual-student level within schools, but the intervention was a pulled-out, small-group (3-5 students) supplemental reading program delivered by dedicated instructors outside the regular classroom, matching the personal-teaching/tutoring exception that allows student-level randomisation.
Primary outcome measures were well-established, widely used standardized instruments (WLPB-R, CTOPP/TOPP-S, DIBELS, TOWRE), even though one supplementary experimental spelling measure was also used.
Because the stronger Year Duration criterion (Y) is met, the weaker Term Duration criterion is automatically considered met as well.
The comparison group's demographics, baseline scores, and the amount/type of any supplemental instruction they received outside the study are documented in detail across four results tables and the narrative text.
Randomisation was conducted at the individual- student level within participating schools, not at the level of entire schools.
The same research team that designed the Proactive Reading/Lectura Proactiva curricula also delivered training, oversaw fidelity, collected the outcome data, and analyzed the results, with no independent external evaluator described.
Outcomes were measured in May after an intervention and assessment window running from October to May, covering roughly 7-8 months of an approximately 9-10 month academic year (i.e., well over 75%).
The additional instructional time given to intervention students is the explicit treatment variable under investigation (comparing supplemental reading intervention versus business-as-usual core instruction), so the lack of an equivalent researcher-provided add-on for the comparison group does not violate the balance requirement.
The paper describes itself as replicating earlier Vaughn-team studies using a new, nonoverlapping sample, but this replication was conducted by the same research group rather than an independent team; an internet search for external citing/related work found no independent replication of this specific study by a different research team.
Outcomes were measured only in reading, phonological processing, and oral language; no other core subjects (e.g., mathematics, science) were assessed, and no vocational/specialised-education rationale is offered to justify this narrow focus.
Follow-up ended with the Grade 1 posttest in this paper; internet search located same-team follow-up papers extending tracking to Grade 2 and to Grade 4-5, but neither documents tracking through actual graduation/completion of the students' school stage.
No pre-registration of the study protocol, hypotheses, or analysis plan is mentioned anywhere in the paper, and no evidence of registration was found via internet search; education-research pre-registration platforms (e.g., AEA RCT Registry, OSF Registries, SREE Registry) were established years after this study was conducted.
Two studies of Grade 1 reading interventions for English-language (EL) learners at risk for reading problems were conducted. Two samples of EL students were randomly assigned to a treatment or untreated comparison group on the basis of their language of...
Randomisation was conducted at the whole-school level, with each school having only a single third-grade classroom, so school-level assignment automatically satisfies the weaker class-level requirement.
The study measured oral English proficiency using the WMLS-R, a norm-referenced, standardized instrument with documented reliability and validity, not a custom-built test.
The intervention and post-test spanned a full 25-week period, which exceeds the minimum one-term (roughly 3-4 month) requirement.
The Control/Comparison group's curriculum, instruction time, and school/student demographics are described in detail and compared against the treatment conditions.
Entire schools, not individual students or single classes within a shared school, were the unit of random assignment to Control, Treatment A, or Treatment B.
No statement documents that the study was conducted by an evaluator independent of the intervention designers; staff from the implementing Costa Rica Multilingue Foundation are themselves co-authors.
The tracked intervention and follow-up period was only 25 weeks, well under the 75%-of-an-academic-year threshold required by the Year Duration criterion.
The Control group's observed instructional time (150 minutes/week) met or exceeded that of both Treatment A (67 minutes/week) and Treatment B (127 minutes/week), so no imbalance favoring the intervention groups is present; the CALL technology is also the explicit treatment variable being tested.
No independent replication of this specific Project EILE trial by a different research team was found in the paper or via external search of citing literature.
The study measured only oral English language proficiency via WMLS-R subtests, with no assessment of other core school subjects such as mathematics or science.
Criterion Y (Year Duration) is not met, and the paper reports only first-year, 25-week findings with no follow-up located through graduation.
No statement anywhere in the paper references a pre-registration platform, registration ID, or registration date for the study protocol, and no such registration was located via external search.
This study presents first-year findings of a 25-week longitudinal project derived from a two-year longitudinal randomized trial study at the elementary school level in Costa Rica on effective computer-assisted language learning (CALL) approaches in an English as a foreign language...
Randomization occurred at the teacher level, with all of each teacher's sections assigned together, exceeding the requirement for class-level randomization.
The study administered and reported results for several widely used, standardized, norm-referenced tests alongside researcher-created measures.
Outcomes were measured from October 2008 to May 2009, roughly seven months after intervention start, well beyond one academic term.
The control group's size, demographics, baseline scores, and business-as-usual instruction are described in detail throughout the Method and Results sections.
Randomization occurred at the teacher level within each school (blocked by school), not at the level of whole schools.
The same research team that designed the ALIAS intervention also trained the program specialists, supervised data collection, and conducted the analysis; no independent evaluator is described.
The pretest-to-posttest interval (October 2008 to May 2009) covers roughly 75-85% of a typical academic year, and the paper describes the program as delivered over the course of one academic year.
Core ELA instructional time was explicitly held comparable across conditions; the additional teacher support the treatment group received is integral to implementing the specific curriculum being tested.
The paper contains no reference to an independent replication of this specific trial by another research team, and none could be located in the wider literature.
Outcomes were assessed only in vocabulary, morphology, reading comprehension, and writing (all within English language arts), with no assessment of other core subjects such as mathematics or science.
The authors explicitly state that outcomes were measured only immediately after the intervention and that longer-term follow-up was not conducted; no later graduation-tracking paper by these authors on this cohort was located.
No statement of pre-registration, a registry platform, or a registration date appears anywhere in the paper, and no matching pre-registration record was located online.
We conducted a randomized field trial to test an academic vocabulary intervention designed to bolster the language and literacy skills of linguistically diverse sixth-grade students (N = 2,082; n = 1,469 from a home where English is not the primary...
The study used a two-stage clustered design in which entire remedial-math class sections (not individual students within a class) were randomly assigned to the tutoring or control condition, satisfying class-level RCT.
Outcomes were measured with the NWEA MAP mathematics assessment, a widely used, externally validated, standardized computer adaptive test, not a custom instrument.
Outcomes were measured at the end of a full academic year, far exceeding the minimum one-term interval required.
The control group is described in detail, including its size, demographic composition, baseline scores, class size, and the specific instruction it received.
Randomisation occurred within each of the seven schools (class sections assigned to condition), not between schools, so this is a class-level rather than school-level RCT.
The same University of Oklahoma faculty team that designed and delivered the high-impact tutoring program also ran the data collection and statistical analysis, with no external or independent evaluator.
Tutoring ran three class periods per week for a full academic year, and outcomes were measured at the end of that same academic year, satisfying the year-duration requirement.
The control group received the same amount of extra instructional time as the treatment group (a full-year remedial math class); the only difference -- presence of a tutor -- is the explicit treatment variable being tested, so the design is balanced by construction.
No independent replication of this specific study or program is reported or found; the paper only references separate high-impact tutoring studies with different designs and teams.
Only mathematics achievement was measured with a standardized exam; other core subjects were not assessed, and no exception for a specialised program is stated.
Students were tracked only through the end of ninth grade, not through graduation, and the paper reports no follow-up study.
The paper contains no statement of pre-registration, registry link, or registration date anywhere in the text or references.
This study uses a randomized controlled trial designed to examine a university-led high-impact tutoring program at seven high schools. The treatment group (n = 525) participated in high-impact tutoring (i.e., groups of 2:1 or 3:1) while the control group (n...
Randomization occurred at the small-group (2-3 student) level within classrooms rather than at the class level, but the intervention is a pull-out, small-group tutoring program delivered outside the regular classroom, which falls under the ERCT exception for personal/tutoring-style teaching.
Core outcomes (receptive vocabulary, decoding, spelling) were measured with widely used, standardized norm-referenced tests, satisfying the exam-based assessment requirement even though one supplementary curriculum-based measure was researcher-developed.
Outcomes were measured about seven months after the intervention began (fall pretest to spring posttest), exceeding the minimum one-term requirement, and this is reinforced by the stronger Year Duration criterion also being met.
Although this is a two-active-treatment comparison rather than a treatment-vs-no-treatment design, the comparison (IBR) condition is extensively documented, including demographics, sample sizes, and detailed procedures.
Randomization was conducted at the small-group level within classrooms across 13-20 schools, not at the school level, so the stronger school-level RCT criterion is not satisfied.
The Connections intervention materials were designed by two of this paper's own authors (Nelson and Vadasy), who also led this evaluation, so the study was not conducted independently of the intervention designers.
Students were tracked to a first-grade follow-up roughly 14 months after the intervention began (kindergarten fall start to first-grade winter follow-up), exceeding the one-academic-year requirement.
Both experimental conditions (Connections and IBR) received identical amounts of supplemental instructional time and tutor attention, so resources were balanced between the two arms compared.
No independent replication of this specific study or the Connections intervention by a different research team was found; the only prior and follow-on studies were conducted by the same overlapping author team, including the "replication" funding itself.
The study measured only vocabulary, decoding, and spelling outcomes (all within the language arts/literacy domain); no other core subjects such as mathematics were assessed.
Tracking ended at a first-grade follow-up, well short of graduation, with no mention of continued tracking or planned follow-up studies beyond this point, and no such follow-up was located via internet search.
The paper contains no statement or reference indicating that the study protocol was pre-registered before data collection began, and no registry record was located via internet search.
A two-cohort cluster-randomized trial was conducted to estimate effects of small-group supplemental vocabulary instruction for at-risk kindergarten English learners (ELs). Connections students received explicit instruction in high-frequency decodable root words, and interactive book reading (IBR) students were taught the same...
The study randomized entire schools to treatment or control, satisfying the requirement for class‑level RCT.
The study employed official state standardized exams for ELA and math, meeting the exam‑based assessment requirement.
Outcomes were measured in subsequent grades, well after at least one full academic term had elapsed, fulfilling the term‑duration requirement.
The control group’s makeup, baseline statistics, and alternate program are clearly documented, satisfying the requirement for a documented control group.
Randomization at the school level fulfills the school‑level RCT criterion.
The study was implemented and analyzed by the same team that developed the program, so there is no independent evaluation group.
Students’ outcomes were measured over multiple grades, covering at least a full academic year of follow‑up.
The control group’s activities were far less intensive than INSIGHTS, so resource allocation was not balanced.
There is no reference to an external, independent replication of the INSIGHTS trial.
The authors measured only ELA and math; other core subjects were not assessed.
No data are reported beyond sixth grade (middle school entry), so graduation tracking is incomplete.
No pre‑registered protocol or registry reference is provided in the paper.
Social‑Emotional Learning (SEL) programs are school‑based preventive interventions that aim to improve children’s social‑emotional skills and behaviors. Although meta‑analytic research has shown that SEL programs implemented in early childhood can improve academic and behavioral outcomes in the short‑term, there is...
The study randomizes at the school level, which is stronger than class-level randomization and satisfies criterion C.
Academic outcomes were measured using standardized instruments from a national framework developed with ACER and the Afghanistan Ministry of Education.
The paper states that post-intervention measurement occurred after a minimum four-month intervention period, meeting term-duration follow-up.
The paper clearly defines the control condition as usual teaching with no intervention and presents baseline characteristics for both groups.
Schools were randomized to intervention and control conditions, meeting the school-level RCT requirement.
The authors trained and supervised intervention delivery, and the paper does not document an external independent evaluator leading the study.
Post-intervention measurement is described as occurring after about 3–4 months, which is below the ERCT requirement of at least 75% of an academic year.
The intervention adds substantial time/training/support relative to the control, but these inputs are integral to the intervention package being tested against business-as-usual schooling.
No independent peer-reviewed replication by a different research team was identified, and the paper itself only discusses replication as a future possibility.
The study reports standardized academic outcomes for reading/general knowledge and mathematics, but it does not assess all core subjects.
The study does not track participants until graduation, and ERCT rules require Y to be met for G to be met.
The trial was registered in ISRCTN in late 2024, after enrolment and the study period described in the paper, so it is not pre-registered.
Background: Conflict and crises have long-lasting and dramatic consequences on the mental health of children. We aimed to investigate the effectiveness of a psychosocial intervention on child mental health in Afghanistan. Methods: A two-arm cluster-randomized controlled trial was conducted in...
Randomization occurred at the school (above-class) level, satisfying the class-level RCT requirement.
Outcomes were measured using the standardized ITBS reading test.
Reading outcomes were assessed after 12–20 weeks (roughly one academic term).
Control group conditions and baseline characteristics were thoroughly described.
No entire school was solely a treatment or solely a control site; randomization was within schools.
An independent evaluation team (with an external data center) conducted the study, separate from the program’s creators.
Outcomes were measured only midyear (half-year), with no full-year follow-up.
The Reading Recovery group got extra daily tutoring that the control group did not receive, resulting in unbalanced time/resources.
No evidence was found of an independent replication of this Reading Recovery study by another team.
Only reading was tested; no standard exams in other core subjects were reported.
The study did not track participants through to graduation.
No pre-registration or registry listing was provided for the study.
Reading Recovery (RR) is a short-term, one-to-one intervention designed to help the lowest achieving readers in first grade. This article presents first-year results from the multisite randomized controlled trial (RCT) and implementation study under the $55 million Investing in Innovation...
Two intact classes were randomly assigned as experimental and control groups, so randomisation occurred at the class level rather than among individual students within one class.
The pre- and post-tests were taken from the writing sections of past TEM-4, a standardised national English proficiency test for English majors in Chinese universities, rather than a test specially designed for the study.
The intervention began in Week 1 and outcomes were measured with a post-test in Week 18, an interval of roughly 4.5 months that exceeds one full academic term.
The control group's size, gender composition, baseline proficiency (Gaokao and TEM-4 scores), and business-as-usual condition are all clearly documented.
Randomisation involved only two intact classes within a single independent college; no schools or institutions were randomised.
The authors designed the GDPF instruction, delivered it, collected the data, and the researcher himself served as one of the essay raters, with no independent third-party evaluation reported.
Outcomes were measured 18 weeks (about 4.5 months) after the intervention began, which is well short of 75% of a full academic year.
Both groups received identical in-class instruction, wrote the same essays on the same schedule, and the additional online dialogic peer feedback activity in the experimental group is itself the treatment variable being tested rather than a separable resource add-on.
No independent replication of this specific GDPF study exists; the paper was published in late 2024/early 2025 and an external search found no peer-reviewed replication by a different team.
Only argumentative writing in English was assessed; no other main subjects were measured and the paper offers no explicit rationale for a specialised-intervention exception.
Measurement stopped at the Week 18 post-test with no follow-up tracking of students to graduation, and criterion Y is not met, which also rules out G.
The paper contains no mention of any trial registry, pre-registration ID, or pre-registered protocol, and no registration for this study was found in an internet search.
While extensive research exists on peer feedback and its effects on writing, there are few experimental studies that rigorously investigate the effects of guided dialogic peer feedback on students' argumentative writing performance. This study, adopting a mixed-methods approach, examined the...
Randomization occurred at the teacher/classroom level (not within a single classroom by student), meeting the class-level RCT requirement.
The study used the standardized Woodcock–Johnson IV Tests of Achievement (WJ-IV) as an outcome measure.
Outcomes were measured at the beginning and end of the school year, exceeding a one-term minimum.
The BAU control condition and its participants are documented with group sizes, demographics, and descriptions of BAU instruction.
Randomization occurred at the teacher/classroom level rather than assigning entire schools to SIWI vs BAU.
The intervention developers/authors appear to have led SIWI delivery supports (PD and coaching), so the study was not independently conducted, despite limited external oversight.
Outcomes were collected at the beginning and end of the academic year, which meets the year-duration requirement, even though COVID-19 reduced instructional exposure for some participants.
SIWI teachers received substantial additional resources (paid PD, ongoing coaching, on-site visits, and provided technology) that were not matched for BAU, and these extra inputs were not framed as the treatment variable.
No independent (different-team) replication of this specific RCT was found; the paper is itself a replication of earlier SIWI work but conducted by the SIWI research team.
The study used a standardized writing/language assessment (WJ-IV), but did not assess impacts across all core school subjects.
The study reports outcomes only within an academic year and does not report tracking students until graduation.
The paper provides IRB information and a data availability DOI but does not provide a pre-registration registry/ID and pre-data- collection registration date.
This study reports findings from a nationwide replication and the second randomized controlled trial (RCT) of Strategic and Interactive Writing Instruction (SIWI), a linguistically responsive framework for teaching writing to deaf students. A total of 50 teachers and their 294...
Cohorts (cluster groups of roughly 20–30 students) were randomly assigned to treatment or control within each school and grade, satisfying class-level cluster randomization.
Outcomes were measured using NWEA MAP, a standardized assessment administered as part of the schools’ regular assessment program.
The intervention began in late January 2023 and the follow-up MAP assessment occurred in mid-May 2023, which is approximately a full academic term after the intervention start.
The control condition is described as continuing Rocketship’s regular Learning Lab reading supports, and the paper reports clear treatment/control sample sizes and baseline equivalence checks.
Randomization occurred within schools at the cohort level rather than assigning entire schools to treatment versus control.
The paper does not explicitly state that the study was conducted by an independent third party separate from the intervention’s designers/provider.
Outcomes were measured from late January 2023 to mid-May 2023, which is far less than 75% of an academic year after the intervention began.
Both groups used the same scheduled Learning Lab time, and the added tutoring/platform resources are the core treatment being tested rather than an unbalanced add-on of extra instructional time.
No independent replication of this specific BookNook Rocketship RCT by a different research team was identified, and none is reported in the paper.
The study uses a standardized reading assessment (MAP Reading) but does not assess outcomes across all core subjects.
The study does not track students through graduation, and per ERCT rules graduation tracking cannot be met when year-duration (Y) is not met.
The paper provides no registry identifier or link to a pre- registered protocol, and no external pre-registration record was found for this study.
This paper describes a 12-week cluster randomized controlled trial that examined the efficacy of BookNook, a virtual tutoring platform focused on reading. Cohorts of first- through fourth-grade students attending six Rocketship public charter schools in Northern California were randomly assigned...
The study randomized at the section (class) level within schools, which meets the requirement for class-level randomization.
The study used the SIMCE, which is the Chilean national standardized exam.
The intervention lasted approximately seven months, which exceeds the minimum one-term duration requirement.
The control group is clearly documented with demographic data and a description of their "business as usual" condition.
Randomization was performed at the section (class) level within schools, not at the school level.
The study was not conducted independently; the lead author developed the program and the author team managed the implementation.
The intervention duration was seven months, which is less than the full academic year (typically 9-10 months) required by the criterion.
The extra time and resources were an integral part of the "bundled" intervention explicitly being tested against business-as-usual, satisfying the exception for this criterion.
There is no evidence provided of an independent replication of this study by a different research team.
The study assessed Math and Language but did not assess Science, which is stated as a subject taught by the teachers.
The study tracked students only until the end of the intervention period (Grade 4), not until graduation.
There is no evidence in the text that the study protocol was pre-registered before data collection began.
This paper presents results from a randomized evaluation of a bundled program employing an external coordinator to aid 4th grade teachers with the integration of a math learning platform that partially replaced regular school math instruction in Chile. Students in...
Randomization includes student-level lotteries, but because the intervention is tutoring, the ERCT tutoring exception applies.
Outcomes are measured using established standardized assessments such as state tests, NWEA MAP, i-Ready, STAR, and PSAT/SAT.
Interventions and end-of-year outcome measurement occur over at least an academic-term scale from intervention start to testing.
The BAU control condition is defined and detailed baseline balance tables document control group characteristics.
Randomization is at student, classroom, teacher, or grade level, not at the school level.
Researchers co-designed the tutoring models with partner districts, so evaluation was not fully independent of intervention design.
Several sites implemented tutoring only in spring 2024 (or for about 12 weeks), which is shorter than a full academic year.
The study explicitly tests the effect of providing additional tutoring resources relative to business-as-usual.
No independent peer-reviewed replication of this specific PLI 2023-24 study design was identified.
Primary outcomes are standardized tests in the tutored subject, not across all core subjects.
This interim report reports end-of-year outcomes and does not track students to graduation; additionally, Y is not met.
An OSF link is provided, but the pre-registration timestamp could not be verified to precede the study start.
This report summarizes the ongoing work by the Personalized Learning Initiative (PLI) research team to understand whether and how scaling high dosage tutoring (HDT) works in the post-pandemic environment. The study involved a large-scale randomized controlled trial with eight partners...
Student-level randomization is acceptable here because the intervention is delivered as small-group tutoring (2-4 students).
The study used widely recognized standardized reading assessments (TOWRE-2, TOSREC, GMRT) with standard scores and reported reliability.
Outcomes were measured after an intervention period running from October to February, which exceeds one academic term.
The Business-as-Usual control is described and the paper reports control-group demographics and details of services received.
Randomization occurred at the student level within schools rather than at the school level.
The authors developed the Engaged Learners program and the study was researcher-implemented with coaching by the first author.
The intervention and measurement window (October to February) is under a full academic year.
The intervention adds substantial instructional time and staffing, but these added resources are integral to what the study is testing against Business-as-Usual.
No peer-reviewed independent replication by other authors was found, and the paper itself calls for future replications.
The study measured reading and attention outcomes only, not standardized outcomes across all core subjects.
The study does not track students to graduation, and Criterion G cannot be met because Criterion Y is not met.
The authors explicitly state that the study was not pre-registered.
We investigate the efficacy of a reading intervention integrated with Engaged Learners, a program that applies behavioral and cognitive principles to increase student behavioral attention and reduce distractions during instruction. Using a three-arm randomized controlled trial, we randomized 159 Grade...
Randomization is at the student level, but the intervention is an individualized AI personalized learning platform, fitting ERCT's personal teaching exception.
The study reports using a standardized test bank (LCME) with reliability and validity evidence rather than a bespoke study-created exam.
Outcomes were measured at Week 12 after the intervention began, matching a term-length (approximately 12 weeks) follow-up window.
The control condition is described in detail and baseline characteristics are reported for both groups.
Randomization occurs at the student level rather than assigning whole schools or sites.
The paper does not provide evidence that the evaluation was conducted by an independent team distinct from the intervention designer.
Despite stating a one-year study period, outcomes are reported at Week 12 rather than after a full academic year of follow-up.
The intervention intentionally adds an AI platform as the treatment, so the control remains business-as-usual by design under ERCT's resource treatment exception.
No independent replication study of this specific intervention trial was found in available sources at the time of this ERCT check.
Outcomes are limited to a single course or domain rather than standardized exams across all main subjects.
The study does not track participants until graduation and, since Year Duration is not met, Graduation Tracking cannot be met under ERCT rules.
The paper does not report a pre-registered protocol or registry entry that can be verified as registered prior to data collection.
This study aims to evaluate the comprehensive impact of an artificial intelligence (AI)-driven personalized learning platform based on the Coze platform on medical students' learning outcomes, learning satisfaction, and self-directed learning abilities. It seeks to explore its practical application value...
Randomisation occurred at the classroom (section) level within each teacher, satisfying the class-level RCT requirement.
The study used the GRADE, a widely recognised norm-referenced standardised reading test, as a formal outcome measure alongside researcher-developed assessments.
Outcomes were measured after a roughly 15-week intervention period, which is consistent with one academic term.
The control group's size, demographic makeup, baseline equivalence, and standard instruction received are documented in detail.
Randomisation occurred within teachers at the section level, not at the level of whole schools.
The intervention's designers directly trained and supervised all data collectors and observers; no independent third-party evaluator is described.
The study's intervention and outcome tracking spanned only about 15 weeks, far short of the 75%-of-a-year threshold required for Y.
Instructional time was equal across arms, and the extra materials/professional development given to treatment sections are the integral treatment variable being tested against a business-as-usual control.
No independent replication of this specific study by a different research team was found in the paper or via internet search; all related follow-up work is by overlapping members of the same author team.
Only vocabulary/academic language and science were assessed; no other core subjects (e.g. mathematics, social studies) were measured.
Criterion Y is not met, and no follow-up publication tracking this study's cohort toward graduation was found via internet search.
No mention of a pre-registered protocol, registry identifier, or registration date is found anywhere in the paper, and internet search found no registry record either.
The goal of this study was to assess the effectiveness of an intervention—Quality English and Science Teaching 2—designed to help English language learners (ELLs) and their English proficient classmates develop academic language in science, as required by the Common Core...
Randomization occurred at the classroom level, satisfying the class-level RCT requirement.
A standardized reading test (Gray Silent Reading Test) was used for outcome measurement.
Post-tests were administered after about 6–7 months of intervention, exceeding a single term.
The control group’s composition and baseline performance were clearly documented and comparable to the intervention group.
Randomization was done at the class level within schools, not at the school level (no whole-school assignment).
The researchers who developed the intervention also implemented the study (no independent evaluators were involved).
The study spanned about 7 months of one school year, with no outcomes tracked for a full year or longer.
Both groups had equal instructional time and curricular resources; ITSS replaced part of the normal class time rather than adding extra time.
The study has not been independently replicated by an unrelated research team.
Outcomes were limited to reading comprehension; no other core subjects were tested.
Participants were not tracked beyond the immediate post-test in 4th grade (no long-term follow-up through graduation).
No pre-registered study protocol was identified for this trial.
Reading comprehension is a challenge for K‑12 learners and adults. Nonfiction texts, such as expository texts that inform and explain, are particularly challenging and vital for students’ understanding because of their frequent use in formal schooling (e.g., textbooks) as well...
Two intact parallel classes (not individual students within one class) were randomly assigned to intervention and comparison conditions, satisfying class-level randomisation.
The writing tests were taken from the database of TEM-4, a nationally standardised Chinese exam for English majors with reported high validity and reliability, and were scored with the well-established Jacobs et al. rubric.
Outcomes were tracked across a full 16-week semester, with the delayed posttest administered about 14 weeks (approximately 3.5 months, one academic term) after the intervention began in week 2.
The comparison group's size, gender, age, language background, baseline scores and exact activities (identical structured task with individual planning) are documented in detail.
Randomisation involved only two intact classes within a single university; no schools or institutions were randomised.
The authors designed the intervention, administered all tests, and analysed the data themselves; independent blind raters scored the texts, but there was no independent conduct or third-party oversight of the study as a whole.
Tracking from intervention start to the final delayed posttest lasted about 14 weeks (one semester), far short of 75% of a full academic year.
Both groups received identical instruction, tasks, and time (20 minutes planning plus 40 minutes writing on the same structured task); the only difference was collaborative versus individual planning, so inputs were fully balanced.
No independent published replication of this specific study was found after a fresh internet search; related collaborative-prewriting studies are either earlier work, by the same authors reusing the same dataset, or different designs rather than replications.
Only English argumentative writing was assessed; no other main subjects were measured and no specialised-context justification per the standard's exception applies.
Measurement ended with the delayed posttest in week 16 of the semester; participants were not tracked to graduation, no follow-up publication tracking this cohort further was found, and criterion Y is not met, which per the standard's dependency rule also rules out G.
The paper contains no mention of any pre-registration or registry entry, and a renewed search of the full text and external sources found no registration record for this study.
Prior studies have reported inconsistent findings with regard to the effects of small-group student talk on developing individual students' English-as-a-foreign-language (EFL) writing ability. To further explore the question under discussion, we designed a quasi-experimental study that included a pretest, a...
Intact classes were the unit of random assignment to the three feedback conditions, satisfying class-level randomisation.
Outcomes were assessed with standard IELTS writing tasks scored via the official IELTS Writing Band Descriptors, a widely recognised standardised assessment.
The intervention ran for the language center's entire two-and-a-half-month semester with the post-test at its end, covering one full term in this context.
The comparison (teacher e-feedback) group's size, conditions, and baseline equivalence via OPT and writing pre-test are clearly documented.
Randomisation was at the class level within a single language center, not among schools or sites.
The authors designed the feedback protocols and ran and analysed the trial themselves; only essay scoring was independent and blinded, so independent conduct is not established.
The full interval from intervention start to final measurement was only about 2.5 months, well below 75% of an academic year.
All three active arms had matched class time, curriculum, and instructor, and the hybrid group's additional dual-source feedback is integral to the treatment being tested.
No independent peer-reviewed replication of this specific three-arm feedback RCT was found, and the paper itself frames such evidence as scarce.
Only EFL writing was assessed; no other subjects or language skills were measured and no explicit exception rationale is provided.
Measurement ended at the end-of-semester post-test with no follow-up tracking to graduation, no subsequent cohort-tracking publication was found, and prerequisite Y fails.
The paper provides no pre-registration reference, registry ID, or registration date, and no registration record was found in IRCT or elsewhere.
Effective feedback plays a critical role in enhancing the writing skills of English as a Foreign Language (EFL) learners. This study examines the comparative effectiveness of three feedback approaches--Teacher e-feedback, AI-based feedback, and a hybrid model--in enhancing the writing performance...
Randomization was performed at the school level, exceeding the class-level requirement.
The study used custom-designed tests rather than a recognized standardized exam.
Outcomes were measured about 15 months after the intervention start, exceeding one academic term.
The control group’s size and treatment condition are clearly described, fulfilling documentation requirements.
Entire schools, rather than individual classes, were randomized to treatment and control.
The evaluation was performed by an independent team (IDB and academic partners), distinct from the OLPC Foundation designers.
Outcomes were measured 15 months post-start, satisfying the full academic year requirement.
Additional resources (laptops and training) were the treatment variable being tested, so the control condition appropriately remained business-as-usual.
An independent replication in Uruguay has confirmed the results.
Only math and language outcomes were assessed, not all main subjects.
A long-term follow-up study tracked student outcomes through graduation, meeting this criterion.
No statement of pre-registration is provided.
This paper presents results from a large-scale randomized evaluation of the One Laptop per Child program, using data collected after 15 months of implementation in 318 primary schools in rural Peru. The program increased the ratio of computers per student...
The study utilized a clustered randomized controlled trial design randomizing at the school level, which satisfies the requirement for class-level or higher randomization.
The primary innovation outcomes rely on custom-developed measures rather than standardized exams, although standardized tests were used for secondary academic outcomes.
The intervention spanned two full academic years, significantly exceeding the one-term duration requirement.
The control group's condition (self-directed preparation) and baseline characteristics are clearly documented and compared to the treatment group.
The study randomized 80 schools to treatment or control conditions, satisfying the school-level randomization requirement.
The intervention was implemented by an independent NGO (Inqui-Lab), while the evaluation was conducted by an academic researcher from Stanford.
The study tracked students over two full academic years, exceeding the one-year duration requirement.
The intervention provided additional resources (kits, training) that were integral to the treatment being tested, while educational time was balanced across groups.
No independent replications of this specific intervention were found in peer-reviewed literature.
The study assesses Math and Science but does not assess other core subjects like Language Arts or Social Studies using standardized exams.
The study tracked students through Grade 9 but did not track them until graduation.
The study was pre-registered with the AEA Trial Registry prior to the start of data collection.
Innovation fuels long-run economic growth, yet education systems in developing countries often overlook the skills required for innovation. This paper provides the first experimental evidence that students can learn core innovation-related skills. I conduct a large-scale clustered randomized controlled trial...
Randomisation at the school level satisfies the requirement for a class-level RCT.
The assessments were study-designed instruments, not recognised standardized exams.
Measurement occurred after five terms, satisfying at least one full academic term of follow-up.
The paper provides detailed baseline characteristics and conditions for the control group in Table 1.
Entire schools, not just classes, were randomly assigned to treatment or control.
The same research team and ICS officers designed, implemented, and analyzed the intervention without independent oversight.
The study tracked outcomes for more than an academic year, satisfying the Year Duration requirement.
The additional teacher is the treatment variable, so business-as-usual resourcing in the control group is acceptable.
Multiple independent studies have replicated this finding.
Only math and reading outcomes were measured, failing to cover all core subjects.
No follow-up through to primary school graduation is reported.
The study was not pre-registered before data collection.
Some education policymakers focus on bringing down pupil–teacher ratios. Others argue that resources will have limited impact without systematic reforms to education governance, teacher incentives, and pedagogy. We examine a program under which school committees at randomly selected Kenyan schools...
Randomization was primarily at the student level (not whole classes or schools), so the class-level randomization requirement is not satisfied.
The primary achievement outcome uses California’s SBAC state math assessment, a standardized exam.
Outcomes were measured long after the intervention started (e.g., grade 11 SBAC after starting in grade 9), exceeding a term-long follow-up.
The control condition is clearly described (business-as-usual placement into Algebra Readiness or standard Algebra tracks) and the paper provides baseline/control-group descriptive detail.
Randomization occurred within campuses and often at the student level, not by assigning entire schools to treatment/control.
The paper describes a district-led reform evaluated using district administrative data and administrator-run assignment, indicating the evaluators were not the intervention delivery team.
The study tracks outcomes from grade 9 through grade 12, which is longer than 75% of an academic year.
The treatment includes substantial added teacher supports and planning resources that are explicitly integral to the tested intervention package, so the imbalance is by design.
No independent replication by a different research team in another context was found in the paper or via internet search.
The study’s standardized exam reporting is not a comprehensive all-core-subject exam set (it centers on math; other subjects are not uniformly covered with standardized exams).
The study follows students through grade 12 and includes an on-time graduation indicator, satisfying graduation tracking.
The paper mentions preregistered questions but does not provide a registry link/ID and no verifiable pre-data-collection registration timing was found.
This random-assignment partnership study examined an innovative district-level reform—the Algebra I Initiative—that placed ninth grade students with prior math scores below grade level into Algebra I classes coupled with teacher training instead of a remedial pre-algebra class. We found that...
Schools (a stronger unit than classes) were randomized, meeting the class-level-or-stronger RCT requirement.
The outcomes are teacher belief measures (questionnaires), not standardized exam-based student achievement assessments.
The first post-intervention measurement in spring 2014 occurs well after the summer 2013 PD began, exceeding one academic term.
The comparison condition is described and both groups’ baseline characteristics and sizes are reported in tables.
Schools were randomized to the CGI or comparison condition, satisfying the school-level RCT requirement.
The authors state they neither developed nor delivered the CGI PD program and acted as a third-party evaluation team.
Measurement in spring 2014 after a summer 2013 start spans most of an academic year, meeting the 75% year-duration threshold.
The added PD time and supports are integral to the CGI PD intervention being tested against business-as-usual.
No independent replication by a different author team of this specific school-randomized CGI-PD-on-beliefs trial was found.
Because standardized exam-based outcomes are not used (E is not met), the all-subject standardized exam requirement is not met.
Neither this paper nor identified follow-up materials track any participant cohort through a graduation endpoint.
No prospective pre-registration is documented in the paper, and the identified registry entry is explicitly retrospective.
Teacher beliefs about mathematics teaching and learning are thought to exert a major influence on instructional practice and student learning. More than 200 teachers in 22 schools participated in a randomized controlled trial of the first two years of a...
Randomization occurred at the school level, which meets or exceeds class-level randomization.
The outcomes are violence and mental-health measures (e.g., SDQ), not standardized academic exams.
The intervention and follow-up span far longer than an academic term (multiple grades plus adult follow-up at age 34).
The control group is clearly described, including its size and that it received no intervention while being followed over time.
Schools (paired sets of schools) were randomized to intervention versus control, satisfying school-level randomization.
The paper discloses that key investigators are developers of the program/curricula, without clear documentation of an independent evaluation lead.
The intervention runs from Grade 1 through Grade 10 and includes long-term follow-up, far exceeding 75% of an academic year.
The intervention adds substantial time and resources, and these added inputs are integral to the Fast Track treatment package being tested against a no-intervention control.
No peer-reviewed independent replication by a different research team in a different context was identified.
Because standardized educational exams are not used (criterion E is not met), all-subject standardized exam coverage is also not met.
The Fast Track cohort is tracked through Grade 12 and into adulthood, which exceeds the graduation horizon.
The Fast Track study began in 1991, but the trial record was first submitted in 2012, so it was not pre-registered.
Domestic violence mechanisms are frequently transmitted across generations, representing a global issue demanding particular attention. This study investigates the intergenerational transmission of intimate partner violence (IPV) and parent-to-child violence (PCV) and whether participating in a multilevel preventive intervention (Fast Track)...
Randomization occurred at the school level, which meets (and exceeds) the class-level RCT requirement.
Academic outcomes were based on teacher assessments of meeting expectations, not standardized exam scores.
Outcomes were measured at least through the end of the school year (and beyond), meeting the term-duration requirement.
The wait-list control group is clearly described with sample sizes, exposure status, and baseline covariates.
The unit of randomization was the school, satisfying the school-level RCT requirement.
The RCT is described as designed and led by the government of Manitoba rather than the intervention developer.
Participants were followed from Grade 1 through Grade 9, far exceeding 75% of an academic year.
The added training/materials are integral to delivering PAX-GBG, and classroom delivery occurs during regular school activities rather than adding extra instructional time.
No independent peer-reviewed replication of this specific Manitoba First Nations administrative-data clustered RCT was found.
Criterion E is not met, so All-subject Exams cannot be met; additionally, outcomes are not standardized exams across all core subjects.
The paper reports follow-up through Grade 9 rather than tracking the cohort through graduation, and no later cohort-follow-up paper to graduation was found online.
No pre-registration registry entry (ID and prospective registration date) for this trial was found or reported.
Purpose PAX Good Behaviour Game (PAX-GBG), a school-based mental health promotion approach, has been shown to improve children’s mental health and academic outcomes. Given that these effects have yet to be shown in Indigenous populations, a partnership with First Nations...
Randomization occurred at the kindergarten-center (cluster) level, which meets or exceeds the class-level randomization requirement.
Outcomes were measured via teacher-report questionnaires rather than standardized exam-based assessments.
Outcomes were measured at least six months after baseline, which exceeds one academic term from intervention commencement to the latest follow-up measurement.
The paper documents the control condition and provides control group sample size and demographics.
Entire kindergarten centers (sites) were randomized, satisfying the school-level RCT requirement.
The paper does not clearly document that the evaluation was conducted independently from the intervention's designers.
The study includes measurements spanning baseline and two later timepoints totaling about 12 months, meeting the year-duration requirement.
Although the intervention required additional teacher training and materials, these appear integral to the intervention package and child sessions were embedded in the regular schedule.
No peer-reviewed independent replication of this specific trial's findings was found or documented.
Because Criterion E is not met, the all-subject standardized exam requirement cannot be met.
The study follows children only through the transition into the first year of school and does not track participants until graduation.
The trial was prospectively registered on ANZCTR on 30/09/2019, before child recruitment at the start of 2020.
Active music and movement engagement has been widely integrated in human socialization across history and cultures, and is particularly prevalent in early childhood play and learning. For clinical populations, music therapy is known to support social skills and wellbeing for...
The study randomized individual students via admissions lotteries, not classes or schools.
Outcomes are measured via course take-up and course passing, not via a standardized exam score.
Outcomes are measured across multiple grade levels (9th-11th), exceeding the one-term follow-up requirement.
The paper documents the control group’s composition and baseline characteristics using a detailed characteristics table and narrative.
Randomization is done via student admission lotteries, not by randomizing schools.
The intervention model is supported by a separate organization, while the study is conducted by researchers affiliated with research institutions.
The paper reports outcomes spanning grades 9-11, satisfying a one-year duration requirement.
The intervention is explicitly described as a comprehensive reform model whose supports and resource-intensive elements are integral to what is being tested.
Independent researchers (AIR) report a follow-up study of Early Colleges based on admission lotteries, providing replication evidence for the model.
The study reports mathematics course outcomes only and does not use standardized exam outcomes across all core subjects.
Follow-up publications by the same research program report high school graduation outcomes and longer-term outcomes after high school.
No public pre-registration record (with a registration date prior to study start) is identified in the paper or via registry searches.
This mixed methods experimental study examined the impacts of the Early College High School model on students' college readiness in mathematics measured by their success in college preparatory mathematics courses in the 9th through 11th grades, and disaggregated for academically...
The original RCT randomized at the school level, which satisfies the class-level (or stronger) RCT requirement.
The long-term follow-up cognitive outcome uses a custom rapid math test rather than a standardized exam.
The intervention ran for eight months, exceeding a full academic term.
The paper documents the control group with baseline and follow-up descriptive statistics in Table 1.
Randomization is at the school level with 34 schools as clusters.
The evaluation team is not the Kumon organization and the paper declares no conflict of interest.
Outcomes are measured in a follow-up conducted six years after the original RCT period.
The intervention adds time and materials, but these added resources are integral to the treatment being evaluated.
No independent replication by a different research team was found for this specific Kumon RCT.
Only mathematics was assessed (and not via standardized exams), so the study does not provide all-subject standardized exam outcomes.
The paper does not show systematic tracking of participants through graduation for the full cohort.
The paper cites an AEA RCT Registry ID but does not provide (and we could not verify) a registration date before the study start.
The COVID-19 pandemic and associated school closures exacerbated the global learning crisis, especially for children in developing countries. Teaching at the right level is gaining greater importance in the policy arena as a means to recover learning loss. This study...
The study randomized at the school level, satisfying the ERCT class-level-or-higher randomization requirement.
Outcomes rely on internal school grades and questionnaires, not a standardized exam-based assessment.
The interval from intervention start after the January pre-test to the May/June post-test exceeds one academic term.
The paper documents the PAU control group size, characteristics, and what support it could receive.
Schools were the unit of randomization, meeting the school-level RCT criterion.
The intervention was evaluated with substantial involvement of the intervention developers and author-led trainer training.
Outcomes were tracked from the December/January pre-test to an October/November follow-up, spanning roughly 9-10 months.
The extra time and staffing are integral to the intervention being tested (PLOS-extra), so PAU as the control is acceptable under the ERCT criterion B exception.
No independent replication of this specific PLOS-extra effectiveness trial was located.
Criterion A is not met because criterion E (standardized exam-based assessment) is not met.
The trial followed students for six months, not through graduation, and no graduation follow-up paper was located.
The paper reports prospective preregistration in REES before data collection, but the registry entry date could not be independently verified without login access.
In secondary education, many students have difficulties planning their schoolwork. These difficulties may not only lead to short-term consequences such as lower grades, but also to long-term psychosocial, professional and financial challenges. To support students with planning problems, we developed...
Randomization was at the individual child level (lottery seats), not at the class or school level, and no tutoring-style exception applies.
Outcomes were measured using widely used standardized assessments (e.g., Woodcock-Johnson, HTKS, digit span), not custom tests created for the study.
The study tracked outcomes from baseline in fall 2021 through spring 2024, far exceeding one academic term.
The control group is clearly defined as lottery non-winners, with sample sizes and alternative preschool enrollment described.
Randomization occurred within lotteries at the child level rather than by random assignment of schools to conditions.
The intervention studied was a business-as-usual Montessori program model not designed by the research team.
Outcomes were tracked from fall 2021 through spring 2024, spanning multiple academic years.
The study explicitly evaluates the real-world Montessori program package (including its resource structure) against typical alternatives, making resource differences part of the treatment definition.
The paper reports that key findings replicate across multiple Montessori preschool RCTs, including at least one independent RCT in another context.
The study does not assess effects across all core school subjects via standardized exam batteries; it reports a selected set of academic and nonacademic outcomes.
Outcomes are tracked only through the end of kindergarten, and the authors explicitly note that longer-run impacts are unknown.
The paper reports registration on REES but does not provide a registration date relative to the study start, so pre-registration before data collection cannot be verified here.
The study uses competitive admission lotteries at 24 oversubscribed U.S. public Montessori schools to estimate impacts of being offered a Montessori PK3 seat on end-of-kindergarten outcomes. The authors report positive impacts on reading and several cognitive outcomes, plus a cost...
Schools (matched sets of schools) were randomly assigned, so the unit of randomization was at least class-level and in fact school-level.
Outcomes were measured with mental health instruments (e.g., SDQ), not standardized exam-based academic assessments.
The intervention began in Grade 1 and outcomes used in this paper were assessed decades later (age 34), exceeding one term.
The control group is described as receiving no intervention and the paper reports control-group sample sizes and pre-treatment comparisons.
Randomization/assignment occurred at the school (matched school set) level.
The paper discloses that key investigators are developers of the Fast Track curriculum/PATHS curriculum, and it does not clearly document independent third-party evaluation.
Outcomes considered in this paper were assessed many years after intervention start, far exceeding 75% of an academic year.
The intervention provided substantial additional programming, but those resources are integral to the Fast Track treatment being tested versus a no-intervention/business-as-usual control.
No peer-reviewed independent replication by a different research team (with clearly documented reproduction of this study) was identified.
Because exam-based academic assessment (Criterion E) is not met, the all-subject standardized-exams requirement is also not met.
Although the target paper focuses on violence/mental health outcomes, a follow-up publication on the same Fast Track RCT explicitly ascertained whether participants graduated from high school or received a GED.
The ClinicalTrials.gov registration was first submitted in 2012, while the study start date is in 1991, so the protocol was not pre-registered before the study began.
Background: Domestic violence mechanisms are frequently transmitted across generations, representing a global issue demanding particular attention. This study investigates the intergenerational transmission of intimate partner violence (IPV) and parent-to-child violence (PCV) and whether participating in a multilevel preventive intervention (Fast...
Random assignment was at the individual applicant (family/ student) level via lotteries, not by whole class or whole school.
ELL designation is based on the standardized WIDA-ACCESS Placement Test (W-APT) with nationally normed thresholds.
Outcomes were tracked from pre-K into kindergarten and through grades 1–3, far exceeding a one-term minimum.
The control condition (half-day, business as usual) is clearly defined and baseline/outcome descriptives are reported by treatment assignment in multiple tables.
Randomization was conducted among applicants within school-site lotteries, not by randomly assigning whole schools to treatment or control.
The full-day expansion was designed and implemented by the district, while the evaluation activities (e.g., baseline assessments and analysis) were carried out by a research team rather than the intervention designers.
Outcomes were followed from pre-K entry through grade 3, which spans multiple academic years and exceeds the 75% of a year requirement.
Although full-day pre-K provides substantially more time than half-day, the paper explicitly defines this time/dosage increase as the treatment contrast and states content was the same aside from time.
No peer-reviewed, independent replication of this specific study’s key outcome finding (K–3 ELL designation reductions) was identified.
Although standardized language assessments underlie ELL designation, the study does not assess impacts across all core academic subjects using standardized exams.
The paper reports follow-up only through grade 3, and no follow-up paper tracking this cohort through graduation was identified.
The paper provides a data/code repository link but does not provide a protocol preregistration entry (with a verifiable pre-data-collection date) for this study.
This study estimates the causal effect of randomized offers of full-day versus half-day pre-K on students’ likelihood of having English language learner (ELL) designations in early elementary grades. We leverage a randomized, controlled trial in a Colorado district serving primarily...
Pupils rather than whole classes were randomised.
A validated standardised exam (PhAB‑2) was used.
Post‑test occurred four months after the December start.
Control demographics and baseline scores are clearly provided.
Randomisation took place within, not between, schools.
Researchers were not affiliated with the program’s developer.
Only six months of data – under one school year.
Extra resources constituted the treatment itself, making balance proper.
The study’s findings were later replicated by an independent team.
The study measured literacy only, not all subjects.
Follow‑up ended two months after the block, not at graduation.
The paper provides no evidence of pre‑registration.
Background. Many school‑based interventions are delivered without evidence of effectiveness. Aims. This study evaluated the Lexia Reading Core5 program with 4‑ to 6‑year‑olds in Northern Ireland. Sample. One hundred and twenty‑six pupils were screened; ninety‑eight below‑average readers were randomised to an 8‑week block of...
Schools were the unit of randomisation, meeting (and exceeding) the class-level RCT requirement.
Outcomes relied on adapted self-report items and a bespoke performance test rather than widely recognised standardised exams.
Outcomes were measured again at week 16 (three months after the intervention), which is roughly one academic term after baseline.
The control group is described as curriculum-as-usual and baseline characteristics are reported by condition.
Schools were randomised to condition, satisfying the school-level RCT criterion.
The trial protocol explicitly distinguishes the Guardian Foundation as developer and universities as evaluators and states the trial was conducted by an independent academic team.
The longest follow-up described is week 16 (around four months), which is far short of 75% of an academic year.
The intervention provides additional structured learning and supports, and these added inputs are integral to the NewsWise treatment package being evaluated.
No independent, peer-reviewed replication of this specific NewsWise RCT was identified.
Because standardised exams were not used (E not met), the all-subject standardised exam requirement is not met.
The study does not track pupils through graduation, and because Y is not met, G cannot be met under the ERCT rules.
The study is registered (ISRCTN13350949), but the registry shows registration occurred after the registry’s date of first enrolment.
Developing children’s news literacy skills to critically engage with news is crucial to their development as informed and engaged citizens. Yet, primary school children remain largely overlooked by research and practice, which have focused primarily on older cohorts. This article...
Student-level randomization is acceptable because the intervention is one-to-one tutoring.
Outcomes rely on self-reported grades rather than standardized exam-based assessments.
The first follow-up occurs about six months after randomization, exceeding one academic term.
The paper clearly describes the control condition and reports baseline balance between groups.
Randomization occurred at the student level, not at the school level.
The tutoring program is run by an external nonprofit, while the evaluation is conducted by academic researchers.
Outcomes are tracked from early 2022 to late 2023, exceeding one academic year.
Any additional resources (free tutoring access) are the treatment variable, so a business-as-usual control is acceptable under ERCT.
No independent replication of this specific Lern-Fair RCT by another team was found.
Criterion E is not met, so Criterion A is automatically not met.
The study follows participants for about 18 months, not until graduation, and no follow-up graduation-tracking paper was found.
The study cites an AEA RCT Registry ID but no publicly accessible record with the registration date could be retrieved to verify pre-registration timing.
Tutoring programs for low-performing students, delivered in-person or online, effectively enhance school performance, yet their medium- and longer-term impacts on labor market outcomes remain less understood. To address this gap, we conduct a randomized controlled trial with 839 secondary school...
Student-level random assignment is clearly described.
Outcomes rely on self-reported grades rather than standardized exams.
Follow-up measurement occurs after a substantial interval post- intervention start.
Treatment and control conditions are clearly described and baseline balance is shown.
Randomization is not conducted at the school level.
Key outcomes are self-reported, not independently assessed.
Outcomes span more than one academic year, including a second follow-up in late 2023.
Extra instructional time is the intervention itself and is explicitly tested.
No independent replications were found in available sources.
Not applicable because exam-based assessment is not used and outcomes are not across all core exams.
No evidence of tracking outcomes until graduation for the full cohort was found.
The study reports preregistration in the AEA RCT Registry and analyzes outcomes as registered.
Tutoring programs for low-performing students, delivered in-person or online, effectively enhance school performance, yet their medium- and longer-term impacts on labor market outcomes remain less understood. To address this gap, we conduct a randomized controlled trial with 839 secondary school...
The trial randomized whole Early Years Settings, which satisfies the class-level (or stronger) randomization requirement.
The study used NRDLS, described as a standardized and validated assessment, as the primary outcome measure.
The study included follow-up about 9 weeks after an 8-week intervention, providing roughly term-long tracking from start to T3.
The study removed the treatment-as-usual control arm, leaving no business-as-usual control group.
Entire Early Years Settings were randomized, meeting the school-level RCT requirement.
Intervention developers (paper authors) trained and supervised delivery, so the evaluation was not independent.
Outcomes were tracked only to a 9-week post-test follow-up, far short of a full academic year.
The two arms were explicitly matched on dosage/delivery and both included comparable homework resources, balancing time and inputs.
No peer-reviewed independent replication of the 2025 BEST vs A-DLS trial was found.
The study measured language and communication only, not standardized outcomes across all main subjects/domains.
The study followed children only for weeks to a few months post- intervention, with no tracking to graduation.
The ISRCTN registry record shows registration before first enrolment, indicating prospective pre-registration.
Children's language abilities set the stage for their education, psychosocial development and life chances across the life course. Aims: To compare the efficacy of two preschool language interventions delivered with low dosages in early years settings (EYS): Building Early Sentences...
The study randomized assignment at the school level, which satisfies and exceeds the class-level requirement.
The study employed custom-developed 21-item assessments aligned with the curriculum rather than widely recognized standardized exams.
The study tracked outcomes from August 2022 to May 2024, covering almost two full academic years.
The control group's baseline characteristics, including demographics and test scores, are fully documented and compared in Table 1.
The study randomized 42 schools into treatment and control groups, satisfying the school-level randomization criterion.
The study was conducted by authors affiliated with the Asian Development Bank, which also funded and supported the implementation of the intervention.
The intervention and data collection spanned two full academic years, exceeding the one-year requirement.
The study explicitly tests the provision of hardware and software resources (tablets and modules) as the primary treatment, justifying the resource difference with the control group.
This is an original study and no independent replication of this specific intervention is reported.
The study measured only Math and English outcomes, excluding Science which was part of the intervention content, and failed the standardized exam prerequisite.
The study stopped tracking students once they graduated and did not collect or report graduation data.
The paper does not provide any reference to a pre-registered study protocol or registry ID.
Although Asian economies have increased access to education, students' learning often trails grade level expectations. In the Philippines, learning worsened through prolonged classroom closure during the coronavirus disease (COVID-19) pandemic. Together with the Department of Education, we conducted a 42-school...
Randomization was conducted at the individual student level (blocked by school and grade), not at the class or school level, and the small-group format does not qualify for the personal-tutoring exception.
The study measured reading outcomes using TOWRE, GRADE, and Michigan's MEAP, all of which are widely recognized standardized assessments, not custom tests.
Outcomes were measured after roughly a full academic year (2010-11) of intervention, well beyond the one-term minimum required by this weaker criterion.
Table 1 provides detailed demographic and baseline reading/motivation data for the control group, including sample sizes and equivalence tests.
The study randomized individual students (blocked by school and grade), not entire schools, so the stronger school-level RCT criterion is not satisfied.
The evaluation was conducted by SRI International researchers, independent of the University of Kansas team that developed the Fusion Reading curriculum.
The reported outcomes reflect roughly a full academic year (2010-11) of intervention, satisfying the 75%-of-a-year threshold.
The extra Fusion class period is the vehicle for the strategy-instruction content being tested, and the control group occupied an equivalent scheduled activity slot ("business-as-usual"), so time allocation is reasonably balanced under the criterion B decision tree.
No independent replication of this specific study by a different research team was found in the paper or in external literature searches.
The study measured only reading-related outcomes (TOWRE, GRADE, MEAP reading, CAIMI reading motivation); no other core subjects were assessed.
The study reports only first-year outcomes with no tracking to graduation, and no follow-up publications tracking this cohort were found via internet search.
No pre-registration statement, registry ID, or registration date is provided in the paper, and none was found via external search of trial registries.
This study estimates the effect of one year of Fusion Reading implementation, a multistrategy supplemental reading intervention for struggling adolescent readers in grades 6-10. The authors conducted a randomized controlled trial in seven Michigan middle and high schools. Students in...
The study randomized whole schools (not individual classes or students) to conditions, which is stronger than class-level randomization and therefore automatically satisfies this criterion.
The paper explicitly states the assessments were researcher-designed ("designed to evaluate students' language and literacy abilities") and "adjusted each year," only loosely modeled on "EGRA-type tasks" -- this is a study-specific custom instrument, not administration of an established, standardized, validated exam.
Outcomes were measured roughly three years after the intervention began, far exceeding the one-term minimum; this is subsumed by the stronger Y (Year Duration) criterion being met.
The control group's business-as-usual PD conditions, baseline balance, and characteristics relative to the treatment schools are documented in detail.
Randomization was conducted at the level of the school, with 180 whole schools allocated across the three arms, directly satisfying the school-level RCT requirement.
The same research team that helped design and implement the intervention with the Department of Basic Education also conducted the data collection and formal analysis, with no statement of an independent third-party evaluator overseeing the study.
The intervention and its measured outcomes spanned three full academic years, far exceeding the 75%-of-a-year minimum.
The additional resources (coaching, lesson plans, extra training days, tablets) are the explicit treatment variable being tested (on-site vs. virtual coaching, each vs. business-as-usual), so the control group appropriately receives standard PD support rather than matched extra resources.
Neither the paper nor further internet search surfaced an independent replication of this specific virtual-vs-on-site coaching comparison by a different research team in a different context.
Criterion E (Exam-based Assessment) is not met, which per the ERCT specification means criterion A cannot be met either; additionally, the mathematics assessment is described only in passing with no detail confirming it is a standardized instrument.
Tracking stopped at the end of grade three (November 2019); there is no indication that students were followed through to graduation from primary or secondary school, and the identified follow-up paper addresses teacher-effect persistence with new cohorts rather than tracking this cohort to graduation.
The AEA RCT Registry record for this trial (10.1257/rct.5148-1.0) shows an initial registration date of December 13, 2019 -- almost three years after the intervention began (March 2017) and after the three-year intervention/data-collection period had already concluded (November 2019) -- so the protocol was not pre-registered before data collection began.
Background: Information Communication Technology (ICT) holds the promise of enabling low-cost teacher professional development at scale. An expert coach, for example, could reach far more teachers virtually, thus reducing salary and transport costs. But the benefits of in-person interaction --...
Randomisation was conducted at the individual student level within schools rather than at the class level, leading to potential contamination across students in the same class.
The primary outcome is based on post‑intervention grade point averages from school records, not a standardized exam‑based assessment.
Outcomes (end‑of‑year GPAs) were measured at the end of the academic year, at least one full term after the intervention.
The control condition and its baseline data are clearly described, including content, fidelity, and demographics.
Randomisation occurred at the student level within schools, not at the school level.
Independent professional research companies conducted data collection and processing, separate from the intervention designers.
Outcomes were tracked through the end of ninth grade, covering a full academic year.
Both intervention and control groups received equivalent session time and attention, balancing educational inputs.
No independent replication of this national study by a different team is reported.
No standardized exam-based assessments across all core subjects; the study relies on administrative GPAs.
Participants were only tracked through ninth grade; no graduation tracking is reported.
The analysis plan and moderation hypotheses were pre-registered on OSF prior to data analysis.
A global priority for the behavioural sciences is to develop cost-effective, scalable interventions that could improve the academic outcomes of adolescents at a population level, but no such interventions have so far been evaluated in a population-generalizable sample. Here we...
Randomisation was at the individual household level, but the intervention is one-on-one telementoring/tutoring, so the personal-teaching exception applies and student-level RCT is acceptable.
Outcomes were measured with tests created by the researchers to mirror textbook content, not with a widely recognised standardised exam.
The intervention began in early September 2020 and outcomes were measured in January 2021 (about 4.5 months later) and again in December 2021, exceeding one full academic term.
The control group is thoroughly documented with baseline demographics, baseline test scores, balance tests and a clear statement that it received no support or alternative learning opportunities.
Randomisation was performed at the individual household level, not at the school or institutional level.
The co-authors designed the intervention, trained the mentors, and led the implementation and analysis themselves; blinded enumerators help data quality but there was no independent evaluation body.
Outcomes were re-measured in December 2021, about 15 months after the intervention began in early September 2020, exceeding 75% of an academic year of tracking.
The additional tutoring/mentoring time is itself the treatment variable being tested against a business-as-usual control during school closures, so the unbalanced control is by design.
No independent team has replicated this specific telementoring study in another context in a peer-reviewed journal; the paper's own reproducibility check is a code/data verification, not a field replication, and related phone-tutoring trials are distinct studies rather than replications.
Although all four core subjects were assessed, the assessments were custom researcher-made tests, so the prerequisite criterion E fails and criterion A fails with it.
Participants in grades 1-3 were tracked for only about one year after the intervention, not until graduation from primary school.
The trial was registered at the AEA RCT registry (AEARCTR-0006395) on 21 September 2020, confirmed directly against the registry record, before endline outcome data collection began in January 2021, so the criterion is met.
Using a randomised experiment in 200 Bangladeshi villages, we evaluate the impact of an over-the-phone learning support intervention (telementoring) among primary school children and their mothers during Covid-19 school closures. Post-intervention, treated children scored 35% higher on a standardised test,...
Randomisation was conducted at the school level (127 schools randomly assigned), which is stronger than class-level and therefore satisfies the class-level RCT criterion.
The English test was constructed, validated and piloted by the research team for the study rather than being a widely recognised, standardised national exam.
Outcomes were measured about nine months after the intervention began (September 2013 baseline to June 2014 endline), which exceeds one full academic term.
The control group's size, conditions and baseline characteristics are documented in detail, with balance tests across demographic and baseline performance variables.
Entire schools were the unit of randomisation in both studies, satisfying the school-level RCT criterion.
The same research team designed the CAL/CAI programs, software, protocols and training and also conducted the surveys and analysis, with no independent third-party evaluator.
The program and outcome tracking spanned one full academic year (September 2013 to June 2014), a little over nine months.
The CAL/CAI sessions ran during existing computer class time against business-as-usual control schools with the same computer facilities, and the software, protocol, training and teacher stipends were integral parts of the treatment being tested.
No independent replication of this specific study by a different research team was found; a citation search turned up only two unrelated CAL/CAI evaluations in India by other teams and a methods paper, none of which replicate this Qinghai school-level CAL/CAI implementer study.
Only English outcomes were measured with a researcher-constructed test, so neither all main subjects nor the prerequisite standardised-exam criterion E is satisfied.
Measurement stopped at the June 2014 endline when students were in grades four and five, with no tracking of the cohort through primary school graduation, and no follow-up publications tracking this cohort were located.
The report contains no mention of a trial registry, registration ID or pre-registered protocol with a date before data collection, and no registry entry for this study was found.
Integrating Information and Communication Technology (ICT) is considered a promising educational input to help disadvantaged students. The overall goal of this paper is twofold: (1) to evaluate the impact of a government-implemented computer-assisted learning (CAL) program by comparing it to...
Teachers (and thus their entire classes) were the unit of random allocation, satisfying class-level randomisation.
The student outcome used a DIME instrument specifically designed for the study, not a widely recognised standardised exam.
Student outcomes were measured about seven and a half months after exposure began, exceeding one academic term.
The control group's size, demographics and baseline proficiency are documented in detail in Tables 2 and 4.
Randomisation was at the individual teacher level, not at the school level.
The IDB evaluators were independent of the Worldfund/Dartmouth team that designed and delivered the intervention.
The average seven-and-a-half-month tracking exceeds 75% of the 40-week academic year defined in the paper.
Teacher training is the explicit treatment variable tested against a business-as-usual control, with students in both arms receiving equivalent class time.
No independent peer-reviewed replication of this specific teacher-training RCT is reported or found.
Only English was assessed, no other core subjects, and criterion E (standardised exam) is not met.
Students were measured only at the end of the school year, with no tracking to graduation, and no follow-up study tracking the cohort was found.
The paper contains no reference to any pre-registration registry, ID, or pre-registered analysis plan, and no such record was found through internet search.
In-service teacher training aims to improve the supply of public education. A randomized experiment was conducted in Mexico to test whether teacher training could increase teacher efficiency in public secondary schools. After seven and a half months of exposure to...
Randomization was carried out at the school level across 29 schools, which satisfies the class-level requirement.
Outcomes were measured with custom research-team items rather than a recognised standardised exam.
The interval from intervention start to the analysed one-year follow-up far exceeds one academic term.
The control group's size and baseline demographic characteristics are documented in Table 1.
Entire schools were the unit of randomisation across 29 vocational high schools.
The same research team designed, implemented, collected, and analysed the trial, with no independent external evaluator.
Outcomes were tracked to a one-year follow-up, covering a full academic year after the intervention began.
The additional educational time and content is the CSE program itself, the explicit treatment variable tested against business-as-usual controls.
No independent replication of this specific study has been reported.
Criterion E is not met, and only SRH outcomes (not all core school subjects) were assessed.
Tracking stopped at a one-year follow-up rather than continuing to graduation.
The parent trial was registered, but this gender-focused secondary analysis was explicitly not pre-registered.
Background: Understanding gender differences in adolescent sexual and reproductive health (SRH) has implications beyond immediate health outcomes, potentially shaping future gender relations, relationship dynamics, and demographic patterns. This study examines patterns of gender divergence in SRH dimensions among adolescents and...
Randomization was conducted within each classroom at the individual-student level, not at the class or school level, and the group intervention is not a personal tutoring exception.
The study measured outcomes with several widely recognized standardized tests (WISC-IV, KTEA-3, YARC, TOWRE-2, WIAT-III) adapted to Romanian, not solely custom-made instruments.
The intervention ran for more than 12 months, and outcomes were measured immediately after and 8 months later, far exceeding one academic term.
The control group's size, ethnic composition, baseline scores, demographics, and business-as-usual condition are documented in the text and in Table 1.
Randomization occurred within classrooms at the student level, not among whole schools, so school-level randomization was not performed.
The intervention was designed by the authors, who also performed the statistical analysis and drew the conclusions; only the test administration was carried out by independent, blinded assessors.
The intervention lasted more than 12 months, with outcomes measured immediately after and 8 months later, exceeding 75% of an academic year.
Both groups received small-group language instruction of comparable format and time, with the study deliberately matching group size so the manipulated variable was the type of instruction rather than added time or budget.
This is described as the first such study in a Romanian-speaking, predominantly Roma sample, and no independent replication of this specific trial exists.
Only language and literacy outcomes were assessed; performance in other core subjects such as mathematics or science was not measured.
Tracking stopped 8 months after the intervention (5th grade), with no follow-up through primary or later graduation and no follow-up publication found.
The study was pre-registered in the Registry of Efficacy and Effectiveness Studies before data collection, and the analyses were conducted in line with the pre-registration.
This pre-registered randomized controlled trial tests the impact of a language comprehension intervention on the language and literacy of 448 upper elementary Roma and non-Roma students from socioeconomically disadvantaged communities. The students from the intervention group participated in 132 language...
Randomization was done at the Head Start site (center) level, which satisfies or exceeds class-level randomization.
The study measured child-level academic outcomes via teacher ratings, not a standardized, exam-based assessment of each child.
The paper reports that the intervention was conducted over an entire Head Start academic term (fall to spring), meeting the term duration criterion.
The control group’s business-as-usual setting is clearly described, including their staffing support and how it differed from the intervention.
Whole Head Start sites (equivalent to schools) were the unit of randomization, fulfilling the school-level RCT requirement.
The study does not mention any external evaluators. The intervention appears to have been evaluated by its own designers, lacking independent oversight.
The intervention spanned an entire preschool year (approximately 9 months), satisfying the one-year duration criterion.
The intervention group received extra training and mental health consultation services, whereas the control group did not receive comparable resources or attention.
No independent replication by other researchers is reported; this was a single-site study carried out by one team.
Academic performance was only assessed in language/literacy and math (via teacher-rated scales), rather than covering all core subjects with standardized exams.
The original CSRP participants were followed up in later years. A subsequent study by the same research team collected data on these students’ outcomes in high school, fulfilling the graduation tracking criterion.
No pre-registered analysis plan or study registration is mentioned. There is no evidence that the trial was registered before data collection.
The role of subsequent school contexts in the long-term effects of early childhood interventions has received increasing attention, but has been understudied in the literature. Using data from the Chicago School Readiness Project (CSRP), a cluster-randomized controlled trial conducted in...
Randomization was performed using classes as the sampling units, satisfying the class-level randomization requirement.
Outcomes were measured using a self-report anxiety scale (GAD-7), not standardized exam-based educational assessments.
Outcomes were measured about 12 weeks after intervention start, which is approximately a full academic term.
The control condition is clearly described (standard PE curriculum) and baseline characteristics and group sizes are reported.
Randomization was conducted at the class level within campuses, not by assigning entire schools/campuses to conditions.
The paper does not clearly document that the evaluation was led by an independent third-party team separate from the intervention designers/managers.
Outcomes were measured after about 12 weeks, far short of 75% of an academic year.
The interventions were delivered within regular PE lesson time (content substitution), so student instructional time appears comparable to control; extra teacher training/monitoring is part of implementing the intervention.
No independent replication of this specific RCT by other authors was found.
Because exam-based assessment (E) is not met, the all-subject standardized exam requirement (A) is also not met.
The study did not track participants to graduation, and no follow- up publication reporting graduation tracking was found.
The trial was registered under NCT05561192 before the reported study start date of February 1, 2023.
Introduction: When manifested in a dysfunctional manner, anxiety constitutes a mental health concern. This study assessed the effects of incorporating 3 different interventions into parts of physical education classes on anxiety symptoms (AS) in high school students. In addition, secondary...
Teachers (i.e., classroom/teacher clusters) were randomly assigned to treatment vs. control, meeting the class-level RCT requirement.
Outcomes are measured using surveys of coding attitudes (ESCAS), not standardized exam-based achievement assessments.
Outcomes were measured from the start to the end of the 2023-2024 school year, exceeding the one-term minimum.
The control condition is described as business-as-usual and the paper reports control-group counts, demographics, and baseline differences.
Randomization was at the teacher/class level rather than by assigning entire schools to treatment vs. control.
The paper describes research-team-delivered PD and support, but it does not clearly state that evaluation/data collection and analysis were conducted independently from the intervention designers.
Outcomes were measured from the start to the end of the school year, satisfying the year-duration requirement.
Treatment received substantially more CS/CT instructional time (and teacher support) than control, and the paper notes it cannot separate curriculum effects from increased instruction time.
No independent peer-reviewed replication of this specific RCT was found in the paper or via external search.
Because standardized exam-based outcomes are not used (criterion E is not met), the all-subject standardized-exams criterion is also not met.
The study reports outcomes only through the end of one school year, and no follow-up publication was found that tracks the cohort through graduation.
The paper reports pre-registration on REES prior to collecting any outcome data, but the registry record could not be independently accessed without login.
Integrating literacy and computational thinking (CT) can broaden computer science education participation, especially for multilingual learners. This study examined how the Computing and AI for All (CAIforALL) Act 1 Curriculum, an ELA-integrated Scratch-based CT curriculum, impacts coding attitudes among elementary...
Classes (not individual students within a class) were used as the randomization unit, meeting the class-level RCT criterion.
Outcomes were measured using a mental-health questionnaire (GAD-7), not standardized exam-based educational assessments.
Outcomes were measured about 12 weeks after intervention start, which meets the minimum term-length follow-up requirement.
The control group is described as standard PE and baseline characteristics are reported by group, including the control.
Randomization was conducted at the class level within campuses, not by randomly assigning whole campuses (schools/sites).
The paper does not clearly document an independent external evaluation team separate from the study team and implementers.
The intervention and outcome measurement covered about 12 weeks, far short of 75% of an academic year.
Student instructional time appears comparable because activities were inserted within standard PE lessons, and extra supports (training/monitoring) are part of the intervention delivery.
No independent replication of this specific trial was found in the paper or via targeted literature searching.
Criterion E is not met, so the all-subject standardized exam requirement cannot be met.
The paper reports no post-intervention follow-up, and criterion Y is not met, so graduation tracking cannot be met.
The trial is reported as registered (NCT05561192) and registry information indicates posting occurred before the study start.
This study assessed the effects of incorporating 3 different interventions into parts of physical education classes on anxiety symptoms in high school students. A parallel, 4-arm randomized clinical trial was conducted with technical high school students from 2 campuses in...
Randomization was clustered (school-level or class-level), meeting the class-level-or-stronger ERCT randomization requirement.
Outcomes were measured with well-being questionnaires/scales, not standardized exam-based educational assessments.
Outcomes were measured at a 26-week follow-up, exceeding a term-length (3–4 month) tracking window from baseline/intervention start.
The paper clearly describes both active and inactive control groups and reports group sizes and measurement timing.
The paper states that randomization was performed at the school level and later reiterates school-level execution of randomization.
The paper does not clearly document independent external evaluation; key trial activities (data collection and analysis) were performed by the author team.
The longest follow-up reported in this paper is 26 weeks, which is below the ERCT threshold of at least 75% of a typical academic year.
The active control closely matches the intervention on session duration and home-practice recommendations, and the extra time/resources are part of the intervention packages being tested.
No peer-reviewed independent replication of this specific trial by a separate research team was found.
Because the study does not use standardized exam-based academic assessments (Criterion E is not met), the all-subject exams criterion is automatically not met.
Graduation tracking is not reported, and ERCT rules also require G to be not met when Y (year duration) is not met.
The trial registration is described as retrospective and is dated after data collection began.
Mental health disorders often emerge during adolescence. Mindfulness interventions may support adolescents’ well-being. However, the evidence supporting the effectiveness of universal mindfulness interventions for adolescents’ well-being is limited and hampered by methodological weaknesses. The present study is the first large-scale...
Schools were randomized at the school level, which is stronger than class-level randomization and therefore satisfies Criterion C.
The outcomes are leadership, wellbeing, behavior, and physical activity/fitness measures, not standardized academic exams.
The intervention began in Term 2 and outcomes were assessed at intervention end in late Term 3, which is at least one term later.
The wait-list control is described as continuing usual practice during Terms 2 and 3, with baseline characteristics reported for both groups.
The unit of randomization was the school, satisfying Criterion S.
Randomization involved an independent researcher, but the paper does not show a fully independent third-party evaluation team separate from the intervention developers.
The intervention began in Term 2 and outcomes were assessed at intervention end in late Term 3, which is well under 75% of an academic year.
The intervention provided substantial additional resources (workshop, equipment, implementation support) that were not matched in the control condition, and these extra inputs were not framed as the treatment variable.
No independent replication by a separate research team was found for this specific Learning to Lead cluster RCT.
Criterion E is not met (no standardized academic exams), so the all-subject standardized exam requirement cannot be satisfied.
The study does not track participants through graduation, and Criterion Y is not met which also forces Criterion G to be not met.
The trial was prospectively registered (ACTRN12621000376842) with a registration date preceding first participant enrolment.
Schools provide an ideal context for developing students’ leadership skills; however, most leadership opportunities (e.g., serving as class president) are typically offered to students who already demonstrate leadership qualities. The aim of our study was to evaluate the effects of...
Randomisation was at the individual-student level within classrooms, not at the class or school level, and the small-group (non-tutoring) nature of the intervention does not qualify for the personal-tutoring exception.
The study used ReadBasix, an externally developed and psychometrically validated standardized assessment battery, for its reading comprehension and language skills outcomes.
Outcomes were measured at the end of an approximately 12-week (about 3-month) intervention period, which approximates the standard's definition of one academic term.
The paper provides detailed, tabulated demographic and baseline performance data for both comparison groups, confirming they were comparable at pretest.
Randomisation occurred at the individual-student level within classrooms; no school-level assignment took place.
The same research team that designed the CLAVES curriculum also conducted and analyzed this trial, with no independent third-party evaluator described.
The intervention and outcome measurement spanned only about 12 weeks, far short of the 75%-of-an-academic-year threshold required for this criterion.
Both comparison groups received identical curriculum content and dosage from the same teacher, so resources were balanced between conditions.
No independent replication of this specific grouping experiment by a different research team was found in the paper or in a search of subsequent citing literature.
Only language and literacy outcomes were assessed, with no coverage of other core subjects and no applicable exception.
There is no evidence of any follow-up tracking beyond the immediate posttest for this cohort, a search of subsequent publications by the author team found no graduation follow-up of this sample, and the prerequisite Year Duration criterion is also not met.
The study was pre-registered with a documented registry ID, and its pre-specified hypotheses and analyses are referenced throughout the paper; the registry itself could not be independently queried to confirm the exact date.
In this preregistered within-teacher randomized controlled trial (n = 84), we tested the effects of grouping English learners (ELs) in homogeneous groups (all ELs) versus heterogeneous groups (ELs and non-ELs) on language, reading comprehension, and argumentative writing. Findings indicated no...
Randomisation was conducted at the individual student level, not at the class or school level, and the intervention is not a one-to-one tutoring exception.
Outcomes are based on ordinary course grades (A-F) assigned by instructors in the courses being studied, not a standardised, widely recognised exam.
The intervention runs a full semester and outcomes are tracked for one and two years after it begins, far exceeding one academic term.
The paper documents the control group's composition, baseline covariates, and the specific standalone DE course it received in detail, including a formal covariate-balance table.
Randomisation occurred among individual students within each college, not among schools/colleges, so the stronger school-level criterion is not met.
The corequisite models were designed and already in place at the colleges themselves; the study authors are independent academic researchers who evaluated pre-existing programs rather than their own intervention.
Outcomes are tracked for one and two full years after the intervention begins, well beyond 75% of an academic year.
The additional instructional time in the corequisite (treatment) condition is integral to the very reform being tested (immediate college-course placement plus concurrent support versus sequential standalone DE), and total instructional time between the two arms is roughly comparable rather than grossly imbalanced.
The paper explicitly identifies itself as the first experimental study of reading/writing corequisite remediation, and no independent replication of this specific RCT by a different research team is reported in the paper or found via subsequent literature.
Criterion E is not met, and the study only measures English/reading-related outcomes, not all main subjects.
The study explicitly tracks outcomes only for one to two years, and no subsequent publication by the same authors tracking this cohort through graduation or degree completion was found.
No statement of pre-registration, registry ID, or pre-registration date is found anywhere in the paper, and no registry record was located via additional search.
This is the first study to provide experimental evidence of the impact of corequisite remediation for students underprepared in reading and writing. We examine the short-term impacts of three different approaches to corequisite remediation that were implemented at five large...
The study randomized entire schools (a stronger unit than class-level), which automatically satisfies the class-level RCT requirement.
Reading comprehension and oral proficiency were measured with MAP and WIDA, both widely used, standardized assessments, not custom-built tests.
The intervention and its outcome measurement spanned only about five weeks, far short of the one-term minimum required by the standard.
The control group's demographics, baseline scores, and business-as-usual instructional condition are documented in detail, including a dedicated table.
Randomization occurred at the school level (via randomized blocks of schools), satisfying the stronger school-level RCT criterion directly.
The same research team that designed the MORE intervention and its predecessor studies also implemented, supported, and scored the current trial, with no independent evaluator described.
Because the Term Duration (T) criterion was not met, the stronger Year Duration criterion cannot be met either, per the standard's cascade rule.
Treatment classrooms spent more time on science and social studies content, but this extra content-area time is integral to the content-integrated literacy design being tested, not an unrelated add-on.
The paper is a self-replication and extension by the same research team of its own prior work; no independent replication by a separate team was found in the paper or via internet search.
Outcomes were limited to reading, writing, vocabulary, and oral proficiency in science and social studies content; no other core subjects (e.g., mathematics) were assessed.
Because the Year Duration (Y) criterion was not met, Graduation Tracking cannot be met either, per the standard's cascade rule; internet search of subsequent publications also found no evidence of graduation-level follow-up of this EL cohort.
The paper states the study design was preregistered but provides no registry name, ID, link, or date to verify that registration preceded data collection, and internet search could not locate the registration record.
The current study replicated and extended the previous findings of content-integrated literacy intervention focusing on its effectiveness on first- and second-grade English learners' (N = 1,314) reading comprehension, writing, vocabulary knowledge, and oral proficiency. Statistically significant findings were replicated on...
The study randomized individual students within schools rather than assigning entire classes or schools to conditions.
The study relied on school-assigned GPAs and grades rather than widely recognized standardized exam-based assessments.
Outcomes were measured over the final three quarters of the school year, which is significantly longer than one academic term.
The control group's demographics, baseline data, and specific activities (neutral writing exercises) are well-documented.
Randomization occurred at the student level within schools, not at the school level.
The study was conducted by the authors themselves, who are also associated with the design of the intervention in previous studies.
Outcomes were measured across the full academic year (Terms 2, 3, and 4) following the start of the intervention in September.
The control group received a placebo activity that matched the treatment group in terms of time and resources.
The replication was conducted by the same authors as the original study, not by an independent research team.
Although the study covered core subjects, it relied on GPA rather than standardized exams, failing the prerequisite Criterion E.
The study tracked students only through the end of the sixth-grade year, not until graduation.
The study protocol was preregistered on OSF before the start of the study.
Recent randomized studies suggest brief social-psychological interventions can help students reappraise common social and academic worries during the difficult transition to middle school and, in turn, improve school performance. We conducted a preregistered student-level randomized controlled trial to assess the...
Randomisation was at the individual applicant level via an enrollment lottery for group ESOL classes, not at the class or school level, and the tutoring exception does not apply.
Outcomes are voter registration, voter participation, and employer-reported earnings from administrative records; no standardised exam-based educational assessment was used as an outcome.
Outcomes were measured from two up to ten years after the first lottery (intervention start), far exceeding the one-term minimum tracking interval.
The control group (lottery non-winners) is documented in detail, including size, baseline demographics, baseline earnings, and the conditions they experienced.
Randomisation was among individual applicants within a single adult education program, not among schools or institutional units.
The evaluation was conducted by university researchers who did not design or operate the FAESL+ program, using administrative lotteries run by program staff and outcome data from independent state agencies and vendors.
Participants were tracked for up to ten years after the lottery (6.9 years on average), far exceeding 75% of an academic year from intervention start.
The additional educational input (access to ESOL classes) is itself the treatment variable tested against a documented business-as-usual control, so the resource imbalance is by design.
The authors state this is the first random-assignment study of business-as-usual US adult ESOL services; internet searches of the paper's citation record (40 citing works, as of July 2026) found related international RCTs on refugee language training but no independent replication of this specific study.
Criterion E fails (no standardised exam outcomes at all), so the all-subject exam requirement automatically fails as well.
Long-run follow-up tracked earnings and voting, not educational progression; no publication tracking this cohort to program completion, credential attainment, or graduation was found in the paper or in internet research.
No pre-registration of hypotheses, methods, or analysis plans is mentioned in the paper or its published journal version; the study retrospectively reconstructs lotteries held in 2008-2016.
While current debates center on whether and how to admit immigrants to the United States, little attention has been paid to interventions designed to help immigrants integrate after they arrive. Public adult education programs are the primary policy lever for...
Entire schools were randomly assigned to conditions, which satisfies the class-level (or stronger) randomization requirement.
The primary outcomes relied on researcher-developed tasks rather than a widely recognized standardized exam, which the authors themselves flag as a limitation.
The intervention ran from August to December 2024 (about five months), and post-test outcomes were measured at the end, exceeding one academic term.
The control group's size, demographics, baseline scores, and business-as-usual condition are documented in the text and in Tables 1 and 2.
Whole schools, not classes or students, were the unit of random assignment to conditions.
The same research team designed, delivered, and evaluated the intervention, and a co-author is associated with the GL-Rime/GraphoGame tool, with no independent evaluator.
The intervention plus measurement spanned only about five months with no follow-up, well under 75% of an academic year.
The phonics instruction and GL-Rime are themselves the explicit treatment variables tested against a business-as- usual control, so the resource difference is integral by design.
No independent research team has replicated this specific RCT; related prior work is by the same overlapping authors.
Only English literacy/reading outcomes were assessed, with no other core subjects, and criterion E is also not met.
Measurement stopped at post-test with no follow-up or graduation tracking, and criterion Y is also not met.
The paper contains no reference to a pre-registered protocol, registry ID, or registration date.
This randomized control trial examined the contribution of teacher-led phonics instruction and GraphoLearn-Rime (GL-Rime) to the development of English literacy and reading skills of Grade 1 level children in a low-income context. Students (N=234) were randomly allocated into three groups....
Randomization was conducted at the individual student level rather than by class or school, failing the class-level RCT requirement.
Outcomes were measured via course grades and credits, not through a recognized standardized examination.
The intervention courses ran for a full 16-week semester, meeting the term duration requirement.
The control group’s composition, consent rates, and baseline covariates are documented in tables and text, satisfying documentation requirements.
The study randomized individual students rather than entire schools, failing the school-level RCT requirement.
The study was conducted by researchers independent from the designers of the corequisite models, satisfying the ERCT requirement for Criterion I.
The corequisite support lasted one semester; the Year-long intervention requirement is not satisfied.
The DE support hours are integral to the corequisite intervention, so the control’s business-as-usual condition is appropriate.
No independent replication of this RCT is reported in the paper.
Only reading and writing outcomes were measured, failing the all-subject exam requirement.
A follow-up study by the same research team tracked the original student cohort through graduation, satisfying the ERCT requirement for Criterion G.
There is no indication that the study protocol was pre-registered before data collection.
This study provides experimental evidence on the impact of corequisite remediation for students underprepared in reading and writing. We examine the short-term impacts of three corequisite models implemented at five large urban community colleges in Texas. Results indicate that corequisite...
Randomisation was performed at the individual participant level rather than by class or school.
The study used a self-developed, non-standardised rating instrument rather than a recognised standardised exam.
Outcomes were measured 12 months after the intervention began, well beyond one full term.
The control group's demographics, baseline scores, and training received are clearly documented in the text and Tables 1 and 2.
Randomisation was at the individual level, not at the school or institutional level.
The same author team designed, conducted, and analysed the trial, so conduct was not independent despite blinded outcome raters.
Outcomes were measured 12 months after the intervention began, covering a full academic year.
Both arms received active, standard-adherent ALS training with comparable resources, with training frequency/format itself being the treatment variable.
No independent replication of this specific trial by a separate research team exists.
Criterion E is not met and only a single specialised domain was assessed, so all-subject coverage fails.
Measurement ended at 12 months with no tracking of participants to any graduation endpoint.
The trial was prospectively registered (DRKS00024822, 26 March 2021) before enrolment began (1 April 2021), confirmed against the DRKS registry record.
Background Skill decay in advanced life support (ALS) is well documented, yet optimal training frequency remains unclear. This trial compared low-dose, high-frequency ALS training with annual full-day training regarding simulation-based resuscitation performance after one year. Methods In this randomized, controlled,...
Whole schools, not students within a classroom, were allocated to the education and control arms, which satisfies the class-level requirement.
Outcomes came from a custom questionnaire the authors assembled from the literature, not from any recognised standardised exam.
The post-test was administered three months after the intervention, which reaches the one-term follow-up requirement.
The control arm's size, demographics, baseline knowledge, and business-as-usual condition are fully tabulated and described.
Ten whole high schools, five per arm, were the unit of allocation, which the Limitations section confirms as school-level randomisation.
The same two authors designed, delivered, measured, and analysed the intervention with no external or third-party evaluation.
The only follow-up was three months after the intervention, far short of 75% of an academic year.
The added PowerPoint education sessions are themselves the treatment variable tested against a business-as-usual curriculum, so the passive control is permitted by design.
Internet searching found no independent replication of this specific trial, which was published only in March 2026.
Criterion E fails and only drug-abuse knowledge was measured, with no school subject assessed at all.
Criterion Y fails, measurement stopped at a three-month post-test, and no follow-up publication tracking the cohort to graduation was found.
The paper reports institutional ethics approval only, with no trial registration number, registry, or pre-registered protocol, and none was found in public registries.
Background: Drug abuse remains a major public health concern, exacerbated by the increasing availability of substances and adolescents' reliance on informal sources of information. Although school-based education is recognized as a key preventive strategy, evidence regarding patterns of drug abuse...
Allocation was made at the kindergarten level with all children in a kindergarten sharing one condition, which satisfies the class-level randomisation requirement.
The paper states outright that every outcome measure was custom-built for the study, so no standardised exam-based assessment was used.
The intervention ran for 15 weeks (about 3.5 months) with outcomes measured at its end, meeting the one-term interval requirement.
The control group's size, demographics, baseline equivalence, and business-as-usual condition are all documented in detail, including fidelity observations.
Whole kindergartens - the preschool institutions implementing the programme - were the units of random assignment, though the paper's interchangeable use of "kindergarten" and "kindergarten classroom" leaves the exact number of institutions unclear.
The authors designed the intervention, built the measures, trained the testers and fidelity coders, and analysed the data, with no independent evaluator reported.
The intervention ran 15 weeks with outcomes taken immediately at its end and no longer-term follow-up, well below 75% of an academic year.
Children's instructional time and content domains were comparable across arms, and the extra 30-hour teacher training and weekly coaching are integral to the ADMIN package being tested.
ADMIN is presented as a first-of-its-kind programme and internet searching identified no independent replication by a different research team.
Criterion E was not met and outcomes covered only language and emergent literacy, with no assessment of other main curriculum domains such as numeracy.
Criterion Y failed, only immediate post-intervention outcomes were collected, and no published follow-up tracks the cohort to graduation.
The paper reports ethics and funding approvals but contains no pre-registration reference, and no registry entry for this trial could be found online.
Arabic-speaking children grow up in diglossia where different language varieties serve complementary functions; They use Spoken Arabic (SpA) for everyday communication but acquire literacy in Standard Arabic (StA), the only variety with a conventional orthography. Research shows that while the...
Randomisation was conducted at the club (cluster) level rather than on individual athletes, satisfying the class-level-or-stronger requirement.
Outcomes were assessed with an author-developed research questionnaire, a diet-adherence index, and self-reported food records rather than any standardised exam.
Outcomes were re-measured at a follow-up roughly four to five months after the intervention started, exceeding the one-term minimum.
The control group's size, demographics, baseline values on all outcomes, and study conditions are documented in detail across Table 2 and Tables 3-6.
Four whole handball clubs, the units implementing the intervention, were randomly allocated 1:1 to intervention and control, satisfying institution-level randomisation.
The same two-author team designed the programme, delivered it, collected the data and analysed the results, with no independent or third-party evaluator involved.
The interval from intervention start to the final follow-up assessment was roughly four to five months, well below 75% of an academic year.
The extra dietitian sessions and materials are themselves the treatment variable under test, so the business-as-usual wait-list control is balanced by design.
The only precursor study is the same authors' own pilot, and an internet search found no independent team that has replicated this trial.
Criterion E is not met and outcomes were limited to sports nutrition knowledge alone, with no other subjects assessed.
Criterion Y is not met and tracking ended about three months after the programme, with no graduation outcome measured, planned, or reported in any later publication.
Despite the paper's claim of prospective registration, the ClinicalTrials.gov record shows first submission on 18 February 2026, a year after the trial started and after data collection had ended.
Background/objectives: Adolescent female athletes often present insufficient dietary intake and limited sports nutrition knowledge (SNK). This study aimed to evaluate the effects of a short-term nutrition education intervention on SNK, dietary intake, and body composition in adolescent female handball players....
Randomization was conducted at the teacher (whole classroom) level, satisfying the class-level RCT requirement.
Outcomes were measured with a custom project-developed observational coding form, not a standardized exam.
The interval from intervention start to posttest spans about four months, satisfying the one-term duration requirement.
The BAU control group's size, demographics, baseline data, and business-as-usual condition are clearly documented.
Randomization occurred at the individual teacher level, not at the school level.
The intervention was designed and delivered by the same research team that conducted and analyzed the study, with no independent external evaluator.
The interval from intervention start to posttest was about four months, well below 75% of an academic year.
The additional professional development and coaching are the explicit treatment variable tested against a business-as-usual control, so the design is balanced.
This specific trial has not been independently replicated; the prior pilot was by the same team and cited studies are not replications of it.
Criterion E is not met and outcomes were individualized observational measures, not standardized all-subject exams.
Criterion Y is not met and the study had no follow-up beyond posttest, so no graduation tracking occurred.
The paper reports ethics approval but provides no pre-registration statement, registry link, or registration date.
The increasing number of children with disabilities served in inclusive preschool classrooms has heightened the need for instructional approaches that support learning during naturally occurring classroom activities and routines. Embedded instruction (EI) is a naturalistic teaching approach that allows teachers...
Randomization was at the individual participant level (not class- or school-level), and no one-to-one tutoring exception applies.
The outcome exams were developed by the course director and instructors for this course rather than being a widely recognized standardized exam.
The paper reports the course running from September 15, 2022, through December 31, 2023, which exceeds a term from start to the end-of-course outcome assessments.
The comparator groups (including the lecture-based condition) and baseline characteristics are clearly documented (e.g., Table 1 and the LBL schedule description).
The study was conducted within one institution and randomized individual participants, not schools/sites.
The paper does not document independent third-party conduct of the intervention evaluation; the author team and course staff appear to have delivered and assessed the course.
The reported course period from September 15, 2022, through December 31, 2023, exceeds 75% of an academic year from start to end-of-course measurement.
The groups received the same preparatory materials, cases, and session count; staffing/simulation differences appear integral to the teaching models being tested rather than separable add-ons.
No independent, peer-reviewed replication of this specific MDTBL vs PBL vs LBL pulmonary-nodule course RCT was identified.
Criterion A is not met because Criterion E is not met and because outcomes are limited to a course-specific pulmonary-nodule test rather than all core subjects.
The study reports only within-course testing (pre/mid/final) and does not track participants until graduation; no follow-up paper reporting graduation outcomes was identified.
The study explicitly states it was retrospectively registered in 2023, after the course start date (September 15, 2022), and no public pre-registration record was found.
Background This study aims to compare multi-disciplinary team-based learning (MDTBL), problem-based learning (PBL) and lecture-based learning (LBL) on the learning outcomes and experiences of medical students. Methods A randomized controlled study was designed to recruit 30 medical students with a...
Randomization occurred at the class (cluster) level, which meets the class-level RCT requirement.
Outcomes were measured with a study-specific 15-item MCQ test rather than a widely recognized standardized exam.
The primary outcome was measured immediately after a single 130-minute session, not at least one academic term later.
The paper clearly documents the comparison (control) group, including group sizes and baseline characteristics.
Randomization was at the class (cluster) level within one medical school, not across multiple schools.
While some procedures were blinded and separated from teaching, there is no clear third-party independent evaluation team separate from the intervention designers.
Because term-duration follow-up is not met and outcomes were measured immediately after the session, year-duration tracking is also not met.
Both groups received comparable instructional time and core materials, and the delivery mode (Spatial WBVE vs classroom) was the intended treatment difference.
No peer-reviewed, independent replication of this specific trial was identified as of the ERCT check date.
Because a standardized exam was not used (E not met), the all-subject standardized exam requirement is not met.
There is no follow-up through graduation, and since year duration (Y) is not met, graduation tracking (G) is also not met.
The trial reports a public registry ID and states registration occurred before participant enrollment, and an earlier preprint reports a registration date that precedes the study start date.
Background: Foundational knowledge of anesthesia techniques is essential for medical students. Team-based learning (TBL) improves engagement. Web-based virtual environments (WBVEs) allow many learners to join the same session in real time while being guided by an instructor. Objective: This study...
Students were randomized to tutoring-style small groups and then those groups were randomized to conditions, which satisfies the ERCT tutoring/personal-teaching exception to class-level randomization.
The outcomes include standardized, commercially published reading assessments (TOWRE-2, TOSWRF-2, TOSREC, DIBELS ORF), not only researcher-developed tests.
The intervention lasted 10 weeks and outcomes were measured at posttest; this does not reach a full academic term and no term-later follow-up is reported.
The comparison condition (Decoding) and the treatment condition (Decoding+Spelling) are clearly described, and baseline demographics and baseline equivalence are reported in tables.
Randomization was not at the school level; students were grouped and then groups were randomized to conditions across multiple schools.
The paper shows the authors led tutor training and project leads monitored adherence, but it does not document an independent third-party evaluation team.
The intervention and measurement window was 10 weeks, far short of 75% of an academic year; additionally, ERCT rules imply Y cannot be met when T is not met.
The Decoding+Spelling condition intentionally added time, and the paper explicitly frames this added time as integral to the tested contrast (adding spelling as implemented in practice), so the resource imbalance is by design.
No independent replication of this specific RCT was found in the paper, and an internet search did not identify an external replication study as of the ERCT check date.
Outcomes focus on literacy-related skills only and do not include standardized assessments across all core school subjects.
The study reports only pretest and posttest around a 10-week intervention and provides no evidence of tracking outcomes through graduation; additionally, ERCT rules imply G cannot be met when Y is not met.
The paper provides no registry link/identifier or dated statement showing the protocol was pre-registered before data collection.
Integrating spelling with word reading instruction is commonly recommended. However, few studies have determined the unique effect of spelling activities on word reading skills for students with reading difficulties. Considering the limited time available for intervention and the fact that...
Randomization was performed at the intact class level (not within classes), satisfying the class-level RCT requirement.
The study used performance-based tasks and questionnaires rather than widely recognized standardized exam-based assessments.
The intervention ran for one academic quarter (12 weeks) with post-testing, meeting the minimum term-duration requirement.
The control group is clearly defined (size and condition) and baseline equivalence is documented.
Randomization was implemented at the classroom level within schools rather than by randomizing schools.
The paper does not document an independent external evaluation team separate from the intervention implementers/researchers.
The study measures outcomes through a 12-week intervention with a short follow-up, which is far less than 75% of an academic year.
The resource differences (CBL platforms and teacher onboarding) are integral to the treatment being tested, and the intervention is delivered within the regular Informatics course structure.
No independent peer-reviewed replication of this specific study was found, and the paper does not report a completed external replication.
Because criterion E is not met, criterion A is not met; the study also focuses on Informatics/cognitive measures rather than all core subjects using standardized exams.
Graduation tracking is not reported, and because criterion Y is not met, criterion G is not met.
The paper does not provide a pre-registration registry/ID and date demonstrating registration before data collection.
Despite growing attention to the digitalization of education and the development of AI-supported learning in Kazakhstan, as well as broader international agendas, the empirical evidence on whether cloud-based learning (CBL) strengthens adolescents’ cognitive development and under what conditions and through...
Randomization occurred at the classroom level, preventing contamination across students in the same class.
The study employs bespoke KA Lite assessments, not established standardized exams.
The study measures outcomes after approximately six weeks, not a full term.
The control group’s makeup and activities are thoroughly described.
The study randomized individual classrooms, not whole schools.
The same research team designed and evaluated the intervention.
The intervention and measurement occur within ~12 weeks, not a full year.
Treatment and control groups received identical time and resources.
No independent replication study is mentioned.
Only mathematics outcomes are measured, not all core subjects.
The study ends after the units and does not follow students to graduation.
The protocol was pre‑registered in the AEA registry before implementation.
This randomized experiment implemented with school children in India directly tests an input incentive designed to increase effort on learning activities against both an output incentive that rewards test performance and a control. Students in the input incentive treatment perform...
Teachers (clusters) were randomly assigned to conditions, satisfying the class-level (or stronger) randomization requirement.
The study used aimswebPlus oral reading fluency with standardized passages and standardized directions, indicating a standardized assessment.
The longest stated start-to-measurement interval is about 11 weeks, which is shorter than a typical academic term (about 3–4 months).
The comparison condition procedures, group sizes, and baseline equivalence at pretest are reported, providing sufficient control/comparison group documentation.
Randomization occurred at the teacher (cluster) level rather than by assigning schools to conditions.
The authors created key materials and trained teachers, and testing was teacher-administered, so independent third-party conduct is not clearly documented.
The study duration (about 11 weeks from start to latest follow-up) is far shorter than 75% of an academic year, and because T is not met, Y is also not met.
Both groups received the same schedule and session length, used the same passages and repeated-reading frequency, and answered the same comprehension questions; remaining differences are integral to the instructional contrast being tested.
No independent replication of this specific clustered RCT was found, and the paper itself only recommends that future studies replicate it.
Outcomes were limited to oral reading fluency (WCPM) rather than standardized exams across all core subjects.
The study reports only short follow-up (4 weeks post intervention), and because Y is not met, G is also not met; no later graduation-tracking follow-up publication was found.
The paper links to OSF for shared data but does not state protocol pre-registration, provide a registry ID for a pre-registration, or give a registration date showing it occurred before data collection.
Fluency is a multidimensional construct that requires automaticity with foundational skills. Fluency is not an end in itself but serves as a bridge between decoding and reading comprehension. However, many secondary students struggle with proficient reading and fail to attain...
Participants were randomized at the individual child level rather than by class (and the intervention is not described as one-to-one tutoring), so the class-level RCT requirement is not satisfied.
The study includes a widely recognized standardized test of decoding (TOWRE), satisfying the exam-based assessment requirement.
The study measures outcomes through follow-up over about 16 weeks from baseline, which is approximately one full academic term.
The paper clearly defines both control conditions and reports baseline comparability across groups, satisfying the documented control group requirement.
Schools/language units were recruitment sites, but randomization was conducted at the child level rather than by school, so the school-level RCT criterion is not satisfied.
GraphoLearn is described as an externally developed program rather than one designed by the study authors, supporting independent conduct relative to the intervention’s designers.
The study spans about 16 weeks from baseline to follow-up, which is far shorter than 75% of an academic year.
The active control matches the game dosage, but the TAU condition is not described as receiving comparable time and resources, and the paper does not clearly state that added resources/time are the treatment variable.
No independent replication of this specific RCT could be identified, and the paper itself frames the trial as the first RCT of this kind in this population.
The study assesses reading/decoding outcomes rather than standardized exam outcomes across all main school subjects, so the all-subject exams criterion is not met.
The study ends after a short follow-up (weeks) and does not track students until graduation; additionally, ERCT specifies that failing year-duration (Y) implies failing graduation tracking (G).
The trial is registered (NCT05295472), but registry dates indicate it was first posted after the study start and after data collection had begun, so it is not pre-registered.
Individuals with developmental language disorder (DLD) often struggle with reading, both with reading comprehension and decoding. Although decoding difficulties are common in the population with DLD, studies investigating phonics-based interventions for these individuals are sparse. This study investigated the effect...
Randomization was performed at the individual student level, not at the class (or higher) level, and the intervention was not a one-to-one tutoring intervention.
The primary educational outcome used a study-specific set of multiple-choice questions rather than a widely recognized standardized exam.
Outcomes were measured up to a twelve-month follow-up after the intervention began, which exceeds the minimum term-length requirement.
The control group is clearly described (no intervention) and the paper reports group sizes and baseline characteristics for all groups (e.g., Table 1).
The trial was not randomized at the school level (it was conducted in a single school with individual-level assignment).
Allocation concealment involved an independent researcher, but the paper does not show that the intervention evaluation was conducted by an external, independent team.
The study tracked outcomes from February 2024 to February 2025, which is approximately one full year and exceeds 75% of an academic year.
The intervention arms received additional instructional inputs (training time and VR equipment), while the control arm received no training; this imbalance is explicitly the intended treatment comparison (training vs no training).
No independent replication of this specific trial was identified, and the paper itself does not report being replicated.
Because the study does not use standardized exams (Criterion E is not met), it also cannot satisfy the all-subject standardized exam requirement.
The study includes a twelve-month follow-up but does not report tracking participants until graduation; no follow-up graduation paper by the same authors was identified.
The paper reports ethics approval and data-collection dates but does not state that a protocol was pre-registered or provide a registry ID and registration date.
Objective: This study evaluated the effectiveness of a virtual reality (VR)-based cardiopulmonary resuscitation (CPR) training program compared with traditional theoretical instruction and a non-intervention control group, on improving and retaining knowledge, attitudes, and self-efficacy of secondary school students. Design Randomized...
The unit of randomization is the teacher (classroom), meeting the class-level RCT requirement.
The RCT’s primary outcomes are platform usage logs, not standardized exam scores; any test-score analysis is non-experimental.
Outcomes are tracked from the May 30 start through November 2022, exceeding one academic term.
The control group is clearly defined (no messages) and baseline characteristics for treatment and control are documented.
Randomization was performed at the teacher level, not at the school level.
The lead author holds IP rights to the software being promoted, so the study is not independent of the intervention’s owner/designer.
The study tracks outcomes for less than a full academic year (March–November), falling short of year-duration tracking.
The only additional resource is the WhatsApp messaging itself, which is the treatment being tested; the control group is business-as-usual by design.
No peer-reviewed independent replication of this WhatsApp-to-teachers messaging RCT was found in a web search as of the ERCT check date.
Only mathematics outcomes are discussed, and criterion E is not met, so the all-subject standardized exam requirement is not satisfied.
Students are followed only within the 2022 school year and not until graduation; criterion Y is also not met, which implies G cannot be met.
No pre-registration identifier or registry entry could be located for this RCT in the paper or via a web search.
The use of self-led educational technologies holds significant potential for improving student learning at scale, but sustaining student engagement with these platforms remains a challenge. We present results from an experimental evaluation implemented following the scale-up of a math platform...
Children were randomly assigned to treatment conditions individually within mixed classes, not by entire class or school, and no tutoring-style exception is invoked.
The study administered the standardised, nationally normed PPVT-III (and Spanish TVIP-H) alongside its custom storybook measure, satisfying the standardised exam-based assessment requirement.
The intervention and its retention testing spanned roughly 12-16 weeks from start to final assessment, approximating the lower bound of a full academic term.
The comparator (English-language home-reading) arm's demographics, size, and baseline performance are extensively documented and confirmed equivalent to the primary-language arm at pretest.
All children were drawn from a single preschool site, with randomisation at the individual-child level rather than across schools.
The sole author designed the intervention, materials, and measures, and personally trained the classroom teachers who delivered it; no independent evaluation team was involved.
Because criterion T is not met, criterion Y cannot be met; independently, the 12-week program is far short of 75% of an academic year.
Both arms received identical storybooks, home-reading schedules, and classroom instruction; only the language of the home-reading books differed, so no extra resources were given to either group.
No replication of this study is mentioned in the paper, and a search did not identify an independent peer-reviewed replication of this specific home storybook-reading crossover design.
Because criterion E is not met, criterion A cannot be met; the study also only assessed vocabulary, not other subject areas.
Because criterion Y is not met, criterion G cannot be met; the study ends at the 12-week posttest with no tracking toward graduation and no identified follow-up publication.
The paper contains no statement of a pre-registered protocol, registry identifier, or registration date.
This study examined how providing either primary- or English-language storybooks for home reading followed by classroom storybook reading and vocabulary instruction in English influenced English vocabulary acquisition. Participants in the study were preschool children (N = 33), from low socioeconomic...
Randomisation was conducted at the individual student level within one course at one university, not at the class or school level, and no tutoring exception applies.
The study used the Cambridge Business English Speaking Test, a widely recognised standardised assessment scored by accredited external examiners.
Outcomes were measured after a 14-week (semester-long) intervention, exceeding the minimum one-term requirement.
The control group's size, demographics, baseline scores and instructional condition are explicitly documented and compared with the treatment group.
Randomisation occurred among individual students within a single university course, not among schools.
Aside from independent exam scoring, there is no explicit statement that the trial's design, data collection, or analysis were conducted independently of the researchers.
The 14-week (roughly one-semester) tracking period is far shorter than 75% of a full academic year.
The extra ASR/AWE resources given to the experimental group are the explicit treatment variable being tested, so the control group's business-as-usual condition is an acceptable comparator.
No independent replication of this specific study was found or reported, confirmed via internet search.
Only English speaking outcomes were assessed, with no other subjects measured or justification given for this narrow scope.
Measurement stopped at the immediate posttest, with no tracking toward course completion or graduation, and criterion Y is not met.
No pre-registration of the study protocol is mentioned anywhere in the paper, and none was found via internet search.
This study investigates how the combined use of automatic speech recognition (ASR) and automated writing evaluation (AWE) influences Chinese English as a foreign language (EFL) learners' speaking anxiety and speaking performance. Adopting a mixed-methods experimental design, the study compares an...
Randomization was at the individual participant level within one course, not at the class (or stronger) level, and no one-to-one tutoring exception applies.
Outcomes were measured using exams and questionnaires developed by the course team for this study rather than a widely recognized standardized exam.
The paper reports that participants took part in the course from September 15, 2022 through December 31, 2023, which exceeds a full academic term.
The comparison groups are described with group sizes and baseline characteristics (gender, age, clinical experience) reported and compared.
Randomization occurred among individual course participants, not among schools or equivalent educational institutions.
The paper indicates the authors developed MDTBL and designed the assessments, with no documented independent third-party evaluation team.
The reported participation/course window (September 15, 2022 to December 31, 2023) exceeds 75% of an academic year.
Core instructional materials, cases, and exams were matched across groups; the additional staffing and simulated patient in MDTBL are integral parts of the MDTBL intervention definition.
No independent replication study of this specific MDTBL trial (or a clearly equivalent MDTBL intervention) by a different team was found.
Because the assessment is not a standardized exam (E not met), the study cannot meet the all-subject standardized exam requirement.
The study reports only within-course assessments and recommends future long-term follow-up, and no follow-up papers tracking the cohort to graduation were found.
The study was retrospectively registered in 2023, after the course/study began in September 2022.
This study aims to compare multi-disciplinary team-based learning (MDTBL), problem-based learning (PBL) and lecture-based learning (LBL) on the learning outcomes and experiences of medical students. A randomized controlled study was designed to recruit 30 medical students with a minimum of...
Randomisation was conducted at the school (cluster) level across nine schools, which satisfies the class-level (or stronger) RCT requirement.
The outcome measure was a physical throwing-accuracy task (radial error), not a standardised academic exam; no exam-based assessment was used.
Outcomes were measured only 48 hours after the last training session, and the full study (pretest through retention) ran about six weeks, short of a full term.
The control group's size, condition (no illusion), and baseline demographic/performance data are clearly documented and compared against the two illusion arms.
Nine public elementary schools were the unit of random assignment, satisfying the school-level RCT requirement.
Only the randomisation sequence was generated externally; the same research/training team designed, delivered, supervised, and assessed the intervention.
The entire intervention and follow-up spanned about six weeks; criterion Y cannot be met because criterion T (term duration) was not met.
All three arms received an identical amount of practice (60 trials over six blocks); only the visual illusion condition, the treatment variable itself, differed between groups.
The authors state this is the first study of its kind, and an internet search for independent replications by other research teams found none.
Only a single motor-performance measure (throwing accuracy) was assessed, and criterion E was not met, so criterion A cannot be met either.
Follow-up stopped 48 hours after training with no further tracking, criterion Y was not met, and an internet search found no follow-up publication by the same authors tracking this cohort.
No pre-registration of the study protocol, registry name, ID, or date is mentioned in the paper, and an internet search of trial registries found no matching registration.
Children with Developmental Coordination Disorder (DCD) have deficits in visual perception and motor planning, which negatively impact motor learning. To address this, paradigms that manipulate visual perception during training have been proposed. This study investigated the influence of a visual...
Randomisation was carried out at the individual student level within a single cohort at one university, not at the class or school level, and no personal-tutoring exception is invoked.
The vocabulary tests were drawn from the New TOEIC Official Test-Preparation Guide III, a widely recognised, standardised, ETS-produced test, not a custom instrument.
Outcomes were measured only 4 weeks after the intervention began, with retention tested 2 weeks later, so the total interval (6 weeks) is far shorter than one academic term.
The control group's size, gender composition, baseline-matching procedure, and detailed app-usage statistics are clearly documented in the text and Table 8.
Randomisation occurred among 20 individual students within a single university cohort, not among schools.
The mobile app (the intervention) was developed by a separate commercial studio, and the academic research team conducting the trial had no described role in designing the app, supporting independent conduct of the evaluation.
As the Term Duration criterion (T) is not met, the stronger Year Duration criterion cannot be met either; the total tracked interval was only about six weeks.
The only substantive difference between conditions is the gamified assessment/ranking feature itself, the explicit treatment variable, while both groups used the same core app functions, word list, and minimum weekly usage expectation.
Neither the paper itself nor an internet search for citing or follow-up literature indicates that this specific study has been independently replicated by a different research team.
Only English vocabulary outcomes were assessed; no other core school subjects were measured, and the paper does not provide the kind of upper- secondary/vocational justification the standard requires for a narrower exception.
Tracking ended two weeks after the four-week intervention with no further follow-up; since the Year Duration criterion (Y) is not met, this criterion cannot be met either, and no subsequent graduation-tracking papers by these authors were found.
The paper contains no statement of pre-registration on any registry platform, nor any registration date or protocol reference, and no internet search of registry platforms located a matching entry.
Many studies have demonstrated that vocabulary size plays a key role in learning English as a foreign language (EFL). In recent years, mobile game-based learning (MGBL) has been considered a promising scheme for successful acquisition and retention of knowledge. Thus,...
Randomisation was conducted at the kindergarten (school) level rather than within a single class, which is a stronger design than the class-level requirement and therefore automatically satisfies it.
The primary target-vocabulary outcome measures were researcher-developed tasks created specifically for this study, rather than standard, widely recognised exams.
The intervention and its posttest measurement together spanned only about 8 weeks, well short of a full academic term.
The control group (kindergartens receiving Implicit Vocabulary Instruction) is clearly documented in terms of size, demographics, materials used, and who delivered the instruction.
Randomisation was conducted at the level of whole kindergartens, satisfying the school-level RCT requirement.
The researchers who designed the intervention also designed the control condition, trained all instructors, and oversaw fidelity checks, with no independent or third-party evaluator described.
Because criterion T (Term Duration) is not met, and the actual intervention/measurement period of about 8 weeks is far short of a full academic year, criterion Y is not met.
Both the experimental and control conditions received identical instructional time and identical storybook materials, differing only in the explicitness of the vocabulary teaching, so no extra resources were present to begin with.
No evidence of an independent replication of this specific study by a different research team was found in the paper or through internet searches.
Because criterion E (Exam-based Assessment) is not met, and only vocabulary and phonological awareness were assessed rather than all main subjects, criterion A is not met.
Because criterion Y (Year Duration) is not met, and no subsequent paper tracking this specific cohort to graduation could be found, criterion G is not met.
The paper contains no statement or reference indicating that the study protocol was pre-registered before data collection began, and no registry entry could be found online.
The introduction of second language (L2) education in kindergartens is ubiquitous in many places globally; nevertheless, research in these settings is scarce compared with that on older learners. L2 vocabulary development is especially germane to these very young learners, rendering...
The study randomized clusters at the class level, meeting the class-level RCT requirement.
Knowledge was measured with a study-specific 15-item MCQ rather than a standardized exam.
Outcomes were measured immediately after a single session rather than at least one academic term after intervention start.
The comparator condition and baseline group characteristics are documented, including a baseline characteristics table by group.
Randomization occurred at the class (cluster) level rather than the school/site level.
The paper does not document that an external independent team conducted the study separate from the intervention team.
Outcomes were measured immediately after a single session, so the study does not meet the year-duration requirement.
Time and instructional content are matched across groups; the delivery-mode differences (including technology) are integral to the treatment being tested.
No independent peer-reviewed replication of this specific trial was found.
The study does not use standardized exams, so it cannot meet the all-subject standardized exam requirement.
The study does not track participants to graduation, and year-long tracking is not present.
The paper reports registration in the Thai Clinical Trials Registry before participant enrollment, but the registry entry posting date could not be independently verified here.
Background: Foundational knowledge of anesthesia techniques is essential for medical students. Team-based learning (TBL) improves engagement. Web-based virtual environments (WBVEs) allow many learners to join the same session in real time while being guided by an instructor. Objective: This study...
Randomisation was at the individual student level, not at the class or school level, and the intervention was classroom instruction rather than one-to-one tutoring.
Writing outcomes were measured with IELTS academic writing tasks scored using official IELTS band descriptors, a widely recognised standardised assessment.
Outcomes were measured after a 12-week (three-month) intervention period, which covers approximately one academic term.
The control group's composition, baseline pre-test scores, and business-as-usual treatment are documented in detail.
All participants came from one university and were randomised individually, so there was no school-level randomisation.
The authors themselves designed, implemented, rated, and analysed the study with no independent third-party evaluation.
The tracking interval was only 12 weeks, far below the required 75% of a full academic year.
The active control received the same duration, teacher, materials, and verified-equivalent practice time, with only the feedback source (ChatGPT vs teacher) differing as the treatment variable.
The authors explicitly call for future replication, and an internet citation-index search found no independent peer-reviewed replication of this specific study.
Only English writing outcomes were measured, not all main subjects studied by the participants.
Measurement ended at the 12-week post-test, the authors explicitly state no long-term follow-up was conducted, and no graduation-tracking follow-up paper was found online.
The paper contains no mention of a pre-registered protocol or any trial registry entry, and none was found via internet search.
This mixed-methods study evaluates the impact of AI-assisted language learning on Chinese English as a Foreign Language (EFL) students' writing skills and writing motivation. As artificial intelligence (AI) becomes more prevalent in educational settings, understanding its effects on language learning...
The study used cluster randomization at the class (scheduled session) level.
Knowledge was assessed with a study-specific 15-item MCQ rather than a widely recognized standardized exam.
Outcomes were measured immediately after a single session rather than at least one academic term after the intervention began.
The paper documents the control condition and provides baseline characteristics and denominators by group.
The study was conducted at a single medical school and randomized scheduled-session clusters, not schools.
The paper does not document a third-party independent evaluation team; implementation occurred within the same instructional setting.
Outcomes were measured immediately post-session, so the study does not include at least 75% of an academic year of outcome tracking.
The study explicitly kept instructional time, content, and materials the same across groups, varying only delivery mode.
No independent replication study by a different research team could be identified.
Because the study does not use standardized exams (Criterion E is not met), it also cannot meet the all-subject standardized exam requirement.
The study does not track participants to graduation and does not meet the year-duration prerequisite.
The trial reports registration in the Thai Clinical Trials Registry before participant enrollment, and an earlier preprint version states a registration date preceding study conduct.
This randomized, controlled, assessor-blinded trial compared a web-based virtual environment (Spatial) with face-to-face delivery of an otherwise identical team-based learning session on anesthesia techniques for fifth-year medical students at Khon Kaen University (Thailand). Knowledge was assessed with a 15-item multiple-choice...
The study randomized whole classes within schools, satisfying the class‑level RCT requirement.
The tests were custom assemblies of items from exam books, not formal standardized exams.
Student performance was assessed at the end of the fall semester, meeting the term‑duration requirement.
The control classes’ makeup, treatment conditions, and baseline data are clearly reported.
The trial randomised classes within schools rather than entire schools.
The authors who developed the CAL were also responsible for its implementation and assessment.
Tracking ceased at the semester’s end, not over a full academic year.
The additional CAL sessions are the treatment itself, so the control group’s business‑as‑usual status is appropriate.
The paper contains no mention of independent replication by a different research team.
The study assessed only math and Chinese; other core subjects were omitted.
Student outcomes were not monitored beyond the semester, so no graduation tracking occurred.
There is no evidence the trial was pre-registered before data collection.
The education of the disadvantaged population has been a long-standing challenge to education systems in both developed and developing countries. Although computer-assisted learning (CAL) has been considered one alternative to improve learning outcomes in a cost-effective way, the empirical evidence...
Randomisation was performed on individual adult learners recruited by email across three universities, not on whole classes or schools, and the collaborative, instructor- delivered nature of the conditions means the tutoring exception does not apply.
Primary proficiency outcomes were measured with the TOEFL iBT speaking, listening and reading tests together with IELTS-aligned writing tasks scored by Criterion E-Rater, all of which are widely recognised standardised instruments.
The intervention ran for a standardised 12 weeks with a further 8-week delayed posttest, giving roughly 20 weeks (about five months) from intervention start to final measurement, which exceeds one academic term.
The static control condition's content, delivery platform, activities and baseline equivalence with the other arms are described, although group-specific sample sizes and demographics are not tabulated.
Allocation was to individual learners via covariate-adaptive minimisation, with no randomisation of schools, campuses or any other institutional unit.
The sole author is the originator of NDLLT and of the CDS model being tested and also performed the conceptualisation, investigation and analysis, so the evaluation was not independent of the intervention designer.
The tracked period runs about 20 weeks from the start of the 12-week intervention to the 8-week delayed posttest, roughly five months, well under 75% of an academic year.
The control arm was an active, content-matched digital curriculum, time-on-task did not differ across groups, and the extra hardware and adaptive AI given to the treatment arms is the treatment variable under test rather than an unbalanced add-on.
No independent replication of this trial by an unrelated research team was found; a citation-index search located six citing works, five self- or co-authored by Bahari and one unrelated systematic review, none of which is an independent RCT replication.
All outcomes lie within English-as-a-foreign-language proficiency and its neurocognitive and affective correlates; no other curricular subject was assessed and no rationale for that restriction is offered.
Follow-up ended with an 8-week delayed posttest, no graduation or degree-completion data were collected, criterion Y is not met (which also precludes G), and a citation search found no follow-up publication tracking the original cohort.
The paper reports only an institutional ethics approval number and names no trial registry, registration identifier, or pre-registration date preceding data collection; an internet search found no matching registration.
Grounded in a critical-realist ontology and a pragmatic-constructivist epistemology, this study operationalizes Nonlinear Dynamic Language Learning Theory (NDLLT) in AI-mediated EFL classrooms and empirically examines motivation as a fluctuating, history-dependent system. A 12-week randomized controlled trial (N = 784; CEFR...
Randomization was conducted at the individual student level within institutes, not at the class or school level, and the classroom-delivered group intervention does not qualify for the one-to-one tutoring exception.
Speaking outcomes were measured with the IELTS speaking examination scored with the official IELTS Speaking Band Descriptors, a widely recognized standardized exam.
The intervention and outcome tracking spanned 13 weeks from the start of instruction to the posttest, which covers approximately one academic term (about three months).
The control group's size, demographics, baseline scores, and business-as-usual traditional speaking course are documented in detail, with baseline equivalence tests reported.
Randomization occurred at the individual student level within institutes; entire schools or institutes were not randomly assigned to conditions.
The researchers themselves designed the study, delivered the instruction with teachers they trained, and one author personally co-scored the speaking tests, with no independent third-party evaluation.
The study lasted only 13 weeks from intervention start to posttest, far short of 75% of an academic year.
Both groups received an equal amount of instructional time in comparable contexts, with the control group taking an active traditional speaking course, so inputs were balanced.
No independent, peer-reviewed replication of this specific Duolingo-based RCT with Chinese EFL learners is reported or known; related AI-speaking studies are different designs by different teams, not replications of this study.
Only English speaking skills (plus self-regulation) were assessed; no other main academic subjects were measured with standardized exams and no specialization rationale covering this is given.
Measurement stopped at the 13-week posttest with no tracking of participants to graduation, and criterion G also fails automatically because criterion Y is not met.
The paper contains no mention of a pre-registered protocol, registry platform, registration ID, or registration date, and no external registry record for this study was found.
This study investigated the effectiveness of artificial intelligence-based instruction in improving second language (L2) speaking skills and speaking self-regulation in a natural setting. The research was conducted with 93 Chinese English as a foreign language (EFL) students, randomly assigned to...
Two intact classes were randomly assigned to the experimental and control conditions, satisfying class-level randomisation.
Pronunciation outcomes were measured with a researcher-built read-aloud task and in-house ratings, and the IELTS-format speaking assessment was administered unofficially by the research team, so no recognised standardised exam was used as an official outcome measure.
The intervention and pre/post measurement spanned a 14-week course, which is approximately one academic term.
The control group's size, demographics, baseline scores, and the teacher-led instruction it received are well documented.
Randomisation was of two classes within one language training center, not of schools or institutions.
The same researcher designed, oversaw, rated, and analysed the study with no independent third-party evaluation.
The study spanned only a 14-week course, far less than 75% of an academic year.
Both groups received matched lesson time and structure with an active teacher-feedback control, the ASR tool being the integral treatment variable.
No independent replication of this specific study is reported; the paper frames itself as filling a research gap, and a targeted internet search found no such replication.
Only L2 pronunciation and speaking were assessed; no other subjects or skills were measured.
Measurement ended at the close of the 14-week course with no follow-up tracking until graduation, and no such follow-up publication was found via internet search.
The paper mentions only ethics approval and contains no reference to a pre-registered protocol or registry ID; no registration record was found via internet search.
This study employed an explanatory sequential design to examine the impact of utilizing automatic speech recognition technology (ASR) with peer correction on the improvement of second language (L2) pronunciation and speaking skills among English as a Foreign Language (EFL) learners....
The study is a self-described quasi-experimental non-equivalent group design with individual students, not classes or schools, allocated to conditions, so class-level RCT is not met.
Outcomes were rated with the CAFIS research rubric on study-designed speaking tasks rather than any widely recognised standardised exam, so criterion E is not met.
Outcomes were measured 12 weeks (about three months) after the intervention began, which just reaches the one-term threshold, so criterion T is met.
The control group's size, composition, baseline scores, and exact activities are documented in detail, so criterion D is met.
All participants came from a single university and were allocated individually, so there was no school-level randomisation and criterion S is not met.
The evaluated intervention (the commercial ELSA Speak app) was not designed by the authors, who declare no competing interests, and scoring was done by raters blind to group assignment.
The study tracked outcomes for only 12 weeks, far less than 75% of an academic year, so criterion Y is not met.
Both groups received identical class time and matched 30-minute daily practice, differing only in practice mode (the treatment itself), so criterion B is met.
This recently published study has not been independently replicated, and earlier ELSA Speak studies are not replications of it, so criterion R is not met.
Only pronunciation was measured with a non-standardised rubric (E fails), so the all-subject exams criterion is not met.
Measurement stopped at the week-12 post-test with no tracking of the second-year students to graduation, so criterion G is not met.
No pre-registration, registry ID, or registration date is mentioned anywhere in the paper or found via internet search, so criterion P is not met.
Pronunciation plays a pivotal role in oral fluency and overall communicative effectiveness in English as a Foreign Language (EFL) learning. However, pronunciation instruction remains a persistent challenge in higher education contexts, particularly where learners have limited opportunities for individualized feedback...
Randomisation was carried out at the class level (three intact classes randomly assigned to conditions), meeting the class-level RCT requirement.
Outcomes were single researcher-administered essay tasks scored by iWrite and Coh-Metrix rather than an official standardised exam administration, so the exam-based assessment requirement is not satisfied.
The intervention ran from week 3 to week 14 with the posttest in week 15 of a semester-long study, so outcomes were tracked for roughly one full academic term after the intervention began.
The control class (n=30) is documented with baseline pretest scores, population details, and a clear description of the teacher-only feedback condition it received.
Randomisation was at the class level within one university department, not at the school or institution level.
The intervention designers themselves conducted, collected, and analysed the study with no independent evaluators or third-party oversight.
The whole study, from pretest to posttest, spanned only one 15-week semester, far short of 75% of an academic year.
All groups had the same teacher, lectures, tasks, and iWrite training, and the extra AWE/peer feedback in the experimental groups was the explicit treatment variable being tested, so the design is balanced.
The study is framed as novel gap-filling research, was published very recently (May 2025), and a July 2026 internet search found no independent peer-reviewed replication of this specific three-arm feedback design.
Criterion E is unmet and outcomes covered only English writing quality, not all main subjects, so the all-subject exams requirement fails.
Measurement ended at the week-15 posttest in the students' sophomore year with no follow-up to graduation, prerequisite criterion Y also fails, and no follow-up publications tracking this cohort were found online.
The paper contains no reference to any pre-registration, registry ID, or published protocol, and none was found through an external registry search.
Nowadays, Chinese EFL learners are increasingly using automated writing evaluation (AWE) to provide feedback on their writing. However, AWE is relatively inadequate for providing elaborate feedback on the global aspects of writing that require human judgment. Thus, how to combine...
Randomisation was performed at the class level, with one class of each matched pair randomly assigned to the experimental condition, satisfying the class-level RCT requirement.
The outcome was measured with a researcher-generated vocabulary test drawn from the app's own word pool, not a recognised standardised exam.
The intervention and tracking spanned one full academic semester, from the first to the last class of the term, meeting the one-term duration requirement.
The control group's size, demographic composition, selection, baseline scores, and business-as-usual conditions are documented in the text and tables.
Randomisation occurred among classes within a single university, not among whole schools, so the school-level RCT requirement is not satisfied.
The same person designed the app and the study and conducted and analysed the experiment, with no independent or third-party evaluation.
The study lasted only a single semester, which is far short of the 75% of a full academic year required.
Both groups received the identical 3,402-word list and same textbook; the smartphone app itself was the explicit treatment variable, with the control given the equivalent printed materials as business as usual.
No independent replication of this specific study by another research team was found after internet searching.
Only English vocabulary was assessed, using a non- standardised custom test, so with criterion E unmet the all-subject requirement fails.
Outcomes were measured only at the end of one semester with no follow-up to graduation, and with criterion Y unmet this criterion fails.
No pre-registration of the study protocol on any registry is mentioned anywhere in the paper, and none was found through internet searching.
The researcher designed a smartphone app to help college students to learn English (L2) vocabulary. The app contained 3,402 English words that were compiled into an alphabetic wordlist with each word displayed on three features; namely: spelling, pronunciation and Chinese...
Randomization was conducted at the class level with intact classes assigned to each condition.
The study employed custom-designed tests of graphing and slope problems rather than a recognized standardized exam.
The study measured outcomes after approximately three months, satisfying the term-duration requirement.
The control group’s size and baseline comparability (NAEP scores) are documented in detail.
Randomization was at the class level within one school, not across multiple schools.
The study was conducted and scored by the authors, with no independent external evaluation.
Outcomes were measured within three months, not tracked over an academic year.
Both groups received the same number of assignments, problems, and review sessions, ensuring balanced time and resources.
There is no evidence of an independent replication study by a different research team that confirms these findings.
The study only assessed mathematics graphing and slope problems, not a full range of subjects.
Follow-up ended at 30 days post-review, with no tracking until graduation.
The paper does not report a pre-registration or registry before the intervention began.
A typical mathematics assignment consists primarily of practice problems requiring the strategy introduced in the immediately preceding lesson (e.g., a dozen problems that are solved by using the Pythagorean theorem). This means that students know which strategy is needed to...
Randomisation was performed on intact classes rather than on individual students, explicitly to avoid within-class contamination, satisfying the class-level RCT requirement.
All outcomes relied on locally administered course examinations and questionnaires developed specifically for this study, with no widely recognised standardised exam used.
The intervention spanned a full 16-week semester with outcomes measured at its end and again one month later, exceeding the one-term requirement.
The control arm's size, demographics, baseline pre-test scores and exact instructional conditions are all documented in the text, Table 1, Table 2 and the CONSORT diagram.
Randomisation was at the class level within a single medical college, so no school- or institution-level randomisation occurred.
The same Cangzhou Medical College team designed the teaching model, delivered it, developed the measures and analysed the data, with independence reported only for the allocation procedure.
Tracking spanned only a 16-week semester plus a one-month follow-up, roughly five months, which is well short of 75% of an academic year.
Contact time and instructional structure were identical across arms with active control activities in every instructional link, and the only added resource, the AI digital tutor and knowledge graph, is the explicit treatment variable being tested.
No independent replication was found in the paper or by internet search; the study is framed as novel, calls for future validation, and was published only two months before this check.
Outcomes were confined to the single intervention course with no measurement of other curriculum subjects, and criterion E was not met, which independently precludes criterion A.
Tracking stopped one month after the final examination for first-year students years away from graduation, no follow-up publication was found by internet search, and criterion Y was not met.
The paper reports only institutional ethics approval and contains no trial registry identifier, link or registration date, and no registration record was found by internet search of Europe PMC or clinical trial registries.
Background: Human Anatomy and Histology & Embryology are foundational courses in nursing education but are often challenging for students due to their complex spatial structures, dense knowledge points, and high cognitive load. Recent advances in generative artificial intelligence (AI) and...
Randomization was at the individual student level, but the intervention is individualized simulator skill training with no shared-classroom contamination, so the personal-teaching exception applies and student-level assignment is acceptable.
Outcomes were custom study-specific measures (needle tip disappearance frequency, system-derived oscillation amplitude, and a blinded three-level safety rating), not a widely recognized standardized exam.
The entire intervention and outcome measurement occurred within a single short session (5-minute practices and immediate tests), far shorter than one academic term.
The control group's size, age, sex, baseline (no prior UGRA experience), and condition (identical training without feedback) are documented.
This was a single-center trial with randomization at the individual student level, not at the school/institution level.
The same team designed and developed the feedback system and also conducted the trial; although the outcome assessor was blinded, there was no independent third-party evaluation team.
Term Duration (T) is not met, so Year Duration cannot be met; the study lasted a single short session, nowhere near a year.
Both groups received identical instructional videos and equal 5-minute practice time; the only difference is the real-time feedback, which is the explicit treatment variable being tested, so time and resources are balanced.
This is described as the first implementation of the system and a pilot study; an internet search found no independent replication by any other research team.
Criterion E is not met and the study measured only a single procedural-skill domain, so All-subject Exams cannot be met.
Criterion Y is not met and there was no follow-up beyond the immediate post-practice test; an internet search found no subsequent graduation-tracking papers by the same authors.
The trial was prospectively registered with the UMIN Clinical Trials Registry (UMIN000055602) on October 1, 2024, before data collection began on November 1, 2024.
Background: Proper needle visualization is a major technical challenge for novices learning ultrasound-guided regional anesthesia (UGRA). We developed a 'You Only Look Once version 5' (YOLOv5)-based system that records quantitative needle trajectory data and delivers real-time visual and auditory feedback...
Although randomization was not at the class level, the intervention was a fully individualized, at-home, computer-based program, satisfying the personal tutoring exception in the ERCT standard.
The study employed the standardized TOWRE subtests for reading fluency, meeting the exam‑based assessment requirement.
Outcome assessments occurred after at least 16 weeks of training (and within a 6‑month participation period), exceeding the one‑term minimum requirement.
The study lacked a separate control group, relying on a within‑subject baseline, thus failing the documented control group requirement.
The study randomized individual children rather than entire schools, so the school‑level RCT requirement is not met.
The intervention was conducted and monitored by the same team that designed it, failing the independent conduct requirement.
Participants were observed for about 6 months, not a full academic year, so the year‑duration requirement is not met.
The control condition had no training or additional support, so resources were not balanced.
No independent research team has published a replication of this trial, so reproducibility is not established.
The study measured only reading skills without assessing other core subjects, so the all-subject exams requirement is not met.
Participants were followed for about 6 months, with no data collection continuing through graduation, failing the graduation tracking requirement.
The trial was prospectively registered long before participants were enrolled, satisfying the pre-registered protocol requirement.
Given the importance of effective treatments for children with reading impairment, paired with growing concern about the lack of scientific replication in psychological science, the aim of this study was to replicate a quasi‑randomised trial of sight word and phonics...
Individual students (not classes or schools) were randomized, and the intervention is not a tutoring-style exception.
Outcomes were measured via scenario-based checklists and manikin software rather than a widely recognized standardized exam.
Outcomes were assessed up to one year after the intervention, exceeding a term-length follow-up.
Both conditions are clearly described and baseline characteristics are reported, providing sufficient documentation for comparison.
Randomization occurred among individual students rather than among schools (or equivalent institutional units).
Key trial activities were performed by the authors/their department rather than a clearly external independent evaluator.
Outcomes were assessed one year after the last educational intervention, satisfying the year-duration requirement.
The added time/instructor input is the intervention contrast being tested (additional practice with continuous assessment vs a practical examination), so unequal resources are integral rather than a confounding add-on.
No independent replication of this specific RCT was found as of the ERCT check date.
Because standardized exams are not used (Criterion E not met), the all-subject standardized exam criterion cannot be met.
Follow-up ends at one year and does not track participants through graduation from their educational stage.
The paper states "Clinical trial number Not applicable" and does not provide a pre-registration record.
Background Proper basic life support (BLS) skills are crucial for laypeople and health care professionals to increase the survival of cardiac arrest patients. A practical examination at the end of a BLS course may be beneficial for prolonging skill retention....
Randomisation was at the individual student level, not the class or school level, and the study is not a one-to-one tutoring intervention.
The study used the standardised, externally developed MIRA readiness assessment and the state SBAC test rather than a custom-built measure.
The intervention and outcome measurement spanned only three weeks, far short of the one-term requirement, with no term-long follow-up.
The comparison groups' demographics, baseline scores, and conditions are documented in the text and in Table 1.
Randomisation was at the individual student level, not at the school or site level.
The same authors who designed the RAMPUP intervention also conducted and analysed the evaluation, with no independent evaluator.
The study lasted only three weeks, far short of a full academic year, and criterion T was not met.
District A's comparison was a time-comparable competing summer program, and District B tested the summer program's added time as the integral treatment variable against a business-as-usual baseline.
This newly developed RAMPUP program and its trials have not been independently replicated by another research team.
Only mathematics outcomes were measured, with no assessment of other main subjects.
Outcomes were measured only at the end of the three-week program, with no follow-up or graduation tracking, and criterion Y was not met.
Both trials were pre-registered at the Registry of Efficacy and Effectiveness Studies with specific IDs, and the paper discusses adherence to the pre-registered plan.
We report results from the implementation and impact evaluations of a three-week summer bridge mathematics program. This program was designed to challenge and support rising ninth grade students and in particular Multilingual Learners. We report the extent to which implementing...
Randomisation occurred at the school level, which satisfies the class‑level RCT requirement.
The authors used a bespoke nine‑item quiz rather than a recognised standardised exam.
The intervention lasted only a few hours over several weeks, not a full term.
The control group’s composition, baseline performance, and lack of intervention are clearly documented.
Randomisation at the school level fulfills the school‑level RCT criterion.
The intervention was designed, delivered, and assessed by the authors’ team with no third‑party evaluation.
The experiment ran under one academic year, so year‐long criterion is unmet.
The control group received no equivalent time or resources, so control was not balanced.
No independent replication of this intervention has been reported.
Only financial literacy was assessed, so all‐subject exam criterion is unmet.
The study tracked outcomes only up to seven weeks post‑intervention, not through graduation.
The trial was pre‑registered in the AEA RCT Registry before data collection.
This paper provides causal evidence on the effects of parental involvement on student outcomes in a financial education course based on two randomised controlled trials with a total of 2,779 grade 8 and 9 students in Flanders. Using an experimental design...
Randomization was conducted at the class (tutorial) level, satisfying the class‑level RCT requirement.
The study used author‑designed pre/post tests rather than a standardized exam.
Outcome measures were collected within days, not after a full academic term.
The handout control group is well described with baseline tasks and scores.
Randomisation occurred at the tutorial‑class level, not at the school level.
The authors’ own team both developed and evaluated the intervention.
Follow‑up lasted only days, not a full academic year.
The AR intervention entailed multimedia app access not matched by the handout.
Independent teams reproduced similar AR‐enhanced learning gains in other contexts.
Only sewing/textiles skills were assessed, not all core subjects.
No follow‑up beyond the immediate post‑workshop period.
No evidence of pre‑registration is provided.
This study contributes to enhancing students’ learning experience and increasing their understanding of complex issues by incorporating an augmented reality (AR) mobile application (app) into a sewing workshop in which a threading task was carried out to facilitate better learning...
Randomisation was conducted at the individual student level, not at the class or school level, and the intervention is not a one-to-one tutoring exception.
The study used a custom-built MCQ test plus a program-specific checklist and generic rating scale, not a widely recognised standardised exam.
Outcomes were measured immediately after a single short training course, far short of the required term-long follow-up.
The control group's size, demographics, baseline knowledge, and conditions are clearly documented in the text, Table 1, Table 3, and the CONSORT diagram.
The study was conducted at a single center with individual-level randomisation, so no school-level randomisation occurred.
The same authoring team designed, conducted, and analysed the trial, so the study was not conducted independently despite the use of blinded assessors.
Criterion T is not met and outcomes were measured immediately post-training, so the year-duration requirement is not met.
Instructional time and content were matched between groups, with the control receiving a time-equivalent 2D-atlas activity, and the differing IVR modality is the integral treatment variable.
No independent replication of this specific IVR-for-ETI trial by a different research team has been published.
Criterion E is not met and only a single narrow ETI-related domain was assessed, so the all-subject requirement is not met.
Criterion Y is not met and there was no follow-up beyond immediate post-training, so no graduation tracking occurred.
The trial was prospectively registered (ChiCTR2400093689) on 10 December 2024, before the December 2024-March 2025 data collection, with pre-specified endpoints.
Introduction: This study aimed to compare immersive virtual reality (IVR)-assisted versus conventional anatomy training for teaching endotracheal intubation (ETI) to novice non-anesthesiology residents enrolled in China's Standardized Residency Training program. Methods: A total of 90 non-anesthesiology residents without prior ETI...
Randomisation was conducted at the individual student level within a single university cohort, not at the class or school level.
The study used researcher-developed custom knowledge and skill tests rather than standardised, widely recognised exams.
Outcomes were measured only 7 days after the intervention began, far short of one academic term.
The control group's demographics, baseline scores, size, and conditions are documented in detail in the CONSORT flow and Tables 1-3.
Randomisation occurred at the individual student level within a single institution, not at the school level.
The same authors designed the VSG intervention and also conducted and analysed the study, with no independent third-party evaluator.
The tracking interval was about one week, far short of a year, and criterion T is not met.
The control group received an equivalent 7-day window of review materials, and the VSG is the explicit treatment variable tested against the standard method of education.
The study describes itself as the first of its kind and no independent replication by another team was found.
Criterion E is not met and the study assessed only a single specialised topic, so the all-subject exam criterion fails.
Criterion Y is not met and the study did not track participants beyond a 7-day post-test, let alone to graduation.
The trial was registered on ClinicalTrials.gov (NCT05309317, first posted 4 April 2022) with a published protocol before data collection began on 21 May 2022.
Aim: This study evaluated the effect of a virtual simulation game (VSG) on the development of nursing students' knowledge and skills related to the prevention of Catheter-Associated Urinary Tract Infection (CAUTI). Design: A parallel group randomized controlled trial. The study...
Randomisation was conducted at the cluster (rotation-subgroup) level rather than the individual student level, satisfying the class-level RCT requirement.
The study used a custom, author-developed 30-item achievement test rather than a widely recognised standardised exam, so this criterion is not met.
Outcomes were measured immediately at the end of a short pediatric neurology rotation component, well under one full academic term, so this criterion is not met.
The control group's size, demographics, baseline scores, and conditions are clearly documented, meeting the requirement.
Randomisation occurred among rotation subgroups within a single medical school rather than across multiple schools or sites, so the school-level RCT criterion is not met.
The same team that designed and built the intervention also delivered it, collected data, and analysed results, with no independent evaluator, so this criterion is not met.
Outcomes were measured within a short rotation window, far less than 75% of an academic year, and criterion T is not met, so Y is not met.
The extra resource (the ARLE app) is the explicit treatment variable tested against an identical standard curriculum, so the business-as-usual control is appropriate and the criterion is met.
No independent replication of this specific pediatric neurology ARLE trial exists; only loosely related AR/VR studies in other domains are cited, so this criterion is not met.
The study used a custom single-subject test (criterion E not met) and assessed only pediatric neurology content, so the all-subject exams criterion is not met.
Measurement stopped immediately after the rotation with no graduation tracking, no follow-up publication exists, and criterion Y is not met, so G is not met.
The paper reports ethics committee approval but provides no public trial pre-registration or registry ID, and none was found externally, so this criterion is not met.
Limited and variable case exposure during pediatric neurology rotations may leave some clinically important presentations underrepresented within a single rotation window. We evaluated whether augmented reality–supported learning environments (ARLEs) integrated into the rotation improve medical students' cognitive achievement beyond the...
Randomization was performed at the individual worker level using even/odd random numbers, not at the class or school level, and the intervention is group-based OSH training rather than one-to-one tutoring.
The study used the Q626 questionnaire, a custom instrument developed by the authors specifically for this study, not a widely recognized standardized exam.
The intervention was a single one-day 4-hour course and the outcome was measured only five days later, far short of one full academic term.
The control group is documented in terms of size, baseline demographics, baseline knowledge scores, and its condition (seminar only, no games).
Randomization was conducted at the individual worker level, not at the school or institution level.
The same authors designed the intervention, the games, and the Q626 instrument and also conducted and analyzed the trial, with no independent or third-party evaluator.
The intervention was a one-day course with outcome measured five days later, nowhere near 75% of an academic year, and criterion T is not met.
The control group received the same seminar and the additional gaming session was the explicit treatment variable being tested against a seminar-only business-as-usual baseline.
This specific gamified OSH trial has not been independently replicated by another research team in a peer-reviewed publication.
The study measured only occupational health and safety knowledge via a custom questionnaire, not all main subjects, and criterion E is not met.
The study only measured immediate post-training recall with no follow-up to graduation, and criterion Y is not met.
The study was pre-registered on ClinicalTrials.gov (NCT06060379) in September 2023, before recruitment began in December 2023, and a protocol was published.
Background: Traditional health and safety training often fails to engage workers effectively. Gamification has emerged as a promising strategy to enhance motivation and learning outcomes. This study aimed to assess the effectiveness of a gamified training course in improving occupational...
Randomization was performed at the individual student level within shared classes rather than at the class or school level, so the class-level RCT criterion is not met.
Outcomes were measured with self-report scales and custom, curriculum-aligned domain-knowledge items developed by the study teachers, not a widely recognised standardised exam.
The intervention and outcome measurement spanned only six 45-minute sessions over a few weeks, far shorter than one academic term.
The control group's size, demographics, baseline characteristics, and "standard ChatGPT" condition are documented in detail, including baseline-equivalence testing.
Randomisation was at the individual student level, not at the school level, so the school-level RCT criterion is not met.
The same author team designed the GenAI intervention, developed the prompts and measures, and conducted and analysed the study, with no independent third-party evaluator.
The study lasted only six 45-minute sessions with immediate posttesting, far short of (and because T is not met, automatically failing) the year-duration requirement.
All three conditions received identical instructional content, time, task formats, and technical settings, with the only difference being the chatbot prompts, so time and resources were balanced across groups.
This specific RCT has not been independently replicated by a different research team in a peer-reviewed outlet.
Outcomes covered only the two intervention subjects (physics and English) using custom items, and because criterion E is not met, criterion A automatically fails.
Students were assessed immediately after the intervention with no follow-up to graduation, and because criterion Y is not met, criterion G automatically fails.
The study, including its hypotheses and data-analysis plan, was pre-registered at OSF, and data collection (T1/T2 in April-May 2025) occurred after pre-registration.
This pre-registered randomized controlled trial investigated whether generative AI (GenAI) tools can support secondary school students' self-regulated learning (SRL) when integrated into regular classroom lessons. A total of 371 students (Grades 7-9) were randomly assigned at the individual level to...
Two intact sixth-grade classes were randomly assigned to conditions, making the class the unit of randomisation.
The study used research-developed and researcher-adapted instruments rather than recognised standardised exams.
The intervention and post-test spanned only six weeks, far shorter than one academic term, with no follow-up.
The control group's size, demographics, instruction, and baseline scores are clearly documented.
Randomisation was at the class level within one school, not at the school level.
The intervention designers also collected and analysed the data, with no independent or external evaluator.
The study spanned only six weeks, far short of an academic year, and criterion T was not met.
Both groups received identical instructional time, and the only difference is the integral teaching medium being tested.
No independent replication of this specific study by a different research team could be found in any source.
Only a single science topic and CT were assessed, and the prerequisite standardised-exam criterion E was not met.
There was no follow-up to graduation, and the prerequisite year-duration criterion Y was not met.
The paper reports ethics approval but no public pre-registration of the study protocol.
In the digital era, fostering computational thinking (CT) skills is essential for scientific literacy and engaging students in authentic scientific practices. Block-based programming environments (BBPEs), such as Scratch, effectively integrate these skills into K-12 science by lowering the syntax barrier,...
Randomization was conducted at the class level using intact classes, satisfying Criterion C.
Outcomes were measured with researcher-run EF tasks rather than a standardized exam-based educational assessment.
Outcomes were measured within the same morning rather than at least one academic term after intervention start.
The comparison conditions and baseline assessment are described, and group/class demographics are reported in an appendix.
Randomization occurred within one school at the class level, not by assigning multiple schools to conditions.
The paper provides no explicit statement that the study was conducted by an independent external evaluation team.
Outcomes were assessed acutely within a single morning, which is far shorter than 75% of an academic year.
The intervention does not add extra instructional time or budget, and all conditions are scheduled within the same pre-class time window.
No independent peer-reviewed replication of this specific trial was found in the paper or via external searching.
Criterion A is not met because Criterion E is not met and the study does not use standardized exams across core subjects.
The study does not track students to graduation and, since Criterion Y is not met, Criterion G is not met.
No pre-registration is reported in the paper, and external registry searching did not identify a registration clearly dated before the study period.
Brief pre-class physical exercise may enhance executive function (EF) processes proximal to mathematics learning. However, the effect is likely dependent on the alignment between the physical exercise and subsequent instructional demands. We tested a dual-process framework in authentic classrooms, predicting...
Although randomization was at the individual level, this is acceptable here because the intervention is an individual, personal-training activity rather than a class-delivered program.
Outcomes were measured with a modified skills assessment tool rather than a widely recognized standardized exam-based assessment.
Outcomes were measured immediately after brief training, not at least one academic term after the intervention began.
The paper documents both groups, sample sizes, baseline characteristics, and comparative results.
Randomization was not conducted at an institution or site level; it was individual-level allocation.
The paper does not document an independent external evaluation team; the authors designed, delivered, and analyzed the study themselves.
Outcomes were not measured over at least 75% of an academic year, and criterion T is also not met.
The study appears to keep training exposure (time and peer observation) comparable across groups, while the guidance modality (AI vs expert) is the intended treatment difference.
No independent replication of this specific custom-made software and trial by a different research team was found.
Criterion E is not met, so criterion A is not met, and the outcomes are not all-subject standardized exams.
The study does not track participants to graduation, and criterion Y is not met so criterion G is not met.
The trial registration date is after the study began, so the protocol was not pre-registered before data collection.
Flexible bronchoscopy is an essential tool for airway management and both diagnostic and therapeutic interventions, particularly in critical care. Accurate identification of tracheobronchial structures is crucial but challenging for less experienced clinicians, often leading to prolonged procedures and increased complication...
Entire intact classes were randomly assigned to experimental vs. control conditions, satisfying class-level randomization.
Outcomes were measured with questionnaires rather than a standardized exam-based assessment.
The follow-up occurred after about 10 weeks, which is shorter than a typical academic term (~3–4 months).
The control class is described with sample size, demographics, the control curriculum, and baseline outcome comparability.
Randomization occurred between classes within one institute, not between multiple schools/sites.
The second and third authors conducted data collection and analysis, and no independent evaluator is documented.
Outcomes were measured after about 10 weeks, far less than 75% of an academic year, and T is not met.
Both groups received comparable instructional time and access, and the control condition was an active course of equal duration.
No independent replication by a different research team was found in the paper or in post-publication searching.
The study does not use standardized exams (E is not met), so it cannot meet the all-subject standardized exam requirement.
The study does not track participants until graduation, and it cannot meet this criterion because Y is not met.
No preregistration record (registry name/ID and pre-data-collection date) is reported or could be verified.
This quasi-experimental study investigated the impact of World Englishes education on enhancing EFL learners’ English language identity and their English emotional resonance in China. Data were collected via a survey comprising two questionnaires. Two teaching classes, which initially exhibited no...
Individual-level assignment is used, but the intervention is explicitly framed as one-to-one tutor-like chatbot practice, fitting the ERCT tutoring exception.
The study uses a custom, study-specific grammar task based on participants’ own sentences rather than a standardized exam.
The study measures outcomes after about one month (4 weeks) from intervention start, which is shorter than a full academic term.
The paper clearly defines the two conditions (ICF vs. DCF) and reports group sizes and participant characteristics, documenting the comparison group adequately.
Assignment is at the participant level (across classes), not at the school/site level.
The authors built the chatbot system and also conducted and analyzed the study, with no documented independent evaluator leading the evaluation.
The study lasts about one month (4 weeks), far short of 75% of an academic year; additionally, since T is not met, Y is not met by ERCT rule.
The two conditions are explicitly described as having the same exposure and feedback generation, differing only in the timing of feedback.
No independent replication of this specific experiment by a different research team was found during the ERCT check.
The study does not use standardized exams (E is not met), so all-subject standardized exam assessment (A) is automatically not met.
The study ends after the post-session assessment/forms and does not track participants to graduation; additionally, since Y is not met, G is not met by ERCT rule.
The paper links to OSF/GitHub for data and code but provides no preregistration identifier/date, and no preregistration record could be verified during the ERCT check.
The emergence of Large Language Models (LLMs) has opened new possibilities for language learning through conversational interaction with chatbots. Yet, little empirical evidence exists on how students experience such interactions and how corrective feedback should be provided. Research suggests that...
Participants were randomized at the individual student level (by survey software), not at the class (or higher) level, and the intervention is not one-to-one tutoring.
The outcome is lecture attendance (QR-based monitoring), not performance on a standardized exam-based assessment.
Outcomes were tracked across an entire 11-week teaching semester, which is a complete academic term in this context.
The control condition is clearly described (including what participants did), and baseline comparability and group sizes are reported.
Randomization occurred among individual students within one university programme, not among schools (or similar institutional units).
The paper provides no clear evidence that an independent third party conducted implementation, measurement, or analysis separate from the intervention author team.
Outcomes were tracked only over an 11-week teaching semester, which is far less than 75% of an academic year.
The control condition is an active control receiving comparable materials and a closely matched task, and the extra time in the encoding-facilitation arm is integral to that treatment.
No evidence was found of an independent replication of this lecture attendance VHS RCT by a different research team.
Because the study does not use standardized exam-based assessments (Criterion E not met), it cannot satisfy the all-subject standardized exam requirement.
The study does not track participants to graduation and, because the year-duration criterion is not met, this criterion is automatically not met.
The paper links to an OSF pre-registration but does not provide a registration timestamp, and the OSF record date could not be verified against the study start.
Background: A volitional help sheet (VHS) is an intervention that promotes the formation of implementation intentions. Research has established that VHSs can change a range of behaviours, including increased attendance at online university lectures. However, research has not yet tested...
Participants were randomized individually to VR vs. audio rather than by intact classes (and the intervention is not tutoring).
The main listening comprehension outcome uses a study-developed (tailor-made) test rather than a standardized exam.
Outcomes were measured immediately after a single-session task, not at least one academic term after the intervention began.
The control (audio) condition is described, group sizes are given, and baseline comparability on key measures is reported.
Assignment was at the individual participant level, not by schools (or equivalent sites/centers).
The VR intervention content was taken from a third-party VR application, and the study reports no funding or role of the application developers in the evaluation.
The study is a single-session experiment rather than tracking outcomes for at least 75% of an academic year.
The VR group received extra technology/orientation time, but these resources are integral to the treatment contrast (VR vs. audio), and the control provides a clear alternative listening mode.
No independent replication of this specific study was found in the paper or via internet searching as of 2026-04-13.
Criterion E is not met (no standardized outcome exam), so A is automatically not met; additionally, outcomes focus on L2 listening rather than all core subjects.
Criterion Y is not met (single-session, not year-long), so G is automatically not met; no graduation tracking or follow-up papers were found.
The paper reports IRB approval but provides no preregistration registry/ID/date, and no external preregistration record was found.
Listening in the real world involves both verbal and non-verbal inputs. However, second language (L2) listening activities in the classroom often lack non-verbal inputs and are removed from the situational and cultural contexts where they would naturally occur. Virtual reality...
Randomization is at the student/session level, but the intervention is one-to-one tutoring, which satisfies the tutoring exception for criterion C.
Outcomes are measured using Eedi platform activity rather than widely recognized standardized exams.
Outcomes are measured within a seven-week study window, which is shorter than a full academic term.
The control condition (static hints) is clearly described and quantified, and the paper documents participant and school context.
The study includes five schools, but randomization is at the student level rather than the school level.
The evaluation is conducted by the intervention-affiliated teams, and the paper does not document an independent third-party evaluation team.
The study lasts seven weeks, far less than 75% of an academic year, and Y also fails because T is not met.
The intervention conditions add substantial tutoring resources compared to static hints, but this resource increase is integral to the intervention contrasts being tested.
No independent peer-reviewed replication of this specific RCT was found, and the paper itself only encourages future replication.
The study focuses on mathematics only and does not assess all core subjects using standardized exams; additionally, A cannot be met when E is not met.
The study does not track students to graduation, and G also fails automatically because Y is not met.
The paper reports ethics review but provides no public pre- registration record, registry ID, or registration date.
One-to-one tutoring is widely considered the gold standard for personalized education, yet it remains prohibitively expensive to scale. To evaluate whether generative AI might help expand access to this resource, we conducted an exploratory randomized controlled trial (RCT) with N...
The trial randomized at the class level (16 classes) rather than assigning different conditions to students within the same class.
Outcomes were measured with self-report questionnaires rather than standardized exam-based assessments.
Outcomes were measured within about 7–8 weeks of intervention start, which is shorter than a full academic term.
The control condition is clearly described, sample sizes are given, and baseline equivalence across groups is reported.
Randomization was at the class level within two schools, not at the school level.
The paper does not document that the evaluation was conducted by an independent third party separate from the intervention developers.
Outcomes were measured far short of 75% of an academic year, and criterion T is also not met.
The interventions were delivered during normal instructional time, and the additional materials/attention appear integral and modest rather than a large unmatched resource increase.
No independent replication of this specific RCT was found in the paper or via internet searching as of the ERCT check date.
Because criterion E is not met and outcomes are not standardized exam scores across subjects, the all-subject exam requirement is not met.
The study includes only immediate post-intervention assessment and no long-term follow-up; additionally, Y is not met so G cannot be met.
No trial registry link/ID or dated statement shows that the protocol was pre-registered before data collection began.
Objectives This study aimed to examine the association between gratitude and defending behavior in response to school bullying among Chinese early adolescents and to evaluate the effects of three types of interventions (i.e., gratitude curriculum, gratitude journal, and gratitude visits)...
Randomisation was at the parent (individual) level, not at the class level.
The study used custom and adapted questionnaires rather than standardised exams.
A six‑month follow‑up assessment provided outcome measurement after at least one full academic term.
The wait‑list control group is described with detailed demographics and conditions.
Randomisation occurred at the individual level, with no schools assigned as units.
The authors who designed the program also delivered and assessed it without independent evaluation.
Follow‑up lasted six months, shorter than the full academic year required.
The study explicitly tests additional training resources as the intervention; the control group remained business‑as‑usual.
No independent replication by another research team is mentioned.
Academic outcomes were measured by custom questionnaires, not in all main subjects via standardised exams.
The study conducted only a six‑month follow‑up and did not track to graduation.
The study was registered after the trial began (ACTRN12613000660785), so it was not truly pre‑registered.
This study evaluated the effects of Group Triple P with Chinese parents on parenting and child outcomes as well as outcomes relating to child academic learning in Mainland China. Participants were 81 Chinese parents and their children in Shanghai, who...
Students (not intact classes or schools) were randomized to groups, so the class-level RCT requirement is not satisfied.
Outcomes were measured using researcher-designed instruments rather than a widely recognized standardized exam.
Outcomes were measured at Week 25 following a 24-week intervention, which exceeds one academic term.
The control condition and baseline comparability are described, including a baseline equivalence table with demographics and pre-test scores.
The unit of randomization is students rather than schools, so this is not a school-level RCT.
The paper does not provide evidence that an independent third party conducted the evaluation or analysis.
The study measures outcomes at Week 25 after a 24-week intervention, which is below the ERCT threshold for year-duration tracking.
The paper does not document extra instructional time for the experimental group beyond the school schedule, implying the intervention substitutes for regular class time rather than adding resources.
No independent, peer-reviewed replication of this specific study was found (in the paper or via an internet search as of 2026-03-14).
Criterion E is not met (no standardized exams), and the study does not use standardized exams across core school subjects.
The study does not track students to graduation, and Criterion Y is not met, which prevents meeting the ERCT dependency requirement for G.
No pre-registration registry, ID, or registration date is reported, and no external registry entry was found as of 2026-03-14.
Introduction: This research paper presents an empirical investigation of the effectiveness of early coding instruction in improving problem-solving skills and computational thinking (CT) among primary school students. The primary research question was to determine whether a structured six-month coding intervention...
The paper specifies cluster randomization at the class level (with school as a stratification factor), meeting the class-level RCT requirement.
The primary educational outcome is assessed with study-specific tools (structured interviews and a custom observation framework) rather than a widely recognized standardized exam.
Outcomes are assessed immediately after a 5-week intervention rather than at least one full academic term after the intervention begins.
The comparator is described, but the paper (as a protocol) does not report realized control-group baseline demographics and baseline performance data by group.
Schools are not randomized to conditions; instead, classes within each school are randomized, so this is not a school-level RCT.
The study is designed, supported, and analyzed by the research team, and no independent external evaluation entity is documented.
The start-to-measurement window is approximately five weeks, far less than 75% of an academic year.
Both arms receive the same instructional dosage (2 h/week for 5 weeks) and comparable supports, so additional time/resources are balanced across conditions and the key difference is instructional setting.
No independent peer-reviewed replication of this specific trial was found, and the paper describes the trial as ongoing and calls for future replication.
Because criterion E is not met, criterion A is automatically not met and the study does not use standardized exams across all core subjects.
The study does not track participants to graduation, and because year duration (Y) is not met, graduation tracking (G) is automatically not met.
The trial is registered on ClinicalTrials.gov with registration dates prior to the study start and the paper states registration occurred before student enrollment and student data collection.
Background: Children are increasingly exposed to stressors related to urbanization, climate instability, and biodiversity loss, which may contribute to poor well-being. School-based outdoor education has shown potential to support both well-being and learning, but robust evidence from randomized trials remains...
The study randomized entire schools to treatment conditions, meeting the class-level RCT criterion.
The study used researcher-designed tests, not standardized exams, to measure financial proficiency.
The intervention consisted of four 50-minute lectures, much shorter than a full academic term.
The paper documents the control group's characteristics, size, and conditions in detail, including baseline scores in Table III.
The study randomized entire schools, fulfilling the requirement for a school-level RCT.
The authors conceptualized, designed the methodology, conducted the investigation, and wrote the paper; no independent team was involved.
The intervention involved four 50-minute sessions and a four-week follow-up, falling short of the required full academic year.
The control group did not receive the financial education program or any comparable substitute, creating an imbalance in educational time/resources.
There is no mention in the paper or readily available external evidence of an independent replication of this specific study.
The study only assessed financial proficiency (knowledge and behavior) and did not use standardized exams for other core subjects.
Follow-up was limited to approximately four weeks post-intervention; there was no tracking until graduation.
The study was registered in the AEA RCT Registry (AEARCTR- 0004431), but the registration occurred *after* the intervention and data collection were completed.
Using a computer-based learning environment, the present paper studied the effects of adaptive instruction and elaborated feedback on the learning outcomes of secondary school students in a financial education program. We randomly assigned schools to four conditions on a crossing...
The RCT randomized at the class level (16 classes), meeting the class-level randomization requirement.
Outcomes were measured with self-report scales rather than a standardized exam-based assessment.
The start-to-posttest interval is about 7–8 weeks, which is shorter than one academic term.
The control condition is explicitly described and baseline equivalence across groups is reported.
Randomization occurred at the class level within two schools, not at the school level.
The paper does not clearly document an independent evaluation team separate from the intervention developers/research team.
Outcomes were measured within about 7–8 weeks, far less than 75% of an academic year, and T is not met.
Intervention activities were delivered within normal instructional time and integrated into the standard curriculum, with no clear evidence of added instructional time or major additional resources versus business-as-usual.
No independent replication by other authors could be identified in the paper, and an online search did not find a published independent replication as of 2026-03-13.
Because criterion E is not met (no standardized exams), the all-subject exam requirement is not met.
The study did not (and could not, given its duration) track students to graduation; it reports no long-term follow-up and Y is not met.
The paper does not report a preregistration record (registry, ID, or date), and none could be confirmed via the article record as of 2026-03-13.
Objectives This study aimed to examine the association between gratitude and defending behavior in response to school bullying among Chinese early adolescents and to evaluate the effects of three types of interventions (i.e., gratitude curriculum, gratitude journal, and gratitude visits)...
Randomization is described at the student/participant level rather than at the class (or school) level, and no tutoring exception is stated.
Outcomes are assessed using a researcher-designed CT test and "Raven Progressive Matrices-style" problem-solving tasks rather than a clearly identified standardized exam-based assessment.
Outcomes are measured at least a term after the intervention begins, with a post-test in Week 25 after a 24-week curriculum.
The control condition is described as the regular school curriculum without deliberate coding instruction, and baseline equivalence is documented.
Although multiple schools are mentioned, the paper does not state that schools were randomized to conditions.
The paper does not document independent third-party conduct of the evaluation, and author roles suggest the study was run and analyzed by the same team.
The study timeline includes a delayed follow-up assessment at Week 36, which is consistent with at least 75% of a typical academic year after the intervention begins.
The intervention adds substantial instructional time (48 hours of coding instruction) without a matched active control, and the paper explicitly notes that a passive business-as-usual control cannot separate coding from general enrichment effects.
The paper does not document an independent peer-reviewed replication of this specific study, and no such replication was identified via targeted searches by DOI/title at the time of this ERCT check.
Criterion E is not met, so the requirement for all-subject standardized exams is not met.
The paper does not report tracking participants to graduation, and no follow-up paper by the same authors reporting graduation outcomes was identified at the time of this ERCT check.
The paper provides no registry name/ID or dated statement showing pre-registration prior to data collection.
Introduction: This research paper presents an empirical investigation of the effectiveness of early coding instruction in improving problem-solving skills and computational thinking (CT) among primary school students. The primary research question was to determine whether a structured six-month coding intervention...
Randomization was at the individual student level (not class- or school-level), and the paper does not describe a tutoring/personal teaching exception.
Outcomes are simulator-based performance tasks rather than a widely recognized standardized exam-based assessment.
The study reports four one-hour training sessions with testing at the beginning/midpoint/end of the study, without documenting a term-long (3–4 month) follow-up window.
The control group condition is explicitly described and baseline characteristics are documented and reported as balanced.
The study was conducted within a single institution and assigned individual students, not multiple schools/sites randomized to conditions.
The evaluated training modalities are third-party tools and the authors disclose no conflicts of interest or financial ties.
The paper does not document outcome measurement at least 75% of an academic year after intervention start, and T is not met.
The intervention adds structured training time by design, and that additional time is the treatment variable being tested against a no-training control group.
No independent replication of this specific 2026 randomized study by a different author team was found.
Criterion A is not met because criterion E is not met and outcomes are limited to simulator tasks rather than all-subject exams.
The study does not track participants to graduation, and G is not applicable under ERCT because Y is not met.
The paper provides ethics approval information but no dated public pre-registration record (registry, ID, and registration date).
Background Simulation-based training is an important component of modern surgical education. While virtual reality (VR) simulators, box trainers, and serious games are all used in laparoscopic training, comparative data on their effectiveness, transferability, and the role of individual cognitive learning...
Randomization was at the individual student level rather than at the class (or school) level, and the intervention was not one-to-one tutoring.
The main learning outcome was measured using a researcher-prepared knowledge test rather than a widely recognized standardized exam.
Outcomes were measured over roughly 6 weeks from the start of the study, which is shorter than one academic term.
The paper documents what the control group received and reports baseline demographics and baseline scores for both groups.
The study randomized individual students within one university program rather than randomizing schools or sites.
The authors appear to have implemented and evaluated the study themselves, with no stated independent third-party conduct.
The study duration is far shorter than 75% of an academic year, and criterion T is not met.
The intervention changes the learning method, but the paper describes substantial training for both groups and does not show a clear unbalanced addition of time/budget to the intervention group.
No independent peer-reviewed replication of this specific study was found or cited as of the ERCT check date.
Criterion E is not met, so criterion A is not met; additionally, the outcomes are not assessed across all core subjects via standardized exams.
The study follow-up ends after a 4-week retention test, and no evidence of tracking participants through graduation (or any graduation-linked endpoint) was found.
The study reports a ClinicalTrials.gov registration (NCT06412835), and registry dates show submission before the study start date.
Background: Collaborative learning is one of the important interactive teaching methods in teaching nursing practices. This study aimed to examine the impact of the collaborative learning approach on nursing students’ knowledge levels and self-directed learning skills related to enteral nutrition....
Students (not whole classes or schools) were randomized, so the unit of randomization was below the class level and risks contamination.
The primary educational outcome used a researcher-developed knowledge test rather than a widely recognized standardized exam.
Outcomes were assessed after a short intervention with only a 4-week follow-up, and the stated study period spans far less than an academic term.
The control condition is clearly described (standard instruction), and baseline demographic comparability between groups is documented.
Randomization occurred at the student level within one university program, not by random assignment of schools or sites.
The study does not report an independent third-party evaluation; key study elements were developed and analyzed by the author team.
The study duration and follow-up are far shorter than 75% of an academic year, and ERCT rules also imply Y is not met when T is not met.
The paper reports equal instructional session duration across groups, and the main resource difference (concept maps/cooperative structure) is the intervention being tested rather than an unmatched time or budget boost.
No independent replication by a different research team was found for this specific study/intervention as of the ERCT check date.
The study does not use standardized exams across all core subjects, and because E is not met, A is not met as well.
The study does not track participants until graduation, and because Y is not met, G is not met as well.
Trial registration information indicates the study was registered on ClinicalTrials.gov before the study start date.
Background: Cancer cases are increasing every day, which makes oncology nursing education very important. Nursing students need to learn how to manage the many symptoms caused by cancer and its treatments. This study aimed to examine the effect of a...
Classes were randomly assigned to conditions, satisfying a class-level RCT.
The outcome measures were study-specific tests rather than standardized exams.
Outcomes were measured in a short pre-post design around a 90-minute intervention, not one term after the start.
The control group is described and the paper provides group composition and background characteristics.
Randomization occurred at the class level, not by random assignment of schools.
The intervention designers and university-affiliated staff were involved in implementation and study conduct.
The study did not track outcomes for a full academic year, and criterion T is not met.
The intervention did not add extra time or budget relative to control; all groups used comparable digital environments.
No independent replication study was identified in the paper or in an external literature search.
Outcomes were not assessed with standardized exams across all main subjects, and criterion E is not met.
There is no evidence of tracking participants to graduation, and criterion Y is not met.
The paper does not report pre-registration, and no matching registry record was identified in an external search.
Students' measurement estimation skills require benchmark knowledge (about measures of known objects) and estimation strategies (ways to compare with benchmarks). While students' estimation skills have been assessed and unpacked in several empirical studies (for length and area but less for...
Randomization was within classes, but the intervention is parent-child one-to-one home teaching, which fits the ERCT tutoring/personal teaching exception for Criterion C.
Primary outcomes are measured with researcher-developed numeracy tasks, not with widely recognized standardized exams.
Outcomes were measured about 6 weeks after intervention start, which is shorter than a full academic term (about 3-4 months).
The control condition is described in detail and baseline characteristics are reported with descriptive tables.
Randomization occurred within classes rather than assigning whole schools to intervention vs control.
The paper does not report that an independent third-party evaluation team conducted the study; implementation and evaluation appear to be run by the research team.
The intervention and outcome measurement span only about 6 weeks, not a full academic year.
The study uses a matched active control with equivalent materials, engagement, and support, isolating numeracy content as the main difference.
No independent replication by a different research team is identified, and the paper frames the work as novel.
Criterion E is not met and the study does not measure standardized outcomes across all core subjects.
Year-long tracking is not present and the paper reports no delayed posttest; no follow-up-to-graduation publications were identified.
The paper states an OSF pre-registration link, but the ERCT check could not verify a time-stamped registration date before data collection.
The home numeracy environment is suggested to influence children's numerical development, but causal evidence for this assertion remains limited. Addressing this gap, we randomly assigned 117 predominantly White 4- to 5-year-olds (M = 4.68 years, SD = 0.2, 47% girls)...
Randomization was at the individual student level (not class- or school-level), and the tutoring/personal-teaching exception does not apply.
Outcomes were measured using study-specific MCQ questionnaires rather than a standardized, widely recognized exam-based assessment.
The study ran from 15/10/2024 to 15/01/2025 with assessment at the end of the sessions, which is approximately one academic term.
The control (PBL) group is described with sample size, baseline characteristics, and a clear description of what the control condition received.
The unit of randomization is students within one institution, not schools or equivalent educational sites.
The paper does not document independent third-party conduct of implementation and/or evaluation separate from the study team.
The study period (15/10/2024 to 15/01/2025) is about three months, which is well below 75% of an academic year, and the paper notes a short follow-up duration.
Both conditions had the same in-session duration (4 h) and similar facilitation, and the additional TBL elements (pre-class preparation materials and readiness assurance tests) are integral to the TBL treatment being compared to PBL.
No independent replication of this specific University of Tabuk TBL-vs-PBL RCT by a different team in a different context was found.
Criterion E is not met (no standardized exam), so criterion A is automatically not met; additionally, the study assesses only surgery-specific knowledge rather than all main subjects.
The study does not track learners through graduation and notes a short follow-up and lack of long-term retention assessment; also, because Y is not met, G cannot be met.
No pre-registration registry, registration ID, or registration date is reported, and no external protocol record was identified.
Objectives: The primary objective measured in our study is to determine whether Team-Based Learning (TBL) is a superior pedagogical approach compared to Problem-Based Learning (PBL) or not. We focused on our secondary objectives, which include promoting problem-solving, facilitating independent learning,...
The study randomized students in small peer-instruction groups for a tutoring-style intervention, which satisfies the ERCT tutoring exception.
The study used custom lesson-specific pre- and post-tests rather than widely recognized standardized exams.
The study ran across two consecutive weeks with immediate post-tests, so it did not measure outcomes at least one academic term after the intervention began.
The in-class active learning control condition and group characteristics are documented, including comparable demographics and baseline physics background knowledge.
Randomization occurred within one university course, not at the school level.
The AI tutor was designed by the authors and the paper does not describe independent third-party conduct of the study.
Outcomes were measured within a two-week crossover design, not after a full academic year.
The AI tutor condition substituted for the in-class lesson and did not add unmatched time or resources for the intervention group.
No independent peer-reviewed replication of this specific study by another research team was found.
Outcomes were limited to physics topics and the study did not use standardized exams across all main subjects (and criterion E is not met).
The study reports only immediate post-lesson outcomes and does not track students until graduation (and criterion Y is not met).
The paper reports IRB approval but provides no evidence of a pre-registered protocol, and no matching public registry entry was found.
Here we report a randomized, controlled trial measuring college students' learning and their perceptions when content is presented through an AI-powered tutor compared with an active learning class. The novel design of the custom AI tutor is informed by the...
Randomization was at the individual child level, but the intervention was delivered one-to-one, fitting the personal teaching exception.
Outcomes were measured with study-specific tasks on novel letters and pseudowords, not standardized exams.
Outcomes were assessed within three consecutive days, far shorter than one academic term.
The paper documents participant characteristics and reports no baseline differences across groups in control measures.
The study sampled from a single school and randomized children, not schools.
No independent evaluation organization is described in the paper.
The study duration and measurement window were days, not an academic year.
The paper states that exposure and duration were made comparable across conditions, balancing time-on-task and practice quantity.
No independent peer-reviewed replication of this specific experiment was found.
This criterion is not met because Criterion E is not met and outcomes are not standardized exams across subjects.
The study does not track participants to graduation, and the one-year prerequisite is not met.
The paper provides an OSF link for materials and data but does not state or verify preregistration with a pre-data-collection date.
Recent research has revealed that the substitution of handwriting practice for typing may hinder the initial steps of reading development. The current experiment investigated the impact of graphomotor action and output variability in letter and word learning using a variety...
The study randomized individual students within a single institution rather than randomizing at the class level.
The study utilized the California Critical Thinking Skills Test (CCTST), a widely recognized standardized assessment.
The intervention lasted only 8 weeks, which is shorter than the required full academic term (typically 3-4 months).
The study documents the control group's size, demographics, and baseline performance scores in detail.
The study was conducted within a single nursing college with randomization at the student level, not the school level.
The study was conducted by the authors themselves; only the randomization process was handled by an independent researcher.
The study duration was 8 weeks, which does not meet the requirement of one full academic year.
The study tests the integration of LLMs as the variable; both groups received equal course time (16 hours), satisfying the balance requirement.
No independent replication of this specific study was found in peer-reviewed journals.
The study measured only critical thinking skills and the specific course test score, not all main subjects taught in the school.
Outcomes were measured immediately after the course ended, with no tracking until graduation.
The paper mentions following CONSORT but does not provide a registration number or evidence of pre-registered protocol.
Background: The integration of Large Language Models (LLMs) into nursing education presents a novel approach to enhancing critical thinking skills. This study evaluated the effectiveness of LLM-assisted Problem-Based Learning (PBL) compared to traditional PBL in improving critical thinking skills among...
No random assignment of classes or students to conditions is described anywhere in the paper; two intact classes were selected because their baseline scores were similar.
The outcome measure was a custom vocabulary test built by the researchers themselves, not a standardised, widely recognised exam.
The 15-week intervention period plus the following posttest stage covers an interval consistent with at least one academic term.
The control group's size, baseline score, demographic profile, and exact instructional treatment at each stage are documented in the paper.
Only one institution was involved, and there is no school-level (or any-level) randomisation described.
The same research team designed the intervention, supervised its delivery, and analysed the results, with no independent evaluator involved.
The total study duration of about 19 weeks falls well short of 75% of an academic year.
The control group received a closely matched substitute activity (handout review of the same words, same time and content) for the same instructional steps, isolating the review medium as the tested variable.
No independent replication of this specific study is reported or referenced in the paper, and none was found through internet citation searches.
Only vocabulary outcomes in one subject (English) were measured, and criterion E (a prerequisite) is not met.
The authors themselves state that only short-term outcomes were measured, with no long-term or graduation follow-up; criterion Y is also not met, which independently rules out criterion G.
No pre-registration statement, registry link, or registration date is mentioned anywhere in the paper, and none was found through internet search.
Understanding the role of English language acquisition in developing the socioeconomic status of Vietnam forms the backdrop for this study, which seeks to shed light on the potential benefits of technology-assisted vocabulary learning. Based on this context, this study employed...
Randomization was performed at the school level, satisfying the requirement for class-level assignment.
The study utilized specific assessments developed by the research group (CSA and TechCheck) rather than widely recognized standardized exams.
The paper specifies the number of lessons but does not provide specific dates or a duration interval to confirm the intervention spanned at least one full academic term.
The control group's demographics are tabulated, and their "business as usual" condition is clearly described.
The study randomized entire schools rather than classes or students, satisfying the school-level RCT requirement.
The study was conducted and analyzed by the same researchers who developed the curriculum.
The study tracked outcomes only until the end of the 24-lesson curriculum, not for a full academic year.
The control group did not receive resources, time, or professional development equivalent to the treatment group.
Previous evaluations were conducted by the same authors/research group; no independent replication is cited.
The study limited assessment to coding and computational thinking, ignoring other core subjects.
Tracking ended immediately after the curriculum implementation.
The paper mentions IRB approval but does not cite a public pre-registration of the study protocol.
Background and context: Early childhood computer science (CS) education is a high-priority focus worldwide, but early childhood CS tools are primarily developed and researched within the United States and Europe. As an example, the Coding as Another Language ScratchJr (CAL-ScratchJr)...
Assignment was made to three intact academic departments (whole cohorts of ~40 students each) rather than to individuals within one shared classroom, satisfying class-level randomisation, although the abstract confusingly labels the design "quasi-experimental."
The reading test was a custom instrument built from a single course textbook unit, not a widely recognised standardised exam.
The intervention and measurement together spanned only about two months (42 hours of teaching), well short of a full academic term.
The control (Lecture Method) group's department, sample size, and pre/post-test descriptive statistics are clearly reported.
Randomisation occurred among departments within a single university, not among separate schools.
The same researcher(s) designed the intervention, delivered/coordinated the teaching, conducted interviews, and analysed the data, with no independent third-party evaluator involved.
Because the Term Duration criterion (T) is not met, the stronger Year Duration criterion is also not met; the study covered only about two months.
All three groups received the same course content, teaching duration, and identical pre/post-tests; only the pedagogical method (not the amount of time, materials, or budget) differed between groups.
The authors explicitly state this is the first study of its kind, and a dedicated internet search found no independent replication.
Only reading skills in English were assessed, no other core subjects were measured, and criterion E (a prerequisite) is not met.
Since criterion Y is not met and no follow-up beyond the immediate post-test is reported or found in later publications, there is no graduation tracking.
No pre-registration of the study protocol is mentioned anywhere in the paper, and an internet search of registry platforms found no matching registration.
This study investigates the impact of Grammar Translation Method (GTM) and the Direct Method (DM) on the development of reading skills of the first-year medical students at 21 September University for Medical and Applied Sciences in Sana'a. A total of...
Randomisation was conducted at the individual student level among students pooled from the same three intact classes, not at the class or school level, and no explicit tutoring exception was invoked by the authors.
The speaking test was adapted from GEPT Kids, a nationally recognised standardised assessment developed by Taiwan's Language Training and Testing Center, using its official elementary speaking rating scale.
Outcomes were measured immediately after a three-week summer intervention, far short of the one-academic-term minimum required by the standard.
The control (No-Bot) group's size, demographics, and the parallel topic-matched activities they completed are clearly documented in the paper, including in Table 1.
The study involved students from a single summer program taught by one teacher, with randomisation at the individual student level rather than across schools.
The same research team that designed and built CoolE Bot also implemented, supervised, and analysed the study, with no independent third-party evaluator.
The study duration (three weeks) is far below the year-long requirement, and since criterion T is not met, criterion Y is automatically not met.
All three groups received equal instructional time and topic-matched worksheet materials, differing only in the conversational partner (chatbot vs. teacher/peers), so resource allocation was balanced.
No independent replication of this study or its custom-built CoolE Bot intervention was found; the paper reports a novel, recently published intervention with no prior replication.
Only English-speaking outcomes were assessed; no other core subjects were measured, and no exception rationale is provided.
Tracking ended at the immediate post-test with no long-term follow-up, and since criterion Y is not met, criterion G is automatically not met.
No statement of pre-registration on any public registry, nor any registration date, is present anywhere in the paper, and no matching registry record was found online.
Generative artificial intelligence (GAI) and automatic speech recognition (ASR) have ushered in promising tools for foreign language learning, notably GAI chatbots. This study investigated the impact of GAI chatbots on elementary school English as a foreign language (EFL) learners' speaking...
Two entire intact classes, not individual students within one class, were randomly assigned to the control and experimental conditions.
The reading comprehension tests were adapted from a Paideia lesson-plan website rather than a standardised, validated exam, and no standardised exam was used for the anxiety outcome either.
The intervention and outcome measurement together spanned only six weeks, far short of a full academic term.
The control group's regular-instruction procedures, activities, and shared demographic/baseline characteristics with the experimental group are described in reasonable detail.
Randomisation occurred between two classes within a single school, not between multiple schools.
The authors themselves acknowledge that one of the researchers personally taught both the control and the experimental group, so the study was not conducted by an independent, third-party team.
Because the Term Duration criterion (T) is not met, the stronger Year Duration criterion is automatically not met as well.
Both groups received the same amount of instructional time on the same texts and curriculum; only the teaching method (Paideia Seminar dialogue versus regular comprehension instruction) differed, so no extra time or budget was given to the experimental group that needed to be matched.
No independent replication of this specific study by a different research team, published in a peer-reviewed journal, was reported in the paper or found through internet searches.
Because Criterion E (standardised exam-based assessment) is not met, and only reading/language arts outcomes were measured (not all core subjects), Criterion A is not met.
Because Criterion Y (Year Duration) is not met, and no follow-up beyond the six-week study period is reported or found, Criterion G is not met.
The paper contains no statement of pre-registration on any registry platform, and no pre-registration record for this study was found online.
This article reports the results of an experimental study on the relative effectiveness of the Paideia Seminar in improving the comprehension of poetry and decreasing reading anxiety. The participants (n = 50) were English as a foreign language (EFL) ninth...
Random assignment was described as being carried out at the level of individual student participants rather than at the level of entire classes.
Writing performance was scored using the rating criteria of CET-4, a nationally recognized standardized English test, based on prompts adapted from that same test.
Outcomes were measured at the end of an eleven-week intervention, which is shorter than the roughly one-term (3-4 month) interval required, with no further follow-up.
The control group's size, demographics, baseline scores, and business-as-usual treatment (peer feedback only, no AI/AWE tool) are clearly documented.
Randomization/selection occurred among individual students from three classes at a single university, not among multiple schools.
The study's own authors (the "first researcher") scored the writing outcomes themselves; no independent third party conducted data collection or analysis.
Since criterion T (Term Duration) was not met, and the study's total duration of eleven weeks is far short of an academic year, criterion Y cannot be met.
Access to the AWE/ChatGPT tools is the explicit treatment variable being tested against a business-as-usual (peer feedback only) control, so the resource asymmetry is integral to the study design.
No evidence of independent replication was found; the study is newly published (online Feb 2025) and internet searches found no other research team reporting a reproduction of this specific three-arm ChatGPT/AWE/ control comparison.
Only English writing performance was assessed; no other core subjects were measured.
Since criterion Y (Year Duration) was not met, and there was no tracking of participants beyond the eleven-week post-test, criterion G cannot be met.
No pre-registration of the study protocol, hypotheses, or analysis plan is mentioned anywhere in the paper, and no registry entry was found via internet search.
The affordances of ChatGPT in language learning and teaching have gained increasing traction. While studies began to investigate the potential of ChatGPT as a feedback provider, little attention was given to ChatGPT's potential impact on students' writing performance and the...
Randomisation was conducted at the level of whole intact classes, satisfying the class-level RCT requirement.
The study used custom, researcher-developed domain-referenced tests rather than a widely recognised standardised exam.
The intervention and outcome measurement together spanned only 6 weeks, well short of one academic term.
The control group's size, teachers, content, and baseline performance are clearly documented and shown to be comparable to the treatment group.
Randomisation was at the class level within the same two school branches, not at the level of separate schools.
The researchers who designed the STAD intervention also trained teachers and directly supervised the study's implementation, with no independent evaluator.
The study lasted only 6 weeks, far short of a full academic year, and criterion T was already not met.
Both groups received the same content, teachers, and duration; only the instructional method (cooperative vs. individualistic) differed, so no resource imbalance is present.
No evidence of an independent replication of this specific study was found in the paper or via external search.
Only ESL rules and mechanics were assessed, and since criterion E is not met, criterion A cannot be met.
No follow-up beyond the immediate post-test was conducted or found via external search, and criterion Y was already not met.
No mention of a pre-registered protocol or registry is found anywhere in the paper or via external search.
This article reports the results of an experimental investigation of the effect of cooperative learning on the acquisition of English as a second language (ESL) rules and mechanics. Four fourth-grade, four fifth-grade, and four sixth-grade intact classes (n = 318...
The paper self-identifies as a quasi-experimental study using convenience-sampled intact classes taught by the researcher himself, with no described randomisation procedure for assigning the two classes to conditions.
The outcome measure was the EFL Achievement Test, an instrument developed by the researcher himself for this study, not a widely recognised standardised exam.
The intervention spanned the full 15-week fall term, and the primary outcome (post-test) was measured at the conclusion of that term, which constitutes at least one full academic term.
The control group's size, gender composition, baseline scores, and instructional conditions (traditional lecture, same syllabus) are clearly documented with supporting tables.
Randomisation (such as it was) occurred at the level of two intact classrooms within a single university, not across multiple schools.
The same researcher designed the intervention and the assessment instrument, personally taught both the experimental and control classes, and conducted the interviews and analysis, with no independent or external party involved.
The intervention and outcome tracking lasted only one 15-week academic term, far short of the 75% of an academic year required by this criterion.
Both groups followed the same syllabus over the same 15-week period with no described extra time or budget given to the experimental group beyond the change in instructional method.
No evidence of an independent replication of this specific study by a different research team is present in the paper or discoverable externally.
Only EFL performance and its sub-skills were assessed; no other core subjects were measured, and criterion E (a prerequisite for A) is not met.
Tracking stopped at a within-term durability test; there was no follow-up until graduation, and criterion Y (a prerequisite for G) is not met.
The paper mentions only Institutional Review Board approval, with no reference to a public pre-registration of the study protocol before data collection began.
This pre-test post-test quasi-experimental study was grounded in a mixed method embedded design to delve into the quality and efficiency of flipped classroom model in enhancing university prep students' overall academic performance in EFL and that in its sub-skills in...
Whole intact classes, not individual students within one class, were randomly assigned to the three study conditions.
All outcome measures (a forced-choice perception test, two production tests, and a teacher rating scale) were custom-built by the author for this study, not standardised exams.
Outcomes were measured about four weeks after the intervention began, far short of one academic term.
The control group's size, demographics, and the instruction it received (comparable duration and content, minus the FFI focus) are described in detail.
Randomisation occurred among classes within a single language institute, not among separate schools or sites.
The sole author designed the materials, trained the teachers, personally monitored lesson delivery, and coded the interaction data himself, with no independent evaluator involved.
Because the term-duration criterion T is not met (total tracking was only about four weeks), the year-duration criterion is automatically not met.
The control group received the same duration and general type of communicative instruction as the treatment groups, differing only in the absence of the FFI focus on /r/, which is itself the treatment variable being tested.
No independent replication of this specific study by another research team was reported in the paper or found in external follow-up literature.
Because the exam-based-assessment criterion E is not met, this criterion automatically fails as well; in any case only /r/ pronunciation was assessed, not other subjects.
Because the year-duration criterion Y is not met, this criterion automatically fails; tracking stopped about two weeks after the lessons ended with no long-term follow-up.
No statement of pre-registration, registry identifier, or pre-registration date appears anywhere in the paper.
The current study examines in depth how two types of form-focused instruction (FFI), which are FFI with and without corrective feedback (CF), can facilitate second language speech perception and production of /r/ by 49 Japanese learners in English as a...
Randomization was explicitly conducted at the student level within shared classrooms, which does not meet the class-level (or stronger) requirement and is not covered by the one-to-one tutoring exception.
The primary outcome instrument is an author- developed pilot measure (P-CTOPPP), not a widely recognized standardized exam.
Outcomes were measured in May/June, several months after the mid-November intervention start, well beyond the one-term minimum.
The control group's size, demographics, and baseline literacy scores are documented in detail in Tables 1 and 2, and its business-as-usual treatment is explicitly stated.
Only a single preschool program (site) was involved, with randomization at the student level within its classrooms, not at the school level.
Two of the three study authors (Lonigan and Farver) co-designed the Literacy Express curriculum being tested, and the third author personally supervised its delivery, so there is no independent conduct.
The documented intervention-to-measurement interval of roughly five to six months does not clearly meet the 75%-of-a-year threshold required for Y.
The extra small-group instructional time is the explicit treatment variable being tested, so the control group's business-as-usual curriculum is an acceptable baseline under the ERCT exception.
The authors describe this as the first study of its kind, and neither the paper nor a targeted internet search identified an independent replication of this specific trial by a different research team.
Criterion E is not met, and independently the study assessed only literacy-related outcomes, not all core subjects.
Criterion Y is not met, and the authors explicitly state no longitudinal follow-up was conducted, with no tracking toward graduation found in the paper or in a search for follow-up publications.
No pre-registration statement, registry ID, or date appears anywhere in the paper, and none was found via internet search.
Ninety-four Spanish-speaking preschoolers (M age = 54.51 months, SD = 4.72; 43 girls) were randomly assigned to receive the High/Scope Curriculum (control n = 32) or the Literacy Express Preschool Curriculum in English-only (n = 31) or initially in Spanish...
The study uses a randomized cross-over design at the peer-group level (student level), which is acceptable under the ERCT exception for personal tutoring interventions.
Outcomes were measured using custom pre- and post-tests designed for the specific lessons, not standardized exam-based assessments.
Outcomes were measured immediately following two single-lesson interventions, falling far short of the one-term duration requirement.
The control group (in-class active learning) is well-documented, including pedagogy, student demographics, and baseline knowledge.
The study was conducted within a single university course, not randomized across multiple schools or institutions.
The study was designed, conducted, and analyzed by the authors, including the course instructors, without independent third-party conduct.
Outcomes were measured over a two-week period, not tracked for a full academic year.
The intervention replaced the control activity without adding extra time; in fact, the intervention group spent less time on task than the control group.
No independent peer-reviewed replication of this specific AI tutoring intervention was found.
The study only assessed physics content knowledge, not all main subjects.
The study tracks learning only for the duration of the lessons and does not follow students to graduation.
The study mentions IRB approval but does not provide evidence of a pre-registered protocol on a public registry.
This study reports a randomized, controlled trial measuring college students' learning and their perceptions when content is presented through an AI-powered tutor compared with an active learning class. We find that students learn significantly more in less time when using...
Intact classrooms, not individual students within a classroom, were the unit of randomisation.
The primary vocabulary outcome measure was a custom, study-specific test built directly from the taught words, not a standardized exam.
The intervention and its outcome measurement spanned only about eight weeks, far short of a full academic term.
The control (PAD) group's size, demographics, baseline scores, and instructional content are described in detail.
Only a single school participated, so randomisation could not occur at the school level.
The same research/development team designed, delivered, and assessed the intervention, with no independent evaluator.
Since the term-duration criterion (T) is not met, the stronger year-duration criterion cannot be met either.
All conditions received the identical total amount of instructional time; only the content emphasis within that shared time differed.
No evidence of independent replication of this specific study was found in the paper or via external search.
Only reading-related constructs (vocabulary, phonological decoding) were assessed; no other core subjects were measured.
Criterion Y (Year Duration) is not met, which automatically fails this criterion; additionally, no follow-up or graduation-tracking publication by the same authors was located.
The paper contains no statement or reference to a pre-registered study protocol, and no registry record was found via external search.
This study examined the added value of a vocabulary plus phonological awareness (vocab+) intervention against a phonological awareness (PA only) intervention only. The vocabulary intervention built networks among words through attention to morphological and semantic relationships. This supplementary classroom instruction...
Randomisation was conducted at the class level (30 classes), which satisfies the class-level RCT requirement.
The primary outcome was a custom-assembled "Floating and Sinking" test built largely from prior research items plus new author-written items, not a widely recognised standardised exam.
Outcomes were measured immediately after a five-lesson intervention and again only 6 weeks later, far short of one academic term.
The comparison (monolingual) group's demographics, baseline achievement, and instructional exposure are documented in detail in Table 1 and the Procedure section.
Randomisation was explicitly conducted at the class level, not at the school level, so the stronger school-level requirement is not satisfied.
The intervention was designed and personally taught by the paper's first author, with no independent or external evaluation team involved.
The study tracked outcomes for at most a few months total (intervention plus a 6-week follow-up), far short of 75% of an academic year, and criterion T was already not met.
Instructional time and teaching materials were held constant across both groups; the only difference was the language of instruction plus minor language scaffolding integral to the bilingual condition itself.
No independent replication of this specific study by a different research team was found in the paper or via external search.
Only science (physics) content knowledge was assessed; no other core subjects were measured, and criterion E was not met.
Tracking stopped 6 weeks after the intervention with no further follow-up reported, and criterion Y was not met.
No pre-registration of the study protocol is mentioned anywhere in the paper, and none was found via external search.
Content and language integrated learning (CLIL) has been widely implemented in Europe. This article presents a randomised controlled field experiment on the effects of CLIL on students' science learning. Thirty sixth-grade intermediate-track German secondary-school classes (722 students) were randomly assigned...
Randomisation was done at the individual student level within one university, not at the class or school level, and the intervention was group classroom instruction, not one-to-one tutoring.
Outcomes were measured with a custom researcher-designed elicitation instrument scored on a 9-point Likert scale by raters, not with any widely recognised standardised exam.
The entire study spanned only four weeks from pretest to delayed posttest, far short of the one academic term (roughly 3-4 months) required.
The control group is clearly documented, including its size, baseline scores with confidence intervals, and confirmation that it received only normal coursework with no PI.
Randomisation occurred among individual students at a single university; no schools or institutions were randomised.
The lead researcher designed the study, personally delivered all treatments, and produced the sole rating dataset used in the analyses, with no independent third-party conduct or oversight.
With only four weeks from pretest to delayed posttest, the study falls far short of tracking outcomes for 75% of an academic year; criterion T is also unmet, which entails Y is unmet.
The four treatment groups received identical time and identically structured sessions, and the extra instruction relative to the no-PI control is the explicit treatment variable (provision of PI), with the control serving as a business-as-usual baseline to rule out practice effects.
An independent team (Alghazo, Jarrah & Al Salem, 2023, Frontiers in Education) conducted a peer-reviewed conceptual replication of this study's perception- vs. production-based PI comparison with Jordanian tertiary EFL learners, explicitly citing and using this study's own definitions and reporting a convergent conclusion for delayed-posttest gains.
Only pronunciation accuracy in English was assessed, using a custom instrument; no other school subjects were measured and criterion E is not met, so A automatically fails.
Measurement ended at a delayed posttest two weeks after instruction, with no tracking of participants to graduation and no follow-up publications found; Y is also unmet, which entails G is unmet.
The paper contains no mention of any pre-registered protocol or registry entry; sharing materials on the IRIS database is not pre-registration, and no registration record was found externally.
While research has shown that provision of explicit pronunciation instruction (PI) is facilitative of various aspects of second language (L2) speech learning (Thomson & Derwing, 2015), a growing number of scholars have begun to examine which type of instruction can...
Randomisation was at the individual student level within a single school for group-based classroom instruction, not at the class or school level, and the tutoring exception does not apply.
Outcomes were measured with a speaking test designed by the researchers based on students' textbooks, scored with a locally modified IELTS rubric, which is a custom instrument rather than a widely recognised standardised exam.
The intervention began in February and outcomes were measured in the last week of May of the same semester, an interval of roughly 3.5 months, which covers at least one full academic term.
Both control groups are documented in detail, including sizes, gender and age composition, baseline speaking scores, and precise descriptions of the instruction each control condition received.
The study randomised individual students within a single school, so no school-level (or institution-level) randomisation took place.
The authors designed the intervention, taught the classes as the teacher-researcher, designed and administered the tests, and analysed the data themselves, with no independent third-party evaluation.
The study lasted one semester (mid-February to late May, about 3.5 months), which is well short of 75% of an academic year.
All three groups received active instruction on the same units with comparable class time; the PBL control did identical projects without phones, and the mobile phone use that differentiates the experimental group is the integral treatment variable being tested.
Neither the paper nor an internet search identified any independent, peer-reviewed replication of this specific mobile-assisted project-based learning trial by a different research team.
Only English speaking skills were assessed with a custom test, so no standardised all-subject assessment took place, and the prerequisite criterion E is also unmet.
Measurement ended with the post-test in May at the close of the semester, with no tracking of participants to graduation and no follow-up publications on this cohort; prerequisite criterion Y is also unmet.
The paper contains no mention of any pre-registration or registry, and an external search found no registration record for this trial.
Combining mobile-assisted language learning (MALL) with project-based learning (PBL) might be the potential framework for enhancing EFL learners' speaking skills. However, only a few studies have scrutinised the impact of modern technologies on project work. More importantly, investigating how MALL,...
Randomization was at the student level, but the intervention is explicitly a personalized tutoring program (an AI chatbot acting as a virtual tutor), so the ERCT tutoring exception applies and student-level randomization is acceptable.
The primary outcome was a custom multiple-choice assessment designed by experts specifically for this study, and the secondary outcome was a school-administered term exam, neither of which is a widely recognized standardized exam.
Outcomes were measured about six weeks after the intervention began (sessions 3 June to 11 July 2024, assessment 11-12 July 2024), which is well short of a full academic term.
The control group's size, demographics, baseline test scores, and business-as-usual condition are documented in detail, including a full balance table.
Randomization was conducted among individual students within nine pre-selected schools, not among schools.
The World Bank author team designed the intervention (prompts, toolkits, training) and also conducted the evaluation, with no statement of an independent third-party evaluator.
The whole study, from intervention start to final measurement, lasted about six weeks, far short of 75 percent of an academic year, and criterion T already fails.
The intervention's extra inputs (after-school sessions, computer lab time, teacher facilitation) are the integral treatment package explicitly being tested against a business-as-usual control, which the ERCT standard accepts as balanced by design.
No independent, peer-reviewed replication of this specific Nigerian Copilot tutoring RCT exists; the authors themselves call for replication, and only a computational reproducibility package (not an independent replication) is available.
Criterion E fails, and the study measured only English plus AI and digital skills, not all main school subjects with standardized exams.
Measurement ended immediately after the six-week program with no follow-up toward graduation, and prerequisite criterion Y is not met.
The paper contains no mention of a pre-registered protocol or registry entry, and no pre-registration for this trial was found in an external search.
This study evaluates the impact of a program leveraging large language models for virtual tutoring in secondary education in Nigeria. Using a randomized controlled trial, the program deployed Microsoft Copilot (powered by GPT-4) to support first-year senior secondary students in...
Randomization was performed at the individual student level within institutes, not at the class or school level, and the group-based intervention does not qualify for the one-to-one tutoring exception.
Writing outcomes were measured with sample tasks from the IELTS, a widely recognised standardised examination, scored with the standard IELTS analytic rubric by two independent raters with good inter-rater reliability.
Outcomes were measured at the end of the 12-week instructional period, an interval of approximately three months from the intervention start, which corresponds to one academic term.
The control group's size, demographics, baseline scores, and the instruction it received are documented in detail, including baseline equivalence tests and Table 1 statistics.
Randomisation occurred among individual students within institutes, not among schools or institutions, so the school-level RCT requirement is not satisfied.
The same research team that designed the AWE-based intervention also delivered the training and conducted the pretests and analyses, with no external or third-party evaluation reported.
The interval from intervention start to outcome measurement was only 12 weeks, far short of 75% of an academic year.
The experimental group had 12 weeks of weekly essay practice versus 8 weeks for controls, and the paper is contradictory about whether controls received the matching workshops, so time and resources were not verifiably balanced.
No independent replication of this specific trial is reported in the paper, and an internet search of the citing literature found no independent peer-reviewed replication of this study.
Only English writing skill was assessed, with no measurement of other main subjects and no explicit rationale offered, so the all-subject requirement is not satisfied.
Measurement stopped at the 12-week posttest with no follow-up tracking to graduation, no follow-up publications were found, and the prerequisite Year Duration criterion is also unmet.
The paper contains no mention of pre-registration on any registry platform or of a protocol published before data collection began, and no external pre-registration record was found online.
In the context of the burgeoning field of second language (L2) education, where proficient writing plays an integral role in effective language acquisition and communication, the ever-increasing technology development has influenced the trajectory of L2 writing development. To address the...
Students were randomised individually rather than by class or school, and the classroom-based intervention does not qualify for the tutoring exception.
The study used custom researcher-made pronunciation tasks scored by a single rater on a Likert scale, not a recognised standardised exam.
Outcomes were tracked from intervention start (Week 1) to a delayed post-test at Week 14, an interval of roughly one academic term (about 3 months).
The comparison group's size, baseline scores, and treatment are documented in tables and text, along with demographic details of the sample.
The study randomised individual students within a single university, so no school-level randomisation occurred.
The same author designed the intervention, taught both groups, and analysed the data, with no independent evaluation team.
The study tracked outcomes for only about 13 weeks, far short of 75% of an academic year.
Both groups received identical time, instructor, and content coverage, with only the instructional method differing as the tested variable.
No independent replication of this specific study by a different research team has been identified, confirmed via a citation-database search conducted during verification.
Only custom pronunciation tasks were used, so with criterion E unmet and no other subjects assessed, this criterion fails.
Measurement ended at Week 14 with no tracking of participants to graduation, and prerequisite criterion Y is unmet; no follow-up publication was found via search.
The paper contains no mention of a pre-registered protocol, registry, or registration date, and none was found via internet search.
This study investigates the efficacy of the type of instruction (i.e., perception-based vs. production-based) on second language (L2) pronunciation acquisition in an English as a foreign language (EFL) context. To achieve this objective, 60 tertiary-level Jordanian learners of English were...
Randomisation was at the student level, but the intervention is a one-to-one tutoring-style conversational bot, so the personal-teaching exception applies.
Outcomes were measured with custom speaking and vocabulary tests built by the authors (only loosely adapted from IELTS materials), not a recognised standardised exam.
The intervention and outcome measurement spanned only six days plus a three-week delayed vocabulary test, far less than one academic term.
The control group's condition, baseline measures and treatment are documented in the methods and result tables, meeting the documentation requirement.
Individual students, not schools or institutional units, were randomly assigned to the two interfaces.
The same authors designed the bot and also ran, measured and analysed the evaluation, with no independent evaluators documented.
The study covered only six days plus a three-week follow-up, far below the required 75% of an academic year.
Both groups used software with identical content over the same period; the bot's interactive features are the integral treatment variable, so the groups are balanced.
The study is presented as novel, was published in mid-2025, and an internet citation search found no independent replication by another team.
Only custom English vocabulary and speaking measures were used; no standardised exams across all main subjects, and prerequisite criterion E is not met.
Participants were followed for only three weeks after the six-day intervention, with no tracking to graduation, and no follow-up publication by the authors was found.
The paper contains no reference to any pre-registered protocol, registry ID, or registration date, and an internet search found no such record either.
With the advancement of artificial intelligence, natural language processing, and speech recognition, conversational agents have emerged as promising tools for second language acquisition. This study designed an English conversational bot to help learners of English as a second language among...
Randomisation was performed at the individual student level within one university course, not at the class or school level, and the intervention is whole-class instruction, not one-to-one tutoring, so no exception applies.
Outcomes were measured with a researcher-assembled listening test compiled from a Pearson practice extract and textbook exercises, piloted by the authors themselves, not a recognised standardised exam.
The intervention ran across a full 12-week teaching semester with the post-test administered at the end of the semester, which corresponds to one academic term ("a semester or equivalent") from intervention start to measurement.
The control group's size, demographics, baseline scores and the exact instruction it received are documented in detail in Table 2, Table 6 and the description of the listening-oriented model.
The trial randomised individual students within a single university; no schools or institutional units were randomised.
The two authors designed the instructional model and also ran and analysed the trial through the course teacher, with no external or third-party evaluation team documented.
The whole trial lasted only 12 weeks, far less than 75% of an academic year (about 9-10 months).
Both groups received the same weekly 2-hour class, the same textbook and the same subskill content over 12 weeks, with the experimental group's CMC speaking tools replacing (not adding to) class time as an integral part of the speaking-listening model being tested.
No independent replication of this speaking-listening subskills RCT by another team was reported in the paper or found in a search of citing literature.
Only English listening competence was assessed, with a non-standardised custom test (criterion E fails), so all-subject standardised assessment is clearly not satisfied.
Measurement ended with the Week 12 post-test, no follow-up to graduation is reported or found in later publications, and the prerequisite criterion Y is not met.
The paper contains no mention of any pre-registered protocol, registry platform, or registration date, and no registry entry for this trial was found.
The purpose of this study is to explore the impact of teaching subskills, namely micro- and macro-skills, with a speaking-listening model on the improvement of listening competence. The research included 112 Chinese tertiary students with intermediate English proficiency who were...
The paper explicitly adopts a non-randomized design with two intact sections and never describes any randomisation procedure, so class-level randomisation is not properly documented or implemented.
Outcomes were measured with a researcher-administered composition test scored on an adopted rubric and a researcher-made attitude scale, not with any widely recognised standardised exam.
The intervention ran for 15 weeks (a full university semester) with outcomes measured at the end of the experiment, satisfying the one-term minimum interval.
The control group's size, condition, and baseline writing, intelligence, and attitude data are documented and shown to be statistically equivalent to the experimental group.
The study involved two class sections within a single college of one university, with no school-level randomisation of any kind.
The single author designed the intervention, prepared the instruments, ran the experiment, and scored the tests himself, with no independent third-party conduct or oversight.
The study lasted only 15 weeks (one semester), well short of 75% of an academic year, with no longer follow-up tracking.
Both groups received identical process-based teaching, the same two weekly lessons, and the same assignments, with the only difference being the reflection sheets that constitute the treatment variable itself.
No independent published replication of this specific reflection-supported process-writing study was found; later similar studies elsewhere are thematically related but are not framed as replications of this trial.
Only EFL writing performance and writing attitude were measured, criterion E is not met, and no other core subjects were assessed.
Measurement stopped at the end of the 15-week experiment with no tracking of students to graduation, criterion Y is not met, and no follow-up publication tracking the same cohort was found.
The paper contains no mention of any pre-registered protocol, registry, or registration date, and no external pre-registration record was found.
The study aims at finding out the effect of process-based writing teaching supported by students' reflection on their performance in, and attitude toward writing. It hypothesizes that there is no statistically significant difference between the mean score of the experimental...
Although allocation was at the class level, the three intact classes were "arbitrarily assigned" with no described randomisation procedure, so a properly implemented class-level RCT is not demonstrated.
All three outcome measures were custom instruments designed by the researchers for this study, not widely recognised standardised exams.
Outcomes were last measured about two weeks after the 1-hour, 2-day intervention began, far short of the required full academic term of follow-up.
The control group's size (n = 10), condition (testing-only with normal instruction and no feedback), and baseline test scores are clearly documented in the text and tables.
The study was conducted within one private language school with three classes allocated to conditions, so there was no school-level randomisation.
The intervention was designed, delivered, and evaluated by the same research team, with one of the researchers personally teaching the treatment sessions and no independent evaluator involved.
The full tracking period was about two weeks from intervention start, nowhere near 75% of an academic year (and criterion T is already not met).
The two treatment groups received identical amounts of instruction (1 hr each), no extra time or budget was added beyond regular schooling, and the testing control simply continued its normal instruction during the same period.
A different, independent research team (Paraskeva & Agathopoulou, 2022) published a peer-reviewed study replicating the recasts-versus-metalinguistic-feedback contrast on the same target structure (past tense -ed) in a different context; Sauro (2009) offers weaker, conceptual support with a different target structure.
Criterion E is not met and the study assessed only one English grammar structure with custom tests, so all-subject standardised assessment is absent.
Measurement stopped roughly two weeks after the intervention, with no tracking of participants to graduation or course completion, confirmed by internet search to find no follow-up publications on this cohort.
The paper contains no reference to any pre-registered protocol or trial registry, only ethics approval and grant funding, and no registry record was found via internet search.
This article reviews previous studies of the effects of implicit and explicit corrective feedback on SLA, pointing out a number of methodological problems. It then reports on a new study of the effects of these two types of corrective feedback...
Two pre-existing intact classes were assigned to conditions with no described randomisation, so a class-level RCT is not documented.
Outcomes were measured with a custom researcher-made writing test scored by the researchers, not a recognised standardised exam.
Outcomes were measured 12 weeks after the intervention began, which approximately equals one full academic term.
The control group's size, population, routine instruction, and baseline pre-test scores are documented in the text and Table 1.
Only two intact classes in one setting were used; there was no randomisation of schools.
The same researchers designed, taught, scored, and analysed the intervention with no independent evaluation team.
The study lasted only 12 weeks from start to final measurement, far below 75% of an academic year.
Both groups had the same class time and received writing instruction, differing only in the teaching method being tested, which is integral to the intervention.
No independent replication of this specific study is reported in the paper, and an internet citation search found none published since.
Only EFL writing was assessed with a custom test, so neither the standardised-exam prerequisite nor all-subject coverage is satisfied.
Tracking ended at the week-12 delayed post-test with no follow-up to graduation, and none was found via internet search.
The paper contains no reference to pre-registration or any trial registry, and none was found via internet search.
One of the most challenging aspects of foreign language learning is writing. Writing is the most demanding and complicated aspect of language system. Writing requires the collective effort of orthographic, graphomotor and other linguistic skills with the inclusion of semantics,...
Randomization occurred at the individual trainee level rather than by intact classes (or equivalent teaching units), so the ERCT class-level randomization requirement is not satisfied.
Outcomes were assessed with study-assembled case sets and rating scales rather than a widely recognized standardized exam.
The stated study period spans from July 18, 2024 to March 31, 2025, which exceeds a typical academic term length.
The control condition is described, but the paper does not provide control-group baseline demographics and baseline performance suitable for ERCT documentation expectations.
This is a single-site residency program trial with individual randomization, not a school- (or institution-) level RCT.
Although assessors were blinded for communication skills, the paper does not document that implementation and evaluation were conducted by an independent third party separate from the intervention designers.
The reported study window (July 2024 through March 2025) spans about 8.5 months, which exceeds 75% of a typical 9–10 month academic year, meeting the ERCT year-duration threshold.
The experimental condition includes added radiologist involvement that is integral to the teaching model being tested against conventional training, so the resource difference is the treatment rather than an unacknowledged confound.
No independent replication is reported in the paper, and no independently published replication of this specific trial was identified in an external literature search.
Because the study does not use standardized exam-based assessments, ERCT criterion A is automatically not met; in addition, outcomes focus on vascular training competencies rather than standardized exams across all core subjects.
The study reports a short follow-up and provides no tracking to a graduation or program-completion endpoint; no follow-up papers by the same authors reporting graduation tracking were identified.
The paper provides no registry identifier or dated statement showing that a protocol was pre-registered before data collection.
This randomized controlled trial evaluates an innovative interdisciplinary teaching model co-led by radiologists and vascular surgeons within China’s standardized residency training program. Forty trainees were randomized into two groups: one receiving collaborative teaching, which included joint lectures, radiologist-attended ward rounds,...
Randomisation was carried out at the individual student level (82 students randomly split into two groups) rather than by randomising entire pre-existing classes, and the paper is not a one-to-one tutoring intervention that would qualify for the exception.
Outcomes were measured with the national College English Test Band 4 (CET-4), a widely recognised standardised exam, with Gaokao English scores used as the standardised baseline.
The paper never states the intervention start date, its length or the interval to the CET-4 measurement, and the authors themselves describe the experimental period as relatively insufficient, so a term-long interval cannot be verified.
The control group's size (41 students), cohort, baseline Gaokao score distribution and mean, and its business-as-usual traditional instruction are all clearly documented.
The study randomised individual students within a single university programme; no schools or institutions were randomised.
The authors, English instructors at the university, designed the CBI intervention and conducted the teaching, testing and analysis themselves with no independent evaluator mentioned.
Criterion T is not met and the paper documents no intervention timeline, so tracking covering at least 75% of an academic year cannot be established.
Both groups took the same Metallurgy Engineering English course in regular class time and differed only in teaching method (CBI 6-T versus traditional), with no extra time, materials or budget given to the experimental group.
No independent replication of this specific CBI metallurgy-English experiment is reported or found; related prior CBI studies involve one of the same authors or different designs and contexts, and an internet citation search found zero papers citing or replicating this study.
Only English proficiency (CET-4) was assessed with a standardised exam; the professed professional-knowledge outcome and other subjects were never measured with standardised tests and no exception rationale is given.
Measurement stopped at the post-experiment CET-4 with no follow-up, criterion Y is not met, and an internet search for subsequent papers by the same authors tracking this cohort found none.
The paper contains no mention of any trial registry, pre-registration ID or pre-specified protocol, and no such registration was found through internet search.
This study investigates the impact of the Content-Based Instruction (CBI) theme-based instructional model on metallurgy major English classrooms. By integrating the CBI theme-based instructional model and employing the 6-T teaching method, the objective is to enhance students' language proficiency and...
Randomisation was done at the individual student level within a single university course, not at the class or school level, and no tutoring exception applies.
Outcomes were measured with researcher-adapted questionnaires and rubric-based teacher/machine ratings on the authors' own platform, not with a widely recognised standardised exam.
The intervention and outcome measurement spanned a full 16-week semester (early September to late December 2022), which covers at least one academic term.
The control group (G1, n = 26) is clearly documented, including its assessment condition, sample demographics, and baseline scores compared statistically with the experimental group.
Randomisation occurred among individual students within one course at a single university, so there was no school-level randomisation.
The E-platform intervention was developed by a team led by the first author, and the same authors conducted, funded, and analysed the study without any independent third-party evaluation.
The study lasted only 16 weeks (one semester), well short of 75% of a full academic year.
Both groups received the same course, the same platform, and three assessment modes each (self and teacher plus either peer or automated), so time and resources were balanced and the assessment mode itself was the treatment variable.
No independent replication of this study by a different research team is mentioned in the paper, and an internet citation search found no such replication published since.
Criterion E is not met, and only English public speaking outcomes were measured, with no standardised assessment of other core subjects.
Measurement ended with the Week 16 post-test in December 2022, criterion Y is unmet, and no internet-searchable follow-up publication tracks this cohort toward graduation.
The paper contains no mention of pre-registration of the study protocol on any registry before data collection, and no registration record was found via internet search.
This quasi-experimental research investigates the employment of a formative assessment platform aided by artificial intelligence in an English public speaking course. The platform integrates deep learning, automatic speech recognition, and automatic writing evaluation. It provides automated assessment and immediate feedback...
Two intact classes (not individual students within one class) were randomly assigned to the experimental and control conditions, satisfying the class-level randomisation requirement despite the very small number of clusters.
Outcomes were measured with a researcher-made 70-item multiple-choice vocabulary test built from the study's own target words, not a widely recognised standardised exam.
The intervention consisted of ten application activities run in August 2021 with the post-test given immediately upon completion of the treatment, far short of the one-term interval between intervention start and outcome measurement.
The control group's size, gender composition, baseline equivalence on the pre-test, and the traditional instruction it received are all explicitly documented.
Randomisation involved only two intact classes within a single university; no schools or institutions were randomised.
The same two authors developed the EVP application and conducted, analysed, and reported the evaluation themselves, with no independent evaluation team or third-party oversight.
The study ran only a short set of activities in August 2021 with immediate post-testing, nowhere near 75% of an academic year, and criterion T already fails.
Both groups received vocabulary instruction on the same 70 target words during the study period; the control had an active traditional teacher-led condition, and the EVP application itself was the integral treatment variable being tested against that business-as-usual method.
No independent replication of this specific EVP study by a different research team is reported in the paper or found via internet citation search of the literature.
Only English vocabulary was assessed, with a custom test, so neither the all-subject coverage nor the prerequisite standardised-exam criterion E is satisfied.
Measurement stopped immediately after the treatment with no follow-up at all, let alone tracking of participants to graduation, and the prerequisite criterion Y fails.
The paper contains no mention of any pre-registration of the study protocol on a registry before data collection, and no matching registration was found via internet search.
With the emergence of new technology, computers have been used for language learning. This experimental study aimed to develop English Vocabulary with a Picture Application (EVP) for improving students' daily English vocabulary memorization and examine the effectiveness of EVP used...
Treatment was randomly assigned to intact class sections (Section 1 vs Section 2), so randomisation was at the class level rather than among students within one class.
Outcomes were measured with self-made, researcher-designed pre- and post-tests, not a recognised standardised exam.
The post-test was given immediately after a brief audiobook listening session within one school quarter, so no term-long interval between intervention start and measurement exists.
The control group is documented with its size (20 Grade 7 students, Section 2), baseline pre-test mean and SD, and the conventional listening condition it received.
The trial involved only two class sections within one school (Colegio de Tablas), so there was no school-level randomisation.
The same researchers who created the audiobook and the tests also delivered the intervention and collected and analysed all data, with no independent evaluator.
The entire study, including the immediate post-test, took place within one school quarter, far short of 75% of an academic year (and criterion T is also not met).
The control group was an active control doing a conventional listening activity under the same procedure and time, so inputs were balanced and the audiobook medium was the treatment contrast.
This 2024 study of a novel, locally developed audiobook intervention has no independent replication by another research team in the peer-reviewed literature, confirmed by internet searches during verification.
Only listening comprehension was assessed with a custom test (criterion E fails), so no all-subject standardised assessment took place.
Tracking ended at the immediate post-test and the authors defer long-term effects to future research, so no graduation tracking occurred, confirmed by internet search during verification.
The paper contains no mention of any protocol pre-registration or trial registry, confirmed by internet search during verification.
This study investigates the impact of the Lib-RUNGOG Audiobook intervention on the listening comprehension skills of Grade 7 students. Employing a randomized pre-test post-test control group design, the study compares the performance of an experimental group exposed to the audiobook...
Allocation followed pre-existing course cohorts chosen by convenience sampling, and no genuine class-level randomisation procedure is described.
Outcomes were measured with a researcher-developed final achievement test, not a widely recognised standardised exam.
The intervention ran for a full semester with outcomes measured at the end of the term, satisfying the one-term minimum.
The GTM comparison group's size, demographics, baseline Nelson and LLOS scores, and instructional condition are documented in detail.
The entire study took place at one university with allocation among internal course cohorts, so no school-level randomisation occurred.
The same authors designed the intervention, taught the courses, built the outcome test, and analysed the data with no independent evaluators.
Tracking lasted only one semester (about 3-4 months), well below 75% of an academic year.
Both groups had identical class time, textbook, and course structure, and the CBI group's extra authentic content materials are integral to operationalising the CBI treatment itself, despite the authors framing the asymmetry as a study limitation.
No confirmed independent peer-reviewed replication of this specific CBI-versus-GTM trial was found in the paper or via external search of the citation graph.
Criterion E fails and only English was assessed with a custom test, so the all-subject exam requirement cannot be met.
Measurement ended with the end-of-semester post-test, and no follow-up toward graduation is reported in this paper or in any subsequent publication by the same authors found via search.
The paper contains no reference to any pre-registered protocol, registry, or registration date, and none was found through external search.
This paper investigated the effect of Content-based Instruction (CBI) on students' English language learning. In so doing, two methods of teaching English, that is, the CBI and the Grammar Translation Method (GTM) were compared with regard to the students' achievement...
Randomization was at the individual nurse (participant) level, not at a class/site level, and no tutoring-style exception applies.
Outcomes were measured using researcher-developed and self-reported instruments rather than standardized exams.
Outcomes were measured immediately at program end (or after about two weeks for controls), not at least one academic term after intervention start.
The control group is clearly described (no education), with group sizes and baseline characteristics reported in tables.
Randomization was not at the school/site level; individual nurses were randomized.
The research team developed and implemented the program and the PI generated the randomization sequence, with no external evaluator documented.
The study ran only from November to December 2023, far less than 75% of an academic year; per ERCT rules, Y is also blocked because T is not met.
The intervention added substantial instructional time/materials, but these resources are integral to the treatment being tested (receiving the CCOM program vs no program).
No independent replication of this specific CCOM program trial was found.
Criterion E is not met (no standardized exams), so A is not met, and outcomes were not broad core-subject exams.
There is no tracking to any graduation milestone, and per ERCT rules G is also blocked because Y is not met.
The protocol was registered in UMIN-CTR, and the registry shows a registered date before the trial's anticipated start and the paper's data-collection period.
Objective: This study aimed to evaluate the effects of the care coordination of pleural mesothelioma program (CCOM program), an educational program that we developed for nurses to improve their knowledge, attitude, and confidence on the care coordination of pleural mesothelioma...
Whole classes (eight classes in four matched pairs) were randomly allocated to conditions by an independent researcher, satisfying class-level randomisation.
Outcomes were measured with researcher-developed experimental tasks (delayed word repetition and rhyme judgment), not with any widely recognised standardised exam.
The intervention was a single one-hour session in mid-April with the posttest in May-June, an interval of only about one to two months, which is shorter than a full academic term.
The passive exposure control group is documented in detail, including its size, demographics, baseline matching, and the exact alternative instruction it received.
Randomisation was carried out at the class level within schools (classes paired within the same school), not by randomly assigning whole schools to conditions.
The authors designed the intervention, delivered it (Bassetti herself taught all sessions), administered assessments, and analysed the data, with no independent third-party evaluation team despite internal blinding.
Outcomes were measured only about one to two months after the one-hour mid-April intervention began, far short of 75% of an academic year, and criterion T is already not met.
The passive exposure control received a matched one-hour session from the same teacher with the same structure, the same words, and the same amount of spoken and written input and practice, differing only in the GPC content being tested.
The authors state this is the first study of its kind, no replication is mentioned in the paper, and an internet citation-database search of works citing this paper found no independent published replication.
Only L2 English phonology was measured with custom tasks; no standardised exams across all main school subjects were used, and criterion E is already not met.
Tracking ended with the May-June posttest; even the planned delayed posttest was cancelled, and an internet search for follow-up publications tracking this cohort found none, while prerequisite criterion Y is also not met.
The paper contains no mention of pre-registration of the trial protocol on any registry, and an internet check confirmed the OSF links in the paper are unregistered materials repositories, not a pre-registration.
The orthographic forms (spellings) of second language (L2) words and sounds affect the pronunciation and awareness of L2 sounds, even after lengthy naturalistic exposure. This study investigated whether instruction could reduce the effects of English orthographic forms on Italian native...
Randomisation was at the student level, but the intervention is one-to-one personalized AI instruction akin to tutoring, so the personal-teaching exception applies.
The outcome tests are only vaguely described as "standardized" with no test names, sources, or evidence of wide recognition, so use of a genuine standardised exam cannot be confirmed.
The interval from intervention start to outcome measurement was only 6 weeks, far shorter than one academic term.
The control group's composition, baseline pre-test scores (Table 1), and the exact condition it received are documented.
Individual students within one university program were randomised, not whole schools or institutional units.
No independent evaluators or third-party oversight are mentioned; the authors appear to have designed the platform and conducted the evaluation themselves.
The study spanned only 6 weeks from start to final measurement, far short of 75% of an academic year.
Both groups received equivalent online study time and content coverage; the only difference, AI-driven personalization, is the treatment variable being tested.
No independent replication of this specific trial is mentioned in the paper or identifiable via internet search (OpenAlex citation count = 0), as the study was only published in 2025.
Only English grammar and vocabulary were assessed, with no standardised exams across other subjects, and prerequisite criterion E is not met.
Outcomes were measured only immediately post-intervention, with no tracking of participants until graduation, and no follow-up publications were found via internet search.
No pre-registration, registry ID, or protocol statement appears anywhere in the paper.
Artificial Intelligence (AI)-driven personalized language learning holds great promise for customizing content and feedback according to learners' needs and proficiency levels. This study explores the potential of AI, particularly large language models (LLMs), for enhancing personalized language learning experiences. The...
Randomisation was at the individual learner level rather than by class or school, and no one-to-one tutoring exception applies.
The study used only the FLCAS self-report anxiety questionnaire and no standardised academic exam-based assessment.
Outcome tracking spanned about seven months from intervention start (10 weekly sessions plus a five-month follow-up), exceeding one academic term.
The control group's size, condition, and baseline FLCAS scores are documented with equivalence checks, despite demographics being pooled across groups.
Individual learners, not schools or language institutes, were the unit of randomisation.
The single author designed, conducted, and analysed the trial himself with no documented external or independent evaluation.
Total tracking of roughly seven months without specific dates falls short of a confirmable 75% of a full academic year.
Both groups had comparable conversation classes, and the chatbot access (with home practice) was the integral treatment variable tested against business-as-usual.
The study is framed as novel, citation databases show zero citing works, and no independent published replication of this specific trial exists.
No standardised exams in any subject were used (criterion E fails), so all-subject exam coverage is impossible.
Tracking ended at a five-month follow-up with no graduation endpoint defined, reached, or found in follow-up literature, and criterion Y also fails.
No pre-registration statement, registry ID, or registration date is provided in the paper, and none was found in the IRCT or ClinicalTrials.gov registries via internet search.
Purpose: This study aimed to evaluate the effectiveness of integrating intelligent educational chatbots into English conversation-based classrooms in reducing speaking anxiety among adult language learners. Methods and Materials: A randomized controlled trial was conducted with 30 adult intermediate-level English learners...
Randomisation was conducted at the individual level (with small friend/family clusters) within each delivery centre, not at the class level, and the intervention was group classes rather than one-to-one tutoring, so no exception applies.
English proficiency was measured with bespoke tests developed by the English Speaking Board specifically for this project, not a widely recognised standardised exam.
Outcomes were measured at the end of the 11-week intervention, an interval shorter than a full academic term of approximately 3-4 months.
The waiting-list control group is thoroughly documented, including its size, socio-demographic profile, baseline proficiency scores and the fact it received no provision during the trial period.
Randomisation was stratified within each of the 22 delivery centres, splitting participants inside every centre between arms, so no school/site-level randomisation took place.
The intervention was developed and delivered by Manchester Talk English under MHCLG, while the trial procedures, randomisation, data collection and analysis were carried out by independent evaluators (Learning and Work Institute, BMG Research and ESB assessors) with external oversight.
The trial tracked outcomes only over the 11-week course, far short of 75 per cent of an academic year, and criterion T is also unmet.
The extra 66 hours of English provision is itself the treatment variable being tested against a waiting-list business-as-usual control, so the deliberate resource difference is integral to the research question.
No independent replication of this CBEL trial by a different research team has been published; the report itself describes the trial as novel evidence-building, and a further internet search found no such replication.
Only English language proficiency was assessed, criterion E is unmet, and no other subjects were measured, so the all-subject exams requirement fails.
Measurement stopped at the end of the 11-week course with no longer-term or completion/graduation tracking, prerequisite criterion Y is unmet, and no follow-up publications tracking this cohort were found online.
A trial protocol exists as an annex to the report, but there is no evidence of registration on a public trial registry before data collection began; no ISRCTN or similar registry entry could be found online.
This report presents findings from a randomised controlled trial (RCT) of a Community-Based English Language (CBEL) intervention aimed at people with very low levels of functional English proficiency. The intervention consisted of 66 hours of guided learning and support delivered...
Randomisation was performed at the individual student level within a single university cohort, not at the class or school level, and the paper does not frame the intervention as one-to-one tutoring that would qualify for the exception.
Outcomes were measured with researcher-administered, CEFR-aligned pre/post listening tests and the FLCAS anxiety scale, with no named, widely recognized standardized exam.
The intervention ran 16 weeks (about four months, a full semester) with post-tests within one week of its end, so the interval from intervention start to outcome measurement covers at least one full academic term.
The control condition, its size, materials, instruction, and baseline comparability (technology familiarity, stratified gender/major, pre-test covariates) are documented, despite some internal inconsistencies in reported sample sizes.
Randomization was at the individual student level within a single university, not at the school or institution level.
The same two authors designed the AI system, designed the trial, collected the data, and analyzed the results, with no external or third-party evaluation team.
The interval from intervention start to final measurement was about 16-17 weeks (roughly 4 months), well short of 75% of an academic year, and no delayed follow-up was conducted.
Both groups received identical instructional dosage (2 x 45-minute sessions/week over the same period) with an active, well-specified control condition, and the AI system's added technology is integral to the treatment being tested.
This December 2025 study reports no replication of itself, and no independent peer-reviewed replication of this specific AI listening system trial exists or is cited; an internet search of citing literature confirms no replication study was found.
Only EFL listening proficiency (plus anxiety and strategy use) was measured, no other core subjects were assessed, and the prerequisite criterion E is not met.
Tracking stopped at an immediate post-test within one week of the 16-week intervention, with no follow-up to graduation, and the prerequisite criterion Y is not met; no follow-up publications tracking this cohort were found online.
The paper reports IRB approval and informed consent but contains no mention of prospective trial registration on any registry; an internet search of common registries found no matching pre-registration record.
This study addresses a critical gap in second language acquisition (SLA) research: the lack of integration between implicit comprehensible input and explicit metacognitive strategy training in AI-driven EFL listening instruction. With a mixed-methods design, it conducted a 16-week randomized controlled...
Randomization was at the individual student level, but the intervention is individualized, self-directed one-to-one pronunciation practice with an AI tutor, so the personal teaching/tutoring exception applies and student-level randomization is acceptable.
Outcomes were measured with a researcher-designed 30-word read-aloud task scored by three teachers, not a widely recognized standardized exam.
The intervention lasted three weeks with a delayed post-test about five to six weeks after the start, far shorter than one full academic term.
The control group's size, composition, baseline scores, and exact conditions (electronic dictionary practice with the same schedule and logging) are clearly documented.
Randomization was at the individual student level within a single private language institute, not across schools or institutions.
The authors designed the intervention and also collected and analyzed the data themselves, with no independent third-party evaluation team, so independence is not established despite blinded raters.
The full study window from intervention start to the delayed post-test was about six weeks, far below 75% of an academic year; also criterion T is not met, which precludes Y.
The control was an active condition with the same word lists, practice schedule, and logging, and the AI tool contrast (plus its brief orientation) is the integral treatment variable being tested against dictionary-based practice.
No independent replication of this 2025 study by a different research team was found in the paper or in an internet literature search conducted during verification.
Only pronunciation accuracy in English was measured, with a custom instrument, so neither the standardized-exam prerequisite nor all-subject coverage is satisfied.
Measurement ended at a delayed post-test two weeks after the intervention, with no tracking of participants to the end of their course of study, and criterion Y is not met, which precludes G.
The paper mentions ethics approval but contains no pre-registration statement, registry name, ID, or registration date, and none was found via registry search.
The integration of artificial intelligence (AI) into language education is rapidly transforming instructional practices and learner engagement. Within the domain of second language acquisition, pronunciation plays a crucial role in achieving communicative competence and intelligibility. Recent advancements in AI technologies...
Randomisation was conducted at the individual‑student level rather than at the class level, so the Class‑level RCT criterion is not satisfied.
The study used researcher‑designed custom tests rather than a standardized, widely recognized exam, so the Exam‑based Assessment criterion is not satisfied.
Outcomes were measured after a 4.5‑month intervention period, which covers at least one term, satisfying the Term Duration criterion.
The study provides detailed baseline characteristics and assessment outcomes for the control group, fulfilling the Documented Control Group criterion.
Randomisation was done at the individual‑student level, not at the school level, so the School‑level RCT criterion is not satisfied.
The same team that designed Mindspark also carried out the trial and analysis, so the Independent Conduct criterion is not satisfied.
Participants were followed for only 4.5 months rather than an academic year, so the Year Duration criterion is not satisfied.
The intervention’s extra instructional time is integral to the treatment, so the Balanced Resources criterion is satisfied.
No independent replication of the study is reported, so the Reproduced criterion is not satisfied.
Only mathematics and Hindi were assessed, so the All‑subject Exams criterion is not satisfied.
Participants were only followed until the endline test, with no graduation tracking, so the Graduation Tracking criterion is not satisfied.
The trial was registered only after data collection began, so the Pre‑registered Protocol criterion is not satisfied.
We study the impact of a personalized technology‑aided after‑school instruction program in middle‑school grades in urban India using a lottery that provided winners with free access to the program. Lottery winners scored 0.37σ higher in math and 0.23σ higher in...
The study randomizes intact sections (all students in a meeting) to flipped or lecture for each lesson, avoiding within‑session mixing and satisfying class‑level assignment.
The study used instructor‑designed course exams instead of standardized external assessments.
Learning outcomes were measured at end of semester, satisfying the term duration requirement.
The control (lecture) condition lacks detailed documentation of participant characteristics and baseline outcomes.
Randomization occurred within course sections rather than entire schools.
Authors who designed the intervention also implemented and analyzed the study.
Outcomes were measured only through one semester, not a full academic year.
Flipped lessons included mandatory pre‑class videos and in‑class exercises not equated by the control group.
An independent replication of the flipped classroom experiment by another research team has been published, satisfying this criterion.
Only econometrics outcomes were assessed, with no broad subject coverage.
The study did not track participants until graduation.
No evidence of pre-registration of study protocols.
Despite recent interest in flipped classrooms, rigorous research evaluating their effectiveness is sparse. In this study, the authors implement a randomized controlled trial to evaluate the effect of a flipped classroom technique relative to a traditional lecture in an introductory...
Randomisation was carried out at the individual student level within a single school rather than by class or school, and the intervention is a self-study mobile system, not one-to-one tutoring, so no exception applies.
Outcomes were measured with 50-item multiple-choice grammar tests custom-designed by two EFL teachers for this study, not a widely recognised standardised exam.
The intervention ran for 16 weeks (a full semester) from the pre-test at the beginning of the semester to the post-test after the treatment, which covers at least one academic term.
The control group's size, gender proportion, year level, English-learning history, baseline test scores, and exact conditions (system access limited to weekly assignments) are clearly documented.
The study took place in a single senior secondary school with randomisation of individual students, so there was no school-level randomisation.
The same authors designed the system, conceived the study, collected the data, and analysed the results, with no independent third-party evaluation or oversight reported.
The tracking period was only 16 weeks (about 4 months), well short of 75% of an academic year.
Although the experimental group received substantially more system access and daily practice time (at least 15 min/day), this added access and usage is the personalised m-learning intervention itself being tested against a near business-as-usual control that used the same system only for weekly assignments, so the imbalance is integral to the treatment.
No independent replication of this specific study by a different research team is mentioned in the paper or found via an internet citation search, and the authors describe it as among the first of its kind.
Only English grammar was assessed with a custom test, so with criterion E unmet and no other core subjects measured, all-subject standardised assessment is absent.
Measurement stopped at the post-test immediately after the 16-week treatment, criterion Y is not met (a prerequisite for G), and no follow-up publication tracking this cohort to graduation was found via internet search.
There is no mention of any pre-registration, registry platform, or protocol registration anywhere in the paper, and no matching registration record was found via internet search.
This study developed a personalized mobile-assisted system with a self-regulated learning (SRL) mechanism to facilitate English-as-a-foreign-language (EFL) students' learning of grammar. A quasi-experimental design, involving an experimental group (n = 278) and a control group (n = 320), was adopted...
Randomisation was performed at the individual student level within one university cohort, not at the class or school level, and the intervention is not one-to-one tutoring, so the class-level RCT requirement is not met.
The primary educational outcome was measured with the IELTS Listening test, a widely recognised standardised exam format, rather than a researcher-made instrument.
Outcomes were measured at the end of an eight-week intervention and at a three-week follow-up, about 11 weeks after the start, which is shorter than a full academic term of roughly 3-4 months.
The control group's size, demographics, baseline scores, and the instruction it received are documented in detail, including baseline equivalence tests against the experimental group.
The trial randomised individual students within a single university, so there was no school-level randomisation.
The sole author designed the study and conducted the data analysis, with no external or third-party evaluation team, so independent conduct is not demonstrated.
The total tracking window of about 11 weeks from intervention start is far below 75% of an academic year, and criterion T is already not met.
Both groups received identical instructional time, sessions, tasks and self-study loads, with the only difference being the AI tool and its immediate feedback, which are the integral treatment variable being tested.
No independent replication of this specific 2025 trial by a different research team was found via internet search; only unrelated studies using different tools and designs exist.
Only English listening comprehension was assessed with a standardised exam; no other main academic subjects were measured and no explicit rationale for a specialised- intervention exception is given.
Tracking ended three weeks after the intervention with no follow-up to graduation, and the prerequisite criterion Y is not met.
The paper reports IRB ethical approval but no pre-registration of the study protocol on any trial registry before data collection.
This randomized controlled trial explored the effects of employing AI-driven methodologies on enhancing listening comprehension, flow experience, and alleviation of listening anxiety among English as a foreign language (EFL) learners. A cohort of 84 Chinese university students, primarily associated with...
Whole intact classes were randomly selected and each entire class received one condition, so allocation was at the class level rather than within classes.
Outcomes were measured with custom researcher-made error correction and picture writing tests, not widely recognised standardised exams.
Outcomes were measured three hours and two weeks after a 90-minute intervention, far short of the required one-term follow-up.
The control group's size, demographics, baseline equivalence and no-treatment condition are documented in the text and tables.
The trial was conducted among four classes within a single middle school, so there was no school-level randomisation.
The authors designed the intervention, the tests and the analysis themselves with no external or independent evaluation team.
The full study lasted about four weeks, far below the required 75% of an academic year of outcome tracking.
Treatments took place within regularly scheduled lessons with matched duration and intensity across groups, so no extra time or budget favoured the intervention groups.
The paper itself replicates earlier work by others, and internet search found no independent team having replicated this specific study.
Only English passive-voice knowledge was tested with custom instruments, so neither standardised exams nor all-subject coverage is present.
Tracking ended two weeks after the intervention, with no follow-up of students until graduation found in the paper or via internet search.
The paper contains no reference to any pre-registered protocol or registry entry, and none was found online.
This study investigates how different form-focused instruction (FFI) timing impacts English as a foreign language (EFL) learners' grammar development. A total of 169 Chinese middle school learners were assigned to four conditions randomly: control, before-isolated FFI, integrated FFI, and after-isolated...
Randomization was at the individual student level (not by class or school), and the intervention is not one-to-one tutoring.
Outcomes were measured with a researcher-developed knowledge test rather than a widely recognized standardized exam.
The intervention is described as a three-session program with a 4-week follow-up, which is shorter than one academic term from intervention start.
The paper documents the control group’s size, condition (standard instruction), and baseline comparability information.
The study was conducted in a single university setting and randomized individual students rather than randomizing schools or sites.
The paper does not document independent third-party conduct separate from the intervention designers and study team.
Outcomes were not tracked for at least 75% of an academic year, and since T is not met, Y is not met by rule.
The paper reports equal instructional time across groups, and the main added inputs (concept maps and cooperative activities) are integral to the intervention rather than a separable resource confound.
No independent peer-reviewed replication of this specific trial was found.
E is not met (no standardized exams), and the outcomes are not all-subject standardized exams.
The study does not track participants until graduation, and since Y is not met, G is not met by rule.
The paper reports a ClinicalTrials.gov registration with a stated registration date before the study period, supporting prospective registration.
Background Cancer cases are increasing every day, which makes oncology nursing education very important. Nursing students need to learn how to manage the many symptoms caused by cancer and its treatments. This study aimed to examine the effect of a...
Randomization was carried out at the individual student level within a single course cohort, not at the class or school level, and the tutoring exception does not apply.
Academic achievement was measured with a researcher- developed, non-validated 15-item multiple-choice quiz rather than a standardised, widely recognised exam.
The intervention consisted of brief instructional sessions with an immediate post-test, far shorter than one full academic term.
Baseline demographics, sample sizes, and pre-instruction scores for the comparison (face-to-face and online) groups are documented in detail in Table 1.
The trial was conducted in a single course at one university with student-level randomization, not randomization of whole schools or institutions.
The same research team designed the metaverse intervention and conducted and analysed the study, with no independent third-party evaluation of the trial.
Because the term-duration criterion (T) is not met and the intervention lasted only brief sessions, the year-duration requirement is also not met.
All three groups received the same instructor, content, materials, and planned duration, differing only in delivery mode, so instructional time and resources were balanced across groups.
No independent replication of this specific trial by a separate research team is reported or identifiable.
Criterion E is not met, and outcomes were limited to a single palpation-related topic, so the all-subject exam requirement is not satisfied.
Criterion Y is not met and there was no follow-up beyond the immediate post-test, so graduation tracking is not satisfied.
Registry verification confirms ClinicalTrials.gov NCT07166705 was first submitted 2 Sept 2025, before the recorded study start of 5 Sept 2025, indicating prospective registration before data collection.
Background: Conventional synchronous online education may be less engaging than face-to-face instruction in physiotherapy education, particularly for practice-oriented content. Metaverse-based environments may offer a more interactive alternative, but comparative evidence in undergraduate physiotherapy students remains limited. This study compared metaverse-based,...
Allocation was at the intact division (class) level via a lottery rather than at the individual-student level, meeting the minimal class-level requirement despite a weak description.
The study used custom, researcher-adapted football skill-performance tests validated only by an expert panel, not a recognised standardised exam.
The intervention-to-measurement interval is not clearly documented and the reported window (under three months) does not clearly establish a full academic term.
The control group's size, baseline scores and command-method condition are documented, with baseline equivalence reported in Table 1.
The trial was run within a single school with allocation at the division level, so no school-level randomisation occurred.
The same researchers designed, delivered, measured and analysed the intervention, with no independent evaluator.
The study spanned under three months, far below 75% of an academic year, and the prerequisite T criterion is not met.
The only extra resource (home video content delivery) is integral to the flipped classroom strategy being tested, and in-class time is comparable, so the groups are balanced.
No independent replication of this specific trial exists; the cited works are related prior studies, not reproductions.
Only two football skills were measured with custom tests and the prerequisite E criterion is not met, so A fails.
Outcomes were measured only with immediate post-tests, with no graduation tracking, and the prerequisite Y criterion is not met.
The paper provides no evidence of any pre-registration of the study protocol on a public registry.
The study identifies the effect of the flipped classroom strategy and the command method in learning certain basic skills in football for second-grade students in the middle stage, and the preference of the effect between the flipped classroom strategy and...
Randomisation was performed at the individual nurse level, not at the class or school level, and no one-to-one tutoring exception applies.
The primary knowledge measure was a researcher-developed custom test, and the other instruments are standardized self-report disposition/perception scales rather than standardized exam-based achievement tests.
Outcomes were tracked with follow-up assessments at 1 month and 6 months after the training, exceeding one full academic term.
The control group's demographics, baseline characteristics, size, and the (traditional) training it received are documented in the text and in Table 1.
Randomisation occurred at the individual nurse level within a single hospital, not at the school/institution level.
The same authors developed the intervention program and also acquired and analysed the data, with no genuinely independent third-party evaluation of the intervention.
The longest follow-up was 6 months after training, which is below 75% of a full academic year.
This is an active-control design in which the control and placebo groups received comparable-time structured training programs, so educational time and resources were balanced across groups.
No independent replication of this specific RCT by a different research team is reported in the paper or identifiable through internet search.
Criterion E is not met and outcomes were confined to the patient-safety training domain, so all-subject standardized exam assessment is not satisfied.
Criterion Y is not met, and the study tracked working nurses only to a 6-month follow-up with no graduation-based long-term tracking.
There is no reference to a pre-registered protocol on any trial registry before data collection began, and none was found by internet search.
Objective: The aim of this study was to determine the effect of the Patient Safety Education Program based on the Constructivist and Traditional Learning Models on nurses' knowledge levels, critical thinking tendencies, and perceptions of clinical decision-making regarding patient safety....
Allocation was by intact whole school rather than by individual students within a class, satisfying the class-level randomisation requirement.
Outcomes were assessed with topic-specific validated scales and self-report questionnaires, not a standardised exam-based assessment.
Outcomes were measured immediately after a single ~45-minute session, far short of the required one-term interval.
The control group's size, demographics, baseline scores, and no-intervention condition are clearly documented.
Only two schools (one per arm) were used, which does not constitute a properly implemented school-level RCT.
The same authors designed, delivered, and evaluated the intervention, with no independent third-party conduct.
Outcomes were measured immediately after a single session, far short of a full academic year (and T is not met).
The extra educational session is itself the treatment variable under test, so the business-as-usual control is appropriate and the criterion is met.
No independent replication of this specific study by a different team has been reported or found.
Only a single epilepsy topic was assessed and the prerequisite criterion E is not met, so A fails.
Outcomes were measured only immediately after the session, with no graduation tracking (and Y is not met).
The paper reports ethics approval only, with no pre-registration on a public trial registry before data collection.
Background & Objective: Misinformation and misconceptions regarding epilepsy in society contribute to the increasing stigma surrounding the condition and its patients. This study aimed to examine the impact of education provided to high school students about epilepsy on their awareness...
Students were randomized at the individual level rather than by classroom, so cross-group contamination could occur.
The learning outcomes were measured using custom-built tests, not standardized exams.
Outcomes were measured immediately after two class meetings, not after a full academic term.
The control condition is clearly described with baseline group characteristics and identical materials.
Randomisation occurred at the student level, not at the school level.
The same research team designed and conducted the intervention, with no third-party evaluator.
The study spans two sessions with no year‑long follow‑up.
Time and materials were identical for both conditions, with only active engagement toggled.
The study has been independently replicated by a different research team.
Only physics learning was assessed, not all core subjects.
The study ended after the course, without tracking students to graduation, and no follow-up by the authors provided such data.
No pre-registration or protocol registry is mentioned.
We compared students’ self-reported perception of learning with their actual learning under controlled conditions in large-enrollment introductory college physics courses taught using active instruction and passive lecture. Both groups received identical content and handouts, and students were randomly assigned without...
Eight intact classes were randomly allocated as clusters using a computer-generated, sealed allocation sequence, satisfying class-level randomisation.
Outcomes were measured with a study-specific ski-technique rating rubric and self-report psychological scales rather than any widely recognised standardised exam.
The whole programme lasted three weeks with final measurement taken immediately at its end, far short of one academic term.
The control group's size, participant characteristics, baseline scores on all outcomes, and the conventional instruction it received are documented in detail and verified by fidelity observation.
Randomisation was among eight intact classes within one single university, not among schools or institutions.
The same research team adapted the protocol, trained the instructors, monitored fidelity, and analysed the data, with no external or third-party evaluator involved.
The study tracked outcomes for only about three weeks, nowhere near 75% of an academic year, and criterion T was also not met.
Time, technical content, practice volume, and instructor qualifications were explicitly matched across arms so that interaction style was the only substantive difference, with the brief autonomy-support instructor training being integral to the intervention tested.
Internet searching found no independent replication of this specific beginner-ski cluster RCT; the authors themselves list replication as future work.
Only ski technique was assessed, using a non-standardised study-specific rubric, and criterion E was not met, so the all-subject requirement fails.
Measurement stopped immediately after the three-week programme with no tracking towards graduation, no follow-up publication was found, and criterion Y was also not met.
Neither the paper nor any searched trial registry records a pre-registration for this study; only institutional ethics approval and CONSORT reporting are reported.
Background: Instructional interactions in beginner skiing classes may influence both objective skill performance and psychological outcomes. However, longitudinal intervention evidence in high-risk skill-learning contexts remains limited, and objective skill performance and psychological outcomes have often been examined separately. Objective: This...
Randomisation was carried out on individual students rather than on intact classes or schools, and the tutoring exception does not apply.
Outcomes were measured with psychometric creativity tasks (GAU, RAT) and a custom survey rather than with a recognised standardised exam.
The post-test was administered 16 weeks after the intervention began, which equals roughly one full academic semester.
The control group's size, attrition, baseline scores on all outcomes and business-as-usual condition are documented, though demographic detail is thin.
The trial ran in a single institution with individual students as the unit of randomisation, so no school-level assignment took place.
The same author team designed, delivered, measured and analysed the intervention, with no external evaluator or third-party oversight.
Outcomes were measured 16 weeks after the start, roughly one semester and well under 75 percent of an academic year.
The extra pre-class materials are integral to the flipped model being tested and reallocate rather than add class time, with class hours and instructor matched across arms.
This 2026 study presents itself as filling a research gap and internet searching found no independent replication of its specific design.
Criterion E is not met and the study measured only creativity constructs, with no attainment assessed in any curriculum subject.
Measurement stopped at the immediate post-test with no tracking to graduation in this or any locatable follow-up paper, and criterion Y was not met.
The paper reports institutional ethics approval only, with no trial registry identifier, pre-registration link or registration date.
Introduction: Creativity is widely acknowledged as an essential competency in art and design education, but approaches for developing it through pedagogical interventions have not been sufficiently explored. This study employed a 16-week self-regulated flipped classroom intervention to examine creativity outcomes,...
Randomisation was performed at the individual student level within a single college, not at the class or school level, and no tutoring exception applies.
Outcomes were measured with a custom multiple-choice knowledge test and a study-adapted checklist with bespoke scoring, not an externally administered standardised exam.
The final outcome was measured only one month after a single training session, far short of a full academic term (~3-4 months).
The control group's size, demographics, baseline performance, and business-as-usual conditions are clearly documented.
Randomisation was at the individual student level in a single college, not at the school level.
The intervention tools were third-party products (the AATE VR CPR Simulator software and AHA guideline content), the outcome was scored by blinded independent assessors, and the authors declared no conflicts of interest.
Outcomes were measured only one month after a single session, far below 75% of an academic year, and criterion T is not met.
The delivery modality itself is the treatment variable; practice repetitions and content were matched (control had more clock time), so no confounding extra resources favoured the intervention.
This specific 2026 trial has not been independently replicated; cited VR CPR trials are separate studies, not replications of it.
Only CPR knowledge and skill were assessed (a single specialised domain), and criterion E is not met, so the all-subject requirement fails.
Follow-up ended at one month with no tracking to graduation, and criterion Y is not met.
No trial registry entry or pre-registered protocol is reported; only ethics approval is mentioned.
Background: The rapid decay of cardiopulmonary resuscitation (CPR) skills remains a pervasive challenge in medical education. Although immersive virtual reality (VR) is increasingly used for training, its efficacy in mitigating long-term skill decay compared to traditional methods remains unclear. Methods:...
Randomization was at the individual student level within one cohort, not at the class or school level, and the video intervention is not one-to-one tutoring.
Outcomes used a study-derived CPR skills checklist and sensor metrics, not a standardized widely recognized exam.
Outcomes were measured about six months after training, exceeding the one-term tracking requirement.
The control group's size, demographics, and baseline characteristics are clearly documented in Table 1.
Randomization was at the student level within a single institution, so no school-level RCT was conducted.
The authors designed the intervention and also conducted and analyzed the trial, with no independent evaluator.
The roughly six-month tracking window is shorter than 75% of an academic year, so the year-duration requirement is not met.
The 2-min video refresher is the explicit treatment variable tested against a business-as-usual control, so the design is balanced.
This is a novel trial with no independent replication of the specific study reported.
With criterion E unmet and only CPR skills measured, the all-subject exam requirement is not met.
Tracking stopped at 6 months with no graduation follow-up, and criterion Y is unmet.
The paper reports IRB approval and CONSORT adherence but no public trial pre-registration ID or date.
OBJECTIVES: Cardiopulmonary resuscitation (CPR) skills decline within months after training, particularly among lay college students. We evaluated whether a brief video-based refresher improves 6-month CPR skill retention. METHODS: We conducted a single-blind randomized controlled trial among 2nd-year college students in...
Randomisation was performed at the individual nurse level within the same facilities rather than by class or school, with contamination explicitly acknowledged.
Outcomes were measured with researcher-developed tests and questionnaires, not widely recognised standardised exams.
Outcomes were measured only up to one month after the short intervention, far less than one full academic term.
The control group's demographics, size, baseline scores, and business-as-usual condition are clearly documented.
Randomisation was at the individual nurse level within facilities, not at the school/facility level.
The same authors developed the intervention and also designed, ran, and analysed the study, with no independent evaluator.
Follow-up was only about one month, far short of a full academic year, and criterion T is not met.
The extra resource given to the intervention group (the e-learning program) is itself the treatment variable being tested against a business-as-usual control.
This is described as the first study of its kind and has no independent replication.
Only cancer-cachexia nursing knowledge was assessed, and criterion E (standardised exam) is not met, so A cannot be met.
There was no tracking to graduation; follow-up was one month, and criterion Y is not met.
The trial was registered on UMIN-CTR on 27 November 2023, before enrolment (11 December 2023) and data collection (April 2024) began.
Background: Nurses often face challenges in assessing and managing cancer cachexia owing to the lack of standardized assessment tools and education. We developed the Comprehensive Cancer Cachexia Assessment Tool and an e-learning program to enhance nurses' knowledge and assessment skills...
Randomisation was performed at the individual student level rather than by whole classes or schools, so the class-level requirement fails.
The study used a custom OSCE checklist developed by the research team, not a recognised standardised exam, so the criterion fails.
The intervention was a single short session with outcomes assessed immediately afterward, well under one academic term.
Both groups' demographics, sizes, and received conditions are documented in Table 3 and the CONSORT diagram, satisfying documentation.
The trial randomised individual students within one school, so school-level randomisation is not met.
The same authors designed, delivered, and evaluated the interventions with no independent evaluator, so independence fails.
The intervention lasted a single session with immediate assessment, nowhere near a year, and T is not met.
Both arms received the same theoretical session and comparable short training time, with only the instructional modality differing, so resources are balanced.
The authors state this is the first comparison of these methods and an internet search found no independent replication.
Only a single skill domain was measured and criterion E is not met, so the all-subject requirement fails.
The study measured a single immediate outcome with no graduation tracking or follow-up publication, and criterion Y is not met.
The trial was registered in the IRCT in August 2023, before the September-December 2023 data collection, satisfying pre-registration.
BACKGROUND: Ensuring accurate assessment of patients' health is paramount in nursing interventions. The Glasgow Coma Scale (GCS) plays a crucial role in evaluating patients' conditions and is deemed a fundamental clinical skill for nursing teams. However, errors in GCS assessment...
Randomization was performed at the individual resident (student) level within a single department, not at the class or school level, and no tutoring exception applies.
Outcomes were measured with custom procedural performance metrics recorded by a blinded observer, not with a standardized, widely recognized exam-based assessment.
Training was a single 2-3 hour session and outcomes were assessed within one week, far short of the one full academic term required.
The control (didactic) group's size, novice status, eligibility criteria, and the instruction it received are clearly documented.
This was a single-center study with randomization at the individual resident level, not randomization of schools or institutions.
The same investigators designed the teaching interventions and conducted the trial; only outcome assessment was blinded, with no independent third-party evaluation.
With training in a single session and assessment within one week, the study spanned far less than 75% of an academic year, and criterion T was not met.
The teaching method itself is the explicit treatment variable, and both arms received comparable structured training time, making this an active-control comparison.
No independent replication of this specific trial by a different team is reported or identifiable; cited studies are related but not replications of this study.
Only a single procedural skill was assessed with custom measures, and since criterion E is not met, criterion A cannot be met.
Assessment occurred within one week of training with no follow-up to graduation, and since criterion Y is not met, criterion G cannot be met.
The trial was registered in the Clinical Trials Registry of India with a specific identifier, obtained after ethics approval and, per the paper, before participant recruitment.
Background and Aims: Neuraxial anesthesia is a core anesthetic skill, but traditional patient-based teaching carries ethical and safety concerns. Simulation-based training offers a structured, risk-free environment for novices to gain competence before performing procedures on patients. This study aimed to...
Randomisation was performed at the individual resident level using SPSS, not at the class or school level, and the intervention is group-based teaching rather than one-to-one tutoring.
Outcomes were measured with a study-specific 11-item OSCE checklist aligned to the training syllabus, not a widely recognised standardised external exam.
The final outcome was measured only one month after a very short training session, which is shorter than one full academic term.
The control group is clearly documented with baseline demographics, sample size, baseline scores, and the traditional teaching it received.
This is a single-center trial with randomisation at the individual resident level, not at the school or institution level.
Although outcome scorers and the statistician were blinded and external to the team, the authors who designed the modified method also implemented the trial and performed the analysis and conclusions.
Follow-up lasted only one month, far short of 75% of an academic year, and the prerequisite term-duration criterion T is not met.
Both arms received training in the same skills with no extra time or budget given to the intervention group (which in fact used less instructor time), and the added video/peer components are integral to the method tested.
This specific modified-Peyton wound-dressing/suture study has not been independently replicated; cited related studies are prior work in different contexts, not replications of this trial.
Only a single skills domain (wound dressing and suture removal) was assessed, and criterion E is not met, so the all-subject exam requirement cannot be satisfied.
Participants were tracked only to one month post-training with no graduation follow-up, and the prerequisite year-duration criterion Y is not met.
No pre-registration is reported; the paper explicitly states the clinical trial number is not applicable, with only ethics-board approval described.
Background: Peyton's four-step approach is a common teaching method in medical resident training. However, it has shortcomings such as long teaching duration, low efficiency, and poor long-term retention. This study proposed a modified Peyton's four-step approach and explored its effectiveness...
Randomisation was performed at the individual student level (by lottery), not at the class or school level, and no one-to-one tutoring exception applies.
The study relied on custom, study-specific instruments (a bespoke ART knowledge test, plus study-administered Mini-CEX and OSCE), not a widely recognised standardised exam.
Outcomes were assessed only one to two weeks after a short six-session course, well short of the required term-long tracking from intervention start.
The control group (n = 25, traditional lecture) is clearly documented with baseline demographics and performance in Table 1.
This is a single-centre trial with individual student-level randomisation, so no school-level randomisation occurred.
The same author team designed, delivered, and analysed the intervention; only outcome scoring was delegated to blinded examiners, which does not constitute independent conduct.
Term Duration (T) is not met and follow-up was only one to two weeks, far short of a full academic year, so Y is not met.
Instructional time was comparable across groups and the additional AI resources are the explicit treatment variable being tested, so the balance requirement is satisfied.
The study is a novel single-centre trial with no independent replication by another research team, and none was found in an external search.
Criterion E is not met and outcomes were limited to the single domain of reproductive medicine, so the all-subject standardised-exam requirement fails.
Year Duration (Y) is not met and participants were not tracked to graduation, only assessed one to two weeks post-course, so G is not met.
No public pre-registration on a trial registry is reported; the protocol is only available on request and an ethics approval number is not pre-registration.
Background: Efficient training of reproductive medicine clinicians is critical in the context of declining global fertility and increasing infertility. Traditional lecture-based instruction often fails to sufficiently develop clinical decision-making skills within limited residency rotations. Innovative strategies that integrate artificial intelligence...
Randomisation was performed at the individual student level (lottery within academic strata), not at the class or school level, and the intervention is not a one-to-one tutoring exception.
Outcomes were measured with an author-adapted, author-developed checklist that received only face validation, not a widely recognised standardised exam.
The intervention and outcome measurement spanned only about three weeks, far short of a full academic term.
The control group's composition, demographics, baseline scores, and condition (traditional bedside teaching only) are clearly documented.
Randomisation occurred at the individual student level within a single institution, not at the school or institutional unit level.
The same authors designed the simulation module and the same investigator delivered both arms, with no independent third-party evaluator conducting the trial.
Because criterion T is not met and the study spanned only about three weeks, the year-duration requirement cannot be satisfied.
The additional 27-hour SBL module is the explicit treatment variable being tested against a business-as-usual traditional-training control, so the resource difference is integral to the study design.
No independent replication of this specific trial by a different team is reported or identified.
Outcomes were limited to cardiorespiratory physiotherapy skills using a custom checklist, with no standardised all-subject exam assessment, and criterion E is not met.
Outcomes were measured at three weeks with no follow-up, and criterion Y is not met, so graduation tracking is not satisfied.
No pre-registration of the protocol on any trial registry before data collection is mentioned anywhere in the paper.
Introduction: Simulation-based Learning (SBL) is widely recognised for bridging the gap between theoretical knowledge and clinical practice in health professional education. However, its systematic application in Indian physiotherapy education is limited, and its impact on developing essential cardiorespiratory clinical skills...
The study allocated students to groups via purposeful (non-random) sampling rather than randomisation, so it is not a class-level RCT.
The study used a custom-developed test rather than a widely recognised standardised exam, so criterion E is not met.
The intervention covered only a 6-hour unit with immediate post-testing, far shorter than one academic term, so T is not met.
The control group's size, demographics, baseline scores, and conditions are clearly documented, so criterion D is met.
The study was conducted in a single school with non-random student allocation, so school-level RCT criterion S is not met.
The author designed, delivered, and evaluated the intervention with no independent third-party conduct, so I is not met.
The study lasted only a 6-hour unit, far short of an academic year, and T is not met, so Y is not met.
Both groups received the same electricity unit over comparable time, and the differing pedagogy (with its integral pre-training) is itself the treatment variable.
The intervention is described as novel and no independent replication was found, so criterion R is not met.
Only the single subject of electricity was assessed via a custom test and criterion E is not met, so A is not met.
Outcomes were measured immediately post-intervention with no follow-up to graduation, and Y is not met, so G is not met.
The paper contains no reference to pre-registration on any registry, so criterion P is not met.
This study aims to improve the results of online learning of a simple electrical circuit. It provides a solution to the barriers of synchronous distance learning, particularly in terms of interactivity and reducing the role of the teacher. The proposed...
Randomization was performed at the individual student level within a single department cohort, not at the class or school level, and the intervention is not one-to-one tutoring.
Outcomes were measured with a self-developed Script Concordance Test (SCT-SSDs) created and validated by the same authors, not a widely recognised standardised exam.
The training lasted only 14 days (with at most a 19-day window) and the post-test was administered immediately after training, far short of one full academic term.
The control group is well documented, with its composition, academic-level distribution, baseline SCT scores, prior course performance, and exact "report-only" condition all described.
Randomisation was at the individual student level within a single university department; no schools (or equivalent units) were randomised.
The same authors who designed the training modules and the outcome instrument also conducted the trial and analysed the data; no independent third-party evaluation is described.
Because the term-duration criterion (T) is not met, and the study tracked only ~14 days with an immediate post-test, the year-duration criterion is also not met.
The control group received a comparable active intervention (report-only videos of the same nine cases), and the reasoning-demonstration content was the explicit treatment variable being tested against this matched baseline.
This specific trial has not been independently replicated by a different research team in a peer-reviewed journal; internet searching found no such replication.
Only a single domain (clinical reasoning for SSD assessment) was assessed via a self-developed SCT, not all main subjects with standardised exams; criterion E is also not met.
Students were tested only immediately after the 14-day training, with no follow-up to graduation; criterion Y is also not met, and no follow-up publication tracking these students to graduation was found.
The paper reports ethics-committee approval but provides no pre-registration of the study protocol on a trial registry before data collection; no registration was found online.
Purpose: Training clinical reasoning skills remains a critical challenge in speech-language pathology education. This study aimed to evaluate the effectiveness of two cloud-based, self-directed instructional modules—diagnostic report viewing and reasoning demonstration—in enhancing students' reasoning skills for speech sound disorders (SSDs)...
Treatment was randomised at the individual student level within the same two classes, not at the class or school level, so the class-level criterion is not met.
The study used a child self-report executive function rating scale, not a standardised exam-based assessment, so this criterion is not met.
The 10-week intervention with immediate post-test only is shorter than one full academic term, so the term duration criterion is not met.
The control group's size, baseline characteristics, and conditions are clearly documented, so this criterion is met.
Schools were not the unit of treatment assignment; students were individually randomised, so the school-level RCT criterion is not met.
The same team designed, implemented, and analyzed the intervention with no independent external evaluator, so this criterion is not met.
The roughly 10-week duration is far short of a full academic year and criterion T is unmet, so the year duration criterion is not met.
The extra structured FMS time is the explicit treatment variable being tested against a business-as-usual control, so the criterion is met.
No independent replication of this specific FMS trial by a different research team exists, so this criterion is not met.
Criterion E is unmet and only executive function (not all main subjects via standardised exams) was assessed, so this criterion is not met.
The study had no long-term follow-up or graduation tracking and criterion Y is unmet, so this criterion is not met.
The paper reports ethics approval but no public trial registry pre-registration, so this criterion is not met.
Given the relative scarcity of randomized controlled trial evidence regarding systematic fundamental motor skills interventions for lower primary school students, this study aimed to examine the effectiveness of a 10-week structured fundamental motor skills intervention on the executive function and...
The study is an explicitly quasi-experimental design using total enumerative sampling, with no randomisation described at any level.
Outcomes were measured using a self-report motivation scale and researcher-developed proformas, not a standardised exam-based achievement assessment.
A two-week intervention with only an immediate post-test reported (no three-month data presented) does not demonstrate term-long tracking, so T is not met.
The control group's demographics, size, baseline academic characteristics, and scores are documented in detail across multiple tables.
The study used no randomisation at any level and therefore no school-level random assignment.
The same researcher designed, delivered, and evaluated the intervention, with no independent third-party evaluator.
The longest follow-up was three months, far short of the 75%-of-an-academic-year requirement.
The additional time and resources constitute the motivational package itself, which is the explicit treatment variable tested against a business-as-usual control.
No independent replication of this specific motivational-package study exists; cited works are different interventions, not replications.
Criterion E is not met and the study did not assess all core subjects via standardised exams, so this criterion fails.
Criterion Y is not met and tracking ended at three months with no follow-up to graduation.
The paper reports only ethics-committee clearance and provides no evidence of pre-registration on any trial registry.
Background: In education, academic motivation both intrinsic and extrinsic plays a critical role in student engagement, learning strategies, and achievement. However, challenges such as low engagement, academic stress, and fluctuating motivation continue to hinder outcomes. To address this, the present...
Randomisation was conducted at the individual student level (not class or school), and the intervention was a small-group, not one-to-one, design, so the criterion is not met.
The primary outcome (MSC) was measured with a researcher-adapted indicator exercise test rather than a standardized, widely recognized exam.
The intervention lasted one week and the posttest was given immediately afterwards with no follow-up, far short of one academic term.
The control group's size, demographics, baseline and posttest performance, and the computational thinking course it received are all clearly documented in the text and tables.
Randomisation occurred at the individual student level rather than at the school or site level, so the school-level RCT criterion is not met.
The intervention was developed, conducted, and analysed by the same research team (the authors), with no independent third-party evaluator, so the criterion is not met.
Outcomes were measured about one week after the intervention began, far short of an academic year, and criterion T was not met, so Y is not met.
The study used an active treated control group (a parallel computational thinking course of comparable planned instructional time and structure), so educational time and resources were broadly balanced across conditions.
The study is described as the first experimental evidence of its kind, and no independent peer-reviewed replication of this specific intervention was found.
Only mathematics-domain outcomes (MSC and arithmetic) were measured; no other core school subjects were assessed with standardised exams, so the criterion is not met.
No follow-up beyond the immediate posttest was conducted, no graduation-tracking follow-up paper was found, and criterion Y was not met.
The authors explicitly state that the study and analysis plan were not preregistered, so the criterion is not met.
Background: Mathematical structuring competence (MSC)—the ability to recognize, apply, and build patterns and structures—is a crucial component of mathematics education and an essential characteristic of mathematical talent. However, systematic approaches to fostering MSC, especially for talent promotion, are still lacking....
Randomization was performed at the individual student level within a single university cohort, not at the class or school level, and the intervention is not a one-to-one tutoring exception.
Outcomes were measured with a study-designed 20-item MCQ and a 30-item OSCE checklist developed by the investigators, not with a widely recognized standardized exam.
The primary skills outcome was measured only three weeks after a single-session intervention, far short of one full academic term.
The control comparator (simulation group) is documented with its size, exact intervention received, and baseline pre-test scores reported alongside the video group.
The trial was conducted at a single university with randomisation at the student level; no schools were randomly assigned.
The same investigator team designed the interventions and assessment instruments and conducted the study; only randomisation and OSCE scoring were delegated, with no independent third-party evaluator running the trial.
Because criterion T (term duration) is not met, the stronger year-duration criterion cannot be met; the study tracked outcomes for only three weeks.
Both groups received the identical 20-min lecture and a standardized 40-min modality-specific session, so instructional time and resources were balanced between the two active arms.
This single trial reports no independent replication of its specific intervention, and no independent reproduction of this study was identified.
The study measured only a single domain (basic life support) and did not use standardized exams across all main subjects; criterion E is also not met, which precludes A.
Tracking ended at the 3-week OSCE with no follow-up to graduation, and criterion Y is also not met, which precludes G.
The trial registry record (ClinicalTrials.gov NCT07368452) was first submitted 18 Jan 2026 and first posted 26 Jan 2026, after the study start date of 04 Oct 2025, so registration was retrospective, not before data collection began.
Background: Basic life support (BLS) is a core competency for medical students, yet simulation-based training can be resource intensive. Scalable alternatives such as facilitated interactive video may support learning, but evidence on skills retention is mixed. This study aimed to...
Randomization was conducted at the individual student level (via Excel RAND), not at the class or school level, and the intervention is a group program rather than one-to-one tutoring.
Outcomes were measured with ad hoc weekly multiple-choice course-content tests created by the authors, not a recognized standardized exam.
The intervention phase lasted four weeks within a 12-week (about three-month) study, and outcomes were measured well under one full academic term after the intervention began.
The control group's size, demographics, baseline performance, and business-as-usual conditions are documented in detail in Table 1 and the procedure section.
Randomization was at the individual student level within a single online course, not at the level of whole schools or institutions.
The same authors designed the BE-Social intervention and also conducted the trial, collected data, and analyzed the results, with no independent third-party evaluator.
The intervention and tracking spanned only about 12 weeks, far short of 75% of a full academic year, and criterion T is not met.
The control group received the same baseline inputs as all groups (daily questions, weekly feedback, weekly videos), and the educational components added to the treatment arms are the explicit treatment variables being tested.
No independent replication by a different research team exists; this study is itself a same-team replication and expansion of Tarifa-Rodriguez et al. (2024).
Only course-specific (applied psychology) outcomes were measured with custom tests, no standardized all-subject exams were used, and criterion E is not met.
Tracking ended after the 12-week study with no follow-up to graduation, and criterion Y is not met.
The paper provides no pre-registration statement, registry identifier, or registration date predating data collection.
Few randomized controlled trials have analyzed evidence-based educational practices delivered through a social media environment. This study used a multi-arm randomized controlled trial to evaluate the critical components of an educational intervention package: study self-management skills training delivered through video...
Randomization was performed at the individual student level within a single cohort, not at the class or school level, so the class-level RCT criterion is not satisfied.
Outcomes were measured with study-specific test scores, skills checklists, and satisfaction surveys rather than a recognized standardized exam.
The intervention was a single 40-minute lecture with outcomes measured immediately and at one week, far short of one academic term.
The control group's size, demographics, baseline pre-test scores, and business-as-usual treatment are documented, allowing comparison with the intervention group.
Randomization occurred at the individual student level within a single university cohort, not at the school level.
The same authors designed, delivered, and evaluated the intervention, with no independent or third-party evaluator involved.
Because Term Duration (T) is not met and the intervention/measurement spanned only about one week, the year-duration criterion is not met.
Both groups received an identical 40-minute lecture from the same instructor using the same materials and content, so instructional time and resources were balanced.
No independent replication of this specific BOPPPS AIDS-education trial by a different team is reported or identified.
Outcomes were limited to a single AIDS/infectious disease topic without standardized exams across all core subjects, and criterion E is not met.
Tracking stopped at one week after a single lecture, with no follow-up to graduation, and the prerequisite Year Duration criterion is not met.
No pre-registration of the study protocol on any registry is mentioned anywhere in the paper.
Background: Infectious diseases education faces dual challenges: addressing knowledge sensitivity and bridging the theory-practice gap. Traditional didactic teaching inadequately cultivates both technical skills and humanistic literacy. Objective: To evaluate the tripartite efficacy of the BOPPPS model in AIDS education, focusing...
Randomization was at the student (within-school) level rather than at the class (or school) level, and the intervention is not one-to-one tutoring.
Outcomes were measured via Likert-scale questionnaires rather than standardized exam-based assessments.
The intervention lasted four 45-minute lessons with a posttest immediately after the intervention, which is far shorter than an academic term.
The control condition (SGM) is clearly described, and baseline outcome data by condition are reported, enabling treatment-control comparison.
Schools were not randomized; instead, students within participating schools/classes were randomly assigned to conditions.
The study was implemented and evaluated within the project team, with no clearly independent third-party evaluator for conduct and analysis.
The intervention and measurement occurred over a very short period, not over at least 75% of an academic year.
The two conditions are described as parallelized and appear to match in lesson time and materials; the difference is the instructional activity, not added time/budget.
No independent replication of this specific experimental study was found in the paper or via web search as of 2026-03-03.
The study does not use standardized exam outcomes (and thus fails Criterion E), so it cannot meet the all-subject standardized exam requirement.
The study did not track students until graduation; additionally, because Criterion Y is not met, Criterion G cannot be met under the ERCT rules.
No pre-registration statement or registry ID/date was found in the paper, and no corresponding public pre-registration record was identified online.
According to Self-Determination Theory, it is essential to experience autonomy, competence, and relatedness in order to develop intrinsic motivation and motivational beliefs such as self-efficacy. Posing and solving one’s own modelling problems may support these needs by offering opportunities for...
Randomisation was carried out at the individual student level (block randomisation, block size 3) within a single cohort, not at the class or school level, and the intervention is not one-to-one tutoring.
Outcomes were measured only with self-report Likert scales (digital competence, clinical self-efficacy, AI attitude), not with any standardised exam-based assessment of educational achievement.
The intervention lasted only eight weekly 90-minute sessions and outcomes were measured immediately at the end, far short of a full academic term of follow-up from intervention start.
The traditional-methods control group is documented with size (n=30), demographics, baseline scores, and a clear description of the business-as-usual condition it received.
Randomisation was at the individual student level within a single institution; no schools (or comparable institutional units) were randomised.
The same two authors designed the scenarios, ran the intervention, collected and analysed the data; only the randomisation step was delegated, with no independent external evaluator for conduct or analysis.
Criterion T is not met (eight-week intervention with immediate post-test), and the study covers nowhere near 75% of an academic year, so Y automatically fails.
All three groups received the same eight-week, weekly 90-minute scenario-based course with equivalent time on task; the only difference was the information source (AI vs internet vs traditional materials), which is the treatment variable itself, so educational time and resources are balanced.
No independent replication of this specific study by a different research team in a peer-reviewed journal is reported or identifiable; the paper is a single original study.
Criterion E is not met, and outcomes covered only a single narrow domain (rehabilitation/clinical scenarios) via self-report scales, not standardised exams across all main subjects, so A fails.
Criterion Y is not met, and outcomes were measured immediately after an eight-week programme with no follow-up to graduation, so G fails.
The authors explicitly state the study was not registered in any clinical trial registry, so no pre-registered protocol exists.
Background This study aims to investigate the effects of using patient scenarios generated by artificial intelligence in rehabilitation education through different training models (artificial intelligence-based, internet-supported+traditional, and traditional only) on students' digital competence, clinical self-efficacy, and attitudes towards artificial intelligence....
Allocation was at the individual student level within a single course (by student ID number parity), not at the class or school level, and the intervention is not one-to-one tutoring.
Outcomes were measured with tests developed in-house by the teaching team, not with a widely recognised standardised exam.
The teaching spanned only five sessions with the post-test administered three days after completion, far shorter than one academic term.
The control (LBL) group's size, demographics, baseline scores, and instructional conditions are clearly documented.
Allocation was at the individual student level within a single institution, not at the school level.
The same team designed, taught, and assessed the intervention, with the same instructor delivering both conditions, and no independent third-party evaluator was involved.
Because the Term Duration (T) criterion is not met, the stronger Year Duration criterion cannot be met, and the study lasted only a few weeks.
Both groups received instruction on the same content from the same instructor with broadly comparable instructional time; the difference is the pedagogical method itself, not an unmatched addition of time or budget.
The study has not been independently replicated by a different research team, and no such replication is reported or found.
Criterion E is not met and only the single subject of pulmonary imaging was assessed, so the All-subject Exams criterion fails.
Because Year Duration (Y) is not met and the study had no follow-up beyond a few days post-intervention, there is no tracking through graduation.
The paper reports no pre-registration and explicitly states "Clinical trial number: Not applicable."
Background: Pulmonary imaging plays a critical role in clinical diagnosis, and developing diagnostic reasoning skills is a key objective of undergraduate medical imaging education. Traditional lecture-based learning (LBL) may be insufficient to actively engage students or to cultivate higher-order diagnostic...
Allocation was made at the individual level using the parity of each nurse's work number within a single hospital cohort, not at the class or school level, and no tutoring exception applies.
Knowledge was measured with a researcher-made 50-item closed-book exam and a custom observational attention scale, not a recognised standardised exam.
The intervention and all outcome measurement occurred within a single five-day training program, far shorter than one academic term.
The control group is documented with its size, training condition (traditional instruction only), and baseline demographic and educational characteristics in Table 3.
Allocation was at the individual level within one hospital cohort, with no randomisation of schools or institutions.
The same research team designed the intervention and instruments and ran the trial; the only independent parties were unaffiliated observers for attention scoring, not an external evaluation of the study as a whole.
The study spanned only five days with no year-long tracking, and since Criterion T is not met, Y cannot be met.
Both groups followed the same curriculum, teachers, materials, and class time; the only difference was the short 2-6 minute KGBL quizzes, which are integral to the gamification intervention being tested rather than a separable extra resource.
No independent replication of this specific pilot trial exists; the paper cites prior Kahoot studies as background but none reproduce this study.
Only theoretical knowledge for the nurse-training content was assessed (via a non-standardised exam), not all core subjects, and since Criterion E is not met, A cannot be met.
There was no follow-up beyond the five-day program and no graduation tracking, and since Criterion Y is not met, G cannot be met.
The paper contains no mention of any pre-registration on a trial registry before data collection.
Purpose: Teachers often face the challenges in maintaining the student attention in class. This study aimed to evaluate the effectiveness of the Kahoot! game-based learning (KGBL) platform in enhancing attention span, theoretical knowledge acquisition, and satisfaction among newly graduated nurses...
Randomization occurred at the individual student level rather than by intact classes (and the intervention is not one-to-one tutoring), so the class-level RCT requirement is not satisfied.
The primary outcome is an author-developed, CEFR-aligned proficiency battery rather than a widely recognized standardized exam, so the exam-based assessment requirement is not satisfied.
The paper reports the dosage in sessions (15 × 90 minutes) but does not clearly document calendar start and outcome measurement dates showing at least one full academic term elapsed from start to measurement.
The control condition is clearly described (traditional instruction), with sample size and baseline/posttest descriptives reported in tables, satisfying the documented control group requirement.
Randomization was not conducted at the school (or site) level; participants were randomized individually, so the school-level RCT requirement is not satisfied.
The intervention platform was custom-developed for the study and there is no clear statement that an independent external evaluator conducted the trial, so independent conduct is not established.
The paper does not provide start and measurement dates demonstrating outcome measurement at least 75% of an academic year after the intervention began, and criterion T is also not met.
Instructional time appears matched across arms (15 × 90-minute sessions) and the added technology resources are integral to the intervention being tested versus business-as-usual, so the balanced control requirement is satisfied.
No independent replication by a different research team in a different context could be identified for this 2026 study.
Because the study does not meet criterion E (it uses an author-developed assessment rather than standardized exams), it cannot meet the all-subject standardized exams requirement.
The study does not report tracking participants through to graduation, and because criterion Y is not met, graduation tracking is also not satisfied.
No explicit pre-registration statement or registry identifier is provided showing the protocol was registered before data collection began.
This randomized controlled trial tested the effects of immersive Virtual Reality (VR) enhanced with artificial intelligence on English language development, operationalized as performance on an author-developed, CEFR-aligned language proficiency battery emphasizing grammatical and lexical performance, in undergraduate Chinese EFL learners...
Randomization was done by intact student groups (cluster-style) rather than mixing treatment and control within the same instructional group.
Outcomes were measured using an instructor-created exam derived from a textbook, not a widely recognized standardized exam.
The intervention and primary outcome measurement occurred over weeks, with the exam administered one week after course completion, which is shorter than one academic term.
The paper describes the traditional-teaching control condition and reports baseline demographics for both groups.
Randomization occurred within one institution among student groups, not across multiple schools or institutional sites.
The authors designed the intervention and also collected and analyzed the data; the paper does not document independent external conduct of the study.
Outcomes were measured within weeks rather than over at least 75% of an academic year, and Y is also not met because T is not met.
The WeChat-based PBL condition required extra out-of-class effort and instructor monitoring without clearly documented time- and effort-matched inputs in the control condition.
No independent replication of this specific trial was found in the paper or via a targeted literature search.
The study does not use standardized exams (E not met) and assesses only an ophthalmology topic rather than all main subjects.
The study did not track outcomes through graduation, and G is also not met because Y is not met.
No evidence of a pre-registered protocol (with an ID and a registration date before data collection) was found.
Background: Ophthalmology poses distinct learning challenges for medical students due to the complex anatomy of the eye and the requirement of essential hands-on skills. Problem-based learning (PBL), a student-centered approach, fosters clinical reasoning and self-directed learning. To address the time...
Individual‑level randomization within one class violates the class‑level RCT requirement.
Assessments were custom course quizzes and a final, not a standardized exam.
Outcomes were collected after only five weeks, not a full term.
Demographics and baseline performance for the control group are fully reported.
Randomization did not occur at the school level.
The same team designed, implemented, and evaluated the study.
Study duration was five weeks, not a full academic year.
Control and treatment groups received equivalent emails and incentives; no extra resources favored treatment.
No independent replication of this RCT is reported.
Outcomes are limited to a single STEM course, not all subjects.
No tracking beyond the short 5‑week course was conducted.
No evidence of prospective trial registration is provided.
Time‑management skills are an essential component of college student success, especially in online classes. Through a randomized control trial of students in a for‑credit online course at a public 4‑year university, we test the efficacy of a scheduling intervention aimed...
Participants were randomized at the individual student level within one course, not by class (and no tutoring exception is explicitly stated).
Outcomes were measured with researcher-created writing prompts and rubric-based ratings (adapted from IELTS descriptors), not a standardized exam.
The outcome measurement occurred about two weeks after the intervention started, far shorter than an academic term.
The control group is clearly defined (teacher feedback), with group sizes and baseline performance reported.
Randomization was not at the school (or site) level; the study took place within a single institute setting with student-level assignment.
While raters were independent and blinded, the study does not document an independent external evaluation team conducting the trial overall.
The study’s pre-to-post tracking spans only about two weeks, which is far less than 75% of an academic year.
The study explicitly matched feedback structure and approximate length across groups, indicating comparable resources and dosage.
No independent peer-reviewed replication of this specific study was identified, and the paper itself does not report replication.
Criterion E is not met, so criterion A is automatically not met; additionally, outcomes focus on writing only rather than all core subjects.
Criterion Y is not met, so criterion G is automatically not met; the paper also reports only short-term outcomes and does not track learners to any graduation milestone.
The paper provides no pre-registration identifier or registry link, and it explicitly lists the clinical trial number as not applicable.
Although considerable research has explored the role of feedback in second language writing, limited studies have compared the effects of AI-generated feedback and teacher-written feedback on both academic performance and emotional experience, particularly in EFL contexts. This mixed-methods brief report...
The study allocated individual children (not intact classes or schools) using visit-order block allocation, not class-level randomisation.
Outcomes were measured with study-specific pre/post geography tasks, not with a widely recognised standardised exam.
Outcomes were measured immediately after a single 10-15 minute session, not at least one academic term after intervention start.
The control condition is clearly described and the paper reports group demographics and baseline/outcome descriptives by group in Tables 1-2.
The study did not randomise schools/sites; it recruited children at a museum and allocated individual children to conditions.
The paper does not document independent third-party evaluation; the author team oversaw the study and authors performed data collection.
Outcomes were measured immediately after a single session, so the study does not meet the ≥75% academic-year duration requirement (and T is also not met).
Time-on-task and materials were matched (same videos, same duration); the difference was movement versus seated viewing as the intended treatment contrast.
No independent peer-reviewed replication of this specific study was found in the paper or via web search as of the ERCT check date.
Because criterion E is not met (no standardised exam), criterion A is automatically not met; the study also assesses only geography content.
Criterion Y is not met, so graduation tracking is automatically not met; no follow-up publications tracking this cohort to graduation were found.
The paper provides an OSF link for data/code but does not document a preregistered protocol with a registration date before data collection.
This video-based study explores the feasibility, acceptability, and efficacy effects of an online movement-based intervention in young children. Seventy-five children were assigned to either an embodied cognition group (watched videos and performed simple full-body movements) or a control group (watched...
Randomization is described at the student level, not at the class (or school) level, and the intervention is not a one-to-one tutoring exception.
Outcomes are measured with researcher-designed instruments and puzzle sets (including Raven-style items), not a clearly identified widely used standardized exam.
Outcomes are measured at Week 25 after a 24-week intervention, which is longer than a typical academic term.
The control condition is described and baseline equivalence is reported, providing sufficient documentation of the control group.
The sample comes from three schools, but assignment is described for individual students rather than randomizing schools to conditions.
The paper does not document an independent third-party evaluator; author roles indicate the authors conducted the study’s key activities.
The study’s pre/post window is Week 0 to Week 25 (about six months), which is below the ERCT year-duration threshold (≥75% of an academic year).
The intervention provides substantial additional structured instruction (48 hours) while the control is explicitly described as passive business-as-usual, so time/engagement resources are not balanced.
No independent replications by other research teams were identified, and the paper itself states replication across settings is needed.
Criterion E is not met (no standardized exam), so criterion A is also not met; additionally, the study does not assess achievement across all core subjects.
Criterion Y is not met (not year-long), so criterion G is automatically not met; no evidence of tracking to graduation (or equivalent completion) was found in the paper or in follow-up publications.
No pre-registration registry, identifier, or registration date (prior to data collection) is reported, and no external registry entry was found.
Introduction: This research paper presents an empirical investigation of the effectiveness of early coding instruction in improving problem-solving skills and computational thinking (CT) among primary school students. The primary research question was to determine whether a structured six-month coding intervention...
Random assignment occurred at the individual student level, not by entire class or school.
The study employed custom 22‑item free‑response tests rather than a standardized exam.
Outcomes were measured over a seven‑day period, not a full term.
The negative control group’s composition and baseline data are clearly documented.
Randomization was at the individual student level, not by school.
The intervention was designed, delivered, and assessed by the same team without independent oversight.
The study tracked outcomes over one week, not a full academic year.
The extra evening sessions are central to the intervention and thus the unmatched control is acceptable.
No independent replication of this RCT is reported.
Only stereochemistry outcomes were measured, and E was not met.
Tracking ended after one week; no graduation‑level follow‑up.
No pre‑registration of the study protocol is mentioned.
The use of the flipped classroom approach in higher education STEM courses has rapidly increased over the past decade, and it appears this type of learning environment will play an important role in improving student success and retention in undergraduate...
The study randomized individual participants rather than classes (and it is not a one-to-one tutoring exception).
The outcomes are questionnaires and rubric/expert ratings, not a widely recognized standardized exam.
The study tracks outcomes over four weeks with only a one-week follow-up, which is shorter than one academic term.
The control condition and baseline equivalence are documented, including what the control group did and a baseline table (Table 2).
The trial is single-institution with participant-level randomization rather than school/site-level randomization.
The paper reports some safeguards (independent assistant for randomization and blinded raters), but it does not document that the study was conducted by an independent external evaluation team separate from the intervention developers.
Outcomes are measured over weeks rather than at least 75% of an academic year, and Y is also disqualified because T is not met.
Time and instructor attention are explicitly matched across arms, and the extra technology (smartphone AR overlay) is the integral treatment contrast rather than an unbalanced add-on resource.
No independent replication study by other authors was found in the paper or via internet searching, and the authors instead call for larger, longer trials.
The study does not use standardized exam-based assessments across subjects, and A is disqualified because E is not met.
The study includes only immediate post-test and a one-week follow-up, and no follow-up paper tracking participants to graduation was found; G is also disqualified because Y is not met.
The paper states outcomes and analyses were prespecified but provides no public preregistration record (registry, ID, and preregistration date) showing registration before data collection began.
We conducted a parallel-group randomized controlled trial in routine calligraphy classes, comparing smartphone-based augmented reality with conventional model copying in twenty adults. The instructor, copybook, venue, and contact time were held constant. The primary outcome—calligraphy self-efficacy— improved more in the...
Random assignment was at the individual student level (within one course), not at the class (or higher) level, and the intervention is not a one-to-one tutoring exception.
Outcomes were measured with study-specific instruments (expert ratings via a rubric and a custom pre-test), not a widely recognized standardized exam.
The intervention and outcome measurement happened within a single short session (about 60 minutes), far shorter than one academic term.
The unguided-rubric control condition is clearly described, and baseline and equivalence information for both groups is reported.
Randomization occurred among individual students, not among schools (or equivalent institutions/sites) implementing the intervention.
The paper does not document that the trial was conducted and evaluated by an independent third-party team separate from the intervention designers/authors.
Year-duration tracking is not reported and, since the study is far shorter than a term, it necessarily fails the one-academic-year duration requirement.
The control condition appears balanced because both groups had the same task and interface, and the only difference was the presence of performance level descriptors (no added time or budget/resources to one group are indicated).
No independent, peer-reviewed replication of this specific study was found, and the paper does not report that it has been replicated by an independent external team.
Because the study does not use standardized exams (criterion E is not met), it cannot satisfy the all-subject standardized-exams requirement.
The study does not track participants through graduation, and because criterion Y is not met, criterion G cannot be met under the ERCT dependency rule.
The paper provides an OSF link for materials/data/analyses availability but does not state that the study was pre-registered with an ID and date before data collection began.
Writing from multiple texts is a widespread task in higher education and requires advanced reading, writing and self-regulatory skills. Research shows that learners often struggle to accurately evaluate their performance in such tasks. This study examines whether rubrics can enhance...
Randomisation was performed among individual students rather than by class (or school), so the class-level RCT requirement is not met.
Outcomes rely on a researcher-developed Nutrition Knowledge Test and questionnaire scales rather than a widely recognised standardised exam-based assessment.
The between-group outcome comparison is conducted immediately after an 8-week intervention (shorter than an academic term), and the later follow-up lacks control-group measurement.
The control group’s size, baseline characteristics, and business-as-usual condition (no nutrition education/materials) are clearly described.
The study is conducted within a single primary school and does not randomise at the school level.
The intervention delivery and the research activities were carried out by the study team, with no explicit independent third-party evaluation.
The intervention and follow-up timeline is far shorter than an academic year; additionally, T is not met, so Y cannot be met.
The intervention adds instructional time and materials, but those additional resources are the treatment itself (a nutrition education programme) being tested against business-as-usual.
No independent, peer-reviewed replication of this specific trial was identified in the paper or through targeted searches using the DOI and trial identifier.
The study does not use standardised exam outcomes across core academic subjects, and E is not met, so A cannot be met.
The study does not track students to graduation; additionally, Y is not met, so G cannot be met.
The trial registration date reported in the paper is after the data collection period, so the protocol was not pre-registered before the study began.
Background: This study aimed to evaluate the effects of a dietitian-led, school-based nutrition education programme on primary school students’ nutrition knowledge, attitudes, behaviours, and anthropometric measurements. Methods: A randomised controlled, prospective design was conducted in a primary school in Mardin...
Randomization was described at the individual student level, not at the class (or school) level, and no tutoring exception applies.
Outcomes were measured with self-report questionnaires (PSQI, ASQ, FFMQ) rather than standardized exam-based achievement assessments.
Outcomes were measured after an 8-week program, which is shorter than a full academic term and not documented as term-long follow-up.
The paper documents the control condition ("no training"), group sizes, and baseline comparability information.
Participants came from a single middle school and randomization was not performed at the school level.
The paper does not explicitly state that the intervention delivery and evaluation were conducted independently from the intervention designers.
Outcomes were measured after 8 weeks, which is far less than 75% of an academic year, and criterion T is not met.
The intervention provided additional time and structured activities, but these resources are integral to the MBSR package being tested against a no-intervention control.
No independent replication of this specific trial was found in searches of the DOI/title as of 2026-03-14.
Criterion E is not met and the study does not report standardized exam outcomes across core subjects.
The paper reports no long-term follow-up (it notes durability was ignored), and criterion Y is not met.
The paper does not mention a registry ID or protocol pre-registration date, and no corresponding registration record could be verified.
This study aimed to evaluate the impact of an 8-week Mindfulness-Based Stress Reduction (MBSR) program on sleep quality and academic stress among Chinese school adolescents. A total of 46 students were recruited and randomly assigned to either the experimental group...
Randomization and delivery were organized at the small-group level rather than at the class (or school) level.
The primary educational outcome (CVS comprehension) was measured with a research instrument rather than a widely recognized standardized exam.
The study timeline (eight weeks plus a 2-week post-test) is shorter than an academic term.
The control condition, control sample size, and baseline descriptive statistics by condition are clearly documented.
Multiple schools participated, but assignment was not randomized at the school level.
The paper does not provide a clear statement that an independent third party conducted the evaluation separate from the intervention designers.
Outcomes were not measured for 75% of an academic year, and ERCT rules also make Y not met when T is not met.
The two conditions were designed to be equivalent in time and structure, differing primarily by contextual embedding.
No independent replication of this specific trial was found in the paper or via targeted online searches as of the ERCT check date.
The study does not report standardized exam outcomes across all core subjects, and ERCT rules make A not met when E is not met.
The study does not track students until graduation, and ERCT rules also make G not met when Y is not met.
No pre-registration registry, ID, or pre-registration date is reported, and no public pre-registration record was found in targeted online searches.
Context personalization is an instructional approach aimed at enhancing students’ engagement and cognitive processing by embedding learning content in familiar contexts. Numerous studies explore the benefits of personalized tasks for learning, but few empirically examine cognitive mechanisms underlying the effects...
The study randomized at the intact class-section level (two classes), satisfying the class-level RCT requirement.
Outcomes were assessed with a researcher-created posttest and rubric rather than a standardized exam-based assessment.
Outcomes were measured immediately after a three-session unit, which is far shorter than a full academic term after start.
The control condition is described, but baseline academic performance is not documented because the design is posttest-only and prior knowledge was not measured.
The trial was conducted within a single school and randomized between two classes, not between schools.
The paper indicates the study design, data collection, and analysis were conducted by the author(s) without an independent external evaluation team.
The study spans only three class sessions and therefore does not track outcomes for at least 75% of an academic year; also, T is not met so Y cannot be met.
Both groups used the same scheduled class time and the same Nearpod delivery platform, with no evidence of additional time or budget provided only to the intervention group.
No independent, peer-reviewed replication of this specific trial by other authors was found in the paper or via internet search.
Because E is not met and outcomes are measured only with a budgeting posttest (not standardized exams across core subjects), A is not met.
The paper reports no long-term follow-up and, because Y is not met, G cannot be met; no follow-up graduation-tracking papers by the same authors were found via internet search.
The paper provides no pre-registration link, registry ID, or registration date, and no registry record was found via search.
Financial literacy remains low among U.S. middle school students, while engagement with traditional instruction often declines. This class-randomized, posttest-only trial (two intact sections) compared Project-Based Learning and Standards-Based Learning in a budgeting unit delivered via Nearpod® across three 40-min class...
Randomization was at the individual student level within one cohort (not class- or school-level), and the study is not a one-to-one tutoring exception.
Outcomes relied on researcher-developed knowledge and OSCE-based performance instruments rather than widely recognized standardized exams.
The intervention lasted one week and the longest reported follow-up was one month, which is shorter than a full academic term.
The control group is clearly described (non-interactive app), with sample sizes and baseline demographics reported for comparison.
Participants were randomized as individual students in a single university program, not by school/site.
The authors’ team developed the intervention and ran the study; an independent statistician and blinding help internal validity but do not demonstrate an independent, third-party evaluation.
The reported follow-up is only up to one month, so the study does not approach 75% of an academic year.
Both groups received the same content and schedule; differences in instructor feedback and interaction are integral to the interactivity treatment being tested.
No independent replication of this specific i-BreathGuard vs n-BreathGuard RCT was found in available sources as of the ERCT check date.
Because criterion E is not met (no standardized exams), criterion A cannot be met; additionally, outcomes are specific clinical competence measures rather than all core subjects.
The study’s follow-up ends at one month post-intervention, far short of tracking participants through graduation, and criterion Y is not met.
The paper reports an IRCT registration and registration date, but the IRCT registry entry could not be accessed to verify registration timing relative to the actual start of data collection.
Competence in mechanical ventilation management is a critical component of nursing education. Technology-enhanced learning strategies may support skill acquisition. This randomized clinical trial evaluated the effectiveness of an interactive mobile application compared with a non-interactive version in improving nursing students’...
Participants were randomized as individuals, not as intact classes (and this is not a one-to-one tutoring exception).
Outcomes were measured with a study-constructed multiple-choice test and an OSCE checklist developed for this study, not a widely recognized standardized exam.
Outcomes were assessed immediately around a single 2-hour training session rather than at least one academic term after the intervention began.
The lecture-based control condition is clearly described and the paper reports group sizes and baseline characteristics.
The trial randomized individual trainees rather than randomizing schools (or other institution-level sites) to conditions.
The VR program was developed by the same hospital department that ran the study, and the paper does not document an external independent evaluation team (despite using blinded raters and an independent statistician for some tasks).
Outcomes were assessed immediately after a 2-hour intervention, so year-long tracking is absent; additionally, since criterion T is not met, criterion Y is automatically not met.
Both arms received equal instructional time (2 hours) and the same core curriculum content; the VR hardware/software is integral to the intended treatment contrast.
No independent replication by a non-overlapping author team was found for this specific study/intervention.
Because criterion E is not met, criterion A is automatically not met; additionally, the study does not assess standardized exams across all core subjects.
The paper does not track participants to graduation, and because criterion Y is not met, criterion G is automatically not met.
The study reports a registration date of February 5, 2025, which is after the stated study period (January 2024 to October 2024), so the protocol was not pre-registered before the study began.
Objective This study aimed to evaluate the effectiveness of VR-based training compared to traditional lecture-based training for medical trainees in managing MCIs, specifically focusing on road traffic accidents. The primary assessment was performed using an Objective Structured Clinical Examination (OSCE)...
Randomisation was conducted among individual students rather than by class (or school), so contamination across classmates remains plausible.
The main knowledge outcome (NKT) was researcher-prepared and aligned to the delivered nutrition education, rather than a widely recognised standardised exam.
The main between-group outcome comparison is at the 8-week post-test, which is shorter than one academic term.
The control group is clearly described, including the control condition, sample size, and baseline characteristics reported alongside outcomes.
The study was conducted in a single school and randomised individuals, not schools.
The intervention and analysis were conducted by the study team without clear third-party evaluation independent of the intervention designers.
The study does not measure outcomes for at least 75% of an academic year after intervention start, and term-duration (T) is not met.
Although the intervention adds instructional time and materials, these inputs are integral to the education programme being tested against a business-as-usual control.
No independent replication of this specific trial was identified in the paper or via post-publication searches.
Exam-based assessment (E) is not met, so all-subject exams (A) cannot be met; additionally, outcomes are nutrition/health measures rather than standardised exams across core school subjects.
Year duration (Y) is not met and no graduation tracking (or follow-up paper tracking the cohort to graduation) was identified.
The reported registration date is after the reported data collection period, so the protocol was not prospectively pre-registered.
Background: This study aimed to evaluate the effects of a dietitian-led, school-based nutrition education programme on primary school students’ nutrition knowledge, attitudes, behaviours, and anthropometric measurements. Methods: A randomised controlled, prospective design was conducted in a primary school in Mardin...
Students (not intact classes or schools) were randomly assigned, and the paper does not document class- or school-level randomization.
Outcomes were measured with self-report scales (self-efficacy and fear of negative evaluation), not standardized exam-based academic assessments.
Primary post-intervention outcomes were collected immediately after an eight-week intervention, which is shorter than a full academic term as defined by ERCT.
The control group condition and timing are described, and baseline measurements were collected for both groups.
The paper describes two institutes as study sites but does not document randomization of institutes/schools to conditions.
The paper does not document an independent external evaluation team for implementation, measurement, or analysis.
Outcomes were tracked only through an eight-week intervention and a four-week follow-up, far less than 75% of an academic year.
The paper explicitly states equivalent instructional time and access to materials for both groups, and the extra instructional structure appears integral to the intervention being tested.
No independent peer-reviewed replication of this specific RCT was found in searches, and the paper frames its findings as "first evidence."
Criterion E is not met (no standardized exam outcomes), therefore criterion A is also not met.
The study does not track students until graduation, and criterion G is also automatically not met because criterion Y is not met.
The paper does not mention a registry, registration ID, or a pre-registered protocol prior to data collection.
This randomized controlled trial examined the effectiveness of repeated reading intervention on reading self-efficacy and fear of negative evaluation (FNE) among Arabic-speaking preparatory students. Sixty second-grade students (ages 14-15) from two preparatory institutes in Gharbia Governorate, Egypt, were randomly assigned...
Randomization was performed at the individual participant level, not at the class (or stronger) level, and no tutoring exception applies.
Outcomes were measured using questionnaires and a self-developed knowledge quiz rather than standardized exam-based assessments.
The intervention and post-test occurred within days (data collection within about one month), far shorter than one academic term from intervention start to outcome measurement.
The paper documents control and experimental group sizes and provides baseline demographic characteristics and pre/post outcome descriptives by group.
The unit of randomization was individual participants from an online panel rather than schools (or other educational institutions) randomized as clusters.
The intervention was developed by the authors’ organization, and the paper does not document that the overall evaluation and analysis were led by an independent external team.
Outcomes were assessed within days after intervention, far below 75% of an academic year; additionally, since criterion T is not met, criterion Y is not met.
The intervention required substantial participant time while the control group was a waiting control with no comparable substitute activity, creating an unbalanced input condition.
No independent replication by a different research team in a different context was found or documented.
Criterion E is not met (no standardized exam-based outcomes), so criterion A cannot be met; additionally, outcomes do not cover all core school subjects.
No tracking through graduation is reported, no follow-up papers tracking this cohort to graduation were found, and criterion G cannot be met because criterion Y is not met.
The study reports OSF preregistration dated October 4, 2024, which precedes the stated data collection period beginning October 14, 2024.
Digital health literacy is an important asset to navigate the (digital) world and lead a healthier life. However, digital health literacy levels are insufficient in many segments of society, particularly among adolescents. In response to these deficits, policymakers have called...
The study randomized individual students, not entire classes or schools.
Assessments used researcher-developed quizzes and checklists, not standardized exams.
The intervention period was 6 weeks, shorter than a full academic term.
The control group's demographics, baseline characteristics, and treatment (routine FC) are documented.
Randomization was at the student level within a single university, not at the school level.
The same authors appear to have designed the intervention, conducted the study, and analyzed the data, with no mention of independent conduct.
The study duration, including data collection, was 11 weeks, which is less than a full academic year. Also, Criterion Y was not met.
The intervention group received gamified activities (extra quizzes, points, badges) which constitute additional resources/ engagement time compared to the control group's routine FC, and this difference was not explicitly tested as the treatment variable nor balanced.
The paper does not mention any independent replication of this specific study by another research team.
The study measured only nursing skills competency and related factors, not performance across all core academic subjects. Also, Criterion E was not met.
The study tracked students for 11 weeks, not until graduation. Also, Criterion Y was not met.
The study was prospectively registered on ClinicalTrials.gov before the intervention likely started based on the semester timing.
Background: Flipped learning excessively boosts the conceptual understanding of students through the reversed arrangement of pre-learning and in classroom learning events and challenges students to independently achieve learning objectives. Using a gamification method in flipped classrooms can help students stay...
The unit of randomization is individual participants (between-subjects), not intact classes or schools, and no tutoring exception applies.
Outcomes are measured using a researcher-designed quiz rather than a widely recognized standardized exam.
Outcomes are measured within a single short session (minutes to about an hour), not at least one academic term after the intervention begins.
The control condition and key baseline/balance characteristics are documented, including a balance table and clear control vs treatment descriptions.
Randomization is conducted among individual participants rather than at the school (or equivalent site/institution) level.
The study does not document independent third-party conduct of the evaluation; the authors are Anthropic-affiliated and describe internal review.
The study duration is about an hour rather than at least 75% of an academic year, and because T is not met, Y is necessarily not met.
The only clear resource difference is access to the AI assistant, which is the explicit treatment variable; otherwise tasks and time limits are comparable across groups.
No independent replication by other research teams was found, and the paper frames itself as an initial study that motivates future work.
Because E is not met (custom quiz), A is not met; additionally, the study assesses only Trio/library-specific skills rather than all core subjects.
The study does not track participants to graduation, and because Y is not met, G is necessarily not met.
The paper links to an OSF pre-registration and states it was done before running the experiment, but the registry entry’s date could not be verified here to confirm it predates data collection.
AI assistance produces significant productivity gains across professional domains, particularly for novice workers. Yet how this assistance affects the development of skills required to effectively supervise AI remains unclear. Novice workers who rely heavily on AI to complete unfamiliar tasks...
Random allocation is described at the individual participant level rather than by intact classes (or schools), so class-level randomization is not demonstrated.
The primary outcome assessments are internally assembled (course-derived written exam plus a study-developed questionnaire) rather than a widely recognized standardized exam.
Although an August 2024 start is stated, the paper does not provide a clear end date or duration showing outcomes were measured at least one academic term after the intervention began.
The control condition and demographics are described, but baseline academic performance (pre-intervention achievement) is not reported for the control group.
The study occurs within one institution and randomizes individuals, not whole schools/sites, so school-level randomization is not shown.
The authors who designed the study also carried out data collection and analysis, and no independent external evaluation team is documented.
The paper does not report an outcome measurement time point at least 75% of an academic year after intervention start, and criterion T is also not met.
The intervention explicitly tests a combined 3D+PBL+annotation teaching package where the additional resources are integral to the treatment definition, making a business-as-usual lecture control appropriate.
No independent replication by a different research team in a different context could be identified.
Because standardized exam-based assessment is not documented (criterion E not met), the study cannot meet the all-subject standardized exam requirement.
The study focuses on short-term outcomes and provides no graduation tracking, and criterion G also fails automatically because criterion Y is not met.
The paper reports ChiCTR registration with an ID and registration date that precedes the stated August 2024 study start.
Background: Traditional lecture-based learning (LBL) faces limitations in teaching complex spinal anatomy and surgical procedures. This study aimed to evaluate the efficacy of a novel Problem-Based Learning (PBL) model integrated with three-dimensional (3D) anatomy software and software-assisted annotation in spinal...
Randomization occurred at the individual student level (not by class or school), and the intervention is not one-to-one tutoring, so ERCT class- level randomization is not satisfied.
Outcomes were assessed using study-created and validated MCQs rather than a widely recognized standardized exam, so the exam-based assessment requirement is not met.
The primary post-test occurred one week after instruction, which is far shorter than an academic term, so the term-duration requirement is not met.
The control condition and baseline comparability are described (control received additional study time), with demographic and pre-test evidence of comparable baseline status, so the control group is adequately documented.
The study was conducted at a single institution with student-level randomization, not school/site-level randomization across multiple schools, so the school-level RCT requirement is not met.
The authors appear to have designed, conducted, and analyzed the study themselves without documenting an independent evaluation team, so independent conduct is not established.
Outcomes were measured within weeks (and T is not met), which is far short of 75% of an academic year, so the year-duration requirement is not met.
The intervention’s additional structure and faculty-moderated discussion are integral components of the educational method being tested against an equal-time study control, so the resource difference reflects the treatment rather than an unintended imbalance.
The paper does not report an independent replication, and an internet search did not identify a peer-reviewed independent replication of this specific trial as of the ERCT check date.
Because standardized exams were not used (E not met), the all-subject standardized exam requirement cannot be satisfied, and the assessments focus only on dental materials topics.
The study measures short-term outcomes only and does not track students through graduation (and Y is not met), so graduation tracking is not satisfied.
The paper reports a CTRI registration number and date, but registry timing relative to first enrollment could not be independently verified from the CTRI database in this review, so pre-registration cannot be confirmed.
Background: Dental materials education poses unique challenges due to the complex integration of scientific principles with clinical applications. Traditional teaching methods often fail to promote deep conceptual understanding. This study investigated whether the process of generating multiple-choice questions (MCQs) by...
Randomization is at the child level, but the intervention is delivered in small-group/individual formats with an experimenter, matching the ERCT tutoring-style exception.
The primary outcomes rely on researcher-created homonym measures rather than a widely recognized standardized exam.
Outcomes were measured about one week after completing a short (~2-week) intervention period, which is far shorter than an academic term.
The paper documents the control conditions and reports baseline comparisons between conditions, enabling interpretation of control-group comparability.
Schools were not randomized to conditions; participants were assigned within schools.
The paper does not document an independent, third-party evaluation team; assignment and delivery are described as being done by the study researchers.
Outcomes were measured within weeks rather than at least 75% of an academic year after the intervention began; additionally, because T is not met, Y cannot be met.
Study 1 uses a time-matched active control, but Study 2 provides greater per-child adult support in the intervention (individual delivery) than the control (pairs), creating an input imbalance not framed as the treatment variable.
No independent replication by a different research team was identified; the paper reports two trials by the same author team.
Because criterion E is not met, criterion A cannot be met; the study also does not assess all core subjects via standardized exams.
The paper reports only short-term post-tests and does not track students to graduation; additionally, because Y is not met, G cannot be met.
Study 2 is stated to be pre-registered on OSF, but the paper does not report a registration date that can be verified as occurring before data collection.
Background: Many words have multiple meanings, which present challenges to learning, yet research has yet to identify effective interventions for homonyms. Lexical inference may be a promising strategy. Aim: To evaluate a brief, novel lexical inference intervention for homonyms. Samples:...
Randomization was at the individual participant level rather than at the class (or stronger) level, so class-level isolation is not ensured.
Outcomes were measured using self-report psychological scales, not widely recognized standardized exams.
The intervention lasted six weeks and outcomes were measured immediately after it ended, which is shorter than one academic term.
The paper provides clear control-group description and baseline comparability information (sample sizes, demographics, and baseline equivalence tests).
Randomization assigned individual participants, not whole schools, so the study is not a school-level RCT.
While randomization was done by an assistant and assessors were blinded, the paper does not document an external independent evaluation team separate from the intervention designers/authors.
The study ran for six weeks with immediate post-testing, which is far below 75% of an academic year, and it also fails Criterion T.
Both groups completed writing tasks on the same schedule and in the same setting with similar supervision, so time and implementation resources appear balanced.
No independent published replication of this specific 2026 PPEW RCT was identified, and the paper itself does not report a replication.
Because the study does not use standardized exam-based outcomes (Criterion E not met), it cannot satisfy the all-subject exams requirement.
The study explicitly reports no long-term follow-up, and because it also fails the year-duration requirement, it cannot meet graduation tracking.
The paper provides ethics approval details but no evidence of a publicly pre-registered protocol (registry name/ID and timing).
This study aimed to examine the interventional effects of positive psychology expressive writing (PPEW) on adolescents’ time attitudes and mental health. A total of 285 adolescents from Northwest China (M = 14.13, SD = 1.075; 53.3% female) were randomly assigned...
Randomization occurred at the individual participant level within sessions (not class- or school-level), and the tutoring exception does not apply.
Outcomes are cheating behavior measured via an experimental logic-puzzle paradigm and self-report, not standardized exam- based educational achievement assessments.
The intervention and main outcome were completed within a single short experimental session (minutes), not tracked for at least one academic term after the intervention began.
The control condition is described in detail (content, structure, matching, and group sizes with baseline comparisons), enabling clear treatment-control comparison.
Participants were recruited from a single university and randomized individually; schools/sites were not randomized.
The trial sessions were administered by the study team (a single primary researcher), without documented independent conduct by a third-party evaluator.
Outcomes were measured immediately within a single session (and all cohorts completed within 2 days), far short of 75% of an academic year; also, since T is not met, Y is not met.
The control condition is an active, structurally parallel program matched to the intervention in length, format, and engagement, indicating balanced time and activity inputs across groups.
No independent replication of this specific RCT (growth mindset intervention reducing cheating) by other authors was identified, and the paper itself frames the study as providing first causal evidence.
Because criterion E is not met (no standardized exam-based achievement assessment), criterion A is also not met; the study does not assess all core subjects via standardized exams.
The study provides only immediate/short-term measurement and no tracking to graduation; additionally, because Y is not met, G is not met.
No protocol pre-registration registry/ID (and no registration date showing registration before data collection) is reported.
Academic dishonesty remains a persistent challenge in higher education, highlighting the need for scalable and cost-effective interventions that target internal motivation. Building on mindset theory, the present research tests the impact of a brief growth- mindset intervention on exam cheating...
Randomization was at the individual student level rather than assigning whole classes (and this was not a one-to-one tutoring intervention).
Outcomes were measured with self-report empowerment and confidence scales rather than a widely recognized standardized exam-based assessment.
The intervention lasted 8 weeks and outcomes were measured immediately after training, which is shorter than a full academic term (about 3–4 months) from start to measurement.
The control condition is clearly described (routine curriculum without simulation) and baseline characteristics and group sizes are documented.
The paper randomizes students within a single university department rather than randomizing at the school/site level.
Although randomization and assessment roles include "independent" personnel, the intervention was developed and data collection led by study authors, so the trial is not fully independent of the intervention designers.
Year-duration tracking is not present because outcomes were measured immediately after an 8-week intervention (and ERCT Y also cannot be met when ERCT T is not met).
The intervention’s added materials and personnel (trained standardized patients, scripted scenarios, debriefing and fidelity procedures) are integral to the treatment being tested, and the paper documents a structured routine curriculum control rather than leaving control students without educational engagement.
The paper does not report an independent replication and no independent replication evidence is provided within the paper text.
All-subject standardized exams were not used, and because ERCT E is not met, ERCT A is automatically not met.
The study did not track participants to graduation, and because ERCT Y is not met, ERCT G is automatically not met.
No pre-registration is provided; the paper explicitly states "Clinical trial number Not applicable."
Background Patient aggression is a persistent challenge in mental health settings, and undergraduate preparation in de-escalation remains variable. This study evaluated whether a standardized-patient (SP) de-escalation simulation improves psychological empowerment and confidence in coping with aggression among psychiatric nursing students....
Randomization was at the individual student level (not class or school), so class-level randomization to reduce contamination is not demonstrated.
Outcomes were measured using locally developed Test A/Test B examinations rather than a widely recognized standardized exam.
The intervention and measurement occurred within a 2-week rotation with immediate post-teaching testing, which is shorter than a full academic term.
The control condition and baseline characteristics (including group sizes and demographics) are described with pre-test scores.
Schools (or clerkship sites) were not randomized; students were individually assigned, so school-level randomization is absent.
The research team developed the animations and also collected and analyzed the data, with no clearly independent evaluator.
The intervention and outcome measurement occurred over a 2-week rotation with immediate post-testing, far below 75% of an academic year.
Instructional objectives and in-session teaching time appear matched across groups; the only extra resource (post-session online access) occurred after the immediate post-test, making any resource imbalance negligible for the main outcome.
No independent replication study is reported or identifiable as a direct replication of this specific RCT.
Because standardized exams are not used (criterion E not met), the all-subject standardized exam requirement cannot be met.
The study does not track participants to graduation and explicitly notes the lack of long-term follow-up; additionally, ERCT rules require criterion Y to be met for criterion G to be met.
The trial registration is explicitly described as retrospective and dated after enrollment, so it was not pre-registered before data collection began.
Objectives Teaching renal pathology, characterized by complex spatial relationships and dynamic pathological processes, poses significant challenges. Traditional methods such as static slides often fail to convey these concepts effectively. This study evaluated the efficacy of custom-developed renal pathology three-dimensional (3D)...
Randomisation was conducted at the individual student level (not by intact classes or schools) and no tutoring exception applies.
Learning was measured with a curriculum-based multiple- choice test developed by the authors, not a widely recognised standardized exam.
Outcomes were measured again at a three-month follow-up, which is approximately one academic term after the intervention began.
The paper clearly defines the comparison groups (TT vs. JS), reports group sizes, and reports baseline and follow- up outcome data by group.
The randomisation occurred among individual students, not among schools (or equivalent institutions/sites).
The authors developed key study materials and conducted the randomisation on site, with no quoted evidence of an independent external evaluation team.
The study’s follow-up lasted three months, which is far short of 75% of an academic year.
The JS condition received substantially more total time and broader learning resources than TT, and this resource difference was not framed as the treatment variable nor matched in the control condition.
No evidence is provided that an independent research team has replicated this specific RCT in another context.
Criterion E is not met (the outcome test is not a standardized exam), therefore All-subject Exams cannot be met.
The study only follows students for three months and does not track outcomes through graduation; additionally, Criterion Y is not met, so G cannot be met.
The paper explicitly states that trial registration is not applicable, providing no pre-registered protocol or registry identifier.
Traditional teaching (TT) is lecturer-centred, while student-centred teaching, including Jigsaw (JS), fosters student interaction. However, the results of research on the effectiveness of JS on learning outcomes is inconsistent. This randomised controlled trial compared the effect of JS and TT...
Students (not whole classes) were randomized, so the design is not class-level (or stronger) randomization as required by ERCT C.
Outcomes were measured using a faculty-designed 20-item MCQ rather than a widely recognized standardized exam.
The post-test was administered immediately after the sessions, which is far less than one academic term after intervention start.
The comparison (cadaveric dissection) group and baseline characteristics are described with group sizes and baseline balance checks, meeting the control documentation requirement.
Randomization occurred among individual students, not among schools (or equivalent institutions/sites).
The paper describes blinding and an independent administrator for allocation, but it does not provide evidence of a truly independent external evaluation team separate from those designing/teaching and building the study instruments.
Outcomes were measured immediately after teaching sessions, far short of 75% of an academic year after intervention start.
Instructional time and core teaching inputs were standardized across groups, and the control group received a comparable active alternative (cadaveric dissection) rather than a lower-resource condition.
No independent replication of this specific study was identified in the paper, and a web search did not locate an external replication as of the ERCT check date.
Because criterion E is not met (no standardized external exam), the all-subject standardized exam requirement is automatically not met.
The study does not track participants until graduation, and it also fails the year-duration prerequisite (Y).
The paper provides IRB approval information but does not provide a public pre-registration record (registry/ID) and date demonstrating registration before data collection.
Purposes Anatomy is fundamental in medical education, yet cadaveric dissection faces challenges including limited specimens, high costs, and chemical hazards. Interactive anatomy tables such as the Pirogov system offer innovative alternatives, but evidence from Southeast Asia is limited. Methods In...
Randomization was conducted at the small-group level (and not fully rigorous), not at the class (or higher) level.
The primary educational outcome (CVS comprehension) was measured using a research instrument rather than a widely recognized standardized exam.
The main outcome measurement occurred within weeks (eight-week period and a post-test two weeks later), not at least one academic term after the intervention began.
The control condition is clearly described, and baseline/descriptive information by condition is reported.
The study took place in multiple schools, but randomization was not conducted at the school level.
The paper does not clearly document that an independent third-party evaluation team conducted the study and analyses.
Outcome measurement occurred far short of 75% of an academic year, and criterion T is not met.
The intervention and control conditions appear time- and materials-structured similarly, differing mainly in contextual embedding of otherwise identical tasks.
No independent replication of this specific study was found in the paper or via internet searching as of the ERCT check date.
Criterion E is not met, and the study does not assess all core subjects using standardized exams.
The study does not track students until graduation, and criterion Y is not met.
No pre-registration (registry, ID/link, and pre-data-collection timing) is reported in the paper, and no corresponding entry was found via internet searching.
Context personalization is an instructional approach aimed at enhancing students’ engagement and cognitive processing by embedding learning content in familiar contexts. Numerous studies explore the benefits of personalized tasks for learning, but few empirically examine cognitive mechanisms underlying the effects...
The study randomized at the student level rather than at the class or school level.
Outcomes were measured with app-administered tests closely aligned to the treatment exercises rather than a standardized external exam.
The intervention and outcome window was six weeks, which is shorter than a full academic term.
The control group is clearly described, including its notification condition and baseline characteristics.
Randomization occurred at the student level rather than the school level.
The paper does not provide a clear statement that the evaluation was conducted by an independent third party separate from the intervention team.
The intervention and measurement period was six weeks, not an academic year.
The additional input (messages/notifications) is the treatment being tested, so a no-message control is the appropriate comparison.
No independent replication study by other authors was found for this specific intervention.
The study does not measure standardized outcomes across all core subjects, and criterion E is not met.
The study does not track participants to graduation and also fails the year-duration prerequisite (criterion Y).
The paper provides ethics approval and data registration statements, but no evidence of a pre-registered study protocol.
We examine whether highlighting streaks - instances of repeated and consecutive behavior when completing learning tasks - encourages 4th to 6th grade students in Peru to increase their use of an online math platform and improve learning. 60,000 students were...
The study randomized at the student level rather than the class or school level.
The study used a custom endline test administered through the app rather than a widely recognized standardized exam.
The intervention duration was six weeks, which is shorter than the required full academic term.
The control group is clearly defined as receiving no messages, and their demographics and baseline performance are well-documented.
The study utilized student-level randomization, not school-level randomization.
The study appears to be conducted by the authors who designed the intervention without a stated independent third-party evaluator.
The intervention lasted six weeks, failing the one-year duration requirement.
The intervention specifically tested the impact of "nudges" (messages) as the treatment variable, so the lack of messages for the control group is by design and balanced.
The paper states this is the first experimental study of its kind in this context, and no independent replication is cited.
The study only assessed math achievement and did not measure outcomes in other main subjects like reading or science.
The study followed up immediately after the six-week intervention; no graduation tracking was conducted.
There is no evidence of a pre-registered protocol with hypotheses and analysis plans in a public registry before data collection.
We examine whether highlighting streaks—instances of repeated and consecutive behavior when completing learning tasks—encourages 4th to 6th grade students in Peru to increase their use of an online math platform and improve learning. 60,000 students were randomly assigned to receive...
The paper is a within-subject experiment with no random assignment, so it is not a class-level RCT.
Outcomes are based on observational coding, not on standardized, exam-based assessments.
The study compares two 10-minute sessions separated by 2 weeks, far shorter than a term-long follow-up.
The sample and baseline (uninformed) condition are documented with detailed participant characteristics.
There is no school-level randomization; individual dyads were recruited and all experienced both contexts.
The study does not report being run by an external evaluation organization independent of the authors.
The study spans 2 weeks between sessions, not a full academic year.
Time and materials are essentially balanced across contexts; the main difference is the informational prime and a different toy set to reduce familiarity effects.
No independent replication study was identified for this 2025 paper.
The study does not use standardized exams in any subject, and it does not assess outcomes across all core subjects.
The study does not track participants until graduation and cannot meet G because it also fails the Year Duration criterion.
The authors state the analyses were not preregistered.
Home math interventions often incorporate informational priming-explicit prompts emphasizing parental math input. While effective in increasing math talk, its impact on child outcome is mixed. This study examined how informational priming shapes the content and dynamic of math interactions. In...
Randomization was at the individual participant level rather than at the class (or school) level.
Outcomes were scored via teacher ratings and an internal AI judge, not via a standardized externally administered exam.
Outcomes were measured within short sessions rather than at least one academic term after the intervention began.
The control condition and sample are clearly documented, including what the control group could and could not do.
The study did not randomize at the school (site) level.
The study was proposed, designed, executed, and analyzed by the author team, not by an independent evaluator.
The study does not measure outcomes over a full academic year, and criterion T is not met.
Time-on-task is held constant across groups and the tested difference is tool access, not extra instructional time or budget.
No independent replication of this specific study was identified at the time of this ERCT check.
The study does not use standardized exams across all main subjects, and criterion E is not met.
The study does not track participants to graduation and criterion Y is not met.
The paper reports IRB approval but no public preregistration record was identified.
With today's wide adoption of LLM products like ChatGPT from OpenAI, humans and businesses engage and use LLMs on a daily basis. Like any other tool, it carries its own set of advantages and limitations. This study focuses on finding...
The unit of randomization was the intact class section, which satisfies the class-level RCT requirement.
The outcome measure was a researcher-created posttest rather than a standardized, widely recognized exam.
Outcomes were measured immediately after a three-session unit, which is shorter than one academic term from start to measurement.
Although the control activities are described, the posttest-only design provides no baseline performance data documenting control group comparability.
The trial randomized two classes within one school rather than randomizing across schools.
The paper does not document an independent evaluation team and states the author conducted all study components themselves.
Outcomes were measured after a three-session unit and the authors state long-term outcomes were not measured, failing the 75%-of- year tracking requirement.
Both conditions used the same scheduled class time and the same delivery platform, with no evidence of extra real resources given only to the intervention group.
No independent replication by other authors was identified, and the paper frames this study type as novel in this domain.
The study does not use standardized exams (Criterion E not met), so it cannot satisfy the all-subject standardized exam requirement.
The study does not include graduation tracking and also fails the year-duration prerequisite (Criterion Y).
The paper provides no protocol pre-registration ID or link, and no registry record could be verified from the information provided.
Financial literacy remains low among U.S. middle school students, while engagement with traditional instruction often declines. This class-randomized, posttest-only trial (two intact sections) compared Project-Based Learning and Standards-Based Learning in a budgeting unit delivered via Nearpod® across three 40-min class...
Randomisation was conducted at the individual student level within a single institute, not at the class or school level, and no tutoring exception applies.
The study used a researcher-designed oral task and a generic rubric, not a widely recognised standardised exam.
Tracking from intervention start to the posttest spanned only about nine weeks, well short of a full academic term.
The control group's size, baseline scores, and business -as-usual instructional activities are clearly documented.
Randomisation occurred among individual students at a single institute, not among schools.
The same research team designed, delivered, and assessed the intervention with no independent evaluators mentioned.
Since criterion T is not met and the study duration was only about nine weeks, criterion Y cannot be met.
Both treatment and control groups received the same lessons and class time; no additional resources were given to the intervention groups, so the balance requirement is trivially satisfied.
No evidence of an independent replication of this specific study by a different research team was found in the paper or via external search.
Only speaking fluency was assessed, not all main subjects, and criterion E (a prerequisite) was not met.
Tracking stopped one week after the intervention ended, with no follow-up toward graduation, and criterion Y was not met.
No pre-registration of the study protocol, registry link, or registration date is mentioned anywhere in the paper.
The current study investigated the impact of using two cooperative learning strategies on the development of oral English language fluency among Iranian intermediate EFL learners. First, 72 learners studying at a private English language institute were randomly assigned to two...
The paper calls its own design "quasi-experimental" and only describes random assignment of students to classes by the university, not random assignment of the four classes to treatment conditions.
Outcome measures were custom checklist and multiple-choice tests built around 43 documentary-specific target words, not standardised exams.
The treatment was a single session, with outcomes measured immediately and one week later, far short of one academic term.
The control group's size, baseline vocabulary level, and lack of treatment are clearly documented.
Randomisation occurred among four classes at a single university, not among multiple schools.
The same researchers designed the materials, ran the treatments, and analysed the data, with no independent evaluator mentioned.
The study tracked outcomes for only about two weeks, and criterion T is also not met, so Y automatically fails.
Exposure to the input mode is itself the explicit treatment variable being tested, so the no-treatment control condition is met by design.
No independent replication of this specific study was found; the authors state it is the first study of its kind, and a later similarly designed study by a different team does not present itself as a replication.
Only vocabulary knowledge was assessed, and criterion E is not met, so A automatically fails.
Tracking ended one week after treatment, criterion Y is not met, and no follow-up publications tracking these participants toward graduation were found.
No pre-registration of the study protocol is mentioned anywhere in the paper, and no internet search identified a registry entry for this study.
This study used a pretest-posttest-delayed posttest design at one-week intervals to determine the extent to which written, audio, and audiovisual L2 input contributed to incidental vocabulary learning. Seventy-six university students learning EFL in China were randomly assigned to four groups....
Individual students, not intact classes or schools, were the true unit of random assignment to conditions, with delivery groups formed only afterward.
The study used custom word/sentence-reading and picture-description tasks with acoustic and listener ratings devised solely for this study, not a standardized exam.
Outcomes were measured only about two weeks after a 4-hour, two-week-long intervention, far short of a full academic term.
Detailed demographic data (Table 1) and a clear description of the control group's comparable instructional conditions are provided.
Participants were individually recruited adult volunteers randomized as individuals, not schools or intact institutional units.
The authors designed the intervention and personally supervised instruction, conducted all testing, and performed all data analysis with no independent evaluator.
The tracked duration (about one month) is far shorter than the required academic-year threshold, and the term-duration criterion it depends on is also not met.
All three groups received the identical 4-hour instructional allotment; only the pronunciation-focused content/feedback embedded within that time differed, so no extra resources were provided to the intervention groups.
The authors describe the study as unprecedented and call for future replication; a later adaptation of the design was conducted by the same original authors rather than an independent research team.
Since criterion E is not met and only a single pronunciation feature was assessed, the all-subject exam requirement is not satisfied.
No follow-up beyond a two-week-delayed posttest was conducted, its prerequisite criterion Y is not met, and no later publication tracks this cohort further.
The paper contains no mention of a pre-registered protocol, registry, or date.
Sixty-five Japanese learners of English participated in the current study, which investigated the acquisitional value of form-focused instruction (FFI) with and without corrective feedback (CF) on learners' pronunciation development. All students received a 4-hr FFI treatment designed to encourage them...
The study is explicitly a quasi-experimental design using two pre-existing intact classes, not a true randomised controlled trial, so the class-level RCT criterion is not satisfied.
The vocabulary test used was a modified version of the VKS built around a custom, study-specific list of 40 target words, not a widely recognised standardised exam.
Outcomes were tracked only to Week 12 of the research timeline (about 11-12 weeks from intervention start), which falls short of a full academic term (approximately 3-4 months).
The control group's size, standard-instruction condition, and baseline demographic/academic characteristics are documented and statistically compared with the experimental group.
Both intact classes were drawn from a single technical college, so randomisation (such as it was) occurred at the class level, not across multiple schools.
The researchers themselves (including a lecturer at the participating institution) designed and directly supervised the intervention and testing procedures, with no independent third-party evaluator involved.
Since criterion T (term duration) is not met, criterion Y is automatically not met; in any case the roughly 12-week tracking period is far short of 75% of an academic year.
The intervention did not add extra class time or budget; both groups followed the same standard ESP syllabus during the same period, differing only in teaching method (Kahoot! versus traditional instruction).
No independent replication of this newly published study exists; a citation-database check confirms zero citing works to date.
Since criterion E (Exam-based Assessment) is not met, criterion A is automatically not met; in addition, only ESP vocabulary was assessed, not all main subjects.
Since criterion Y (Year Duration) is not met, criterion G is automatically not met; no follow-up publication tracking this cohort was found either.
No pre-registration of the study protocol is mentioned anywhere in the paper, and none was found via external registry/metadata checks.
Research over the past decade has identified gamification as an emerging approach in language teaching, particularly in updating vocabulary instruction through game-based methods. This study investigates the efficacy of Kahoot! in improving learners' recall and retention of ESP vocabulary within...
Individual students, not intact classes, were randomly assigned to the reading, listening, reading-while- listening, and control conditions, and no tutoring exception applies.
The two dependent measures (Test A matching/recall, Test B meaning recall) were custom-built by the authors specifically for this study's 17 target collocations, not standardised exams.
Outcomes were measured only about six to seven weeks after the intervention began, far short of a full academic term.
The control group's size, composition, baseline scores, and exact activities (no reading/listening, only the same tests) are clearly documented.
Randomisation occurred among individual students within a single college cohort, not among schools.
The same two authors designed the intervention, built the target-item list and both dependent measures, and there is no statement of an independent third-party team conducting the study.
Since Term Duration (T) is not met, Year Duration cannot be met either; the whole study spanned only about 8 weeks.
The extra reading/listening exposure is itself the explicit treatment variable under investigation, so the no-treatment control lacking that exposure is a valid "business as usual" baseline rather than an unbalanced confound.
No independent replication of this specific study (same design, materials, and research questions, conducted by a different team in a different context) was found in the paper or in a subsequent internet-based literature search.
The study measured knowledge of only 17 collocations drawn from one graded reader, not performance across all main subjects, and criterion E (a prerequisite) is not met.
Since Year Duration (Y) is not met, Graduation Tracking cannot be met; the study's longest follow-up was a delayed posttest four weeks after the intervention ended, and no follow-up publication tracking this cohort was found.
There is no mention anywhere in the paper of a pre-registered protocol, registry platform, or registration date, and no internet search evidence of a registration was found.
There has been little research investigating how mode of input affects incidental vocabulary learning, and no study examining how it affects the learning of multiword items. The aim of this study was to investigate incidental learning of L2 collocations in...
Randomisation was performed at the individual student level within a single secondary school, not at the class or school level, and no tutoring exception applies.
The study used a custom-designed Reading Comprehension Skills Test and Attitude Scale created and validated by the researcher, not a standardised, widely recognised exam.
Outcomes were measured immediately after a short, twelve-session intervention lasting only a few weeks, far short of a full academic term.
The control group's size, baseline test scores, and the "business as usual" traditional treatment it received are clearly documented and compared with the experimental group.
Randomisation occurred at the student level within a single secondary school, not among multiple schools.
The study was designed, conducted, and analysed solely by the same researcher, with no evidence of an independent or third-party evaluation team.
Since the term-duration criterion (T) is not met and the intervention lasted only twelve short sessions, the stronger year-duration requirement is also not met.
Both groups received their normal amount of classroom reading instruction; only the teaching strategy (Think-Aloud versus traditional skimming/scanning) differed, so no additional time or budget was given exclusively to the intervention group.
No independent replication of this specific study by another research team was found in the paper or via internet search; other cited Think-Aloud studies target different grade levels and contexts.
Only reading comprehension (and an attitude scale) were assessed; no other core subjects were measured, and criterion E is also not met.
There is no follow-up beyond the immediate post-test, no subsequent tracking papers by the author were found via internet search, and since criterion Y is not met, criterion G cannot be met either.
No pre-registration of the study protocol, hypotheses, or analysis plan is mentioned anywhere in the paper, and no registry record was found via internet search.
The current study's objective examines the effectiveness of using a Think-Aloud strategy in improving Saudi EFL learners' reading comprehension and attitudes towards learning. A quantitative study with a quasi-experimental design was implemented through applying two different instruments: Reading Comprehension Skills...
Randomisation was carried out at the individual student level (60 volunteers split in half), not at the class or school level, and no tutoring exception applies since this is regular classroom reading instruction.
The reading comprehension test was a custom instrument created for the study, and the authors explicitly state no standardised assessment tool was used.
Outcomes were measured about eight to nine weeks after the intervention began, which falls short of a full academic term.
The control group's size, baseline (pre-test) performance, and the nature of the instruction it received are clearly described in the text and accompanying tables.
Randomisation occurred among individual student volunteers at what appears to be a single site, not among schools.
There is no statement that the study was conducted by an independent, third-party evaluation team separate from the researchers who designed and delivered the intervention.
Since criterion T (Term Duration) is not met, the stronger Year Duration criterion cannot be met either.
Both groups received instruction over the same schedule and the same reading passages; the intervention differed in teaching method, not in additional time or budget, so no imbalance in resources arises.
No independent replication of this specific study by a different research team is reported or was found.
Only reading comprehension was assessed, and since criterion E (standardised exam) is not met, this stronger criterion cannot be met either.
There is no follow-up tracking beyond the ten-week study, and since criterion Y is not met, this criterion cannot be met either.
There is no mention anywhere in the paper of a pre-registered protocol, registry platform, or registration date.
This experimental study investigated ESL Reading Comprehension with Task-Based Language Instruction: Emphasizing the Identification of the Main Idea district RYK, Pakistan. Using a pre-test and post-test control group design, sixty ESL students aged 14-16 were randomly assigned to experimental (TBLT-based...
Randomisation in the Phase II training study was conducted at the individual student level within each school, not at the class or school level, and the intervention was group instruction rather than one-to-one tutoring, so no exception applies.
Outcomes were measured with custom, researcher-designed listening and speaking instruments rather than widely recognised standardised exams.
The entire intervention and outcome measurement spanned only about two weeks, far short of the one-term minimum.
The control group's size, composition, and the specific business-as-usual activities it received instead of strategy training are clearly documented.
Randomisation occurred at the individual student level within each school rather than between schools.
The same research team that designed the strategy taxonomy and training curriculum also delivered the intervention and analysed the results, with no independent third-party conduct described.
Because criterion T (Term Duration) is not met, criterion Y is automatically not met; independently, the actual duration was only about two weeks.
Instructional time was explicitly equalised across the metacognitive, cognitive, and control groups, with the manipulated variable being the strategy instruction itself.
No evidence was found of this specific training experiment having been independently replicated by a different research team in a peer-reviewed outlet.
Because criterion E (Exam-based Assessment) is not met, criterion A is automatically not met; the study also assessed only English listening and speaking, not other subjects.
Because criterion Y (Year Duration) is not met, criterion G is automatically not met; no follow-up beyond the immediate post-test is described, and no subsequent tracking papers were located.
No pre-registration of the study's hypotheses, methods, or analysis plans is mentioned anywhere in the paper, and none would be expected given the study predates modern trial registries.
Recent research on cognition has indicated the importance of learning strategies in gaining command over second language skills. Despite these recent advancements, important research questions related to learning strategies remain to be answered. These questions concern 1) the range and...
Randomisation occurred at the individual student level within one TEFL cohort, not at the class or school level, and no tutoring exception applies.
The study used a custom-selected, self-report vocabulary scale rather than a standardised exam, so this criterion is not met.
The five-week intervention plus three-week delayed follow-up totals roughly eight weeks, short of the minimum one-term duration required.
The control group's size, baseline scores, and treatment are documented and confirmed comparable to the experimental groups at baseline.
Randomisation was at the individual student level within one institution, not at the school level.
The same author team designed, delivered, and analysed the study with no independent evaluators involved.
Since Term Duration (T) is not met and the total tracked period is only about eight weeks, Year Duration is not met.
All groups received identical content and schedule; the only extra element (a short app tutorial) is negligible and tied to the medium being tested.
No independent replication of this specific study is reported or evident; related prior studies and 2024-2025 citing papers are not replications of this particular experiment.
Since criterion E is not met and only vocabulary knowledge (a single subject) was assessed, this criterion is not met.
Tracking ended shortly after the intervention with no follow-up toward graduation, Y is not met, and no later graduation-tracking publication was found.
No pre-registration of the study protocol is mentioned anywhere in the paper, and none was found in external registries.
This study systematically evaluates the comparative efficacy of digital flashcards on mobile and computer platforms versus traditional paper-based flashcards in augmenting academic vocabulary knowledge development among Iranian undergraduate university students. A randomized controlled trial was conducted with 112 subjects, allocated...
Randomisation was carried out at the individual student level within a single writing course, not at the class or school level, and no personal-tutoring exception applies.
Writing quality and reflection depth were scored by the researchers' own trained raters using researcher-adapted rubrics, not a widely recognised standardised exam.
The intervention spanned only a short, explicitly acknowledged "short intervention period" covering a single three-stage writing cycle, far below one academic term.
The comparison (individual) group's size, structure, and baseline (prewriting) performance are explicitly documented and shown to be statistically comparable to the collaborative group.
The study was conducted with one cohort in one writing course at a single university, with no school-level randomisation.
There is no indication of an independent, third-party team conducting or overseeing the study; the same institution's researchers appear to have designed, run, and evaluated the intervention.
Because criterion T (Term Duration) is not met, and the study is explicitly a short, single writing-cycle intervention, the year-long duration requirement is not met.
Both arms received identical structured-reflection scaffolding at each writing stage; the only manipulated variable was the collaborative-versus- individual structure itself, which is the intended treatment contrast, so no extra time or budget was given to either arm alone.
No independent replication of this study is reported or plausible, given its very recent publication date, and a dedicated internet search found no such replication.
Because criterion E is not met, criterion A cannot be met; additionally, only ESL writing performance was assessed, not all core subjects.
Because criterion Y is not met, criterion G cannot be met; the study also explicitly stops at the end of the short intervention with no further follow-up, and no subsequent tracking papers by these authors were found.
The paper describes institutional ethical approval but makes no mention of pre-registering the study protocol, hypotheses, or analysis plan before data collection, and no registry entry was found via internet search.
This study investigates how structured reflection embedded within process-oriented writing improves ESL learners' writing quality and metacognitive engagement in both collaborative and individual contexts. Drawing upon Murray's Process Writing Model and Moon's Reflective Thinking Model, a mixed-methods study involving 60...
Randomisation was carried out at the individual student level using a stratified block design, not at the class or school level, and the intervention was a group course rather than one-to-one tutoring.
Outcomes were measured with self-report psychological scales (engagement, self-efficacy, emotions), not with a standardised, widely recognised exam-based assessment.
The intervention lasted only 12 weeks and outcomes were measured immediately at the end of the intervention, so there is no documented interval of at least one full academic term from intervention start to measurement.
The control (NEAI) group's demographics, size, and the content it received are clearly documented, including baseline and post-test descriptive statistics.
Randomisation occurred at the individual student level, not at the school (or institution) level.
The same author who conceived and designed the intervention also performed the experiment and analyzed the data, with no independent external evaluation team running or analyzing the trial.
Since criterion T (Term Duration) is not met and the total study duration (12 weeks) falls far short of an academic year, criterion Y is not met.
Both the AI and NEAI groups received an identical number and length of sessions and the same course content, differing only in whether the Grammarly tool (the treatment variable) was included.
No independent replication of this specific trial by a different research team was found in the paper or through internet searches of citing literature.
Since criterion E (Exam-based Assessment) is not met and only affective/engagement constructs were measured rather than academic performance across subjects, criterion A is not met.
Criterion Y (Year Duration) is not met, and no follow-up publication tracking this cohort toward graduation was found; the study measured outcomes only immediately after the 12-week intervention ended.
The paper mentions IRB ethical review but provides no registry name, ID, or date confirming pre-registration before data collection began, and no internet search surfaced a matching trial registration.
A major challenge in educational technology integration is to engage students with different affective characteristics. Also, how technology shapes attitude and learning behavior is still lacking. Findings from educational psychology and learning sciences have gained less traction in research. The...
Random assignment was at the individual student level rather than at the class (or school) level.
Outcomes were measured with self-report items and diaries rather than standardized exam-based academic assessments.
Outcomes were measured over a seven-day period with a one-week follow-up, which is much shorter than one academic term.
The PTS control condition is described, group sizes are reported, and baseline comparability is discussed.
The unit of randomization is individual students, not schools/sites.
The study procedures describe researcher-led implementation and fidelity checks without stating an independent external evaluator.
The study duration is far shorter than 75% of an academic year, and ERCT rules also require Y to be false when T is false.
The control condition is described as time- and structure-matched, and the additional MCII fidelity feedback appears limited and integral to delivering MCII rather than a large resource imbalance.
No independent replication of this specific RCT was found in the paper or via external searching.
The study does not use standardized exams and therefore cannot meet the all-subject standardized exam requirement (and A fails when E fails).
The study reports only a one-week follow-up and no graduation tracking; additionally, G must be false when Y is false.
The paper explicitly states the study was not preregistered.
Academic procrastination is a pervasive challenge in higher education, particularly within low-structure, digitally mediated learning environments. This study investigates the efficacy of Mental Contrasting with Implementation Intentions (MCII) as an intervention to reduce procrastination among undergraduate students in a private...
Randomisation was carried out at the individual-student level within a single speaking course cohort, not at the class or school level, and no tutoring exception applies.
A researcher-designed speaking test only "adapted from" IELTS descriptors was used, not the standardised IELTS exam itself.
No clear intervention start date or measurement date is given, so the interval to at least one academic term cannot be verified from the text.
The control condition's activities and group size are described, though group-specific baseline demographics are not reported.
Randomisation occurred at the individual-student level within one department, not at the school level.
The same research team designed and delivered the intervention and the lead researcher was directly involved in scoring outcomes; no independent external evaluator is described.
Criterion T is not met, so criterion Y is automatically not met; no evidence of year-long tracking is provided in any case.
Both groups followed the same weekly course structure; the specific tool (YouGlish vs. traditional materials) is the treatment variable being compared, with no evidence of extra time or budget for the experimental group.
The authors present this as the first study of YouGlish in the Kuwaiti EFL context; a targeted internet search (Semantic Scholar, general web) found zero citing works and no independent replication of this specific study.
Criterion E is not met, so criterion A is automatically not met; only a single skill domain (speaking) was assessed in any case.
Criterion Y is not met, so criterion G is automatically not met; the study only reports immediate post-test results with no long-term follow-up, and no subsequent publication tracking this cohort was found.
No mention of pre-registration of the study protocol appears anywhere in the paper, and no registry record was located.
Despite growing interest in digital tools for speaking skills, research on YouGlish's impact on EFL learners' speaking achievement, particularly in Kuwait, remains limited. This study examined the effect of integrating YouGlish in English lessons on Kuwaiti EFL students' speaking achievement...
Participants were randomly assigned as individual students drawn from mixed majors into six conditions, not as intact classes or schools, and this is not a one-to-one tutoring intervention.
The outcome measures (ten custom multiple-choice questions per text and a written-recall/pausal-unit scoring scheme) were researcher-designed for this study, not a widely recognised standardised exam.
The entire experiment, from reading the passage to completing all assessment tasks, was administered in a single session and completed within 45 minutes, far short of one academic term.
The control ("no question") group's size, composition, and procedure are clearly documented, and baseline reading competence was confirmed to be equivalent across all six randomly assigned groups.
Randomisation occurred among individual students at a single institution rather than among schools.
The same author team designed the materials, ran the experiment, and performed the analysis; no independent third-party evaluator is described.
Since the Term Duration (T) criterion is not met (the study was a single 45-minute session), the stronger Year Duration criterion is automatically not met.
All six groups underwent an identical procedure and time allotment (single 45-minute session, same questionnaire, written recall, and multiple-choice test); the only difference between groups was the presence/type of embedded questions in the text, which is the core variable under investigation, so no extra, unmatched time or budget was given to any condition.
No independent replication of this specific study by a different research team was found in the paper or in an internet search of works that cite it.
Since criterion E (Exam-based Assessment) is not met, criterion A is automatically not met; in addition, only English reading comprehension was assessed, not all main subjects.
Since criterion Y (Year Duration) is not met, criterion G is automatically not met; no follow-up or graduation-tracking publication by the same authors was found either.
The paper reports institutional ethics approval but no pre-registration of the study's hypotheses, methods, or analysis plan on a public registry prior to data collection, and no registry entry was found on internet search.
Answering text-related questions while reading is a questioning strategy which is called adjunct questions or embedded questions, the benefits of which have been established in first-language reading as to enhance comprehension. The present study aims to study the effects different...
Randomisation occurred at the individual student level in a laboratory setting, not at the class or school level, and no valid tutoring/personal-teaching exception was invoked.
Outcomes were measured using a custom-built forced-choice identification test created by the researchers for this study, not a standardised, widely recognised exam.
The entire study, from pretest to final (delayed) posttest, spanned only about four weeks, far short of one academic term.
The control group's size, procedure (same training minus CF), and baseline/outcome scores are clearly documented in the text and in Tables 1-3.
Randomisation was conducted at the individual participant level across multiple different institutes, not at the school/institute level.
The same authors who designed the intervention and materials also conducted data collection and analysis themselves, with no independent evaluator involved.
Since criterion T (Term Duration) is not met, and the whole study spanned only about four weeks, criterion Y is not met.
The presence/type of corrective feedback is the explicit treatment variable being tested, so the no-CF control condition is the appropriate "business-as-usual" baseline rather than an unbalanced resource allocation.
No evidence was found of independent replication of this specific study by a different research team; the authors themselves call for future replication.
Criterion E (standardised exam) is not met, so criterion A is automatically not met; additionally, only two specific vowel contrasts were assessed rather than a broad set of subjects/skills.
Tracking ended entirely within about a month of the first training session, with no continued follow-up and no graduation tracking; criterion Y is also not met, which further precludes G.
The paper documents open sharing of materials but contains no statement of pre-registration of the study's hypotheses, design, or analysis plan before data collection.
This study investigated the effects of different types of corrective feedback (CF) provided during second language (L2) speech perception training. One hundred Korean learners of L2 English, randomly assigned to five groups (n = 20 per group), participated in eight...
Two pre-existing ("intact") university classes were assigned to conditions with no described randomisation procedure, so this is a quasi-experiment, not a class-level RCT.
Outcomes were measured via researcher-elicited read-aloud speech samples rated with a custom comprehensibility slider and acoustic/phonetic coding, not a standardised exam.
The interval from the start of the instructional treatment (around Week 5) to the post-test (Week 12) was only about seven weeks, well short of a full academic term.
The control group's composition, baseline proficiency, and the activities they received in place of the intervention are described in detail.
Only two intact classes at a single university were involved, with no school-level (or any) randomisation described.
The intervention was designed and delivered by one of the authors, who also served as the sole instructor for both groups.
Because criterion T (Term Duration) is not met, criterion Y is automatically not met per the ERCT dependency rule.
The control group received meaning-oriented lessons explicitly matched in duration to the experimental group's lessons, isolating the suprasegmental-instruction content as the treatment variable.
No independent replication of this specific study is mentioned in the paper, and an internet search found no subsequent replication by other research teams.
Because criterion E (Exam-based Assessment) is not met, and only suprasegmental/comprehensibility pronunciation outcomes (not other core subjects) were assessed, criterion A is not met.
Because criterion Y (Year Duration) is not met, criterion G is automatically not met, and no follow-up tracking toward graduation is reported or found via internet search.
No statement of pre-registration of the study protocol appears anywhere in the paper, and an internet search found no registry entry for this study.
The current study examined in depth the effects of suprasegmental-based instruction on the global (comprehensibility) and suprasegmental (word stress, rhythm, and intonation) development of Japanese learners of English as a foreign language (EFL). Students in the experimental group (n =...
The paper explicitly self-identifies as a quasi-experimental non-equivalent-groups design with assignment based on prior class results rather than genuine randomisation, so criterion C is not met.
The study used a teacher-made, researcher-designed test rather than a standardised, widely recognised exam, so criterion E is not met.
The paper does not clearly document the intervention start date, and its vague four-month follow-up claim is not linked to any reported data, so the term-duration interval cannot be verified as met.
The control group's size, school, comparability basis and instructional treatment (grammar translation method) are described, satisfying a minimal documentation requirement.
The study used only two sections within a single school, not school-level randomisation, so criterion S is not met.
The same researcher designed, delivered, and evaluated the intervention with no independent third-party evaluator, so criterion I is not met.
Since criterion T is not met and there is no evidence of a near-year-long duration, criterion Y is not met.
No additional class time, materials, or budget beyond normal instruction is described for either group, so the resource-balance requirement is trivially satisfied.
No independent replication of this specific study was found in the paper or via internet search; the cited studies are related but distinct research, not replications.
Only English writing was assessed, and since criterion E is not met, criterion A is not met.
Since criterion Y is not met, criterion G is automatically not met; no follow-up publication tracking this cohort toward graduation was located via internet search.
No mention of pre-registration is found anywhere in the paper, and no trial-registry record was located via internet search, so criterion P is not met.
Cooperative learning may have beneficial impact on classroom teaching. The study aimed to identify the effect of cooperative learning in writing ability of the 7th class students in the subject of English. The population of the study consisted of 68...
Randomisation was performed at the individual student level, not at the class or school level, and the intervention is a group curriculum rather than personal tutoring, so the exception does not apply.
The study relies on self-report psychometric scales (MHLS, DSS, GHSQ, MAKS) rather than a standardised, widely recognised academic exam.
The interval from intervention start (week 4) to the three-month follow-up (week 24) spans roughly four to five months, meeting the one-term minimum.
Detailed demographic and baseline outcome data for the control group are reported in Table 4 and confirmed to be equivalent to the intervention group.
The trial was conducted at a single institution with student-level (not school-level) randomisation.
The same authors who designed the intervention also supervised data collection and analysis; only the randomisation sequence was generated independently.
The intervention-to-final-measurement interval is only about four to five months, well below 75% of an academic year.
The intervention group received substantial extra instructional time and facilitator support unmatched by the waitlist control, and the authors themselves acknowledge this imbalance rather than framing the extra resources as the explicit treatment variable.
The paper is a newly accepted article with no independent replication reported in the text or found via internet search, and its recency precludes any replication having yet occurred.
Criterion E is not met, and the study assesses only mental-health-literacy-related constructs rather than all core academic subjects.
Criterion Y is not met, tracking stopped at the three-month follow-up with no tracking through graduation, and no follow-up publication by the same authors tracking this cohort was found.
The authors explicitly state the study was not registered, and no pre-registration reference is provided anywhere in the paper or found via independent registry search.
Background: Mental health literacy remains insufficient among university students, even as mental health difficulties grow more prevalent in this population. Blended learning approaches that combine digital and face-to-face components may hold advantages over conventional delivery methods, though empirical evidence supporting...
Randomisation was performed at the individual student level into six AI-tool groups and one control group, not at the class level, and no tutoring exception is invoked.
The primary academic-writing outcome was assessed with the IELTS Academic Writing Test, a widely recognised international standardised exam.
No start date, end date, or duration in weeks/months is given anywhere for the intervention, so a term-length follow-up cannot be verified.
The control group is identified only by a group label, sample size, and a generic descriptor ("Regular CALL Course"), without demographic or baseline detail specific to that group.
Randomisation occurred at the individual student level within a single participant pool, not among schools.
The study has a single author who both designed the AI writing-tool intervention and conducted the entire evaluation, with no independent third-party evaluator.
Since criterion T (Term Duration) is not met and no duration is documented at all, criterion Y cannot be met either.
Duration and assessment were standardized across all groups, and the tool used (AI vs. regular CALL) is itself the treatment variable being tested, so the control received a comparable "business as usual" course.
No independent replication of this specific study is mentioned or discoverable; the study was only published in 2025.
Only writing/language outcomes were assessed; no other core academic subjects (e.g. mathematics, science) were measured, and no specialised-subject exception is argued.
Since criterion Y (Year Duration) is not met, criterion G is automatically not met; the study also reports only an immediate posttest with no follow-up.
No pre-registration platform, registry ID, or registration date is mentioned anywhere in the paper.
Purpose: The purpose of this study is to investigate the transformative role of artificial intelligence (AI) tools in enhancing academic writing proficiency among English as a Foreign Language (EFL) learners. By focusing on the balance between syntactic complexity and clarity,...
Randomisation was carried out at the individual student level within a single institution's cohort, not at the class or school level, and the study is classroom-based rather than one-to-one tutoring, so the tutoring exception does not apply.
The reading tests used to measure outcomes were researcher-developed and modified rather than a pre-existing widely recognised standardised exam.
Outcomes were measured within about two months of intervention start (one month of treatment plus a delayed posttest three weeks after treatment ended), well short of a full academic term.
The control group's size, demographic composition, and the instruction/activities it received are clearly documented in the methods section.
Randomisation occurred among individual students within a single institution, not among schools or institutions.
The same two authors designed the intervention, delivered the instruction/oversight, and conducted the analysis, with no independent evaluator described.
Since criterion T (Term Duration) is not met, this stronger Year Duration criterion is automatically not met; in any case the total tracked period was about two months, far short of a year.
The additional ChatGPT-based practice is the explicit treatment variable the study set out to test, with the control group receiving a comparable "business as usual" package of face-to-face instruction and extra quizzes.
No independent replication of this specific study by a different research team was found; the closest related work is a prior study by the same authors.
Since criterion E (Exam-based Assessment) is not met, this stronger All-subject Exams criterion is automatically not met; in any case only reading was assessed.
Since criterion Y (Year Duration) is not met, this stronger Graduation Tracking criterion is automatically not met, and no follow-up beyond the three-week delayed posttest is reported.
No pre-registration of the study protocol, registry link, or registration date is mentioned anywhere in the paper, and no registry entry was found via internet search.
The main objective of this study was to measure the impact of ChatGPT on the reading skills of language learners. Therefore, a total of fifty Omani students with intermediate English proficiency were selected and randomly assigned into two groups, one...
Students (not intact classes or schools) were randomly assigned, so the trial is not a class-level (or stronger) RCT.
Outcomes were measured with Likert-scale questionnaire items rather than a standardized exam-based assessment.
The intervention and outcome measurement occurred over four 45-minute lessons with an immediate posttest, not at least one academic term after the intervention began.
The control condition (SGM) is clearly described and baseline outcome data (pretest self-efficacy) are reported by condition.
The study involved two schools but did not randomize at the school level; students were randomized within schools into conditions.
The paper does not document a clearly independent third-party evaluation; implementation and monitoring were conducted by project- affiliated master’s students.
Because term-duration tracking is not met (and the study is far shorter than 75% of an academic year), the year-duration criterion is not satisfied.
The control condition is parallelized in instructional time and structure (same number/length of lessons and similar teaching methods), so resources are balanced.
No peer-reviewed independent replication of this specific PSOM vs. SGM experiment was found in the paper or via an internet literature search.
Because the study does not use standardized exam-based assessments (Criterion E is not met), the all-subject standardized-exams criterion is not met.
The study does not track students through graduation and (because Year Duration is not met) Graduation Tracking cannot be met.
The paper provides no registry link/ID or dated statement indicating pre-registration before data collection.
According to Self-Determination Theory, it is essential to experience autonomy, competence, and relatedness in order to develop intrinsic motivation and motivational beliefs such as self-efficacy. Posing and solving one’s own modelling problems may support these needs by offering opportunities for...
Randomisation was performed at the individual student level within a single school, not at the class or school level, and no tutoring exception applies.
The outcome assessment is an unnamed 40-item multiple-choice test with no evidence of being a recognised standardised exam.
The intervention and measurement window spanned only eight to twelve weeks, short of a full academic term.
The control group's size, baseline scores, and the "regular classroom teaching" it received are clearly documented and statistically compared to the treatment group.
Only a single school was involved, with individual students allocated to conditions, not a randomisation of multiple schools.
No statement of independent, third-party data collection or evaluation is provided; the authors appear to have designed and conducted the study themselves.
The study lasted only eight to twelve weeks, far short of 75% of an academic year, and criterion T is not met.
The extra resources (laptop, internet access) given to the treatment group are the explicit treatment variable being tested against a business-as-usual classroom control.
The authors propose replication only as future work; no independent replication of this study has been reported or found through internet search.
Only vocabulary outcomes in a single subject area were measured, and criterion E is not met.
No follow-up tracking beyond the immediate posttest is reported or found through internet search, and criterion Y is not met.
No pre-registration of the study protocol is mentioned anywhere in the paper, and none was found through internet search.
Since English has been deemed a foreign language in Indonesian schools, how EFL students learn and acquire English vocabulary has been a hot topic. The purpose of this study is to demonstrate the impact of WLI on Indonesian EFL students'...
The study is a self-described quasi-experimental design using two intact classes assigned partly on scheduling and administrative grounds, with contradictory statements about randomisation, so a properly implemented class-level RCT is not established.
Outcomes were measured with a researcher-designed 2-minute speaking task, a researcher-adapted rubric, and a researcher-developed questionnaire rather than any widely recognised standardised exam.
The intervention lasted eight weeks with post-tests immediately afterwards in Week 9, which is roughly two months and falls short of a full academic term of 3-4 months.
The control group (n = 26) is clearly documented with demographics, baseline comparisons, and a full description of the parallel peer-to-peer condition it received.
The study involved two intact classes within a single university; no schools or institutions were randomised.
The authors designed, delivered, and analysed the intervention themselves, with the researcher also serving as the course instructor and no external evaluation team.
Tracking lasted only about ten weeks in total, far below 75% of an academic year, and criterion T is already unmet.
The control group completed parallel peer-to-peer speaking activities with the same topics, task structures, materials, instructor, and approximate time-on-task, so educational inputs were balanced and the AI tool itself was the treatment variable.
The paper was published in July 2026 and no independent peer-reviewed replication of this specific study exists or is referenced.
Only EFL speaking outcomes were measured with custom instruments; no other core subjects were assessed and criterion E is not met, so criterion A automatically fails.
Measurement stopped at the Week 9 post-test with no follow-up or graduation tracking, and prerequisite criterion Y is not met.
The paper contains no pre-registration statement or registry ID, and no external registration record was found.
Introduction: The growing availability of generative artificial intelligence (AI) tools has created new opportunities to enhance language learning, particularly speaking practice. However, studies on the effectiveness of AI voice-chat applications in enhancing English as a Foreign Language (EFL) learners' speaking...
Randomization was performed at the individual student level within a single online course, not at class or school level, and the intervention is not one-to-one tutoring, so the class-level RCT requirement is not satisfied.
Outcomes were measured with the course's own 42 summative module-test questions, a course-specific instrument rather than a widely recognised standardised exam.
The study was a short self-paced summer MOOC with outcomes collected within the course modules, and no interval of at least one academic term between intervention start and outcome measurement is documented.
The control group (Swedish-medium version) is clearly documented with its size, baseline demographics in Table 1, and its condition, an identical course differing only in language of instruction.
Randomisation occurred at the individual student level within one online course; no schools or institutional units were randomised.
The authors themselves translated/adapted the course, ran the trial, and analysed the data, with no external or third-party evaluation team documented.
Since criterion T is not met and the study was a short summer MOOC snapshot, tracking clearly did not span 75% of an academic year.
Both groups received identical courses with the same time, content, and resources, differing only in the medium of instruction, so inputs were fully balanced.
No independent published replication of this specific randomized EMI-versus-SMI course experiment was found in the paper or via external search; the authors themselves call for future replication.
Only programming/computing knowledge was assessed with a custom course test, so with criterion E unmet and no other subjects measured, the all-subject exam requirement fails.
Measurement ended with the course's own module tests, with no tracking of participants to graduation, and criterion Y is unmet which also fails this criterion.
The paper contains no mention of a pre-registered protocol or registry entry, and an external search found no pre-registration for this trial.
Stakeholders and researchers in higher education have long debated the consequences of English-medium instruction (EMI); a key assumption of EMI is that student's academic learning through English should be at least as good as learning through their first language (usually...
Randomization was at the individual student level (and the two groups were in fact separate academic-year cohorts), not at the class or school level, and the group-based peer feedback intervention does not qualify for the one-to-one tutoring exception.
Outcomes were measured with a course-specific narrative paragraph writing task scored by a researcher-made rubric, not a widely recognised standardised exam.
The intervention and outcome measurement were completed within a single 3-week instructional unit, far short of one academic term.
The control group's size, cohort, condition (OPF plus teacher feedback), and baseline and final writing scores are clearly documented, with sample demographics reported.
The study took place in a single university with allocation at the student/cohort level; no school-level randomisation occurred.
The single author designed, taught, administered, and scored the intervention with no external or independent evaluation team mentioned.
The full study lasted only 3 weeks, nowhere near 75% of an academic year, so the year-duration requirement fails (as does its prerequisite T).
Both groups received the identical 3-week OPF activity and class time, and the only input difference (presence or absence of teacher feedback) is the explicit treatment variable being tested.
No independent replication of this specific study exists; it was published in May 2025, has zero recorded citations as of the verification date, and no external search identifies any replication by another team.
Only EFL narrative paragraph writing was assessed, with a custom rubric rather than standardised exams, so all-subject coverage fails (and prerequisite E is unmet).
Measurement ended with the final draft at week 3 of the unit; no tracking of participants to graduation is reported, and no follow-up publication tracking this cohort was found.
The paper contains no mention of any pre-registered protocol, registry, or registration date, and no external registry record for this study could be found.
Online Peer Feedback (OPF) is a proven and effective peer editing tool in EFL writing classrooms. However, the level of effectiveness varies depending the relationship between peer editors and on the available peer editing tools. This study investigates the impact...
The study is a quasiexperiment in which four intact classes were assigned to conditions without randomisation, so it is not a class-level RCT.
Outcomes were custom researcher-coded oral fluency measures from picture-narration tasks, not standardised exams (TOEIC was used only to describe baseline proficiency).
Outcomes were measured only up to one week after a one-to-two-week training phase, far short of a full academic term.
The control group's size, origin, baseline proficiency and pretest scores, and conditions (tests only) are clearly documented.
Allocation involved four intact classes within one university, so there was no school-level randomisation.
The authors designed, delivered (the second author was the classroom instructor), and analysed the intervention with no independent evaluators.
The study lasted only a few weeks in total, far short of the required 75% of an academic year (and T is not met).
The experimental groups received identical amounts of practice within regular class time, and the task-repetition practice itself was the treatment variable tested against a business-as-usual control.
No independent replication of this specific study by a different research team was found in the paper or via external search (related studies by Kakitani & Kormos, 2024, and Tabari et al., 2025, differ in design/modality or are not independent).
Criterion E is not met and only L2 oral fluency was measured, so all-subject standardised assessment is absent.
Measurement stopped one week after training with no tracking of the first-year students to graduation (Y is not met, and no follow-up publication tracking this cohort was found).
The paper reports an Open Materials badge but no pre-registration of hypotheses, methods, or analyses before data collection; no registry entry was found via search.
To examine the effects of task repetition with different schedules, English-as-a-foreign-language classroom learners performed the same oral narrative task six times under three different schedules. They narrated the same six-frame cartoon story (a) six times consecutively in one class (massed...
Randomisation was at the individual student level within one institute, not at the class or school level, and the intervention is not one-to-one tutoring.
Outcomes were measured with custom researcher-designed tests (UGJT and an adapted MKT), not widely recognised standardised exams.
Outcomes were measured about five weeks after the intervention began, far less than one full academic term.
The control group's size, demographics, baseline scores and exact conditions are clearly documented, enabling proper comparison.
Individual students within a single site were randomised; there was no school-level randomisation.
The authors designed, taught, tested and analysed the study themselves with no independent third-party involvement.
The study spanned only about five weeks from intervention start to posttest, far below 75% of an academic year.
All groups received equal time, identical passages and the same instructor, so inputs were balanced and only the tested instructional technique differed.
No independent published replication of this specific study was identified in the paper or via external search.
Only custom tests of English passive voice were used; no standardised exams across all main subjects were administered, and criterion E is not met.
Measurement ended one week post-treatment with no tracking of participants to graduation, and no follow-up papers by the same authors were found.
The paper contains no reference to any pre-registered protocol or registry entry, and none was found via external search.
The importance of teaching explicit knowledge of grammar has been one of the most controversial issues in L2 instruction. Despite the existence of considerable number of studies on different methods of grammar instruction, very few studies have investigated the role...
Randomisation was performed at the individual student level within three universities, not at the class or school level, and the intervention (video input before an oral task) is not a personal tutoring intervention that would qualify for the exception.
Outcomes were measured with a custom picture-prompted oral narrative task analysed by research tools (HD-D, TAALES, Praat), not with any widely recognised standardised exam.
The whole study spanned only three sessions over a few days, with the repeat oral task administered two days after the immediate task, far short of one academic term.
The no-input baseline group is clearly documented, with its size, condition, demographics, and baseline vocabulary and working memory scores reported and compared across groups.
Randomisation was at the individual student level; no schools or institutions were randomly assigned to conditions.
The same research team designed the intervention, delivered the sessions, transcribed, coded, and analysed all data, with no external or third-party evaluation reported.
With outcomes measured two days after the intervention began, the study falls vastly short of the required 75% of an academic year; criterion T is also unmet, so Y cannot be met.
The extra input exposure given to the treatment groups is itself the treatment variable being tested against a deliberate no-input baseline, and the added video time (about 6-18 minutes) is minimal and integral to the design.
No independent replication of this specific study by a different research team is reported in the paper or found in an external search; the authors themselves call for replication studies.
Since criterion E is not met and only L2 speaking outcomes were measured, with no assessment of other core subjects, the all-subject exams criterion fails.
Measurement ended two days after the intervention with no follow-up, so participants were not tracked to graduation; criterion Y is also unmet, which rules G out.
The paper contains no mention of any pre-registration, registry, or published protocol for the study.
The present study investigates the impact of meaningful input on L2 learners' vocabulary use and their fluency in oral performance (immediate and repeat tasks), as well as whether the effects are mediated by learners' prior vocabulary knowledge and working memory....
The authors explicitly declare a quasi-experimental design using two pre-existing intact classes, so a properly implemented class-level RCT is not demonstrated.
Outcomes were measured with ad hoc, school-made trimester exams, not a widely recognised standardised test.
Only about two months elapsed from intervention start (January 2020) to outcome measurement (early March 2020), which is shorter than one full academic term.
The control group's size, demographics, baseline exam scores and traditional-instruction conditions are documented in detail.
The study involved only two classes within one single school, so there was no school-level randomisation.
The authors designed the intervention, created the instruments and analysed the data themselves, with no independent third-party evaluation.
The study lasted only about two months, far below the required 75% of an academic year.
Both groups had identical class time, teacher and content; the gamified digital activities were integral to the tested methodology, so resources were balanced.
The study is presented as a first exploratory work, and a citation-database search found only one unrelated citing paper and no independent replication.
Only performance in English was measured with ad hoc exams; no other core subjects were assessed and criterion E is not met.
Tracking ended with the March 2020 trimester exam, no graduation follow-up was found in the paper or in the authors' later publication record, and criterion Y is not met.
The paper contains no reference to any pre-registered protocol or registry entry, and none was found online either.
Teaching English as a Foreign Language (EFL) is constantly searching for the most successful learning method. Using active learning methodologies and ICT tools in educational contexts entails some changes in the teaching-learning process. This research work analyzes the effect of...
Randomisation was at the individual student level rather than at the class or school level, and the computerized practice intervention does not qualify for the tutoring exception.
Outcomes were assessed with multimedia discourse completion tasks and a rubric custom-designed for this study, not with a widely recognised standardised exam.
The interval from intervention start to the final delayed posttest was only about five to seven weeks, which is shorter than one full academic term.
The comparison group's size, demographic characteristics, baseline pretest performance, and exact treatment conditions are documented in detail.
The study randomised individual students within one university rather than randomising schools or institutional units.
The single author designed, delivered, and evaluated the intervention herself, with no independent evaluation team or third-party oversight.
The study tracked outcomes for only about five weeks after the intervention, far short of 75% of an academic year, and criterion T is not met.
Both groups received identical instruction, tasks, feedback, and practice time, with only the sequencing of the same 20 tasks differing as the treatment variable.
The paper describes itself as the first study of its kind, and a citation search found no independent peer-reviewed replication of this specific experiment by a different research team.
Criterion E is unmet and the study assessed only English pragmatics outcomes with custom instruments, not standardised exams across all main subjects.
Measurement stopped at a five-week delayed posttest, with no tracking of the sophomore participants through to graduation found in the author's later publications, and criterion Y is not met.
No pre-registration statement, registry identifier, or registration date is mentioned in the paper, and none was found via an external registry search.
A handful of second/foreign language (L2) studies have examined the effects of practice schedules and reported the advantage of interleaved practice (i.e., practice multiple skills simultaneously) over blocked practice (i.e., practice one skill first and then proceed to the next...
Randomisation was done at the individual student level within one university cohort, not at class or school level, and the tutoring exception does not apply.
Outcomes were measured with the widely recognised standardised Cambridge B1 PET listening test (two parallel forms), with only six added recognition items as a minor modification.
Outcomes were measured at the end of a 10-week intervention, which is shorter than the required full academic term of roughly 3-4 months.
The study has no control group at all - all three randomised arms were active treatments - so a documented control group is absent.
Individual students within a single university were randomised, so there was no school-level randomisation.
The single author designed, delivered, monitored, and analysed the intervention herself with no independent evaluation team or third-party oversight.
The study tracked outcomes for only 10 weeks, far less than 75% of an academic year.
All three comparison arms received equivalent practice time, apps, and instructor support, with only the app modality (the treatment variable) differing between groups.
The study is presented as a novel first comparison, and a citation search found no independent peer-reviewed replication of this specific design by another team.
Only English listening was assessed; no other main subjects or skill areas were measured with standardised exams.
Participants were only tracked to the immediate posttest after 10 weeks, with no follow-up until graduation found in the paper or in a citation search of subsequent publications.
The paper contains no reference to a pre-registered protocol, registry ID, or registration date.
Mobile apps are becoming part and parcel of our daily lives. Hence, this study examined the differential effects of app modes on listening comprehension and recognition. From a pool of Egyptian EFL sophomores, 107 students were randomly assigned into 3...
Randomisation was at the individual student level within one course at a single institute, not at class or school level, and the one-to-one tutoring exception does not apply.
Outcomes were measured with custom-designed parallel writing tasks scored on a rubric adapted from IELTS descriptors, not with a widely recognised standardised exam.
The interval from intervention start to outcome measurement was only two weeks (Week 1 to Week 3), far short of the required full academic term.
The control group's size, baseline scores, participant characteristics, and the exact feedback condition it received are clearly documented in the text and Tables 1-2.
The study randomised individual students within one private language institute; no schools or institutional units were randomised.
The single author designed the intervention, ran the study, and analysed the data himself, with only blinded essay raters and no independent evaluation team.
The full study interval was about two weeks, far below 75% of an academic year, and prerequisite criterion T is also not met.
Both groups received structurally matched feedback of the same approximate length under identical task conditions, so time and resources were balanced and only the feedback source differed.
The recently published study presents itself as a novel contribution and no independent peer-reviewed replication of this specific trial exists.
Only EFL paragraph writing was assessed with a custom instrument, and since prerequisite criterion E fails, A fails as well.
Measurement ended at the two-week post-test with no follow-up or graduation tracking, no subsequent tracking paper was found, and prerequisite criterion Y is not met.
No pre-registration, registry ID, or registration date is reported anywhere in the paper, and no external registry record was found.
Although considerable research has explored the role of feedback in second language writing, limited studies have compared the effects of AI-generated feedback and teacher-written feedback on both academic performance and emotional experience, particularly in EFL contexts. This mixed-methods brief report...
Randomisation was done at the individual student level within a single university cohort (and the paper even labels the design quasi-experimental), not at the class or school level, and the classroom-based intervention does not qualify for the tutoring exception.
Outcomes were measured with researcher-constructed gap-filling/multiple-choice tests built for this study, not with any widely recognised standardised exam.
The interval from intervention start to the final delayed post-test is only about eight weeks (a 4-week intervention plus a 4-week delay), well short of one full academic term.
The control group's size, baseline scores, comparability, and the traditional business-as-usual instruction it received are all explicitly documented.
The study randomised individual students within one university; no schools or comparable institutional units were randomised.
The same authors designed the cognitive linguistics-based intervention, delivered it, collected the data, and analysed the results, with no external or independent evaluation team.
The total tracking period was about eight weeks, far below 75% of an academic year, and criterion T is already not met.
Both groups received the same amount of instructional time (4 weeks, two 90-minute classes weekly) on the same phrasal verb content, with the control getting an active traditional programme, so no extra time or resources favoured the intervention group.
This January 2026 study reports no independent replication of itself, and prior similar cognitive-linguistics PV studies it cites are earlier separate experiments, not replications of this trial. An internet search for citing or replicating papers found none published as of this check.
Criterion E failed (custom tests), and only English phrasal verb knowledge was assessed, with no other core subjects measured.
Measurement stopped four weeks after the 4-week intervention, with no tracking of participants to graduation and criterion Y already unmet. No follow-up publication tracking this cohort was found via internet search.
The paper contains no mention of any pre-registered protocol, trial registry, or registration date.
Introduction: This study aims to evaluate the effectiveness of cognitive linguistics-based instruction in enhancing the comprehension and retention of English Phrasal verbs (PVs) among Vietnamese EFL learners. Methodology: Eighty business-majored students from two classes were selected and divided into two...
Randomisation was at the individual student level, not at the class or school level, and the intervention was whole-class teaching, so no tutoring exception applies.
Outcomes were measured with a custom, author-made written discourse completion test rather than any recognised standardised exam.
Outcomes were last measured about three weeks after a short four-lesson intervention, far less than one full academic term of tracking.
The business-as-usual comparison group (PPP-EG) is documented with its size, shared demographic profile, condition received, and baseline pre-test scores.
The trial involved a single school with student-level random assignment, so no school-level randomisation occurred.
A single teacher-researcher designed the intervention, taught all groups, created and scored the tests, and analysed the data without any independent oversight.
The whole study spanned only a few weeks, far short of the required 75% of an academic year of tracking.
All three arms received the same four regular 45-minute lessons from the same teacher with documented comparable time allocations, and remaining structural differences were the treatment contrast itself.
No independent replication of this specific 2024 study by a different research team in a peer-reviewed journal is reported or known; the only citing works found are by the same author.
Criterion E is not met and only L2 pragmatic production was assessed, with no standardised exams covering other main subjects.
Measurement ended three weeks after the intervention with no tracking of students through to graduation, and no follow-up publication tracking this cohort was found.
The paper contains no mention of pre-registration, a registry ID, or a published protocol of any kind, and no external registry record was found.
The present study investigates the effect of different types of task implementation on teaching L2 interactional sequences. 81 EFL learners were randomly assigned to one of three experimental groups. In the first experimental group (T1-EG, n = 27), implicit instruction...
Students were randomised individually within each class (within-class randomisation), not by whole classes or schools, and the intervention was whole-class instruction, not one-to-one tutoring, so no exception applies.
The primary and most outcome measures were researcher-developed or research-specific instruments (affix knowledge test, morphological awareness tasks, Tong et al. reading comprehension task), not widely recognised standardised exams, with only part of the word reading composite (WRAT-4) being standardised.
The intervention lasted only five weeks and outcomes were measured immediately after it ended, far short of one academic term of tracking from intervention start.
The active control (affix knowledge) group is documented in detail, including its instructional condition, baseline equivalence checks, and pretest/posttest descriptive statistics in Table 1.
The study took place in a single school with students randomised within classes; no schools were randomised.
The authors designed the interventions and measures and conducted the evaluation themselves; teacher blinding is noted but there is no independent third-party evaluation team.
The study tracked outcomes for only about five weeks, far below 75% of an academic year, and criterion T is also not met.
An active control group received the same dosage (two 50-min sessions per week for 5 weeks), the same affix content, teacher training, videos, reading passages and games, differing only in the targeted strategy instruction being tested.
The paper is explicitly presented as the first study of its kind, and an internet search of its citation record found no independent replication by a different research team.
Only English reading-related outcomes were measured, no other core subjects were assessed, and prerequisite criterion E is not met.
Measurement ended immediately after the 5-week intervention with no delayed posttest or tracking toward graduation, an internet search found no graduation-tracking follow-up by the same authors, and prerequisite criterion Y is not met.
The paper contains no mention of any trial registry, pre-registration ID, or pre-registered protocol, and an internet search of trial registries found no matching registration.
Previous studies have demonstrated the overall effectiveness of morphological interventions in enhancing word reading and reading comprehension. However, the specific practices within these morphological interventions vary significantly and few intervention studies have investigated how targeted instruction in distinct aspects of...
Randomisation was conducted at the individual student level within one institution, not at the class or school level, and the intervention is not one-to-one tutoring.
Writing outcomes were scored with a researcher-designed 100-point analytic rubric by two teachers, not with a widely recognised standardised exam.
The whole study, from baseline to the final retention task, lasted only six weeks, which is well short of one academic term.
The control (regular AI) group's size, condition, baseline writing scores, and comparability checks are clearly documented.
Randomisation occurred among individual students within a single institution, not among schools or comparable implementation units.
A single author designed the interventions and also conducted, analysed, and reported the study with no external or third-party evaluation.
The entire study lasted six weeks, far short of 75% of an academic year, and criterion T is already not met.
All groups received the same AI feedback and writing tasks; the modest extra training time (FRAC/APCA activities) is the explicit treatment variable being tested against regular AI use.
The paper, published in June 2026, reports no independent replication of this specific FRAC/APCA trial and none could be identified via internet search.
Only English argumentative writing was assessed, with no standardised exams and no measurement of other core subjects, and criterion E is not met.
Tracking ended at the six-week retention task with no follow-up to graduation, and prerequisite criterion Y is not met.
The paper reports ethics approval but no pre-registration of the study protocol on any registry.
Introduction: As generative AI becomes increasingly integrated into writing instruction, the central educational challenge is not only to provide feedback but also to help students interpret, evaluate, and use such feedback critically. This study examined whether two metacognitive interventions--a Feedback...
Randomization was at the individual participant level (dental interns), not at the class (or stronger) level, and no tutoring exception applies.
The outcome was a researcher-developed 15-question MCQ quiz rather than a widely recognized standardized exam.
Outcomes were measured immediately after a brief 15-minute quiz, not at least one academic term after the intervention began.
The control group is clearly defined (baseline knowledge, no AI), its size is reported, and its conditions are described alongside the intervention group and outcome tables.
Randomization occurred among individual interns within one institution, not among schools (or equivalent educational sites).
The paper does not document an independent third-party evaluation; instead, the study team designed the assessment and ran the study.
The study does not track outcomes for 75% of an academic year and does not provide term-length follow-up.
Although the AI-assisted group received additional resources (internet and AI tools), these resources are the explicit treatment being tested against baseline practice, with time-on-task held constant.
No independent replication of this specific study was found, and the paper only cites other related AI education RCTs rather than replications of this experiment.
This criterion is not met because exam-based assessment (E) is not met; additionally, outcomes focus on a single case-based quiz rather than standardized exams across all core subjects.
The study does not track participants until graduation and, since year duration (Y) is not met, graduation tracking (G) is not met.
The paper provides ethics approval details but does not report a public pre-registration record, registry ID, or a registration date before data collection.
Introduction: This study aimed to assess the impact of artificial intelligence (AI) assistance on immediate task performance and evaluate perceived task load and AI acceptance among dental interns in an educational setting. Methods: A pragmatic experiment was conducted among 132...
Randomisation was done at the individual student level, not at the class or school level, and the intervention is not one-to-one tutoring, so the exception does not apply.
Outcomes were measured with a self-report autonomy questionnaire and a researcher-made interview, not with any standardised exam-based assessment.
The interval from intervention start to outcome measurement was only six weeks, which is far shorter than the required full academic term.
The control group's size (37), demographics, baseline OPT proficiency, pretest autonomy scores, and traditional face-to-face condition are all clearly documented.
Randomisation occurred at the individual student level within a single university branch, so there was no school-level assignment.
The same researchers designed, delivered, and evaluated the intervention themselves, with no external or independent evaluation team involved.
The six-week study duration is far shorter than 75% of an academic year, and the prerequisite term-duration criterion is also unmet.
The thrice-weekly SMS messages are the explicit treatment variable integral to the tested intervention, and both groups shared the same course and teacher, so the design is acceptably balanced.
No independent replication of this specific study exists; a citation search confirms the cited similar studies and citing papers are earlier or unrelated works, not replications of this trial.
Criterion E is unmet and only an EFL autonomy questionnaire was measured, so no all-subject standardised exam assessment exists.
Measurement ended immediately after the six-week treatment, with no follow-up tracking of participants until graduation; the one same-author follow-up paper found addresses vocabulary outcomes, not graduation tracking.
The paper contains no mention of a pre-registered protocol, registry platform, or registration date, and an independent search found no registration record.
This paper investigates the efficiency of text messaging as an English as a Foreign Language (EFL) instructional tool to enhance learner autonomy and perception at the Islamic Azad University-South Tehran Branch, Iran. The study considers seventy-four learners to participate in...
Individual students from one kindergarten course list were randomized into two newly formed groups, which is student-level (not class-level) randomization for a group-teaching (non-tutoring) intervention.
Outcomes were measured with a custom pictorial pre/post-test designed by the researchers for this study, not a widely recognised standardised exam.
The intervention lasted only four weeks with the post-test administered immediately after the eighth session, far short of one academic term.
The control (non-RALL) group's size, composition, baseline performance, and conditions (identical lessons and games without the robot) are clearly documented.
Randomisation occurred among 38 individual children within a single kindergarten; no schools or sites were randomised.
The authors designed the intervention, wrote the materials (including the co-authored textbook), programmed the robot, taught both classes, and analysed the data themselves, with no independent evaluators.
The study tracked outcomes for only four weeks, far below 75% of an academic year (and criterion T is also not met).
Both groups received identical lesson plans, games, materials, and eight one-hour sessions; the only difference was the robot itself, which was the explicit treatment variable being tested.
No independent replication of this specific RALL pragmatics study by a different research team was found; the authors themselves call for future replication.
Only English pragmatic performance (two speech acts) was assessed with a custom test; no other subjects were measured, and criterion E is not met, so A automatically fails.
Measurement stopped immediately after the four-week course with no follow-up or graduation tracking, and criterion Y is not met, so G automatically fails.
The paper contains no mention of any pre-registered protocol, registry platform, or registration date, and no registry entry for this study was found online.
Technology, as a source of instruction, has fulfilled various purposes in foreign language learning environments. During the last decade, Robot-Assisted Language Learning (RALL) has attracted teachers' and researchers' attention due to the look and feel of humanoid robots. However, in...
Two intact second-grade classes were randomly assigned to experimental and control conditions, so randomisation was at the class level (albeit with only two clusters).
Outcomes were measured with a researcher-made observation card and a teacher's journal, not with any standardised exam.
Outcomes were observed only during the ten-week treatment, which is shorter than the required full academic term of roughly 3-4 months.
The control group is described only by size and its conventional instruction, with no demographic or baseline performance documentation.
The study randomised two classes within a single Kuwaiti school, so no school-level randomisation took place.
The researchers designed, delivered, observed, and analysed the intervention themselves, with no independent evaluators.
The ten-week study falls far short of the required 75% of an academic year, and the prerequisite term-duration criterion is also unmet.
Both groups had the same regular reading class time, and the Starfall-based instruction itself was the integral treatment variable tested against conventional teaching.
The study is presented as among the first of its kind and no independent published replication of it was found.
No standardised exams were used at all (criterion E fails) and only English reading was observed, not all main subjects.
Data collection ended with the ten-week treatment and no graduation tracking or follow-up of the cohort exists.
The paper contains no reference to any pre-registration, registry ID, or published protocol.
This study examines the effect of a Starfall-based instructional program on second-grade pupils' reading comprehension in a Kuwaiti public school in the first semester of the academic year 2023/2024. The participants were divided into a control group taught conventionally per...
Intact classes (not individual students within a class) were randomly assigned to the two task-repetition conditions, so the unit of randomisation was the class.
Outcomes were custom research measures of utterance fluency derived from manual PRAAT analysis of recorded speech, not any standardised, widely recognised exam.
The entire intervention and all outcome measurement took place within a single scheduled class session, far short of one academic term of follow-up.
There is no business-as-usual control group, and the comparison (PTR) group's demographics and baseline characteristics are not documented separately from the pooled sample.
Randomisation occurred among intact classes within a single private language school, not among schools or sites.
The sole author designed the two carousel versions, implemented the study, and manually performed the fluency analysis herself, with no independent evaluation team.
Since the entire study occurred within one class session, tracking did not remotely approach 75% of an academic year (and criterion T is already not met).
Both randomised conditions received the same amount of class time and an identical activity structure (three storyboard retellings during a normal lesson); the only difference - same versus different stories - is the treatment contrast itself.
No independent, peer-reviewed replication of this specific poster carousel fluency study exists; the paper itself calls for future replications and a targeted citation search found only review/overview articles citing the study, not replications.
Only L2 speaking fluency was measured with custom instruments; no other subjects were assessed and criterion E is not met, which automatically fails this criterion.
Measurement ended within the single training session with no follow-up of participants, let alone tracking to graduation; a search for follow-up publications by the same author found none, and prerequisite criterion Y is not met.
The paper contains no mention of any pre-registration, registry platform, or registered protocol, and a targeted search of registries and citation databases found no pre-registration record for this study.
This paper reports on the impact of an English as a Second Language (ESL) speaking activity - the poster carousel - on English learners' second language (L2) fluency. Two versions of the poster carousel were developed to observe the effect...
Randomization was at the individual student level, but the intervention is one-to-one tutoring by a chatbot positioned as a personal English tutor, so the tutoring exception applies and student-level RCT is acceptable.
Outcomes were measured with a custom assessment built from 25 sentences taken from each participant's own chat utterances, not a standardised, widely recognised exam.
The intervention comprised 12 chatbot sessions over roughly four weeks with outcomes measured immediately after the last session, which is far shorter than one academic term.
The study has no control group by design, and the two feedback-timing arms lack documented per-group demographics and baseline performance, so no properly documented comparison group exists.
Randomisation occurred at the individual student level within four classes at a single institution, not at the school level.
The same research team designed, built, administered, and analysed the chatbot intervention with no external or third-party evaluation.
The whole study, from first session to final assessment, lasted about one month, far below 75% of an academic year.
Both groups received identical chatbot sessions, identical session dosage, and the same system-generated amount of feedback, differing only in feedback timing, which is the treatment variable itself.
No independent replication of this specific chatbot feedback-timing RCT by a different team has been published; related studies exist but do not reproduce this study's design, and a citation-graph search found no replication.
Only English grammar knowledge was assessed, with a custom instrument, so neither the standardised-exam prerequisite nor the all-subjects coverage is satisfied.
Measurement stopped immediately after the final chat session with no follow-up, and since criterion Y is not met this criterion cannot be met either.
The paper contains no mention of a pre-registered protocol or registry entry made before data collection; the only external repository (OSF) was created in January 2023, after data collection had already finished.
The emergence of Large Language Models (LLMs) has opened new possibilities for language learning through conversational interaction with chatbots. Yet, little empirical evidence exists on how students experience such interactions and how corrective feedback should be provided. Research suggests that...
Randomisation was done at the individual student level within a single school, not at the class or school level, and the classroom-based feedback intervention does not qualify for the one-to-one tutoring exception.
Outcomes were measured with researcher-administered essay tests scored on a rubric adapted from the Cambridge B2 scale, not with a widely recognised standardised exam.
The interval from intervention start to outcome measurement was only 8 weeks, which is shorter than one full academic term (approximately 3-4 months).
The control group is clearly documented with its size (n = 30), comparable demographics and baseline proficiency, baseline test scores, and a description of the teacher-only feedback it received.
The study took place in a single school with randomisation at the individual student level, so there was no school-level randomisation.
The authors themselves designed the study, conducted the data collection, and performed the analyses, with no independent third-party evaluation team beyond the two external essay raters.
The study lasted only 8 weeks, far short of 75% of an academic year, and since criterion T is not met this criterion automatically fails as well.
The extra input (immediate AI-generated feedback via Grammarly, within a hybrid AI-plus-teacher feedback package) is the explicit treatment variable tested against business-as-usual teacher feedback, with in-class time, curriculum, and teacher feedback time deliberately equalised across groups.
No independent replication of this specific study exists; earlier similar trials by other teams (e.g., Wei et al. 2023) predate this study and are prior related work, not replications of it.
Only EFL writing was assessed, no other school subjects were measured, and since criterion E is not met this criterion automatically fails as well.
Measurement ended at the Week 8 post-test with no follow-up toward graduation, and since criterion Y is not met this criterion automatically fails as well.
The paper reports ethics approval but contains no mention of any pre-registered protocol on a trial registry before data collection began, and no registration for this study was found via internet search.
This study examines the impact of AI-assisted writing feedback on the writing skills of secondary-level EFL students. A sample of 60 Turkish high school students was divided into an experimental group receiving feedback from an AI writing assistant and a...
Randomisation was carried out at the individual student level within a single participant pool, not at the class or school level, and the intervention was not one-to-one tutoring.
Outcomes were measured with researcher-made form-recognition and meaning-recall MWE tests created for this study, not with any widely recognised standardised exam.
The whole study spanned only four weeks, with the viewing treatment and the unannounced posttest occurring in the same session in Week 4, far short of one academic term.
The comparison (active control) conditions are documented with group sizes, demographics, and baseline measures of proficiency, vocabulary knowledge, and working memory, with statistical checks confirming baseline comparability.
Randomisation occurred at the individual participant level within a single university, so no school-level randomisation took place.
The single author designed the materials, conducted the experiment, and analysed the data as part of his master's dissertation, with no independent third-party evaluation.
The study lasted only four weeks with an immediate posttest, so it falls far short of 75 percent of an academic year, and the weaker term criterion T is also unmet.
All three groups received identical time and materials - the same 30-minute video watched twice in the same session - with only the subtitle format differing, so inputs were fully balanced across conditions.
This 2025 study describes itself as the first of its kind, and a fresh internet search in July 2026 still found no independent published replication of this specific experiment.
Only knowledge of English multiword expressions was measured with custom tests, so neither the exam-based prerequisite (criterion E) nor coverage of all main subjects is satisfied.
There was no delayed posttest or longer-term follow-up of any kind, let alone tracking of participants until graduation, the prerequisite criterion Y is unmet, and a fresh internet search found no follow-up publications on this cohort.
The paper contains no mention of any pre-registration; direct inspection of the linked OSF project confirms it is an unregistered project hosting only video excerpts, an OASIS document, and appendices, not a pre-registered protocol.
Given the accumulating evidence about audiovisual input as a valuable resource from which knowledge of multiword expressions (MWEs) can be built up incidentally, the next inquiry arises as to what can be done to promote MWE uptake from this resource....
Two intact classes were conveniently selected and one was simply "designated" as the experimental group, so no proper class-level (or any) randomisation is described.
Outcomes were measured with researcher-designed rule analysis, grammar, oral, and written tests piloted for the study, not with any recognised standardised exam.
The intervention lasted eight weeks with main post-tests one day later and only a single delayed written test about three months after the start, so a full term of tracking is not clearly reached.
The control group's size, gender mix, proficiency level, baseline equivalence on semester grades, and the instruction it received are all clearly documented.
The study took place in a single high school with only two classes, so there was no school-level randomisation.
The first author, a teacher at the study school, designed the treatment, taught it, and analysed the data herself with no independent evaluation team.
Total tracking from intervention start to the final delayed test was roughly three months, far short of 75% of an academic year.
Both classes received the same amount of instructional time, the same contextualised input, and practice, differing only in the instructional technique being tested.
No independent replication of this specific study is reported; the paper only builds on earlier related research by Fotos and Ellis rather than being replicated itself.
Only English tense/grammar outcomes were measured with custom tests; no other school subjects were assessed and criterion E is not met.
Measurement ended with a delayed written test about a month after the oral test, with no tracking of students to graduation.
The paper contains no mention of any pre-registration, registry, or published protocol.
Two approaches to grammar instruction are often discussed in the ESL literature: direct explicit grammar instruction (DEGI) (deduction) and indirect explicit grammar instruction (IEGI) (induction). This study aims to explore the effects of indirect explicit grammar instruction on EFL learners'...
Students (not intact classes/schools) were randomized to groups, so class-level randomization was not used.
The primary knowledge outcome used a researcher-prepared test, not a widely recognized standardized exam.
Outcomes were measured over about six weeks from start to latest follow-up, which is shorter than a full academic term.
The control condition and baseline characteristics are described, including time/resources for control and demographic comparability.
Randomization was not conducted at the school/site level; it was done among students within one university program.
The authors/researchers carried out core trial activities with no clear independent evaluation team or blinded outcome assessment.
The study duration is far shorter than 75% of an academic year, and T is not met, so Y cannot be met.
Both groups received substantial instructional time and lab practice; the main contrast is collaborative versus individual practice rather than a clear one-sided resource increase.
No independent peer-reviewed replication of this specific RCT was found as of the ERCT check date.
Criterion E is not met and the study does not assess all core subjects with standardized exams.
The study reports only short-term follow-up (4 weeks) and does not track participants until graduation; Y is also not met.
The trial was registered on the date the study began (and not clearly before study start), so it does not meet pre-registration.
Collaborative learning is one of the important interactive teaching methods in teaching nursing practices. This study aimed to examine the impact of the collaborative learning approach on nursing students’ knowledge levels and self-directed learning skills related to enteral nutrition. This...
Randomisation was performed at the individual student level within a single course cohort, not at the class or school level, and no tutoring exception applies.
Academic performance was measured with a researcher-developed 20-item multiple-choice test, not a widely recognised standardised exam.
The intervention lasted only four weekly two-hour sessions with outcomes measured immediately at the end, far short of one academic term.
The control group's size, demographics, baseline scores, and business-as-usual lecture condition are clearly documented.
Randomisation occurred at the individual student level in a single institution, not at the school level.
The intervention was designed and delivered by the authors themselves, with no independent third party conducting the trial.
The study lasted only four weeks with immediate measurement, far below the year-duration requirement; criterion T is also not met.
In-class time, instructor, content, and slides were identical across arms; the additional pre-class engagement is the integral defining feature of the flipped classroom tested against business-as-usual lecture.
No independent replication of this specific trial by a different research team is reported or identified.
Only a single fluids-and-electrolytes module was assessed with a custom test, and criterion E is not met, so all-subject exams cannot be satisfied.
Outcomes were measured immediately after the intervention with no follow-up to graduation, and criterion Y is not met.
The trial was retrospectively registered in December 2025, years after the 2019-2020 data collection, so it was not pre-registered.
Introduction: Nursing students often struggle with the fluids and electrolytes course because of the complexity, extensive detail, and fleeting nature of the content. This study aimed to compare the effects of two teaching approaches, the flipped classroom and traditional lecture-based...
Randomization was performed at the individual student level rather than by class or school, and the digital intervention is not one-to-one tutoring.
The study used a custom author-adapted questionnaire rather than a widely recognised standardised exam.
Outcomes were measured only about three days after baseline, far shorter than one academic term.
The control group's demographics, size, baseline scores, and no-treatment condition are clearly documented.
Randomization occurred at the individual student level within one university, not at the school level.
The same team that designed the intervention also delivered, collected, and analysed the study, with no independent external evaluator.
The study spanned only a few days, far short of a full academic year, and Term Duration is not met.
The additional educational materials are the explicit treatment variable tested against a business-as-usual control, so the imbalance is by design.
This novel single-site study has not been independently replicated by another research team.
Only a single custom knowledge domain was assessed, and criterion E is not met, so all-subject exams fails.
Outcomes were measured days after baseline with no tracking to graduation, and criterion Y is not met.
The trial was registered retrospectively in June 2026, long after data collection ended, so it was not pre-registered.
Background: Inappropriate antibiotic use in companion animals contributes to antimicrobial resistance within the One Health context. Educational interventions targeting non-health companion animal owners, particularly undergraduate students who frequently make day-to-day animal care decisions, have remained limited. Methods: A randomized controlled...
Randomization was performed at the individual student level within a single cohort, not at the class or school level, and the intervention is group-based teaching rather than one-to-one tutoring.
Outcomes were measured with a self-report empathy questionnaire (JSPE) and a study-specific custom checklist (SCCG), neither of which is a standardized exam-based achievement assessment.
The intervention ran about one month and outcomes were measured in the final week of the course, so the interval from intervention start to measurement was far shorter than one academic term.
The control group is well documented, with baseline demographics and baseline outcome scores tabulated and the control curriculum described in detail.
This was a single-center study with randomisation at the individual student level, so no school-level randomisation occurred.
The authors designed the study and intervention and also conducted and analysed it, with no independent external evaluation of the trial as a whole.
Follow-up was about one month and criterion T is not met, so the year-duration requirement cannot be satisfied.
The control was an active, time- and dose-matched curriculum receiving equivalent contact time and instructional resources, so the groups were balanced.
This is a single new trial with no independent replication of it reported or identified.
Only empathy and communication were assessed, not all core subjects, and criterion E is not met, so this criterion cannot be met.
Students were assessed only up to the end of the course with no tracking to graduation, and criterion Y is not met.
The authors explicitly state the study was not prospectively registered.
Background: Empathy and effective communication are core competencies in emergency care, yet structured training explicitly targeting these skills is uncommon in undergraduate curricula. We evaluated whether a Nonviolent Communication (NVC) curriculum, a structured framework centered on observation, feelings, needs, and...
Randomization was performed at the individual student level within a single university cohort, not at the class or school level, and no tutoring exception applies.
Learning outcomes were measured with a researcher-built 26-item knowledge test, not a widely recognized standardized exam.
The intervention and outcome measurement spanned only about one to two weeks, far shorter than a full academic term.
Both study arms, including the comparator group, are documented with size, demographics, and baseline characteristics.
Randomization occurred at the individual student level in a single institution, not across schools.
The same authors designed the intervention, ran the study, and analyzed the data, with no independent or third-party evaluator.
The study covered only about one to two weeks, far short of a full academic year, and criterion T is not met.
Both arms received the same pre-class content and comparable in-class time, with the pedagogical modality itself being the treatment variable compared.
No independent replication of this specific study by a different team is reported or identified.
Only a single domain (hypertension pharmacotherapy) was assessed, and criterion E is not met, so A cannot be met.
Measurement stopped immediately after the in-class session with no follow-up to graduation, and criterion Y is not met.
Only an institutional research-ethics approval is reported; no pre-registered protocol on a trial registry is provided.
Background: Active learning methods like flipped classrooms, case-based learning (CBL), and game-based learning (GBL) are increasingly important in medical and pharmacy education. While studies suggest integrating these methods may improve outcomes, direct comparisons of CBL and GBL within flipped classrooms...
The study explicitly used a quasi-experimental design with no randomisation, so no class-level RCT exists.
The outcome was a locally developed department final exam, not a widely recognised standardised assessment.
The intervention spanned only 12 class hours with no term-long interval or follow-up to the measurement.
The control group's size, demographics, baseline scores, and lecture-based condition are clearly documented.
A single-site quasi-experimental study with no randomisation cannot meet school-level RCT requirements.
The same team designed, delivered, and analysed the study, with no independent external evaluation.
A 12-class-hour intervention is far short of a year, and Term Duration (T) was not met.
Instructional time and resources were matched across groups, and the short pre-class microlectures are integral to the flipped design rather than a confounding add-on.
No independent replication of this specific study exists; the authors only call for future verification.
Only physics was assessed and Criterion E was not met, so the all-subject requirement fails.
No graduation tracking is reported and Criterion Y was not met.
Only a vague mention of pre-specified analysis plans is given, with no registry link, ID, or date before data collection.
This study evaluates the effectiveness of integrating the flipped classroom model with problem-based learning (PBL) in undergraduate physics education in China. A total of 200 first-year students were assigned to either a lecture-based control group or a flipped-PBL group (n=100...
Randomization was performed at the individual student level within a single course, not at the class or school level, and no tutoring exception applies.
The primary knowledge and skills outcomes were measured with instruments the researchers developed themselves, not with a widely recognized standardized exam.
The entire intervention and outcome measurement spanned only about two weeks, far shorter than one full academic term.
The control group's size, demographic profile, baseline scores, and condition (virtual lecture only) are clearly documented and shown to be comparable to the intervention groups.
Randomization occurred at the individual student level within one nursing department, not across schools or institutions.
The same two researchers designed the interventions and the assessment instruments and also conducted the study and analysis, with no independent or third-party evaluator.
With a total intervention-to-measurement span of about two weeks (and term duration already unmet), the study falls far short of the year-duration requirement.
The additional engagement time is inseparable from the Kahoot and telesimulation methods that are themselves the treatment being tested against a business-as-usual virtual-lecture control, though the authors acknowledge exposure time differed across groups.
This specific combined game-based learning plus telesimulation intervention has not been independently replicated; the authors state no comparable study exists.
Only the single topic of postpartum hemorrhage was assessed, and the underlying exam-based criterion (E) is not met.
Outcomes were collected immediately after the intervention with no follow-up to graduation, and the prerequisite year-duration criterion (Y) is not met.
The trial was registered retrospectively, after data collection, so it does not meet the pre-registration timing requirement.
Introduction: Recently, new methods such as game-based learning, mobile learning, and telesimulation have been used in lessons. In this study, we aimed to evaluate the effects of game-based learning and telesimulation methods on the knowledge, skills, motivation, and self-efficacy levels...
Randomisation was at the individual student level by lottery within one cohort, not at the class or school level, and no tutoring exception applies.
Outcomes were measured with a custom 14-item DOPS checklist built from the study's own video content and institutional SOP, not a widely recognized standardized exam.
Outcomes were measured only one week after the intervention, far shorter than one academic term.
The control group's population, size, and conditions (equipment familiarization at T1, delayed video at T2) are documented, along with sample demographics.
This was a single-center trial with randomization at the individual student level, not at the school level.
The same anesthesiology team designed the instructional video and conducted, analyzed, and reported the study, with no independent third-party evaluator.
The tracking interval was one week, nowhere near a full academic year, and criterion T is also not met.
The instructional video is the explicit treatment variable being tested, and the control group received a time-comparable active alternative (equipment familiarization) during the same session.
No independent replication of this specific study by a different team in a peer-reviewed journal is reported or evident.
Only a single procedural skill was assessed, not all core subjects, and criterion E is not met, so A cannot be met.
Participants were followed for only one week with no tracking to graduation, and criterion Y is not met.
The paper reports only an ethics reference, not a pre-registered trial protocol with hypotheses and analysis plan filed before data collection.
The integration of digital formats into undergraduate medical education offers a promising approach to enhance procedural skill acquisition. This study investigates the efficacy of a short, structured instructional video as part of a blended learning curriculum for teaching radial artery...
Randomization was performed at the individual-teacher level within the same workshops, not at the class or school level, and no tutoring exception applies.
Outcomes were measured with researcher-developed pre/post vignettes, not with a standardised, widely recognised exam.
The intervention and both measurements occurred inside a single ~100-120 min PD session, with the post-vignette taken after only 70 min, far short of one term.
The comparison (Open Mode) group and all subgroups are documented with demographics, sample sizes, and baseline pre-vignette performance.
Randomization was at the individual-teacher level, not at the level of whole schools or implementing units.
The authors designed the Mastering Math PD program and also conducted, coded, and analysed the trial, with no independent third-party evaluator.
Because the term-duration criterion T is not met (outcomes measured within a single session), the stronger year duration criterion cannot be met.
Both randomised conditions received the identical session structure and an equal 30-min systematization phase, differing only in mode (discussion vs. expert video), so time and resources were balanced.
This is a first-of-its-kind study with no independent replication by a different research team.
Criterion E is not met, and only multiplication-related teacher practices were assessed, so all-subject exam coverage cannot be satisfied.
Criterion Y is not met and there is no follow-up beyond the single session, let alone tracking to graduation.
The paper contains no pre-registration statement, registry ID, or registration date.
Many teachers seek to foster students' conceptual understanding, e.g., for meanings of multiplication, but are only partially prepared to go beyond surface translations between multiple representations. This concerns out-of-field teachers as well as in-field teachers (i.e., with mathematics teaching certificate)....
Randomisation was performed at the individual student (resident) level via lottery, not at the class or school level, and no tutoring exception applies.
Outcomes were measured with study-adapted scoring tables rather than a widely recognised standardised exam.
Outcomes were measured immediately after a short wet-lab training block, well under one academic term, with no term-long follow-up.
The control group's size, demographics, baseline scores, and business-as-usual treatment are clearly documented.
Randomisation was at the individual student level within a single institution, not at the school level.
The same authors designed and delivered the intervention and conducted the assessments, with no independent or third-party evaluator.
The intervention and its measurement spanned far less than a full academic year, and criterion T is not met.
Both groups were explicitly matched on content, class duration, instruments, consumables, and instructor, and any extra pre-class self-study is integral to the flipped-classroom method being tested.
No independent replication of this specific study exists; the authors present the flipped-plus-TBL combination as novel and previously lacking.
Only pterygium excision surgical skill and related theory were assessed, and criterion E is not met.
There was no follow-up to graduation; only immediate post-training outcomes were measured, and criterion Y is not met.
The paper reports ethics approval but provides no pre-registration of the study protocol in any trial registry before data collection.
Background: Standardized ophthalmology residency training emphasizes competency-based education, with fine surgical skills representing a key challenge. However, the current traditional lecture-based teaching program has a long learning curve due to insufficient active participation and practice time. The flipped classroom combined...
Randomization was at the individual student level rather than at the class (or higher) level.
The outcome exams were locally developed (Test A/Test B) rather than a widely recognized standardized exam.
Outcomes were measured immediately after teaching within a 2-week rotation, not at least one academic term after intervention start.
The control group condition and baseline comparability (including group sizes and baseline demographics/scores) are clearly reported.
The study randomized students, not schools or comparable institutional units, so it is not a school-level RCT.
The authors developed the 3D animations and also performed core evaluation activities (including analysis), with no clear independent evaluator.
The intervention and outcome measurement occurred within a 2-week rotation with immediate post-testing, far shorter than 75% of an academic year; also, Y cannot be met if T is not met.
Teaching time/objectives were standardized across groups and the main difference was the instructional medium; any extra post-session access to animations is unlikely to affect the immediate outcome.
No independent, peer-reviewed replication of this specific RCT was identified.
The study did not use standardized exams (so E is not met), and it assessed only renal pathology/nephrology knowledge rather than all core subjects.
The paper reports no long-term follow-up (including no tracking to graduation), and G cannot be met if Y is not met.
The trial was registered retrospectively on 2025-11-18 after the study period (August–October 2025), so it was not pre-registered.
Objectives Teaching renal pathology, characterized by complex spatial relationships and dynamic pathological processes, poses significant challenges. Traditional methods such as static slides often fail to convey these concepts effectively. This study evaluated the efficacy of custom-developed renal pathology three-dimensional (3D)...
Randomisation was at the individual student level within one school with a shared teacher, not at the class or school level, and no tutoring exception applies.
Outcomes were measured with custom sport-specific motor metrics (jump distance and successful jumps), not a recognised standardised exam.
The intervention and outcome measurement spanned only about six weeks, well short of the one-term minimum.
The control group's size, sex composition, age, baseline performance, and standard-PE condition are clearly documented.
Only one school participated and randomisation was at the individual student level, so school-level randomisation is absent.
The intervention was designed, delivered, and analysed by the same team with no independent external evaluator, so independence is not established.
The roughly six-week duration is far below the year-long requirement, and criterion T is also not met.
Both groups received equal regular PE curriculum time, with the GBG being a game format overlaid on existing sessions rather than added time or budget, so the groups are balanced.
The study is presented as novel with no independent replication of this specific GBG long-jump performance trial.
Only a single long-jump motor skill was measured with non-standardised metrics, and criterion E is not met, so A fails.
Measurement ended at the six-week post-test with no graduation tracking, and criterion Y is not met.
The paper reports only local ethics approval and no public trial pre-registration with a registry ID or date.
Purpose: This study investigated the effect of the Good Behaviour Game (GBG) on targeted motor performance metrics during long jump instruction. Methods: In this randomized controlled trial, 40 middle school students (age: 13.10±.40 years) were equally distributed between experimental and...
Randomisation was performed at the individual student level, not at the class or school level, and no personal-tutoring exception applies.
All outcome instruments were custom tools built by the study team for this study, not widely recognised standardised exams.
The intervention was a single ~60-minute session and final measurement occurred only about nine days later, far short of one academic term.
The comparison groups' demographics, sizes, and conditions are documented in Table 1 and the methods, including confirmation of what each arm received.
Randomisation was at the individual student level in a single institution, not at the school level.
The same team designed the instruments and intervention, delivered the sessions, and scored the outcomes, with no independent or third-party evaluator.
Term duration (T) is not met, and the study spanned only about two weeks, so the year-duration criterion cannot be satisfied.
All three arms received a session of comparable scheduled duration, and the differing instructional support is integral to the teaching modalities being compared, which is the study's treatment variable.
This is a novel single-site pilot with no independent replication reported or found.
Exam-based Assessment (E) is not met, and only a single content domain (choking management) was assessed, so the all-subject criterion fails.
Year Duration (Y) is not met and there was no follow-up to graduation, with tracking ending nine days after the intervention.
No pre-registration of the study protocol is reported; only an IRB exemption is noted, and the authors recommend prospective registration for future work.
This pilot randomized controlled trial compared AI-assisted self-directed learning, lecture-based instruction, and simulation-based education in teaching emergency choking management to pre-professional health students. Twenty students (n = 6-7 per group) enrolled in a two-week preparatory program were randomly assigned to...
Randomisation is asserted only by an "R" symbol in what the paper itself calls a symbolic representation of the Solomon design, while the Method text reports convenience sampling of the school and the four intact classes.
Outcomes were a unit-specific researcher-developed achievement test and a Likert attitude scale, not a widely recognised standardised exam.
Post-tests were administered at the end of an eight-week implementation, well short of one full academic term.
Control group size, instructional condition, and CG1 baseline scores with equivalence tests are documented, although no demographic detail is reported.
All four arms were classes within one conveniently selected school, so no school-level randomisation occurred.
The same two authors designed, implemented, collected, analysed and interpreted the study, with no external evaluator or third-party oversight reported.
The complete study window was eight weeks, roughly a fifth of an academic year, and criterion T was not met.
All arms covered the same unit in the same lessons over the same period, and the flipped group's extra out-of-class video time is integral to the model being tested against business-as-usual instruction.
Internet searching found no independent replication of this single-school Solomon four-group study, which was published only in May 2026.
Only one unit of Social Studies was assessed, with no other core subjects measured, and criterion E was not met.
Measurement ended with post-tests at the close of the eight-week implementation, and no follow-up publication tracking this cohort toward graduation could be found.
No pre-registration, registry identifier, or protocol is reported or independently locatable; only an institutional ethics committee approval is documented.
The aim of this study is to examine the effects of the Flipped Classroom Model and the Argumentation-Based Learning Model on sixth-grade students' academic achievement and attitudes toward the Social Studies course. The study was conducted in the spring semester...
Randomisation was at the individual nurse level within one operating room department, not at class or school level, and no tutoring exception applies.
All outcome measures were custom instruments built or adapted by the research team for this study, not recognised standardised exams.
Primary outcomes were measured just one week after a single 4-hour training session, far short of one academic term.
The control group's size, demographics, professional experience and the exact content of the training it received are all documented in detail.
The trial was conducted in a single operating room department with individual nurses randomised, so no school- or site-level randomisation occurred.
The same authors developed the MR module and the outcome instruments and ran and analysed the trial themselves, with only partial assessor blinding as a safeguard.
Outcomes were tracked for only one week, nowhere near 75% of an academic year, and criterion T was not met.
Both arms received an identical 4 hours of structured training including a shared 2-hour lecture, and the only extra resource, the MR hardware and module, is the treatment variable itself.
The study presents itself as the first RCT in this area and internet searching found no independent reproduction of this specific trial.
Criterion E is not met and outcomes covered only the single narrow domain of ACSS scrub nursing, so the all-subject requirement fails.
Tracking stopped one week after training, no follow-up publication by the same authors was found, and criterion Y was not met.
The paper states the trial was not prospectively registered and no registry entry could be located in any trial registry.
Background: Mastery of instrument sequencing and spatial anatomy is crucial for scrub nurses in anterior cervical spine surgery (ACSS), yet conventional training often fails to provide immersive, three-dimensional practice. Aim: This randomized controlled trial evaluated a novel Mixed Reality (MR)...
Randomisation was at the individual student level within intact sections rather than at the class or school level, and the intervention is not personal tutoring.
All outcomes used bespoke, curriculum-aligned instruments (custom STEAM test, study-specific rubric and kinematic proxy) rather than any recognised standardised exam.
Outcomes were measured at the end of an 8-week intervention with no later follow-up, which is shorter than one full academic term.
The control group's size, baseline demographics and pretest scores (Table 1), and instructional conditions are documented in full, with fidelity checks confirming no STEAM elements leaked in.
The trial was single-site with individual-level randomisation, so no randomisation among schools or implementing institutions took place.
The single author both designed the STEAM-RB curriculum and carried out the investigation, analysis, and reporting, with only partial safeguards rather than third-party evaluation.
The tracking interval was only 8 weeks, far below 75% of an academic year, and criterion T was not met.
Class time, instructor experience, facilities, and content coverage were matched across arms, and the only extra resource (low-cost smartphone video feedback) is integral to the intervention being tested.
Internet searching found no independent replication of this STEAM-Roliball trial; the paper presents itself as filling a gap and lists replication in other sports and sites as future work.
Criterion E was not met, and all outcomes were confined to the intervention's own Roliball/STEAM domain with no assessment of other school subjects.
Measurement stopped at the 8-week posttest with no follow-up toward graduation of these first-year undergraduates, no follow-up publication was found, and criterion Y was not met.
The paper explicitly states the trial was not prospectively registered, and no registry record was found by internet search.
University physical education (PE) is increasingly expected to foster physical literacy by integrating movement competence with cognitive understanding and intrinsic motivation. However, rigorous evidence regarding the effectiveness of interdisciplinary approaches, such as STEAM (Science, Technology, Engineering, Arts, and Mathematics), remains...
Randomisation was performed on individual students within each classroom, not on whole classes or schools, and no tutoring exception applies.
Learning outcomes were measured with researcher-built, self-piloted pre- and posttests rather than a recognised standardised examination.
The whole trial, from pretest to posttest, took place inside a single 90-minute lesson, far short of one academic term.
The control group's size, its exact activity, and its baseline motivational and prior-knowledge scores are reported alongside the experimental group.
No schools were randomised; allocation occurred among individual students inside each participating classroom.
The authors developed the digital learning environment, designed the tests, personally taught the lessons, and analysed the data, with no third-party evaluator.
The study ran for a single 90-minute lesson, so the year-long tracking requirement fails, and criterion T was not met either.
Both conditions received an identical workbook and an identical, pre-fixed time schedule, and the only difference - the digital simulation itself - is the explicit treatment variable under test.
Internet searching found no independent replication of this trial by any other research team; the only related studies are by the same author group.
Only a single narrow topic - the 'part of many wholes' fraction concept - was assessed, and criterion E was not met.
Measurement ended minutes after the intervention, no follow-up publication tracking this cohort exists, and criterion Y was not met.
Neither the paper nor any trial registry record shows a pre-registered protocol for this study.
Fractions are challenging but essential for mathematical learning. Educational technology may support students' acquisition of fraction concepts. This study investigates underlying cause-and-effect mechanisms, i.e., whether features, e.g., authentic simulations, have a motivating effect in learning situations-resulting in an indirect learning-promoting...
Randomisation was carried out on stratified pairs of students within each classroom, with all three conditions present in the same class, so the unit of randomisation was below the class level and contamination was not prevented.
The conceptual understanding outcome was measured with a ten-item pretest and posttest assembled and adapted by the researchers for this study, not with a recognised standardised exam.
The whole intervention and outcome measurement took place inside a single 85-minute session, with the posttest administered immediately after the 40-minute knowledge organization phase, far short of one academic term.
The control condition is documented in detail, including its size (n=80), the exact learning activities and eight prompts it received, and a full table of baseline demographics, language proficiency and pretest scores with statistical confirmation of baseline equivalence.
Schools were not randomised at all; the four participating schools and 16 classes merely supplied the sample, and allocation was performed on student pairs within each classroom.
The same team designed the self-learning environment and the instructional videos in their own prior design-research work, and then ran, coded and analysed the trial themselves, with no external evaluator or third-party oversight reported.
The entire trial ran within a single 85-minute session with an immediate posttest, so the tracking interval falls drastically short of 75% of an academic year, and Criterion T is also not met.
All three conditions received exactly the same instructional time within one 85-minute session and an active, content-matched control that worked through a written worked example of the same task plus transfer practice, so educational time and inputs were balanced.
No independent replication of this trial by a different research team is reported in the paper, and dedicated internet searching found only same-author companion publications rather than any reproduction.
Only a narrow slice of mathematics - conceptual understanding of variables as generalizers and algebraic expressions - was assessed, with a researcher-built instrument, so neither the all-subject coverage nor the Criterion E prerequisite is satisfied.
Measurement ended with the posttest inside the same 85-minute session, no follow-up publication by the same authors tracks the cohort to graduation, and Criterion Y is also not met.
The paper contains no reference to a trial registry entry, pre-registered protocol, or registration date preceding data collection, and no registry record was found online.
Self-learning environments with instructional videos are well-established for procedural skill remediation, but their effectiveness for conceptual understanding of complex concepts has been less examined. This study examines the impact of more or less structured prompts in instructional videos on students'...
Randomisation was at the individual student level within a single cohort, not at the class or school level, and the one-to-one tutoring exception does not clearly apply.
All outcome instruments were custom-developed by the authors for this study and not formally validated, so no recognised standardised exam was used.
The intervention and its outcome measurement occurred within a single same-day workshop, far short of the required one-term interval.
The control group's size, baseline scores, and conditions received are documented in the text and Table 1, satisfying the documentation requirement.
Randomisation was at the individual student level within a single faculty, so no school-level randomisation occurred.
The same team designed, delivered, and analysed the trial, with only internal blinding and no independent third-party evaluator.
Outcomes were measured immediately after a one-day workshop with no year-long tracking, and criterion T is not met.
Practice time and the shared lecture/video were balanced across arms, and the AR technology and guided practice are integral treatment variables tested against a business-as-usual control.
No independent replication of this specific trial exists (confirmed by external search); the authors explicitly call for future replication.
Only a single specialised skill was assessed with custom instruments, and criterion E (the prerequisite) is not met.
There was no follow-up beyond immediate assessment and no graduation tracking (no follow-up paper found), and criterion Y is not met.
The paper reports ethics approval and CONSORT reporting but provides no public pre-registration of the protocol before data collection.
(1) Background: Augmented reality (AR) simulation may accelerate psychomotor skill acquisition in clinical education, but comparative evidence is scarce. This three-arm randomized controlled trial compared AR simulation, basic task-trainer simulation, and lecture-based instruction for urinary catheterization training. We hypothesized that...
Randomisation was at the individual student level within a single cohort, not at the class or school level, and the intervention is group teaching rather than tutoring.
Outcomes were measured with a topic-specific research questionnaire (the SKQ), not a widely recognised standardised exam.
The longest follow-up was six weeks, shorter than the one-term minimum required by the criterion.
The control group's size, demographics, baseline scores, and condition are documented in detail in Tables 1 and 2.
Randomisation was at the individual student level in a single university, not at the school/institution level.
The same team designed, delivered, and analysed the intervention, with no independent or external evaluator.
Follow-up was only six weeks, far short of a year, and criterion T was not met.
Both groups received equal time, materials, and hands-on practice, differing only in the instructional method being tested, so resources were balanced.
The study is described as the first of its kind and, after internet searching, no independent peer-reviewed replication was found.
Only one specialised topic was assessed and criterion E was not met, so all-subject assessment fails.
Tracking ended at a single six-week follow-up with no graduation tracking found after internet searching, and criterion Y was not met.
The trial was registered retrospectively, after data collection (registry start 4 March 2025 vs first posted 31 July 2025), so it was not pre-registered.
Aim: This study aimed to evaluate the preliminary effectiveness of a ChatGPT-integrated educational session on nursing students' knowledge acquisition and retention regarding endotracheal suctioning, and to explore their perspectives and learning experiences. Methods: A pilot pre-test/post-test parallel-group randomized controlled trial...
Randomization was performed at the individual student level within a single nursing school, not at the class or school level, and the intervention was not purely one-to-one tutoring.
The knowledge questionnaire and skill checklist were custom-designed by the researchers, not standardized, widely recognized exams.
The intervention consisted of two 2-hour sessions and the final measurement was only one month later, well short of a full academic term.
The control group (n=20) is documented with baseline demographics, baseline scores, and confirmation that it received no study intervention.
Randomization occurred at the individual student level within one nursing school, not at the school or institution level.
The researchers designed the educational content and also delivered and evaluated the intervention, with no independent third-party conduct of the trial.
The intervention and follow-up spanned only about one month, far short of the required 75% of an academic year, and criterion T is not met.
The additional educational time is the treatment variable being tested against a business-as-usual control, and the two experimental arms received the same educational package and equal time.
No independent replication of this specific trial by a different research team in a peer-reviewed journal is reported or identified.
Only mechanical ventilation knowledge and skills were assessed, using custom tools, so the all-subject standardized-exam requirement is not met.
Follow-up ended one month after the intervention with no tracking to graduation, and criterion Y is not met.
Internet verification of the TCTR registry shows the trial was registered retrospectively, after enrollment and after data collection had been completed.
Introduction: Learning mechanical ventilation principles can reduce patient complications. This study examined the effects of task-trainer simulation and peer education on nursing students' knowledge, clinical performance, and action speed related to basic principles of mechanical ventilation (BPMV). Methods: This three-arm...
The study is explicitly a quasi-experiment with best-effort randomization, mixing class-level and student-level assignment and assigning one whole school to the control without randomization.
Outcomes were measured with a custom selection of four Code.org exercises scored by number of attempts, not a recognized standardized exam.
The intervention ran for about one month with the post-test roughly a month after the pre-test, which is shorter than a full academic term.
The paper documents the size, gender composition, mean age, baseline pre-test scores and treatment (regular STEM, no programming) of the control groups.
Randomization was not conducted at the school level; one whole school was assigned to control and others were split by class or student within the school.
The authors designed the adaptive drill-type intervention and the extended Code.org approach and also conducted and analysed the study themselves, with no independent evaluator.
Term Duration (T) is not met and the intervention spanned only about a month, far short of a full academic year.
The active control (CTR) group received the same total instructional time as the experimental group, with the differentiating drill-type content being the treatment variable under test.
No independent replication of this study by a different research team is reported or found.
Only programming performance was assessed via custom Code.org exercises; no standardized all-subject exams were used and criterion E is not met.
Tracking ended at a post-test about a month after the pre-test, with no follow-up to graduation, and criterion Y is not met.
Only institutional ethics approval is reported; there is no pre-registration of the protocol, hypotheses and analysis plan on a public registry before data collection.
Background and Context: Informatics Education research has often called for including the coverage of debugging skills in education, whose acquisition also fosters the learning of programming. Programming and debugging alike are complex tasks, whose mastering requires operating on multiple interconnected...
Randomisation was at the individual intern level (random number table), not at the class or school level, and the intervention is not one-to-one tutoring.
The outcome measures were instructor-developed theoretical and clinical-skill tests and a Mini-CEX rating tool specific to PHN, not widely recognised standardised exams.
The intervention lasted only four weeks with outcomes measured immediately, far short of the one-term requirement, and no term-length follow-up was conducted.
The control group's size, demographics, baseline course grades, and lecture-based conditions are clearly documented in Table 1 and the text.
This was a single-center trial randomising individual interns, with no randomisation at the school or institutional-unit level.
The same authors designed the intervention, delivered the teaching, and analysed the data, so the study was not conducted independently of the intervention designers despite assessor blinding.
The study lasted only four weeks with immediate outcome measurement, far short of an academic year, and Term Duration (T) is also not met.
Both groups received the same content over the same four-week period from the same instructor, and the AI-assisted PBL-CBL method is the integral treatment variable, so resource allocation is balanced.
No independent replication of this specific PHN AI-assisted PBL-CBL trial exists; the cited studies are related but distinct interventions and contexts.
Outcomes covered only the single topic of post-herpetic neuralgia, not all main subjects across the curriculum.
Outcomes were measured immediately after four weeks with no follow-up to graduation, and Year Duration (Y) is also not met.
No trial pre-registration (registry, ID, or date) is reported or locatable externally; ethics approval and CONSORT adherence do not satisfy the pre-registration requirement.
Background: This study aimed to explore the efficacy of a teaching model integrating artificial intelligence-assisted problem-based learning (PBL) with case-based learning (CBL) in standardized clinical teaching of postherpetic neuralgia (PHN) for medical interns. Methods: A total of 120 interns at...
Randomization was at the individual student level (crossover), not at the class (or higher) level, so ERCT class-level randomization is not satisfied.
Outcomes were measured using study-developed MCQs (even though validated) rather than a widely recognized standardized exam.
The primary outcome was measured one week after instruction, which is far shorter than a full academic term.
The control condition and baseline/demographic information are reported, enabling meaningful comparison.
This was not a school-/site-level randomized trial; it was conducted in one institution and randomized individual students.
The paper does not document independent third-party conduct of the trial’s implementation and evaluation.
Outcomes were measured on the scale of a week, not at least 75% of an academic year (and T is also not met).
The control group received additional study time, and the additional faculty-moderated discussion/support is an explicit and integral part of the intervention package being tested.
No independent, peer-reviewed replication of this specific intervention/trial was identified as of the ERCT check date.
Because the study does not use standardized exams (E not met), it cannot meet the all-subject standardized exam requirement.
The study measures only immediate outcomes and does not track participants to graduation (and Y is not met).
A registry ID and registration date are provided, but it cannot be verified from the paper (or accessible registry data) that registration occurred before the first data collection/enrollment.
Dental materials education poses unique challenges due to the complex integration of scientific principles with clinical applications. Traditional teaching methods often fail to promote deep conceptual understanding. This study investigated whether the process of generating multiple-choice questions (MCQs) by students...
Randomization was at the individual student level within a single course, not at the class or school level, and the intervention is not one-to-one tutoring.
The outcome was measured with an author-developed, study-specific terminology test rather than a widely recognised standardised exam.
The intervention-to-measurement interval was only about two months with no delayed follow-up, shorter than one full academic term (~3-4 months).
The control group's size, demographics, baseline scores, and business-as-usual conditions are clearly documented in Table 1 and Fig. 2.
The study randomized individual students at a single institution, with no school-level randomization.
The same team designed, administered, analyzed, and reported the trial, with no independent third-party evaluation of the intervention they created.
The study covered only about two months from start to measurement, far short of a year, and criterion T was not met.
The intervention group received substantial extra study/gameplay time as a supplementary add-on that the business-as-usual control group did not get in equivalent matched form.
The MedQuiz trial is novel and has not been independently replicated by another team; the authors themselves call for future replication.
Only a single subject (medical terminology) was assessed, and criterion E was not met, so all-subject exams cannot be satisfied.
The study measured only immediate outcomes with no follow-up to graduation, no follow-up papers were found, and criterion Y was not met.
The authors explicitly state the study was not registered in any clinical trial registry, and no registry record was found online, so registry pre-registration is absent.
Healthcare students often struggle with learning medical terminology due to its complexity and abstract nature. This randomized controlled trial assessed the effectiveness of MedQuiz, a digital serious game, in enhancing immediate terminology acquisition and user satisfaction among 60 undergraduate students...
The study is observational and did not randomize at the class level.
The study measures course grades rather than using standardized exams.
There is no intervention with outcomes measured after one academic term.
The paper does not document a distinct control group.
No school-level randomization was performed.
The study was conducted by the authors without an independent evaluator.
There is no intervention tracked for a full academic year.
No attempt to balance class time or resources.
The study's findings have been independently replicated by others.
The study does not use all-subject standardized exams.
No graduation tracking is performed.
No pre-registered protocol is referenced.
We model how class size affects the grade higher education students earn and we test the model using an ordinal logit with and without fixed effects on over 760,000 undergraduate observations from a northeastern public university. We find that class...
No randomisation occurred; groups were formed by year of enrollment, so this is not an RCT at the class level.
The study used an internal course final exam plus a satisfaction questionnaire, not a recognised standardised assessment.
The intervention was a 24-class-hour chapter with outcomes taken at course end, with no documented full-term tracking interval.
The control group's size and conditions are documented, though demographics are pooled and no baseline scores are given.
No school-level randomisation occurred; a single institution and teacher with cohort-based assignment.
The same team designed, delivered, assessed, and analysed the intervention with no independent evaluator.
Criterion T is not met and the intervention covered only a 24-class-hour chapter, far short of a year.
The reform group received additional out-of-class online micro-lectures and pre-class tasks (extra time/resources) not matched for the control group, so B is not met.
No independent replication of this specific study by another team exists.
Criterion E is not met and only one pharmacognosy chapter was assessed, so all-subject exams cannot be satisfied.
Criterion Y is not met and outcomes ended at course completion with no graduation tracking.
No pre-registration, registry ID, or protocol date is mentioned anywhere in the paper.
Background There is a gap between the basic knowledge and practice skill in the traditional teaching of pharmacognosy for pharmacy undergraduates. The outcome-based education (OBE), which emphasizes student-centered and results-oriented teaching strategy, aligns with the cultivation demand for comprehensive pharmaceutical...
The study is explicitly a quasi-experiment using intact classes with non-equivalent groups, not a randomised controlled trial assigning classes or schools to conditions.
Outcomes were measured with a researcher-developed Biology Achievement Test, not a widely recognised standardised exam.
The intervention covered only three topics from one term's scheme with outcomes measured immediately after instruction, and no interval of at least a full academic term from start to measurement is documented.
The control group is identified by size and treatment (conventional method) but lacks documented demographic and baseline characteristics establishing comparability.
The study is a quasi-experiment that assigned intact schools to conditions without randomisation, so it does not constitute a school-level RCT.
The same researcher designed the intervention, developed the outcome instrument, and conducted the study, with no independent or third-party evaluation.
Because criterion T (Term Duration) is not met, the stronger Year Duration criterion is automatically not met, and no year-long tracking is documented.
Both groups received normal class periods on the same content; the only difference was the instructional method itself, and the technology-enabled flipped delivery is integral to the intervention being tested.
There is no independent replication of this specific study; cited works are separate studies in different contexts, not reproductions of this trial.
Only biology was assessed, using a non-standardised researcher-made test, so criterion E fails and consequently the all-subject requirement also fails.
Because criterion Y is not met and the study measured outcomes immediately after instruction with no follow-up to graduation, graduation tracking is not satisfied.
There is no mention of any pre-registration of the study protocol on any registry before data collection.
The study investigated the Effect of Flipped Education (FE) on Senior Secondary Two (SS II) Students' Achievement in Biology. The study adopted a quasi experiment of the pre-test post-test non-equivalent control group design. The population was 856 SS II students...
The paper does not describe any randomization at the class level.
No standardized exam-based assessment is implemented in the paper.
No term-long outcome measurement is reported in the paper.
Control group demographics and baseline data are not provided.
No school-level random assignment is executed as part of this paper.
The study was conducted by the intervention's own authors.
No outcomes tracked over a full academic year are provided.
Treatment classes had extra teacher time; controls did not.
No independent replication of the interventions is reported.
Outcomes measured only in targeted subjects, not across all.
No long-term tracking through graduation is provided.
The RCT was preregistered on OSF prior to data collection.
The effect of a reduced pupil–teacher ratio has mainly been investigated as that of reduced class size. Hence we know little about alternative methods of reducing the pupil–teacher ratio. Deploying additional teachers in selected subjects may be a more flexible...
Students (not intact classes or schools) were randomized to conditions, so the class-level (or stronger) randomization requirement is not met.
Outcomes are measured via course interaction and performance proxies (submissions, immediate success, hint ratings), not standardized exams.
The course and outcome tracking span four weeks, which is shorter than a full academic term.
The control condition is described, but baseline demographics and baseline performance needed to document control comparability are not reported.
The randomization occurs among students within a course run, not among schools or other institutional sites.
The paper does not document independent third-party implementation or evaluation separate from the intervention designers/authors.
The study duration is four weeks and therefore does not meet the academic year tracking requirement; it also cannot meet Y because T is not met.
Reflection prompting adds time/effort only in intervention arms, but this added effort is the treatment being tested against a hints-only control.
No independent replication by a different research team in a different context was found.
Because the study does not use standardized exams (E not met), it cannot satisfy the all-subject standardized exam requirement.
Participants were not tracked to graduation, and the study is far shorter than the year-duration prerequisite (Y not met).
No protocol registry link, ID, or registration date is reported, and no pre-registration record was found.
Generative AI tools, such as AI-generated hints, are increasingly integrated into programming education to offer timely, personalized support. However, little is known about how to effectively leverage these hints while ensuring autonomous and meaningful learning. One promising approach involves pairing...
Participants were randomized individually (not by class/school), and the intervention was not one-to-one tutoring.
Outcomes were assessed with an OSCE developed for this study and a short written test, not a widely recognized standardized external exam.
Outcomes were measured immediately at the end of a short course (one day, plus one week of preparatory access), not at least one academic term after the intervention began.
The study explicitly states key baseline demographics and prior ultrasound experience were not collected, and it does not report baseline performance measures for the control group.
The trial was conducted at a single university site and randomized individual students, not schools (or equivalent institutional sites).
The authors created the blended modules and were involved in course implementation and assessments, with no stated independent external evaluation team.
Outcomes were assessed at the end of the curriculum and the study duration is far shorter than 75% of an academic year; also, T is not met.
The blended and conventional conditions provide comparable supervised hands-on resources and intentionally substitute online preparation for face-to-face theory; there is no evidence of a non-integral resource advantage for the intervention group.
No independent replication of this specific trial by a different research team was found or documented at the time of this ERCT check.
E is not met (no standardized external exams), so A cannot be met; the outcomes are limited to ultrasound OSCE and a short written test rather than all-subject standardized exams.
The paper reports no long-term follow-up, and no follow-up publications by the same authors tracking participants to graduation were found; also, Y is not met so G cannot be met.
The study states trial registration was not applicable and explicitly reports it was not pre-registered in a trial registry.
Background: Ultrasonography is an essential clinical tool, offering rapid, bedside, imaging that supports timely clinical decision-making. Its effectiveness, however, depends heavily on examiner skill, requiring structured, practice-oriented training. Traditional tutor-led ultrasound teaching is limited by personnel and resource shortages. Blended...
Randomization was at the student level within one school, not at the class (or higher) level, and the intervention is not described as one-to-one tutoring.
Outcomes were measured with self-report questionnaires/scales (PSQI, ASQ, FFMQ), not standardized exam-based academic assessments.
Outcomes were measured immediately after an 8-week intervention, which is shorter than a full academic term (typically ~3-4 months).
The control group is clearly described (no training) and the paper reports sample sizes, demographics, and baseline comparability across groups.
This is not a school-level RCT because only students (not schools) were randomized and the study took place in one middle school.
The paper does not document an independent external evaluator; "two independent researchers" double-checking data entry does not establish independence from intervention design/delivery.
Outcomes were measured after 8 weeks, far shorter than 75% of an academic year; additionally, per ERCT rules, if T is not met then Y is not met.
The intervention group received substantial additional structured time and activities (weekly 90-min sessions plus daily practice) while the control group received no comparable substitute activity.
No independent replication of this specific study by other authors was found in the paper or via internet search as of the ERCT check date.
Criterion A is not met because Criterion E is not met and the study does not use standardized academic exams across subjects.
Graduation tracking is not reported; additionally, per ERCT rules, if Y is not met then G is not met, and no follow-up paper reporting graduation tracking was found.
No pre-registration link/ID or registration date is reported, and searches did not identify a public pre-registered protocol for this specific trial.
This study evaluated the impact of an 8-week Mindfulness-Based Stress Reduction (MBSR) program on Chinese adolescents' sleep quality and academic stress. Forty-six students were randomly assigned to an experimental group (n=22) receiving MBSR or a control group (n=24) receiving no...
Randomization was at the individual student level (not class- or school-level), and no one-to-one tutoring exception is stated.
Outcomes rely on manually graded EiPE responses and survey items rather than a widely recognized standardized exam.
The activity and measurement occur in a short time window (late in a semester) rather than at least one full term after the intervention begins.
The paper describes what the control group received and gives group sizes, but does not report detailed control group demographics and baseline characteristics as required by ERCT.
Randomization was not conducted at the school (or site) level; it was conducted at the student level in a single university course.
Key measurement and analysis were conducted by the author team, with no clearly described independent third-party evaluation.
The study is a short activity with immediate post measures and does not track outcomes for at least 75% of an academic year (and T is not met).
The treatment adds transparency information and a quiz as the treatment variable being tested; the extra time/inputs are integral to the intervention rather than a confound.
No independent replication of this specific transparency RCT was found in the paper or via external literature search.
Because the study does not use standardized exam-based assessments (E not met), it cannot satisfy the all-subject standardized exams requirement.
The study does not track participants through graduation; additionally, Y is not met, so G cannot be met under ERCT.
No protocol registry link/ID or dated statement indicating pre-registration prior to data collection was found in the paper or via registry-focused search.
The development of effective autograders is key for scaling assessment and feedback. While NLP based autograding systems for open-ended response questions have been found to be beneficial for providing immediate feedback, autograders are not always liked, understood, or trusted by...
Randomization was at the individual child level (not class- or school-level), and the paper does not frame the intervention as a tutoring-style exception.
The main pre/post outcomes are researcher-created homonym measures rather than standardized exam-based assessments.
Post-testing occurred about one week after a short (~2-week) intervention, far shorter than one academic term from intervention start.
The paper clearly describes what the control groups received and reports baseline comparability checks between conditions.
Participants were drawn from multiple schools, but randomization was not conducted at the school level.
The intervention delivery and assignment were carried out by the research team, and no independent evaluation team is documented.
The intervention and follow-up span only weeks, far below 75% of an academic year, and Criterion Y is not met when Criterion T is not met.
Study 2 likely provides more experimenter time per child in the inference condition (individual) than in the control (pairs), and this resource imbalance is not framed as the treatment variable.
No independent replication by a different author team was found or documented; the paper reports two trials conducted by the same research team.
Standardized exams across all core subjects are not used, and per the ERCT dependency rule, Criterion A is not met because Criterion E is not met.
The study does not track participants to graduation, and under ERCT rules Criterion G is not met because Criterion Y is not met.
Although Study 2 is stated to be pre-registered on OSF, the registration record/date could not be verified and the paper does not provide dates to confirm pre-registration occurred before data collection.
Background: Many words have multiple meanings, which present challenges to learning, yet research has yet to identify effective interventions for homonyms. Lexical inference may be a promising strategy. Aim: To evaluate a brief, novel lexical inference intervention for homonyms. Samples:...
Randomization was at the individual student level (within one course cohort), not at the class or school level, and no tutoring exception applies.
Outcomes were measured via self-report questionnaires (and the OSCE was ungraded with no recorded scores), not via standardized exam-based assessment outcomes.
The intervention-to-outcome interval is within the same semester (week 7 OSCE vs week 12 final exam), which is shorter than a full academic term follow-up from intervention onset.
The control condition is described, but the paper does not provide clearly reported control-group baseline characteristics and/or baseline performance separately by group.
The study randomized individual students within one university course, not multiple schools (or equivalent institutions/sites).
The paper does not report an independent external evaluation team; core study roles were performed by the author group.
Outcomes were measured within a single semester (week 7 OSCE to week 12 exam), far short of 75% of an academic year.
Both groups received the same OSCE resources (same format/stations and duration), and the only difference (timing) is the intended treatment variable, so resource imbalance does not confound the intervention effect.
As of the ERCT check date, no independent peer-reviewed replication of this specific randomized timing study was found.
Criterion E is not met (no standardized exam-based outcomes), therefore criterion A cannot be met under the ERCT rules.
The study measured outcomes only through the final exam within the same semester and (since criterion Y is not met) cannot satisfy graduation tracking; no follow-up-to-graduation papers were found.
No pre-registration statement, registry/platform ID, or timing evidence is provided, and no external preregistration record was found by DOI/title searches.
Background With the introduction of the new dental licensing regulations (ZApprO) in Germany, preclinical teaching time was substantially reduced, particularly affecting practical training. To support students’ learning under these conditions, a formative Objective Structured Clinical Examination (OSCE) was implemented early...
Randomisation was at the individual learner level, not classes.
Outcomes used a course-specific final test rather than a recognised standardised exam.
Outcomes were measured within a short-duration course without term-long follow-up.
The control condition is described, but baseline and full control-group characteristics are not documented for most participants.
Randomisation was not conducted at the school (institution) level.
The paper describes the authors implementing the intervention themselves rather than an independent evaluator.
Outcomes were not tracked for a full academic year after the intervention began.
Any added time from writing responses is the treatment itself and is described as minimal.
No independent replication of this specific RCT is identified.
Because E is not met, A is automatically not met.
Because Y is not met, G is automatically not met; no evidence of graduation tracking was found.
The paper provides an OSF link for data and code, but it does not report a pre-registered protocol with a pre-data-collection date.
This study investigates the effectiveness of brief reflection interventions designed to support self-regulated learning in a short, Massive Open Online Course for in-service teachers. Two types of text-based reflection prompts were tested in a randomised controlled trial with over 5,000...
No quote confirms class- or school-level randomisation; assignment to planning conditions is only described as "matched," and the unit involved was individually recruited university volunteers, not classes or schools.
Outcomes were measured with custom-coded temporal speech fluency metrics (speech rate, pausing, repair) derived from an oral narrative task, not with any standardised exam.
The entire study, from planning to the fourth task performance, took place within a single roughly 35-minute session, far short of one academic term.
The control (NP) condition's treatment is briefly stated, but no explicit sample size, demographics, or baseline data specific to that group are quoted in the main text.
Randomisation, to the extent it occurred, was among individually recruited university volunteers at a single institution, not among schools.
The same research team designed the study, collected the data (with two RAs), transcribed and verified the recordings, and conducted the analysis; no independent evaluators are described.
Since criterion T (term duration) is not met, and the whole procedure occurred within a single roughly 35-minute session, criterion Y is necessarily not met.
The extra planning time given to the L1P, L2P, and L1P/L2P groups is itself the treatment variable under investigation, so the NP control group's lack of planning time represents the intended business-as-usual baseline rather than an unaddressed imbalance.
No independent replication of this specific study is mentioned or was found; internet searches confirm this is a very recent (2025) study extending prior work rather than reporting or being the subject of any published reproduction.
Since criterion E (exam-based assessment) is not met, criterion A is automatically not met.
Since criterion Y (year duration) is not met, criterion G is automatically not met, and no follow-up publications tracking this cohort toward graduation were found through internet searches.
No pre-registration of the study protocol, registry platform, or registration date is mentioned anywhere in the paper, and none was found through internet searches of the publisher's page.
This is an investigation of the interplay between collaborative pre-task planning language, task repetition, and L2 proficiency in oral narrative task performance. A total of 128 EFL learners engaged in paired collaborative planning under one of four conditions: L1, L2,...
Randomisation was carried out at the level of individual students via "stratified random assignment", not at the level of classes or schools, and no tutoring exception applies to this group-based teaching method.
The listening test was a researcher/institution-designed instrument built from course textbook materials and a custom rubric, not a widely recognised standardised exam.
Outcomes were measured immediately at the end of an eight-week intervention window, well short of one academic term (~3-4 months).
The control group's condition is described only in vague, one-line terms, with no demographic detail, and the paper's participant counts are internally inconsistent.
No school-level (or class-level) randomisation is described; assignment was of individual students.
Both authors are staff/researchers at the same institution whose students were studied, and there is no mention of an independent evaluator or external data-collection team.
Since Term Duration (T) is not met, Year Duration cannot be met; total tracking (8 weeks) is far short of 75% of an academic year.
Both groups received instruction in parallel over the same eight-week period, with no indication that the intervention arm received extra time or budget beyond substituting the teaching method.
No independent replication of this specific study is reported or found; other authors' Suggestopedia studies are different studies in different contexts, not replications of this one.
Only English listening comprehension was assessed, and Criterion E (which A depends on) is not met.
Since Year Duration (Y) is not met, Graduation Tracking cannot be met, and there is no follow-up beyond the immediate post-test.
No mention of a pre-registered protocol, registry name, or registration date appears anywhere in the paper.
This study investigated the effectiveness of Suggestopedia in improving English listening comprehension competence among 134 Global Citizenship Program Level 3 and 4 students at Swinburne Vietnam over 8 weeks. The research problem was the need to enhance students' English listening...
The paper is an explicitly quasi-experimental study with only two intact classes, one flipped into each condition, which is not a properly implemented class-level RCT.
Writing outcomes were scored with a researcher-adopted six-aspect rubric rather than any standardised, widely recognised exam.
The whole intervention and outcome measurement occurred within roughly two weeks (Weeks 5-7 of the course), far short of one academic term.
There is no genuine control group, and group-specific documentation is limited to baseline draft scores, with demographics reported only for the pooled sample.
Assignment involved two intact classes within a single university; no schools or institutional units were randomised.
The authors designed, delivered, rated, and analysed the intervention themselves, with a member of the research team even providing the Group B feedback; no independent evaluator was involved.
Since the tracking interval was only about two weeks, the 75%-of-an-academic-year requirement is necessarily unmet.
The extra instructor feedback given to Group B is the explicit treatment variable being tested (combined feedback versus ChatGPT-only feedback), so the resource difference between arms is integral to the design.
No independent replication of this specific study exists; the paper cites related but distinct ChatGPT-feedback studies, not reproductions of this trial, and a targeted internet search found no later replication.
Only academic English writing was assessed, with a custom rubric, so neither the all-subject coverage nor the standardised-exam prerequisite (criterion E) is satisfied.
Measurement ended with the revised draft and a Week 7 questionnaire; no participant was tracked to graduation, no follow-up publication was found, and criterion Y is unmet.
The paper contains no mention of any pre-registered protocol, registry, or registration date, and no matching registration was found in registries such as OSF via internet search.
Generative artificial intelligence (GAI) language models, exemplified by ChatGPT, are significantly helping language learners in writing practice by providing immediate formative feedback. This study investigates whether ChatGPT's writing feedback influences graduate students' academic writing abilities and compares it with combined...
Randomisation was done at the individual student level within a single cohort of English majors, not at class or school level, and the intervention was classroom teaching, not one-to-one tutoring, so no exception applies.
Outcomes were measured with a researcher-made 22-item test on the words "join" and "connect", not with any recognised standardised exam.
Outcomes were measured only 3 and 14 days after the brief instruction, far short of the one-term (roughly 3-4 month) follow-up the criterion requires.
The control group is described only by size, major and an unsupported claim of comparable proficiency, with no demographic detail or baseline performance data reported.
The experiment randomised 46 individual students within a single university; no schools or institutional units were randomised.
The sole author designed the metaphor-based teaching approach, conducted the experiment herself, and analysed the results, with no independent evaluator or third-party oversight mentioned.
The study spanned only about two weeks from instruction to the final 14-day delayed test, nowhere near 75% of an academic year; T is not met, so Y cannot be met.
Both groups received instruction on the same 22 target items, with the contrast being teaching method (cognitive-linguistic vs traditional) rather than any extra time, materials or budget given only to the intervention group.
No independent replication of this specific 46-student experiment was found in a fresh internet search; related conceptual-metaphor vocabulary studies by other teams are separate experiments, not replications of this study.
Only knowledge of two English vocabulary items was tested with a custom instrument; no other subjects were assessed and criterion E already fails, so A fails automatically.
Tracking ended 14 days after instruction with no follow-up to graduation, no subsequent papers by this author were found, and criterion Y is not met, so G fails.
The paper contains no mention of any registry, registration ID or pre-registered protocol, and a fresh internet search found no registration for this study.
The "conceptual metaphor" fundamentally shapes human cognition to a significant extent. This paper explores the connotations, evolution, and distinctive features of different types of conceptual metaphor theory. It proposes the integration of metaphor theory into college English vocabulary teaching to...
Randomization was done at the individual student level within a classroom-based (not one-to-one tutoring) intervention, so the class-level requirement is not satisfied.
Outcomes were measured with a speaking test and questionnaire constructed for this study, not a recognised standardised exam.
The intervention and follow-up lasted only six weeks, with post-tests immediately after, which is far shorter than one academic term.
The control group is described only by size and mean test scores, with no demographic or baseline characteristics, and the paper explicitly labels its own descriptive and inferential results tables as fictional and hypothetical.
Randomisation occurred at the individual student level within a single university, with no schools or institutions being randomised.
A single author designed, delivered, and evaluated the intervention with no external or third-party evaluation team.
The whole study lasted six weeks, which is far below 75% of an academic year; since criterion T fails, Y fails as well.
Both groups received classroom instruction and speaking practice over the same six-week window with no documented extra time or budget for the experimental group; the role-play activity is the treatment variable itself rather than a separable added resource.
No independent, peer-reviewed replication of this specific six-week Jordanian role-play trial was found in the paper or via external search; the study was published only in February 2025, leaving little time for replication.
Only speaking fluency in English was assessed with a custom test; no other core subjects were measured and criterion E is not met, which also fails A.
Measurement stopped at the six-week post-test with no tracking to graduation; criterion Y is not met, which also fails G, and no follow-up publications tracking this cohort were found.
The paper contains no mention of any pre-registered protocol, registry, or registration date, and no external registry record for this study was found.
This research examined whether role-play exercises improved Jordanian EFL students' speaking fluency. Fifty intermediate EFL students were studied for six weeks. The experimental and control groups were randomly assigned. Both groups got classroom instruction and speaking practice; however, the experimental...
Randomisation was done at the individual student level within a single school (60 students paired and randomly assigned to two groups), not at the class or school level, and the intervention was whole-group classroom teaching, not tutoring.
Outcomes were measured with a 35-item MCQ test developed by the researcher specifically for this study, not a widely recognised standardised exam.
The intervention lasted only about 4-6 weeks (22 days total) with the post-test administered immediately afterwards, far short of a full academic term.
Beyond stating that the control group was taught by the Direct Method, the paper reports no baseline (pre-test) scores or demographic breakdown for the control group and the group sizes are inconsistently reported (30 vs 20) without explanation.
The study took place in a single school with individual students randomised to groups, so there was no school-level randomisation.
The researcher himself developed the module and the test, delivered the treatment, and conducted the analysis, with no independent third-party conduct or oversight.
The study spanned only about 4-6 weeks from intervention start to final measurement, nowhere near 75% of an academic year, and criterion T is already not met.
Both groups received the same amount of instructional time (35-40 minute sessions over the same 22-day period) on the same 10th-grade PTB English topics, differing only in teaching method, so no extra time or budget was given to either group and time/resources were balanced.
No independent replication of this specific study is reported or found; prior GTM-vs-DM studies cited in the literature review predate this trial and are not replications of it, and an internet search for later citing or replicating studies by independent teams found none.
Only English outcomes were measured with a custom test, no other school subjects were assessed, and criterion E is not met, which automatically fails criterion A.
Measurement stopped at the post-test immediately after the 4-6 week treatment, with no follow-up tracking to graduation and no follow-up publications found via internet search, and prerequisite criterion Y is not met.
The paper contains no mention of any pre-registration, registry, or published protocol for the study, and no registry entry was found via internet search.
The grammar translation method and direct method were compared in this study to see how they impacted students' English learning outcomes at secondary school. These results were obtained through the use of an experimental pre-test and post-test control group design....
Sixty individual students were randomly divided into two groups within one university, so randomisation was at the student level, not the class or school level, and the intervention was group instruction rather than one-to-one tutoring.
Outcomes were measured with an unnamed 100-point vocabulary pre/post test evidently built for the study; the only standardised exam mentioned (CET-4) was used solely to stratify participants, not as an outcome measure.
The paper reports no intervention start date, end date, session schedule, or interval between pre-test and post-test, so a tracking period of at least one academic term cannot be established.
There is no untreated control group, and the two comparison groups are documented only by a bare assertion of no significant baseline differences without any demographic tables, baseline statistics, or numeric data.
Randomisation was at the individual student level within a single university, with no schools or institutional units assigned to conditions.
The author team designed the learning schemes, ran the experiment, and analysed the data under their own student-innovation grant, with no mention of any external or independent evaluator.
No study duration is reported at all, so tracking over at least 75% of an academic year is not documented; criterion T already fails, which by rule also fails Y.
Both arms were active learning conditions teaching the same target vocabulary, with the instructional mode (explicit rule teaching versus implicit contextual input) being the treatment contrast itself, and no extra time or resources are reported for either arm.
This May 2026 paper reports no independent replication of its specific experiment; a targeted internet search found no citing, follow-up, or replication papers for this study.
Only English vocabulary knowledge was assessed with a custom test, no other subjects were measured, and the prerequisite criterion E is not met.
Measurement ended at the immediate post-test with no follow-up or graduation tracking; the prerequisite criterion Y is not met, and no internet search evidence of a follow-up study was found.
The paper contains no mention of any pre-registration, registry platform, registration ID, or protocol published before data collection, and an internet search found no registry entry for this study.
This study explores the effects of explicit and implicit learning strategies on vocabulary acquisition among learners at different English proficiency levels, with the aim of optimizing college English vocabulary teaching approaches. A total of 60 non-English majors were selected as...
Individual students within a single cohort were randomly assigned to the two groups, not entire classes or schools, and the intervention is group classroom teaching, not one-to-one tutoring, so no exception applies.
Outcomes were measured with a researcher-made pre/post-test ("group interview"), with no named, widely recognised standardised exam.
The paper gives no intervention start date, end date, or measurement date, so a term-long interval from intervention start to outcome measurement cannot be established.
Beyond stating that the control group received traditional instruction and reporting its test-score statistics, the paper provides no demographic or baseline characteristics of the control group and even contradicts itself on sample size and composition.
Randomisation was at the individual student level within a single college; no schools or institutional units were randomised.
The same authors designed the cooperative learning intervention, taught/ran the experiment, and analysed the data, with no external or third-party evaluation mentioned.
Since criterion T (term duration) is not met and no dates or durations are reported, the study cannot demonstrate tracking over at least 75% of an academic year.
The intervention swapped the teaching method (cooperative learning vs. traditional instruction) within normal class teaching, and the control group received comparable instruction with no evidence of extra time, budget, or materials given only to the experimental group.
Neither the paper nor a fresh internet search shows any independent peer-reviewed replication of this specific SUST cooperative learning reading trial by a different research team.
Criterion E is not met (no standardised exam), and only EFL reading was assessed, with no measurement of other main subjects.
Criterion Y is not met, and measurement stopped at the post-test with no follow-up tracking of students to graduation, either in the paper or in later publications by the same authors found via internet search.
The paper contains no mention of any pre-registration, registry platform, protocol ID, or registration date, and no registry entry was found for this study on internet search.
This research article is aiming at investigating the Impact of Using Cooperative Learning Strategy in Improving EFL Students' Reading Skill. Subjects were 40 male university students in the English Department, College of Education, SUST. They were randomly assigned into two...
Randomisation was carried out at the individual student level within a single university cohort, not at the class or school level, and the intervention is peer dyad chat rather than one-to-one tutoring, so no exception applies.
Outcomes were measured by researcher-coded error, recast, and uptake counts from chat transcripts, not by any standardised exam.
The entire study consisted of two chat sessions of about one hour each with outcomes measured immediately from those sessions, far short of a full academic term.
The control condition is described procedurally, but no group-specific demographic breakdown or baseline performance data for the control group is reported, so comparability cannot be assessed.
Randomisation occurred among 20 individual students within a single university course, with no school-level or site-level assignment.
The sole author designed the study and the researcher/teacher also delivered the intervention and coded the outcomes, with no independent third-party evaluation.
The study lasted only two one-hour sessions with immediate measurement, nowhere near 75% of an academic year, and criterion T is already not met.
Both groups performed the same tasks in the same two one-hour SCMC sessions, and the only differences (form-focus instructions and teacher prompts) are integral components of the corrective-feedback treatment being tested.
No independent replication of this specific 2026 study is reported or findable; related SCMC feedback studies by other teams are distinct designs, not replications of this trial.
Only EFL linguistic accuracy/uptake was measured via custom transcript coding; no other subjects were assessed and criterion E is not met, which automatically fails A.
Measurement ended immediately after the two chat sessions with no follow-up of any kind, let alone tracking to graduation, and prerequisite criterion Y is not met.
The paper reports ethics approval but contains no mention of any pre-registered protocol, registry, or registration date.
The use of computer-mediated communication (CMC) in language learning, in a variety of forms, has expanded rapidly since the 1990s. It is well documented that computer-mediated communication (CMC) can create a positive learning environment in which learners engage in authentic...
Randomization was performed at the individual student level within a single department, not at the class or school level, and the intervention is simulation skills practice rather than one-to-one tutoring.
Outcomes were measured with a custom 300-point scoring rubric developed by the study team, not a widely recognized standardized exam.
The whole program lasted only four weeks with final outcomes measured at week 4, far short of a full academic term.
The control group is well documented, including size, baseline demographics, and the traditional mannequin training it received.
This was a single-center trial randomizing individual interns, not schools or institutions.
The same authors designed the VR intervention and conducted the trial; blinding of the assessor does not make the conduct independent of the intervention designers.
The program lasted only four weeks, far short of a full academic year, and criterion T is already not met.
The VR group practiced individually with unlimited repetitions while the control group shared one mannequin in groups of five, an unmatched difference in practice intensity that the authors themselves call the most significant threat to internal validity rather than the treatment variable.
No independent replication of this specific trial is reported or found; it is described as a preliminary study needing replication.
Only a single procedural skill (thoracentesis) was assessed with a custom rubric, and since criterion E is not met, A cannot be met.
Follow-up ended at week 4 with no tracking to graduation, and criterion Y is already not met.
The authors explicitly state the trial was not prospectively registered.
Background: Thoracentesis is an essential clinical procedure, but its teaching is often limited by patient safety concerns and insufficient opportunities for repeated practice. Virtual reality (VR) offers immersive, repeatable simulation-based training; however, its effectiveness for thoracentesis has not been rigorously...
The study is a descriptive survey using a questionnaire, not an RCT, and no randomisation of classes or students to conditions was performed.
Outcomes were measured with a researcher-designed Likert-scale questionnaire of self-reported perceptions, not a standardised exam.
The study is a one-time cross-sectional survey with no intervention and no follow-up interval, so no term-long tracking exists.
The study has no control group at all, as it is a single-group descriptive survey, so no control group documentation exists.
No school-level (or any) randomisation to conditions was conducted; this is a non-experimental survey.
The single author designed the instrument, administered it, and analysed the data, with no independent third-party evaluation; and the study is not an intervention trial.
Since criterion T is not met, criterion Y cannot be met; the study has no intervention and no year-long tracking.
There is no intervention group and no control group, so no balancing of educational time or resources can occur.
The study is a single descriptive survey with no independent replication, and there is no reproduction of this specific study by another team.
Criterion E is not met, and the study assesses only self-reported perceptions rather than standardised exams across all main subjects.
Criterion Y is not met, and the study is a one-time survey with no follow-up or tracking to graduation.
No pre-registration of a study protocol on any registry is mentioned anywhere in the paper.
The study examined the influence of Artificial Intelligence (AI) on the academic performance of students in private secondary schools in Ikwo Local Government Area of Ebonyi State. The study was guided by three objectives: to assess the level of awareness...
The study is not randomized at the class (or school) level; the same cohort experienced all modalities in a fixed sequence.
Outcomes are measured using self-report surveys rather than a standardized exam-based assessment.
Measurements are taken immediately around a one-week rotation and pre/post sessions, not at least one academic term after start.
There is no separate documented control group; the design is a single cohort experiencing all three modalities.
The study is conducted within one institution and does not randomize schools/sites to conditions.
Independent third-party conduct is not documented; the authors report conducting the analysis themselves.
The study does not measure outcomes over 75% of an academic year, and term duration is also not met.
There is no control group and the compared modalities differ in format and resources (e.g., one-on-one vs small group), so balanced inputs cannot be established.
No independent peer-reviewed replication of this specific study was found, and the paper itself reports a single-university sample.
Because exam-based assessment (E) is not met, all-subject standardized exams (A) are also not met.
The study is brief and does not track participants to graduation; no follow-up graduation-tracking publications were found.
The paper reports no clinical trial number and describes the work as retrospective, providing no evidence of prospective pre-registration.
Simulation-based education is a crucial element of medical training, providing safe and realistic environments to develop clinical skills and confidence. This study evaluates the effects of three key simulation methods—Standardized Patients (SP), High-Fidelity Simulators (HFS), and Virtual Reality (VR)—on medical...
Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.
Share your research with us for review
Provide feedback to help us make things better.
If you're the author, let us know about necessary updates or corrections.