Effects of Heterogeneous Versus Homogeneous Grouping on English Learners' Language and Literacy Development: Evidence from a Randomized Controlled Trial

Michael J. Kieffer, C. Patrick Proctor, Andrew W. Weaver, Sasha Karbachinskiy, Qihan Chen, Qun Yu, Gabriella Solano, Aaron Coleman, Shaelyn M. Cavanaugh, Xiaoying Wu, Elise Cappella, Rebecca D. Silverman

Published:
ERCT Check Date:
DOI: 10.3102/00028312251355989
  • reading
  • language arts
  • L2 languages
  • K12
  • US
0
  • C

    Randomisation was at the individual-student level within classrooms, not at the class or school level, and the small-group (non-tutoring) nature of the intervention does not qualify for the personal-tutoring exception.

    "We randomly assigned individual students classified by their schools as ELs (n = 84) to either homogeneous groups (four ELs) or heterogeneous groups (two ELs and two non-ELs)."

  • E

    The study used ReadBasix, an externally developed and psychometrically validated standardized assessment battery, for its reading comprehension and language skills outcomes.

    "English reading comprehension was measured using the ReadBasix assessment program... developed by researchers at the Educational Testing Service (Sabatini et al., 2019)."

  • T

    Outcomes were measured at the end of an approximately 12-week (about 3-month) intervention period, which approximates the standard's definition of one academic term.

    "All participating ELs received ~12 weeks (18 total hours) of CLAVES..."

  • D

    The paper provides detailed, tabulated demographic and baseline performance data for both comparison groups, confirming they were comparable at pretest.

    "As shown in Table 2, there were no statistically significant differences between the heterogeneous and homogeneous groups at pretest on any variables..."

  • S

    Randomisation occurred at the individual-student level within classrooms; no school-level assignment took place.

    "After selection, participants were individually randomly assigned within self-contained fourth and fifth grade classrooms to homogeneous groups... or heterogeneous groups..."

  • I

    The same research team that designed the CLAVES curriculum also conducted and analyzed this trial, with no independent third-party evaluator described.

    "CLAVES was designed over a 2-year period in collaboration with multilingual students and their teachers..." (curriculum designed by co-authors of this study, per Proctor et al., 2021)

  • Y

    The intervention and outcome measurement spanned only about 12 weeks, far short of the 75%-of-an-academic-year threshold required for this criterion.

    "Teachers were asked to implement CLAVES for 30 minutes a day three times a week for 12 weeks, or a total of 18 hours of instruction."

  • B

    Both comparison groups received identical curriculum content and dosage from the same teacher, so resources were balanced between conditions.

    "Importantly, they reported the same dosage for their heterogeneous and homogeneous groups; the median time per day and the pattern of days per week were reported to be the same for the two groups."

  • R

    No independent replication of this specific grouping experiment by a different research team was found in the paper or in a search of subsequent citing literature.

  • A

    Only language and literacy outcomes were assessed, with no coverage of other core subjects and no applicable exception.

    "We evaluated the main effects of grouping on proximal language skills..., an intermediate measure of core analytic language skills..., and more distal measures of reading comprehension and argumentative writing."

  • G

    There is no evidence of any follow-up tracking beyond the immediate posttest for this cohort, a search of subsequent publications by the author team found no graduation follow-up of this sample, and the prerequisite Year Duration criterion is also not met.

  • P

    The study was pre-registered with a documented registry ID, and its pre-specified hypotheses and analyses are referenced throughout the paper; the registry itself could not be independently queried to confirm the exact date.

    "The study was preregistered at the Registry of Efficacy and Effectiveness Studies (https://sreereg.icpsr.umich .edu/sreereg/) as #8020.1v2."

Abstract

In this preregistered within-teacher randomized controlled trial (n = 84), we tested the effects of grouping English learners (ELs) in homogeneous groups (all ELs) versus heterogeneous groups (ELs and non-ELs) on language, reading comprehension, and argumentative writing. Findings indicated no significant main effects of grouping. However, preregistered moderation analyses indicated that heterogeneous groups benefited students with higher English language skills (Hedges' g = 0.27-0.59 or 0.75-1.93 grade equivalents), whereas homogeneous groups benefited students with lower English skills (g = 0.31-0.58 or 1.00-1.55 grade equivalents). Instructional observations indicated that teachers provided more specialized strategies for ELs in homogeneous groups and more authentic questions for students in heterogeneous groups. Findings question the default use of homogeneous grouping and support considering English proficiency when making instructional and policy decisions for EL instruction.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was at the individual-student level within classrooms, not at the class or school level, and the small-group (non-tutoring) nature of the intervention does not qualify for the personal-tutoring exception.
      • "We randomly assigned individual students classified by their schools as ELs (n = 84) to either homogeneous groups (four ELs) or heterogeneous groups (two ELs and two non-ELs)."
      • Relevant Quotes: 1) "In this preregistered within-teacher randomized controlled trial (n = 84), we tested the effects of grouping English learners (ELs) in homogeneous groups (all ELs) versus heterogeneous groups (ELs and non-ELs) on language, reading comprehension, and argumentative writing." (Abstract, p. 909) 2) "Students (n = 84) were individually randomly assigned to within-teacher groups, with teachers providing instruction to both the homogeneous and heterogeneous groups." (p. 912) 3) "We randomly assigned individual students classified by their schools as ELs (n = 84) to either homogeneous groups (four ELs) or heterogeneous groups (two ELs and two non-ELs)." (p. 916) 4) "After selection, participants were individually randomly assigned within self-contained fourth and fifth grade classrooms to homogeneous groups (n groups = 14; n ELs = 56) or heterogeneous groups (n groups = 14; n ELs = 28), as shown in Figure 1." (p. 922) Detailed Analysis: The unit of randomisation in this study is unambiguously the individual student. Students within the same self-contained classroom, taught by the same teacher, were individually randomly assigned to a four-EL homogeneous group or a two-EL/two-non-EL heterogeneous group. This is explicitly a within-classroom, within-teacher design; no whole classes or schools were randomised to conditions. The ERCT exception for Criterion C applies only to interventions "designed for personal teaching like tutoring," where student-level randomisation is acceptable even without class- or school-level assignment. Here, the intervention is delivered to small groups of four students at a time (either four ELs, or two ELs plus two non-ELs), not one-to-one tutoring, so this exception does not clearly apply. Because both conditions being compared are active instructional groups delivered by the same teacher in the same classroom, cross-contamination between "treatment" and "control" (in the classic sense) is less of a concern than in a typical treatment-vs-no- treatment RCT. Nonetheless, per the letter of the ERCT standard, the unit of randomisation must be the class or school (absent the tutoring exception), and here it was the individual student. Final: Criterion C is not met because randomisation occurred at the individual-student level within single classrooms, and the tutoring exception does not apply to this small-group (not one-to-one) intervention.
    • E

      Exam-based Assessment

      • The study used ReadBasix, an externally developed and psychometrically validated standardized assessment battery, for its reading comprehension and language skills outcomes.
      • "English reading comprehension was measured using the ReadBasix assessment program... developed by researchers at the Educational Testing Service (Sabatini et al., 2019)."
      • Relevant Quotes: 1) "English reading comprehension was measured using the ReadBasix assessment program (previously known as the Reading Inventory and Scholastic Evaluation or RISE) developed by researchers at the Educational Testing Service (Sabatini et al., 2019)." (p. 924) 2) "ReadBasix is a web-administered reading skills componential battery designed with vertical scales for children in grades 3-12." (p. 924) 3) "The technical manual for ReadBasix reports an acceptable (IRT) marginal reliability of .753 for grade 4 and .674 for grade 5 of the norming sample (Sabatini et al., 2019)." (p. 924) 4) "The English vocabulary subtest of ReadBasix presents students with a target word and asks them to choose the most closely related word from three options... The technical manual for ReadBasix reports an IRT marginal reliability of .832 to .867 for grades 4 and 5 of the norming sample." (p. 924) Detailed Analysis: The study's key distal outcome, reading comprehension, was measured with ReadBasix, a standardized, norm-referenced battery developed by researchers at the Educational Testing Service and validated with published technical manuals and reliability statistics across grades 3-12. ReadBasix was used at both pretest and posttest for reading comprehension, and it is a widely used, externally developed standardized assessment, not a custom test built solely for this study. Other outcomes (CLAVES language skills factor, argumentative writing rubric, CALS-I) are more proximal/researcher- or field-developed instruments, but the criterion requires that a standard exam-based assessment be used somewhere in the outcome battery, which is satisfied by the ReadBasix reading comprehension and vocabulary/morphology subtests. Final: Criterion E is met because the study used ReadBasix, an externally developed, standardized, psychometrically validated assessment battery, for its reading comprehension, vocabulary, and morphology measures.
    • T

      Term Duration

      • Outcomes were measured at the end of an approximately 12-week (about 3-month) intervention period, which approximates the standard's definition of one academic term.
      • "All participating ELs received ~12 weeks (18 total hours) of CLAVES..."
      • Relevant Quotes: 1) "All participating ELs received ~12 weeks (18 total hours) of CLAVES, with half randomly assigned to homogeneous groups (consisting of four ELs) and half assigned to heterogeneous groups..." (p. 912) 2) "Teachers were asked to implement CLAVES for 30 minutes a day three times a week for 12 weeks, or a total of 18 hours of instruction." (p. 922) 3) "They reported teaching a median of 31 lessons with some variability (SD = 5.20 lessons; minimum = 20; maximum = 35)." (p. 923) Detailed Analysis: The intervention was designed to run for approximately 12 weeks (about 2.75-3 months), with posttest measures administered at the conclusion of this period. The ERCT standard defines a term as "approximately 3-4 months." Twelve weeks sits at the low end of this range and is consistent with a school trimester/quarter-length term in many US districts. There is no explicit statement of a longer gap between intervention start and outcome measurement beyond the instructional period itself, but the approximately 12-week span of active tracking from intervention onset to posttest is reasonably close to the standard's "approximately 3-4 months" definition of a term. Final: Criterion T is met, on balance, because the approximately 12-week (about 3-month) span from intervention start to posttest measurement approximates the standard's definition of one academic term.
    • D

      Documented Control Group

      • The paper provides detailed, tabulated demographic and baseline performance data for both comparison groups, confirming they were comparable at pretest.
      • "As shown in Table 2, there were no statistically significant differences between the heterogeneous and homogeneous groups at pretest on any variables..."
      • Relevant Quotes: 1) "Table 2. Demographics and pretest descriptive statistics, overall and by group (n = 84)" reporting gender, home-language background, reading comprehension, decoding, vocabulary, morphology, English language skills composite, and argumentative writing separately for "Homogeneous mean (SD)(n = 56)" and "Heterogeneous mean (SD) (n = 28)." (p. 921, Table 2) 2) "As shown in Table 2, there were no statistically significant differences between the heterogeneous and homogeneous groups at pretest on any variables, including gender (p = .609), home-language background (all p values > .10), reading comprehension (p = .274), word reading (p = .694), or English language skills (p = .600)..." (p. 928) 3) "Following selection, student participants included 84 fourth and fifth graders classified as ELs (41 girls; see Table 2). On family surveys, the most common home language was Mandarin or Cantonese (61%), followed by Spanish (21%)..." (p. 919) Detailed Analysis: Although this study does not have a classic untreated "control" arm (both homogeneous and heterogeneous groups receive the same CLAVES curriculum), the comparison group in each contrast (e.g., the homogeneous condition when heterogeneous is the focal comparison, or vice versa) is extensively documented. Table 2 provides detailed demographic and pretest performance data (gender, home-language background, reading comprehension, word recognition, vocabulary, morphology, English language skills composite, and argumentative writing) broken out separately by group, along with significance tests confirming baseline equivalence. This level of detail satisfies the intent of Criterion D, which requires that the comparison group's composition, baseline characteristics, and conditions be clearly documented so that comparability can be assessed. Final: Criterion D is met because the paper provides detailed demographic and baseline performance data for both comparison groups, confirming their comparability at the outset of the study.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred at the individual-student level within classrooms; no school-level assignment took place.
      • "After selection, participants were individually randomly assigned within self-contained fourth and fifth grade classrooms to homogeneous groups... or heterogeneous groups..."
      • Relevant Quotes: 1) "We randomly assigned individual students classified by their schools as ELs (n = 84) to either homogeneous groups (four ELs) or heterogeneous groups (two ELs and two non-ELs)." (p. 916) 2) "After selection, participants were individually randomly assigned within self-contained fourth and fifth grade classrooms to homogeneous groups (n groups = 14; n ELs = 56) or heterogeneous groups (n groups = 14; n ELs = 28)..." (p. 922) 3) "Six elementary schools in a large urban school district in the northeast of the United States participated in this research." (p. 916) Detailed Analysis: Six schools participated, but schools themselves were not the unit of randomisation; they were simply the recruitment sites. Within each participating classroom, individual students were randomised to homogeneous or heterogeneous small groups. There is no quote indicating that whole schools (or even whole classes) were assigned wholesale to one condition or the other; instead, every participating classroom (and by extension every school) contained both homogeneous and heterogeneous groups simultaneously. Final: Criterion S is not met because randomisation occurred at the individual-student level within classrooms, not at the school level.
    • I

      Independent Conduct

      • The same research team that designed the CLAVES curriculum also conducted and analyzed this trial, with no independent third-party evaluator described.
      • "CLAVES was designed over a 2-year period in collaboration with multilingual students and their teachers..." (curriculum designed by co-authors of this study, per Proctor et al., 2021)
      • Relevant Quotes: 1) "CLAVES was designed over a 2-year period in collaboration with multilingual students and their teachers, with the goal of creating a language-based literacy curriculum..." (p. 923), citing "Proctor et al., 2021," where C. Patrick Proctor is a co-author of this paper. 2) "Prior evidence suggests that CLAVES is efficacious at promoting ELs' language and reading comprehension (Proctor et al., 2020) as well as writing (Silverman et al., 2021)." (p. 914), both of which are studies by co-authors of the present paper (Proctor, Silverman). 3) "We thank Tyara Dabrio, Rachel Hodes, Audrey McMaster, Dee Perry, Aimee Salgado, Summer Wu, and Mayee Yeh for their essential contributions to data collection." (Acknowledgments, p. 940) 4) "This research was funded by the Institute of Education Sciences, U.S. Department of Education Award No. R305A200069." (p. 940) Detailed Analysis: The CLAVES curriculum used in this study was designed by several of the same authors who conducted this trial (notably Proctor and Silverman, who also authored prior CLAVES efficacy studies). The data collection was carried out by research assistants thanked in the acknowledgments, but there is no indication that these individuals, or any party, constituted an independent third-party evaluation team separate from the authors' own research program. External funding from IES supports financial independence but does not establish independent conduct of the evaluation itself, which the criterion requires. There is no statement anywhere in the paper of an external, independent evaluator or agency handling data collection, analysis, or interpretation apart from the curriculum-developing research team. Final: Criterion I is not met because the same research team that designed the CLAVES curriculum also designed, implemented, and analyzed this trial, with no independent third-party evaluation described.
    • Y

      Year Duration

      • The intervention and outcome measurement spanned only about 12 weeks, far short of the 75%-of-an-academic-year threshold required for this criterion.
      • "Teachers were asked to implement CLAVES for 30 minutes a day three times a week for 12 weeks, or a total of 18 hours of instruction."
      • Relevant Quotes: 1) "All participating ELs received ~12 weeks (18 total hours) of CLAVES..." (p. 912) 2) "Teachers were asked to implement CLAVES for 30 minutes a day three times a week for 12 weeks, or a total of 18 hours of instruction." (p. 922) 3) "This study was conducted amid the COVID-19 pandemic, which constrained recruitment and led to a more modest sample size than we intended." (p. 938) Detailed Analysis: The Y criterion requires outcome tracking spanning at least 75% of a full academic year (roughly 9-10 months). The intervention and its associated outcome measurement here spanned only about 12 weeks (roughly 3 months), dramatically shorter than the required threshold. There is no mention of any longer-term follow-up assessment beyond the immediate posttest at the end of the 12-week implementation period. Final: Criterion Y is not met because the study's active intervention and measurement period (~12 weeks) falls far short of the required 75% of an academic year.
    • B

      Balanced Control Group

      • Both comparison groups received identical curriculum content and dosage from the same teacher, so resources were balanced between conditions.
      • "Importantly, they reported the same dosage for their heterogeneous and homogeneous groups; the median time per day and the pattern of days per week were reported to be the same for the two groups."
      • Relevant Quotes: 1) "Thus, in this study, we controlled for instruction by using CLAVES, a small-group literacy curriculum, and provided it to all students in the heterogeneous and homogeneous groups." (p. 914) 2) "Teachers were asked to implement CLAVES for 30 minutes a day three times a week for 12 weeks, or a total of 18 hours of instruction." (p. 922) 3) "Importantly, they reported the same dosage for their heterogeneous and homogeneous groups; the median time per day and the pattern of days per week were reported to be the same for the two groups. Teachers also reported finishing CLAVES on the same lesson for both groups." (pp. 922-923) Detailed Analysis: Both the homogeneous and heterogeneous conditions received the identical CLAVES curriculum, from the same teacher, for the same reported dosage (median 30-40 minutes per session, same number of sessions per week, finishing at the same lesson). There is no imbalance in instructional time, materials, or teacher attention documented between the two conditions; the only manipulated variable is group composition (EL/non-EL mix), not the amount or type of instructional resources provided. Applying the ERCT criterion B decision procedure: no extra time or budget is present for one condition versus the other (both receive identical CLAVES dosage), so this criterion is trivially satisfied by resource parity between the two arms being compared, without needing to reach the treatment-variable or within-subjects branches of the decision tree. Final: Criterion B is met because both grouping conditions received identical instructional time, curriculum content, and teacher-reported dosage.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific grouping experiment by a different research team was found in the paper or in a search of subsequent citing literature.
      • Relevant Quotes: 1) "In this study, we did not seek to establish further evidence of the efficacy of the curriculum. Instead, we chose CLAVES as a promising baseline instructional model for testing the effects of grouping, because (a) it has preliminary evidence of efficacy (Proctor et al., 2020; Silverman et al., 2021)..." (p. 914) 2) "Kieffer, M. J., & Weaver, A. W. (2024). Classroom concentration of English learners and their reading growth. Educational Researcher..." (References, p. 943) -- a related but distinct correlational study by two of the same authors, not an independent replication of this specific grouping RCT. Detailed Analysis: The paper references prior studies of the CLAVES curriculum's general efficacy (Proctor et al., 2020; Silverman et al., 2021), but these are studies of the curriculum itself, not independent replications of this specific experimental comparison of homogeneous versus heterogeneous grouping. No other research team is cited as having replicated this specific grouping experiment in a different context. An internet search of the citation record for this paper (via OpenAlex, DOI 10.3102/00028312251355989) identified five citing works as of this check: (a) Weaver, Kieffer, Proctor, Cappella, Wu, Cavanaugh, and Karbachinskiy (2025, Reading Research Quarterly), which reuses data from this same RCT to study peer social ties rather than replicate the grouping comparison; (b) Black, Le, Abbady, Romano, Carlson, Miciak, Francis, and Kieffer (2025, Peabody Journal of Education), which shares an author with this study but examines a distinct topic (linguistic course concentration) in unrelated cohorts recruited 2013-2015, not this grouping intervention; (c) Chen, Kuo, Chang, Huang, and Du (2026, Cognitive Development), an unrelated Mind-Mapping intervention study; (d) Acosta (2026), a theoretical paper on linguistic heterogeneity; and (e) Wu, Kieffer, and Proctor (2026, Reading and Writing), a same-author-team study of morphological awareness. None of these is an independent research team's replication of this specific heterogeneous- versus-homogeneous grouping experiment. Final: Criterion R is not met because no independent replication of this specific grouping experiment was found in the paper or in a search of the subsequent citing literature.
    • A

      All-subject Exams

      • Only language and literacy outcomes were assessed, with no coverage of other core subjects and no applicable exception.
      • "We evaluated the main effects of grouping on proximal language skills..., an intermediate measure of core analytic language skills..., and more distal measures of reading comprehension and argumentative writing."
      • Relevant Quotes: 1) "We evaluated the main effects of grouping on proximal language skills (i.e., vocabulary and morphology taught in CLAVES), an intermediate measure of core analytic language skills (Uccelli et al., 2015), and more distal measures of reading comprehension and argumentative writing." (p. 916) 2) "This study has some limitations to note... We anchored our study with a specific curriculum, which allowed us to control for curriculum across groups and increase internal validity but lessened external validity to the extent that findings may be less generalizable to instruction with substantially different curricula." (p. 939) Detailed Analysis: All outcomes measured in this study (CLAVES language skills, core analytic language skills, reading comprehension, argumentative writing) fall within the single domain of English language and literacy. No outcomes in mathematics, science, or other core subjects were assessed. The exception for All-subject Exams applies to "highly specialised interventions in upper secondary or vocational education," which does not describe this elementary (grades 4-5) small-group literacy intervention, so the exception does not apply. Final: Criterion A is not met because the study measured outcomes only in the language/literacy domain, with no assessment of other core subjects and no qualifying exception.
    • G

      Graduation Tracking

      • There is no evidence of any follow-up tracking beyond the immediate posttest for this cohort, a search of subsequent publications by the author team found no graduation follow-up of this sample, and the prerequisite Year Duration criterion is also not met.
      • Relevant Quotes: 1) "All participating ELs received ~12 weeks (18 total hours) of CLAVES..." (p. 912) -- the entire study timeline, from pretest through posttest, is confined to this ~12-week implementation window. 2) "This study has some limitations to note. This study was conducted amid the COVID-19 pandemic, which constrained recruitment and led to a more modest sample size than we intended." (p. 938) -- the limitations section does not mention any longer-term or graduation-tracking follow-up. 3) "Future research, including experimental or quasi-experimental studies with rich qualitative components, could shed further light on the issue of how to group ELs within schools." (p. 940) -- framed as a suggestion for future work, not evidence that follow-up tracking occurred. Detailed Analysis: There is no mention anywhere in the paper of tracking participants beyond the immediate posttest at the end of the 12-week intervention, let alone through graduation. An internet search for subsequent publications by this author team (Kieffer, Proctor, Weaver, Silverman, et al.) that might track this same 84-student EL cohort toward graduation did not identify any such follow-up study. The citing literature located via OpenAlex includes a paper on graduation and college-enrollment tracking (Black et al., in press, Peabody Journal of Education), but that study follows different cohorts (recruited 2013-2015) examining linguistic course concentration, not the participants or intervention from this CLAVES grouping RCT. No follow-up publications tracking this specific cohort were found. Per the standard's dependency rule, because Criterion Y (Year Duration) is not met, Criterion G cannot be met either. Final: Criterion G is not met because there is no evidence of tracking beyond the immediate posttest, no follow-up publication tracking this cohort toward graduation was found in a search of the citing literature, and Criterion Y (a prerequisite) is also not met.
    • P

      Pre-Registered

      • The study was pre-registered with a documented registry ID, and its pre-specified hypotheses and analyses are referenced throughout the paper; the registry itself could not be independently queried to confirm the exact date.
      • "The study was preregistered at the Registry of Efficacy and Effectiveness Studies (https://sreereg.icpsr.umich .edu/sreereg/) as #8020.1v2."
      • Relevant Quotes: 1) "In this preregistered within-teacher randomized controlled trial (n = 84)..." (Abstract, p. 909) 2) "The study was preregistered at the Registry of Efficacy and Effectiveness Studies (https://sreereg .icpsr.umich.edu/sreereg/) as #8020.1v2." (p. 916) 3) "In addition, we explored a priori, preregistered questions about how the effects of grouping differed by students' pretest English language skills..." (p. 916) 4) "As indicated in the preregistration, power analyses indicated a minimum detectable effect size for the main effect of grouping of 0.33." (p. 922) 5) "Given the exploratory nature of this question, we did not specify an adjustment for multiple comparisons in our preregistration." (Note 5, p. 941) Detailed Analysis: The paper explicitly documents pre-registration of the trial's hypotheses, design, and planned analyses at the Registry of Efficacy and Effectiveness Studies, with a specific registration ID (#8020.1v2) provided. Multiple references throughout the paper (confirmatory research question 1, preregistered moderation analyses for research question 2, and preregistered power analysis details) confirm that the design and primary/moderation analyses were specified in advance. No quote in the paper suggests that registration occurred after data collection began. An attempt was made to independently verify the registration date by querying the SREE Registry of Efficacy and Effectiveness Studies directly for entry #8020.1v2; the registry's search interface requires interactive/authenticated access and returned no results through automated web search, so the exact registration date could not be independently confirmed from the registry itself. However, the paper's internal evidence (references to preregistered confirmatory and moderation analyses woven consistently throughout the methods and results, with no indication of post hoc registration) supports that registration preceded data collection. Final: Criterion P is met because the study was pre-registered with a documented registry ID and the pre-specified hypotheses and analyses are referenced throughout the paper, though the exact registration date could not be independently confirmed via the registry itself due to access restrictions.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.