Effects of a University-Led High-Impact Tutoring Program on Low-Achieving High School Students: A 3-Year Randomized Controlled Trial

Daniel Hamlin, Corey Peltier, Stacy Reeder

Published:
ERCT Check Date:
DOI: 10.3102/00028312261450125
  • mathematics
  • K12
  • US
  • digital assessment
1
  • C

    The study used a two-stage clustered design in which entire remedial-math class sections (not individual students within a class) were randomly assigned to the tutoring or control condition, satisfying class-level RCT.

    "we used a two-stage clustered randomization procedure with the high-impact tutoring treatment being administered at the classroom level within schools" (p. 12)

  • E

    Outcomes were measured with the NWEA MAP mathematics assessment, a widely used, externally validated, standardized computer adaptive test, not a custom instrument.

    "Developed by NWEA, the MAP assessment is a computer adaptive test that is used to evaluate mathematics ability and growth throughout the school year. More than 9,700 schools use the MAP assessment in 145 countries" (p. 14)

  • T

    Outcomes were measured at the end of a full academic year, far exceeding the minimum one-term interval required.

    "This assessment was administered again at the end of the academic year, enabling us to determine the effects of high-impact tutoring on students' mathematics test scores during the school year." (p. 14)

  • D

    The control group is described in detail, including its size, demographic composition, baseline scores, class size, and the specific instruction it received.

    "we compared the outcomes of ninth grade students who were randomly assigned to either a remedial mathematics class providing high-impact tutoring (treatment) or to a remedial mathematics class delivered by a classroom teacher only (control)." (p. 7)

  • S

    Randomisation occurred within each of the seven schools (class sections assigned to condition), not between schools, so this is a class-level rather than school-level RCT.

    "eligible eighth grade students were randomized to remedial mathematics class sections at each high school for the following academic year. These remedial class sections were then randomly selected to be either a high-impact tutoring treatment group or a remedial mathematics class control group" (p. 13)

  • I

    The same University of Oklahoma faculty team that designed and delivered the high-impact tutoring program also ran the data collection and statistical analysis, with no external or independent evaluator.

    "Four faculty and four support staff in the College of Education at the University of Oklahoma participated in the high-impact tutoring program." (p. 7)

  • Y

    Tutoring ran three class periods per week for a full academic year, and outcomes were measured at the end of that same academic year, satisfying the year-duration requirement.

    "In the pooled sample, students (n = 525) in the treatment group participated in high-impact tutoring (i.e., groups of 2:1 or 3:1) three class periods per week for an entire academic year." (p. 3)

  • B

    The control group received the same amount of extra instructional time as the treatment group (a full-year remedial math class); the only difference -- presence of a tutor -- is the explicit treatment variable being tested, so the design is balanced by construction.

    "treatment and control group students within schools received approximately the same amount of extra instructional time for the academic year, with class periods typically being 50-55 minutes in duration." (p. 7)

  • R

    No independent replication of this specific study or program is reported or found; the paper only references separate high-impact tutoring studies with different designs and teams.

    "In previous randomized controlled trials, high-impact tutoring at the high school level has exhibited larger positive effects than the effects estimated in this study (de Ree et al., 2021; Guryan et al., 2021)." (p. 24)

  • A

    Only mathematics achievement was measured with a standardized exam; other core subjects were not assessed, and no exception for a specialised program is stated.

    "The MAP assessment was our primary outcome of interest... the MAP assessment is designed to measure students' growth and achievement in the areas of algebraic thinking, numbers and operations, measurement and data, and geometry" (p. 14)

  • G

    Students were tracked only through the end of ninth grade, not through graduation, and the paper reports no follow-up study.

    "Three separate cohorts of ninth grade students at seven high schools participated in this study from 2021 to 2024." (p. 10)

  • P

    The paper contains no statement of pre-registration, registry link, or registration date anywhere in the text or references.

Abstract

This study uses a randomized controlled trial designed to examine a university-led high-impact tutoring program at seven high schools. The treatment group (n = 525) participated in high-impact tutoring (i.e., groups of 2:1 or 3:1) while the control group (n = 438) attended a remedial mathematics course. The treatment group showed a difference of nearly a half-year of learning (0.13 SD) compared with the control group. We also found no evidence that 2:1 student-tutor groups were more effective than 3:1 groups. Although the university-led program produced strong effects, it was delivered at a high cost. Future work is needed to investigate strategies for reducing the cost of high-impact tutoring while maintaining effectiveness.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • The study used a two-stage clustered design in which entire remedial-math class sections (not individual students within a class) were randomly assigned to the tutoring or control condition, satisfying class-level RCT.
      • "we used a two-stage clustered randomization procedure with the high-impact tutoring treatment being administered at the classroom level within schools" (p. 12)
      • Relevant Quotes: 1) "To estimate causal effects, we used a two-stage clustered randomization procedure with the high-impact tutoring treatment being administered at the classroom level within schools (see Figure 1)." (p. 12) 2) "In late spring, eligible eighth grade students were randomized to remedial mathematics class sections at each high school for the following academic year. These remedial class sections were then randomly selected to be either a high-impact tutoring treatment group or a remedial mathematics class control group taught solely by a classroom teacher." (p. 13) 3) "In year 3 of the study, we included an additional treatment condition within the high-impact tutoring course sections by randomly assigning students to either 2:1 or 3:1 student-tutor group sizes." (p. 13) Detailed Analysis: The randomisation unit for treatment/control assignment was the class section, not the individual student. Students were first randomised to remedial-math sections, and those whole sections were then randomised to the tutoring or control condition, which is exactly the class-level design the criterion requires and avoids within-class contamination. (The subsequent 2:1 vs. 3:1 randomisation in year 3 is a secondary sub-assignment within the tutoring arm and does not affect the primary treatment/control randomisation unit.) Because this is also a personal/small-group tutoring intervention, the standard's tutoring exception would in any case permit even student-level assignment, but the study exceeds that bar by randomising at the class level. Criterion C is met because whole class sections, not individual students within a section, were randomly assigned to treatment or control.
    • E

      Exam-based Assessment

      • Outcomes were measured with the NWEA MAP mathematics assessment, a widely used, externally validated, standardized computer adaptive test, not a custom instrument.
      • "Developed by NWEA, the MAP assessment is a computer adaptive test that is used to evaluate mathematics ability and growth throughout the school year. More than 9,700 schools use the MAP assessment in 145 countries" (p. 14)
      • Relevant Quotes: 1) "The MAP assessment was our primary outcome of interest. Developed by NWEA, the MAP assessment is a computer adaptive test that is used to evaluate mathematics ability and growth throughout the school year." (p. 14) 2) "More than 9,700 schools use the MAP assessment in 145 countries, and NWEA has a 40-year history of developing such assessments." (p. 14) 3) "The reliability evidence (test-retest reliability and marginal reliability) for the MAP is strong. For the ninth grade mathematics assessment, test-retest reliability is .0.9, and marginal reliability is 0.96 (NWEA, 2019). External alignment studies offer further evidence to support the validity of the MAP assessment (Egan & Davidson, 2017)." (p. 14) 4) "These studies have shown that 97% of MAP items are aligned with Common Core State Standards (ninth grade mathematics: r = .72)." (p. 14) Detailed Analysis: The primary outcome measure is the NWEA MAP mathematics assessment, an externally developed and independently validated, widely used standardized test administered in thousands of schools internationally, with documented reliability and alignment to curriculum standards. It was not designed by the study authors for this specific intervention. Criterion E is met because the study relied on a widely recognised, externally validated standardized exam (NWEA MAP) rather than a custom-built assessment.
    • T

      Term Duration

      • Outcomes were measured at the end of a full academic year, far exceeding the minimum one-term interval required.
      • "This assessment was administered again at the end of the academic year, enabling us to determine the effects of high-impact tutoring on students' mathematics test scores during the school year." (p. 14)
      • Relevant Quotes: 1) "At the start of the academic year, all ninth grade students took the NWEA MAP assessment, providing us with baseline achievement data in mathematics. This assessment was administered again at the end of the academic year, enabling us to determine the effects of high-impact tutoring on students' mathematics test scores during the school year." (p. 14) 2) "In the pooled sample, students (n = 525) in the treatment group participated in high-impact tutoring (i.e., groups of 2:1 or 3:1) three class periods per week for an entire academic year." (p. 3) Detailed Analysis: Baseline achievement was measured at the start of ninth grade and the primary outcome at the end of the same academic year, an interval far longer than the one-term minimum required by this criterion. Because this interval also satisfies the stronger Year Duration (Y) criterion, the weaker Term Duration criterion is automatically met as well. Criterion T is met because outcomes were measured a full academic year after the intervention began, well beyond one term.
    • D

      Documented Control Group

      • The control group is described in detail, including its size, demographic composition, baseline scores, class size, and the specific instruction it received.
      • "we compared the outcomes of ninth grade students who were randomly assigned to either a remedial mathematics class providing high-impact tutoring (treatment) or to a remedial mathematics class delivered by a classroom teacher only (control)." (p. 7)
      • Relevant Quotes: 1) "In this study, we compared the outcomes of ninth grade students who were randomly assigned to either a remedial mathematics class providing high-impact tutoring (treatment) or to a remedial mathematics class delivered by a classroom teacher only (control)." (p. 7) 2) "Class sections for both study conditions were capped at 22 students. In the control group, the average class size was 19 students over the 3-year period of analysis." (p. 7) 3) Table 3, "Summary Statistics," reports means, standard deviations, and ranges of baseline and end-of-year math RIT scores, GPA, gender, FRL status, race/ethnicity, ELL status, and special education status separately for the control group sample (n = 438). (p. 15) 4) Table 4, "Comparison of the Baseline Characteristics of the Treatment and Control Groups," reports baseline math scores and demographic proportions for the control group in the pooled sample and each of the 3 years. (p. 16) Detailed Analysis: The paper documents the control condition thoroughly: what instruction control students received (a remedial mathematics class led solely by a classroom teacher, without a tutor), its size (n = 438; average class size 19), and detailed baseline demographic and achievement data broken out by year and pooled sample in Tables 3 and 4. This level of detail allows readers to assess comparability of the control group to the treatment group. Criterion D is met because the control group's composition, size, baseline characteristics, and the instruction it received are clearly and quantitatively documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred within each of the seven schools (class sections assigned to condition), not between schools, so this is a class-level rather than school-level RCT.
      • "eligible eighth grade students were randomized to remedial mathematics class sections at each high school for the following academic year. These remedial class sections were then randomly selected to be either a high-impact tutoring treatment group or a remedial mathematics class control group" (p. 13)
      • Relevant Quotes: 1) "In late spring, eligible eighth grade students were randomized to remedial mathematics class sections at each high school for the following academic year. These remedial class sections were then randomly selected to be either a high-impact tutoring treatment group or a remedial mathematics class control group taught solely by a classroom teacher." (p. 13) 2) "Three separate cohorts of ninth grade students at seven high schools participated in this study from 2021 to 2024." (p. 10) 3) "In year 1, 223 ninth grade students were randomized to treatment and control groups... at one high school where there were three treatment class sections and seven control group sections." (pp. 10, 12; the sentence spans a page break around Table 2 on p. 11) Detailed Analysis: Every participating school hosted both treatment and control class sections; the unit that was randomly assigned to condition was the class section within a school, not the school itself. No school was entirely assigned to one condition or the other. This is precisely the class-level (not school-level) design that the ERCT standard distinguishes at Level 2, so the stronger school-level criterion is not satisfied even though the weaker class-level criterion (C) is. Criterion S is not met because randomisation was conducted at the class-section level within each of the seven schools, not at the level of whole schools.
    • I

      Independent Conduct

      • The same University of Oklahoma faculty team that designed and delivered the high-impact tutoring program also ran the data collection and statistical analysis, with no external or independent evaluator.
      • "Four faculty and four support staff in the College of Education at the University of Oklahoma participated in the high-impact tutoring program." (p. 7)
      • Relevant Quotes: 1) "Four faculty and four support staff in the College of Education at the University of Oklahoma participated in the high-impact tutoring program. Three faculty members were mathematics education scholars while program staff comprised undergraduate and graduate research assistants." (pp. 7-8) 2) "The program's tutoring director organized tutor recruitment and hiring, tutor training sessions, school site coordination, and data management." (p. 8) 3) "Program staff also performed observations of tutors at school sites that were followed by formative feedback on weekly tutor training days." (p. 9) 4) "For the main analyses, we estimated the following regression model in Stata, clustering standard errors at the school level" (p. 18) Detailed Analysis: The intervention (recruitment, training, supervision, curriculum pacing, and data management of the tutoring program) was designed and run entirely by University of Oklahoma faculty and their own program staff. The same faculty authors ("we estimated the following regression model") also performed the outcome analysis. No quote anywhere in the paper describes an external, third-party evaluation team independent of the intervention's designers collecting data or analysing results; data management and analysis were handled by the program's own staff and the authors themselves. Criterion I is not met because the intervention designers, program staff, and the researchers who analysed the results are the same team, with no documented independent or third-party evaluator.
    • Y

      Year Duration

      • Tutoring ran three class periods per week for a full academic year, and outcomes were measured at the end of that same academic year, satisfying the year-duration requirement.
      • "In the pooled sample, students (n = 525) in the treatment group participated in high-impact tutoring (i.e., groups of 2:1 or 3:1) three class periods per week for an entire academic year." (p. 3)
      • Relevant Quotes: 1) "In the pooled sample, students (n = 525) in the treatment group participated in high-impact tutoring (i.e., groups of 2:1 or 3:1) three class periods per week for an entire academic year." (p. 3) 2) "one tutor worked with two or three low-achieving ninth grade students for an entire class period three times a week over the course of the school year." (p. 7) 3) "At the start of the academic year, all ninth grade students took the NWEA MAP assessment... This assessment was administered again at the end of the academic year" (p. 14) Detailed Analysis: The intervention was delivered across the full ninth-grade academic year (not a shorter sub-period), and the primary outcome measure was collected at the very end of that year, matching or exceeding the roughly 9-10 month academic-year benchmark the standard requires. Because Y is met, the weaker Term Duration (T) criterion is automatically satisfied as well. Criterion Y is met because both the intervention and the outcome-tracking window spanned a full academic year.
    • B

      Balanced Control Group

      • The control group received the same amount of extra instructional time as the treatment group (a full-year remedial math class); the only difference -- presence of a tutor -- is the explicit treatment variable being tested, so the design is balanced by construction.
      • "treatment and control group students within schools received approximately the same amount of extra instructional time for the academic year, with class periods typically being 50-55 minutes in duration." (p. 7)
      • Relevant Quotes: 1) "Because both groups were enrolled in the same remedial mathematics course, treatment and control group students within schools received approximately the same amount of extra instructional time for the academic year, with class periods typically being 50-55 minutes in duration." (p. 7) 2) "This control condition allowed us to distinguish the effects of high-impact tutoring from the added instructional time that it offers. It also created a higher threshold for observing program effects than studies using business-as-usual control conditions that do not provide instructional time in mathematics." (pp. 3-4) 3) "The same teacher within schools was usually responsible for leading both treatment and control group sections." (p. 7) 4) "the per-student cost of high-impact tutoring was $5,207, with 84% of total costs arising from tutor compensation. The remedial mathematics course was only $467 per student, so the additional per-student cost of high-impact tutoring to achieve a 0.13 SD effect on mathematics RIT scores was $4,740." (p. 23) Detailed Analysis: Applying the criterion B decision procedure: the intervention group does receive an additional resource relative to control -- a dedicated tutor working with 2-3 students -- which carries a substantially higher per-student cost ($4,740 more). However, the authors explicitly designed the study so that instructional TIME is held constant across arms (same remedial course, same class-period length, often the same teacher), and the paper explicitly frames the presence/absence of a dedicated tutor as the core treatment variable under test ("distinguish the effects of high-impact tutoring from the added instructional time that it offers"). This matches the "resources are the treatment being tested" branch of the decision tree: the extra tutor resource is integral to and is precisely what is being evaluated, while the more easily confounding factor (extra time) is explicitly balanced between arms. The control group is not simply business-as-usual; it is an active, time-matched comparison, which is a stronger design than most studies in this literature. Criterion B is met because instructional time was explicitly balanced between groups, and the remaining resource difference (a tutor) is the explicit treatment variable under study, clearly framed as such by the authors.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study or program is reported or found; the paper only references separate high-impact tutoring studies with different designs and teams.
      • "In previous randomized controlled trials, high-impact tutoring at the high school level has exhibited larger positive effects than the effects estimated in this study (de Ree et al., 2021; Guryan et al., 2021)." (p. 24)
      • Relevant Quotes: 1) "In previous randomized controlled trials, high-impact tutoring at the high school level has exhibited larger positive effects than the effects estimated in this study (de Ree et al., 2021; Guryan et al., 2021)." (p. 24) 2) "de Ree et al. (2021) reported large effects on mathematics achievement (0.44-0.72 SD) for a small sample of 49 high school students participating in high-impact tutoring, but the control group in that study did not receive additional minutes of mathematics instruction." (p. 25) 3) "In a larger analysis, Guryan et al. (2021) found strong positive effects on mathematics achievement when tutored high school students were compared with control group students enrolled in an elective course" (p. 25) Detailed Analysis: The de Ree et al. (2021) and Guryan et al. (2021) studies cited are prior, independent tutoring trials, but they are different programs (different country/context, different control-group designs, different research teams and university), not replications of this specific university-led Oklahoma program or its design. The paper is itself a novel 3-year trial and does not report an independent replication of this particular study. Internet Verification (2026-07-28): The article was published online (AERJ OnlineFirst) on 2026- 06-26 at doi.org/10.3102/00028312261450125. Semantic Scholar's record for this DOI shows a citation count of zero as of this check, and general web searches for replications of this specific university-led Oklahoma tutoring design returned no matches. Given the paper's very recent publication, no independent replication currently exists. Criterion R is not met because no independent replication of this specific study was reported or located, including after a fresh internet search.
    • A

      All-subject Exams

      • Only mathematics achievement was measured with a standardized exam; other core subjects were not assessed, and no exception for a specialised program is stated.
      • "The MAP assessment was our primary outcome of interest... the MAP assessment is designed to measure students' growth and achievement in the areas of algebraic thinking, numbers and operations, measurement and data, and geometry" (p. 14)
      • Relevant Quotes: 1) "The MAP assessment was our primary outcome of interest... In ninth grade, the MAP assessment is designed to measure students' growth and achievement in the areas of algebraic thinking, numbers and operations, measurement and data, and geometry (NWEA, 2019)." (p. 14) 2) "In addition to analyzing the NWEA MAP assessment, we investigated students' grade-point average at the end of the year (scale 0-4.0)." (p. 14) 3) "We originally collected data on student tardiness and absences, but we excluded these indicators" (p. 14) Detailed Analysis: The only standardized exam-based outcome analysed is mathematics (NWEA MAP). GPA was also examined, but GPA is a composite of teacher-assigned grades across courses, not a standardized exam-based assessment of other subjects, so it does not satisfy the "all-subject exams" requirement, which itself depends on criterion E-style standardized assessments in each subject. No standardized exams in reading, science, or other core subjects were reported, and the paper does not offer a specialised/vocational-education rationale (the exception under this criterion) for limiting assessment to mathematics; math remediation for a general ninth-grade population is not such a specialised context. Criterion A is not met because only mathematics was assessed with a standardized exam, with no standardized assessment of other core subjects and no stated exception.
    • G

      Graduation Tracking

      • Students were tracked only through the end of ninth grade, not through graduation, and the paper reports no follow-up study.
      • "Three separate cohorts of ninth grade students at seven high schools participated in this study from 2021 to 2024." (p. 10)
      • Relevant Quotes: 1) "Three separate cohorts of ninth grade students at seven high schools participated in this study from 2021 to 2024." (p. 10) 2) "This assessment was administered again at the end of the academic year, enabling us to determine the effects of high-impact tutoring on students' mathematics test scores during the school year." (p. 14) 3) "Future work is thus needed to test whether remedial high school mathematics courses might be a cost-effective alternative to high-impact tutoring programs." (p. 26) Detailed Analysis: The study design comprises three separate annual ninth-grade cohorts (2021-2024), each followed only for its single ninth-grade year, rather than a single cohort followed longitudinally through to high school graduation. The paper's own discussion of "future work" is framed around cost-effective alternatives and program design, with no mention of an ongoing or planned longitudinal follow-up of these students to graduation. Internet Verification (2026-07-28): Because Y is met, this criterion required an internet check for a possible companion follow-up publication. Given the article's OnlineFirst date of 2026-06-26 and a Semantic Scholar citation count of zero, together with general web searches for "Hamlin," "Peltier," or "Reeder" plus "graduation" or "follow-up," no subsequent tracking study of this cohort was found. Criterion G is not met because outcomes were tracked only through the end of the ninth-grade year in which each cohort was tutored, with no follow-up through graduation reported or located.
    • P

      Pre-Registered

      • The paper contains no statement of pre-registration, registry link, or registration date anywhere in the text or references.
      • Relevant Quotes: 1) No sentence anywhere in the Methods, Data Analytic Strategy, Funding, or Declaration of Conflicting Interests sections mentions a trial registry (e.g., OSF, AsPredicted, ClinicalTrials.gov, AEA RCT Registry) or a pre-registration date. 2) "The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This study was funded with generous support from the Randall and Lenise Stephenson Family Foundation." (p. 34) Detailed Analysis: A full review of the paper, including its funding and disclosure statements and reference list, found no mention of a publicly pre-registered protocol, hypotheses, or analysis plan, nor any registry identifier or timing information that would confirm registration occurred before data collection began. Internet Verification (2026-07-28): Searches of the AEA RCT Registry and OSF Registries for trials by Hamlin, Peltier, or Reeder relating to high-impact tutoring in Oklahoma high schools, plus general web searches for a pre-registration record, returned no matching entries. Criterion P is not met because there is no evidence in the text, or found via internet search, of a pre-registered study protocol.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.