Implementation and Impact of a Summer Mathematics Bridge Program for Multilingual Learners: Evidence from Randomized Controlled Trials

Haiwen Chu, Leslie Hamburger, Jill Neumayer DePiper

Published:
ERCT Check Date:
DOI: 10.3390/educsci16050796
  • mathematics
  • K12
  • US
  • digital assessment
0
  • C

    Randomisation was at the individual student level, not the class or school level, and the study is not a one-to-one tutoring intervention.

    "For the impact study, we implemented two experiments with assignment at the student level..."

  • E

    The study used the standardised, externally developed MIRA readiness assessment and the state SBAC test rather than a custom-built measure.

    "The outcome measure for the study was the Mathematics Initiative Readiness Assessment (MIRA), a reliable measure of readiness for high school mathematics..."

  • T

    The intervention and outcome measurement spanned only three weeks, far short of the one-term requirement, with no term-long follow-up.

    "For the intervention group (first cohort), the outcome variable is the post-test on the MIRA which was taken at the end of the three-week program."

  • D

    The comparison groups' demographics, baseline scores, and conditions are documented in the text and in Table 1.

    "Table 1. Baseline characteristics of samples in Districts A and B."

  • S

    Randomisation was at the individual student level, not at the school or site level.

    "We estimate the impact of the intervention, which was individually assigned to students..."

  • I

    The same authors who designed the RAMPUP intervention also conducted and analysed the evaluation, with no independent evaluator.

    "Author Contributions: Conceptualization, H.C. and L.H.; methodology, H.C. and J.N.D.; software, H.C.; investigation, H.C. ..."

  • Y

    The study lasted only three weeks, far short of a full academic year, and criterion T was not met.

    "We report results from the implementation and impact evaluations of a three-week summer bridge mathematics program."

  • B

    District A's comparison was a time-comparable competing summer program, and District B tested the summer program's added time as the integral treatment variable against a business-as-usual baseline.

    "we shifted to a classic RCT in which we compared the summer bridge program to business-as-usual remedial mathematics provided through the Advancement Via Individual Determination [AVID] Algebra Readiness program..."

  • R

    This newly developed RAMPUP program and its trials have not been independently replicated by another research team.

    "Our original experimental design, adapted from an evaluation conducted by Snipes et al. (2015), was a delayed-intervention randomized controlled trial (RCT)."

  • A

    Only mathematics outcomes were measured, with no assessment of other main subjects.

    "We report results from the implementation and impact evaluations of a three-week summer bridge mathematics program."

  • G

    Outcomes were measured only at the end of the three-week program, with no follow-up or graduation tracking, and criterion Y was not met.

    "For the intervention group (first cohort), the outcome variable is the post-test on the MIRA which was taken at the end of the three-week program."

  • P

    Both trials were pre-registered at the Registry of Efficacy and Effectiveness Studies with specific IDs, and the paper discusses adherence to the pre-registered plan.

    "Both studies were pre-registered at the Registry of Efficacy and Effectiveness Studies (the study in District A is registered as 18643.2v1, and the one in District B is registered as 18643.1v1)."

Abstract

We report results from the implementation and impact evaluations of a three-week summer bridge mathematics program. This program was designed to challenge and support rising ninth grade students and in particular Multilingual Learners. We report the extent to which implementing teachers enacted the activities as intended and note challenges that implementing teachers reported with the ambitious program for mathematics learning in a summer setting. For the impact study, we implemented two experiments with assignment at the student level: a classic randomized controlled trial with 37 students in the analytic sample and a delayed-intervention randomized controlled trial with 114 students in the analytic sample. Our impact analysis was a single-level linear regression model with demographic and prior math achievement covariates. While the impact on student math outcomes was positive, ranging from 0.03 to 0.19 standard deviations, these effects were not statistically significant due to sample attrition.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was at the individual student level, not the class or school level, and the study is not a one-to-one tutoring intervention.
      • "For the impact study, we implemented two experiments with assignment at the student level..."
      • Relevant Quotes: 1) "For the impact study, we implemented two experiments with assignment at the student level: a classic randomized controlled trial with 37 students in the analytic sample and a delayed-intervention randomized controlled trial with 114 students in the analytic sample." (Abstract) 2) "We estimate the impact of the intervention, which was individually assigned to students, on the math outcome measure by conducting a linear regression." (p. 8) 3) "For each of these subgroups, we created a random variable and assigned a random number to each of the slots. We then determined the median of the random number variable and assigned the intervention condition ... and the comparison condition ... as students applied by site ... they were placed in the next available slot and assigned treatment status based upon the slot into which they were placed." (p. 8) Detailed Analysis: The ERCT 'C' criterion requires randomisation at the class level (or the stronger school level), unless the intervention is one-to-one tutoring or personal teaching, in which case student-level randomisation is acceptable. The paper states explicitly and repeatedly that assignment was "at the student level" and that the intervention "was individually assigned to students." The RAMPUP program was a classroom-based summer program built around small-group activities and whole-class discussion, not one-to-one tutoring, so the tutoring exception does not apply. Randomising individual students to intervention and comparison conditions within the same summer sites creates the contamination risk the criterion is designed to guard against. Criterion C is not met because randomisation was conducted at the individual student level rather than at the class or school level, and no tutoring exception applies.
    • E

      Exam-based Assessment

      • The study used the standardised, externally developed MIRA readiness assessment and the state SBAC test rather than a custom-built measure.
      • "The outcome measure for the study was the Mathematics Initiative Readiness Assessment (MIRA), a reliable measure of readiness for high school mathematics..."
      • Relevant Quotes: 1) "The outcome measure for the study was the Mathematics Initiative Readiness Assessment (MIRA), a reliable measure of readiness for high school mathematics, which was found to have internal consistency of 0.83 as measured by Cronbach's alpha (Briggs, 2024). The MIRA is a computer-administered test that was given at the beginning and end of the summer session to serve as a pre- and post-test of student mathematics achievement." (p. 8) 2) "and so relied instead on the state standardized math test, the 8th grade Smarter Balance Assessment Consortium (SBAC) math test both as a covariate and to determine whether groups were equivalent at baseline." (p. 9) 3) "Briggs, D. (2024). Mathematics initiative readiness assessment (MIRA) technical report." (References) Detailed Analysis: Criterion E requires a standardised, widely recognised exam-based assessment that was not custom-built for the study. The primary outcome, the MIRA, is a pre-existing readiness assessment with its own published technical report (Briggs, 2024) and reported internal consistency of 0.83; it is a general measure of readiness for high school mathematics rather than a test aligned to the RAMPUP curriculum content (cross-cutting concepts, patterns, graph theory, equivalence). In addition, the study uses the 8th grade Smarter Balanced Assessment Consortium (SBAC) math test, an established state standardised assessment, as a baseline covariate. Both are externally developed, standardised instruments rather than researcher-designed tests. Criterion E is met because the study used the standardised, externally developed MIRA readiness assessment (and the state SBAC test) rather than a custom-built test aligned to the intervention.
    • T

      Term Duration

      • The intervention and outcome measurement spanned only three weeks, far short of the one-term requirement, with no term-long follow-up.
      • "For the intervention group (first cohort), the outcome variable is the post-test on the MIRA which was taken at the end of the three-week program."
      • Relevant Quotes: 1) "We report results from the implementation and impact evaluations of a three-week summer bridge mathematics program." (Abstract) 2) "The MIRA is a computer-administered test that was given at the beginning and end of the summer session to serve as a pre- and post-test of student mathematics achievement." (p. 8) 3) "For the intervention group (first cohort), the outcome variable is the post-test on the MIRA which was taken at the end of the three-week program." (p. 9) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins, or, for naturally short interventions, that there is at least term-long follow-up tracking. Here the intervention was a three-week summer program (13-14 sessions), and the outcome (MIRA post-test) was collected at the very end of that three-week window. The interval from intervention start to measurement is therefore only about three weeks, far short of a full term, and there was no term-long follow-up tracking after the program ended. Criterion T is not met because outcomes were measured at the end of a three-week program, well short of one academic term, with no longer-term follow-up.
    • D

      Documented Control Group

      • The comparison groups' demographics, baseline scores, and conditions are documented in the text and in Table 1.
      • "Table 1. Baseline characteristics of samples in Districts A and B."
      • Relevant Quotes: 1) "Therefore, we shifted to a classic RCT in which we compared the summer bridge program to business-as-usual remedial mathematics provided through the Advancement Via Individual Determination [AVID] Algebra Readiness program (AVID, 2023)." (p. 7) 2) "Within the second cohort of students, the comparison condition was summer learning loss associated with not taking any math classes or participating in summer math learning programs (cf. Snipes et al., 2015)." (p. 7) 3) "Table 1. Baseline characteristics of samples in Districts A and B." listing N, Ethnicity, Gender, SBAC math (M/SD), and MIRA pre- (M/SD) for both Intervention and Comparison columns. (p. 13) Detailed Analysis: Criterion D requires the control group to be well-documented, including demographics, baseline performance, and the conditions it experienced. The paper clearly documents the comparison conditions in both districts: District A's comparison received business-as-usual AVID remedial mathematics, and District B's comparison was a delayed cohort experiencing summer learning loss with no summer math. Table 1 provides comparison-group sample sizes, ethnicity breakdowns, gender, SBAC math scaled scores, and MIRA pre-test scores alongside the intervention group, and Tables 2, 4, and 6 document enrollment and attrition for the comparison groups. Criterion D is met because the control/comparison group's composition, baseline achievement, and conditions are documented in detail in the text and in Table 1.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was at the individual student level, not at the school or site level.
      • "We estimate the impact of the intervention, which was individually assigned to students..."
      • Relevant Quotes: 1) "For the impact study, we implemented two experiments with assignment at the student level..." (Abstract) 2) "We estimate the impact of the intervention, which was individually assigned to students, on the math outcome measure by conducting a linear regression." (p. 8) 3) "Randomization in both districts was conducted following a procedure adapted from Snipes et al. (2015) to create a spreadsheet for assignments" ... assigning individual student "slots" to conditions. (pp. 7-8) Detailed Analysis: Criterion S requires randomisation at the level of the school or the implementing institution/unit. In this study, randomisation was performed at the individual student level; students were sorted into "slots" and assigned to intervention or comparison conditions. No schools, campuses, sites, or clusters were randomly assigned. The two districts were partners chosen for convenience, not randomised units. Criterion S is not met because randomisation occurred at the individual student level rather than at the school or site level.
    • I

      Independent Conduct

      • The same authors who designed the RAMPUP intervention also conducted and analysed the evaluation, with no independent evaluator.
      • "Author Contributions: Conceptualization, H.C. and L.H.; methodology, H.C. and J.N.D.; software, H.C.; investigation, H.C. ..."
      • Relevant Quotes: 1) "we approached the design of a summer bridge program with a more ambitious, socioculturally driven vision of mathematics learning and teaching. Our iterative design, development, and improvement work was further grounded in a design-based research approach..." (pp. 2-3) 2) "For two summers leading up to the experiments described in this study, we intensively studied implementation ... we tested and iteratively improved the usability and feasibility of the materials." (p. 3) 3) "Author Contributions: Conceptualization, H.C. and L.H.; methodology, H.C. and J.N.D.; software, H.C.; investigation, H.C.; writing-original draft preparation, H.C. ... funding acquisition, H.C. and L.H." (p. 18) 4) "Conflicts of Interest: The authors declare no conflicts of interest. Although the RAMPUP curriculum is licensed for reuse ... there is no financial benefit that would accrue to the authors..." (p. 18) Detailed Analysis: Criterion I requires the evaluation to be conducted independently of the team that designed the intervention. Here the same WestEd authors who designed and iteratively developed the RAMPUP curriculum (Chu & Hamburger and colleagues) also conceptualised the study, developed the methodology, wrote the randomisation software, and carried out the investigation and analysis. There is no statement of an external, third-party evaluator or independent evaluation team; the disclosure of no conflicts of interest does not establish independent conduct, since the designers themselves ran the trial. Criterion I is not met because the intervention designers (the WestEd authors) also designed, implemented, and analysed the evaluation, with no independent third-party evaluator.
    • Y

      Year Duration

      • The study lasted only three weeks, far short of a full academic year, and criterion T was not met.
      • "We report results from the implementation and impact evaluations of a three-week summer bridge mathematics program."
      • Relevant Quotes: 1) "We report results from the implementation and impact evaluations of a three-week summer bridge mathematics program." (Abstract) 2) "In District B, the intervention was implemented over 13 sessions, each of which was scheduled to be three hours in length." (p. 7) 3) "For the intervention group (first cohort), the outcome variable is the post-test on the MIRA which was taken at the end of the three-week program." (p. 9) Detailed Analysis: Criterion Y requires that outcomes be measured at least 75% of a full academic year (roughly 9-10 months) after the intervention begins. The intervention here lasted only three weeks over the summer, with outcomes measured at the end of that window. This is far short of a full academic year, and no year-long tracking occurred. In addition, per the criteria-specific rule, because criterion T (Term Duration) is not met, criterion Y cannot be met. Criterion Y is not met because the study spanned only three weeks, far short of a full academic year, and criterion T was not met.
    • B

      Balanced Control Group

      • District A's comparison was a time-comparable competing summer program, and District B tested the summer program's added time as the integral treatment variable against a business-as-usual baseline.
      • "we shifted to a classic RCT in which we compared the summer bridge program to business-as-usual remedial mathematics provided through the Advancement Via Individual Determination [AVID] Algebra Readiness program..."
      • Relevant Quotes: 1) "What was the impact of participating in RAMPUP compared to the comparison condition of: Participating in business-as-usual remedial mathematics instruction for all students? Participating in a delayed cohort who experienced summer learning loss..." (p. 6) 2) "we shifted to a classic RCT in which we compared the summer bridge program to business-as-usual remedial mathematics provided through the Advancement Via Individual Determination [AVID] Algebra Readiness program (AVID, 2023). The AVID program was designed to cover 6th through 8th grade standards in 15 four-hour 'units' ... The summer session in District A, however, was 14 days of instruction in two-hour blocks." (p. 7) 3) "Within the second cohort of students, the comparison condition was summer learning loss associated with not taking any math classes or participating in summer math learning programs (cf. Snipes et al., 2015). In District B, the intervention was implemented over 13 sessions, each of which was scheduled to be three hours in length." (p. 7) Detailed Analysis: Following the criterion B decision tree: the treatment variable being tested is participation in the RAMPUP summer bridge program itself. In District A, the comparison group received a competing summer program (AVID business-as-usual remedial mathematics) over the same 14-day, two-hour-block summer session, so the two groups received broadly comparable instructional time and resources - a balanced head-to-head curriculum comparison. In District B, the comparison condition was an explicit business-as-usual baseline of summer learning loss (no summer math), and the study's stated research question is precisely to test the impact of the summer program against that no-program baseline. A summer bridge program inherently consists of added summer instructional time; that added time is integral to and is itself the treatment variable being tested, so a business-as-usual (no summer math) comparison is appropriate by design rather than a confounding imbalance. Under either district, criterion B is satisfied: District A matches resources directly, and District B tests the added summer instruction as the integral treatment variable against a business-as-usual baseline. Criterion B is met because District A used a time-comparable competing summer program as its comparison, and District B explicitly tested the summer program (its added instructional time) as the integral treatment variable against a business-as-usual summer-learning-loss baseline.
  • Level 3 Criteria

    • R

      Reproduced

      • This newly developed RAMPUP program and its trials have not been independently replicated by another research team.
      • "Our original experimental design, adapted from an evaluation conducted by Snipes et al. (2015), was a delayed-intervention randomized controlled trial (RCT)."
      • Relevant Quotes: 1) "Our original experimental design, adapted from an evaluation conducted by Snipes et al. (2015), was a delayed-intervention randomized controlled trial (RCT)." (p. 7) 2) "This program was designed to challenge and support rising ninth grade students and in particular Multilingual Learners." (Abstract) 3) "Given the ambition of the program, it may be that teachers need more experience implementing the activities in the curriculum." (p. 18) Detailed Analysis: Criterion R requires that this specific study, or its central experimental claim, have been independently replicated by a different research team in a peer-reviewed outlet. RAMPUP is a newly developed intervention, and this 2026 paper reports its first efficacy trials; the authors adapted their design from Snipes et al. (2015) but that is a methodological template, not a replication of RAMPUP. An internet search in July 2026 across WestEd, the National R&D Center to Improve Education for Secondary English Learners, ERIC, and journal sources located only the developers' own RAMP-UP materials, webinars, and reports and found no independent replication of this trial by a different research team; the program is new and this paper reports its first RCTs. The null, underpowered results themselves have not been reproduced elsewhere. Criterion R is not met because this specific intervention and trial have not been independently replicated by a different research team in a peer-reviewed publication.
    • A

      All-subject Exams

      • Only mathematics outcomes were measured, with no assessment of other main subjects.
      • "We report results from the implementation and impact evaluations of a three-week summer bridge mathematics program."
      • Relevant Quotes: 1) "The outcome measure for the study was the Mathematics Initiative Readiness Assessment (MIRA), a reliable measure of readiness for high school mathematics..." (p. 8) 2) "We report results from the implementation and impact evaluations of a three-week summer bridge mathematics program." (Abstract) 3) "RAMPUP was designed to promote conceptual understanding, participation by design, and purposeful focus for English Learners." (p. 6) Detailed Analysis: Criterion A requires that impact be measured across all main subjects using standardised exam-based assessments, not just the intervention subject. This study measured only mathematics outcomes (via the MIRA, with SBAC math as a covariate). No other core subjects (e.g., reading, language arts, science) were assessed, and no justification for a specialised single-subject focus of the type contemplated by the exception (upper secondary or vocational specialisation) is provided. Measuring a single subject cannot satisfy the all-subject requirement. Criterion A is not met because only mathematics was assessed, with no measurement of impact across other main subjects.
    • G

      Graduation Tracking

      • Outcomes were measured only at the end of the three-week program, with no follow-up or graduation tracking, and criterion Y was not met.
      • "For the intervention group (first cohort), the outcome variable is the post-test on the MIRA which was taken at the end of the three-week program."
      • Relevant Quotes: 1) "The MIRA is a computer-administered test that was given at the beginning and end of the summer session to serve as a pre- and post-test of student mathematics achievement." (p. 8) 2) "For the intervention group (first cohort), the outcome variable is the post-test on the MIRA which was taken at the end of the three-week program." (p. 9) 3) "the first module on patterns is adjacent enough to 9th grade content that we plan on modifying it into a week-long introductory unit ... that may be appropriate to test as a bounded ninth grade intervention at the very beginning of the school year." (p. 18) Detailed Analysis: Criterion G requires follow-up tracking of participants until they graduate from the relevant educational stage. Here outcomes were collected only as a pre- and post-test within the three-week summer program, with no follow-up after the program ended and no tracking toward graduation. The authors describe only future plans for modified experiments, not longitudinal graduation tracking. An internet search in July 2026 for subsequent or follow-up publications by these authors tracking this cohort to graduation found none. In addition, per the criteria-specific rule, because criterion Y (Year Duration) is not met, criterion G cannot be met. Criterion G is not met because there was no follow-up beyond the three-week program, no graduation tracking, and criterion Y was not met.
    • P

      Pre-Registered

      • Both trials were pre-registered at the Registry of Efficacy and Effectiveness Studies with specific IDs, and the paper discusses adherence to the pre-registered plan.
      • "Both studies were pre-registered at the Registry of Efficacy and Effectiveness Studies (the study in District A is registered as 18643.2v1, and the one in District B is registered as 18643.1v1)."
      • Relevant Quotes: 1) "Both studies were pre-registered at the Registry of Efficacy and Effectiveness Studies (the study in District A is registered as 18643.2v1, and the one in District B is registered as 18643.1v1)." (p. 7) 2) "For District B, although we had not pre-registered SBAC test scores to establish baseline equivalence, we performed checks where required by attrition." (p. 9) 3) "Institutional Review Board Statement: The study was conducted in accordance with the Declaration of Helsinki and approved by the Institutional Review Board of WestEd (protocol code 2020-10-6, approved 7 May 2024)." (p. 18) Detailed Analysis: Criterion P requires that the study protocol be pre-registered on a recognised registry before data collection begins. The paper explicitly states that both trials were pre-registered at the Registry of Efficacy and Effectiveness Studies (REES), providing specific version-controlled registration identifiers (18643.2v1 and 18643.1v1) for each district. The authors further discuss the substance of what was and was not pre-specified (e.g., acknowledging that SBAC scores were "not pre-registered" for baseline equivalence), which demonstrates a genuine pre-registration that constrained the planned analyses. A direct attempt to open the REES entries to confirm the exact registration dates was blocked by the registry's access restrictions (HTTP 403), but the IRB approval (7 May 2024) and the summer 2024 data collection, together with the paper's explicit use of "pre-registered" and named version IDs, support registration prior to data collection. Criterion P is met because both trials were pre-registered at the Registry of Efficacy and Effectiveness Studies with specific registration IDs, and the paper discusses adherence to and deviations from the pre-registered plan.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.