Effectiveness of a Spanish Intervention and an English Intervention for English-Language Learners at Risk for Reading Problems

Sharon Vaughn, Paul T. Cirino, Sylvia Linan-Thompson, Patricia G. Mathes, Coleen D. Carlson, Elsa Cardenas Hagan, Sharolyn D. Pollard-Durodola, Jack M. Fletcher, David J. Francis

Published:
ERCT Check Date:
DOI: 10.3102/00028312043003449
  • reading
  • L2 languages
  • K12
  • US
1
  • C

    Randomisation occurred at the individual-student level within schools, but the intervention was a pulled-out, small-group (3-5 students) supplemental reading program delivered by dedicated instructors outside the regular classroom, matching the personal-teaching/tutoring exception that allows student-level randomisation.

    "Randomization was carried out within schools without further stratification." (p. 455)

  • E

    Primary outcome measures were well-established, widely used standardized instruments (WLPB-R, CTOPP/TOPP-S, DIBELS, TOWRE), even though one supplementary experimental spelling measure was also used.

    "The Woodcock Language Proficiency Battery-Revised (WLPB-R; Woodcock, 1991...) is a well-standardized instrument with a normative sample of 6,359..." (p. 457)

  • T

    Because the stronger Year Duration criterion (Y) is met, the weaker Term Duration criterion is automatically considered met as well.

    "...administered in both English and Spanish to all participants... prior to the intervention (October) and following its completion (May)." (p. 456)

  • D

    The comparison group's demographics, baseline scores, and the amount/type of any supplemental instruction they received outside the study are documented in detail across four results tables and the narrative text.

    "27 of 46 (59%) comparison group students received one or more types of reading instruction over their core instruction. The amount of time additional instruction was provided to these 27 students ranged from 7.9 to 93.2 hours..." (p. 465)

  • S

    Randomisation was conducted at the individual- student level within participating schools, not at the level of entire schools.

    "Randomization was carried out within schools without further stratification." (p. 455)

  • I

    The same research team that designed the Proactive Reading/Lectura Proactiva curricula also delivered training, oversaw fidelity, collected the outcome data, and analyzed the results, with no independent external evaluator described.

    "We used Proactive Reading (Mathes et al., 2005; Mathes, Torgesen, Wahl, Menchetti, & Grek, 1999)... In designing the reading intervention, Lectura Proactiva (Mathes, Linan-Thompson, Pollard-Durodola, Hagan, & Vaughn, 2001)..." (pp. 459, 460)

  • Y

    Outcomes were measured in May after an intervention and assessment window running from October to May, covering roughly 7-8 months of an approximately 9-10 month academic year (i.e., well over 75%).

    "Eight different intervention instructors provided the Spanish or English intervention... from October through May." (p. 461)

  • B

    The additional instructional time given to intervention students is the explicit treatment variable under investigation (comparing supplemental reading intervention versus business-as-usual core instruction), so the lack of an equivalent researcher-provided add-on for the comparison group does not violate the balance requirement.

    "Both the Spanish and English interventions were designed to supplement and extend instructional time in reading and oral language. The experimental nature of the interventions justified this design, as opposed to providing these interventions instead of standard classroom instruction." (p. 464)

  • R

    The paper describes itself as replicating earlier Vaughn-team studies using a new, nonoverlapping sample, but this replication was conducted by the same research group rather than an independent team; an internet search for external citing/related work found no independent replication of this specific study by a different research team.

    "However, the current study was an attempt to replicate previous results in a nonoverlapping sample of students so that sample-specific findings could be discerned from those that can be generalized or strengthened by occurring independently in two different samples of students." (p. 454)

  • A

    Outcomes were measured only in reading, phonological processing, and oral language; no other core subjects (e.g., mathematics, science) were assessed, and no vocational/specialised-education rationale is offered to justify this narrow focus.

    "It included measures of letter knowledge, phonological awareness, rapid naming, language proficiency, reading decoding, reading comprehension, reading fluency, and spelling." (p. 456)

  • G

    Follow-up ended with the Grade 1 posttest in this paper; internet search located same-team follow-up papers extending tracking to Grade 2 and to Grade 4-5, but neither documents tracking through actual graduation/completion of the students' school stage.

    "We hope to follow these students through third grade to determine the extent to which significant differences in reading between intervention and comparison students are maintained..." (p. 482)

  • P

    No pre-registration of the study protocol, hypotheses, or analysis plan is mentioned anywhere in the paper, and no evidence of registration was found via internet search; education-research pre-registration platforms (e.g., AEA RCT Registry, OSF Registries, SREE Registry) were established years after this study was conducted.

Abstract

Two studies of Grade 1 reading interventions for English-language (EL) learners at risk for reading problems were conducted. Two samples of EL students were randomly assigned to a treatment or untreated comparison group on the basis of their language of instruction for core reading (i.e., Spanish or English). In all, 91 students completed the English study (43 treatment and 48 comparison), and 80 students completed the Spanish study (35 treatment and 45 comparison). Treatment students received approximately 115 sessions of supplemental reading daily for 50 minutes in groups of 3 to 5. Findings from the English study revealed statistically significant differences in favor of treatment students on English measures of phonological awareness, word attack, word reading, and spelling (effect sizes of 0.35-0.42). Findings from the Spanish study revealed significant differences in favor of treatment students on Spanish measures of phonological awareness, letter-sound and letter-word identification, verbal analogies, word reading fluency, and spelling (effect sizes of 0.33-0.81).

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation occurred at the individual-student level within schools, but the intervention was a pulled-out, small-group (3-5 students) supplemental reading program delivered by dedicated instructors outside the regular classroom, matching the personal-teaching/tutoring exception that allows student-level randomisation.
      • "Randomization was carried out within schools without further stratification." (p. 455)
      • Relevant Quotes: 1) "Of these students, 94 were randomized into intervention (n = 42) and comparison (n = 52) groups... Randomization was carried out within schools without further stratification." (p. 455, Spanish study) 2) "For the English study, 362 students were screened, and 108 (30%) met the criteria; of these students, 96 were randomized into intervention (n = 46) and comparison (n = 50) groups." (p. 455) 3) "Eight different intervention instructors provided the Spanish or English intervention (or both interventions) to treatment participants in small groups of 3 to 5 students for 50 minutes a day, 5 days a week, from October through May... The intervention was provided in addition to the students' core reading lessons and did not conflict with their daily scheduled reading time." (pp. 461, 464) Detailed Analysis: The quotes make clear that randomisation was performed at the level of the individual student within a school, not at the class or school level. Under a strict reading of Criterion C this would normally fail, since students in the same classroom could be split between intervention and comparison conditions, risking contamination (e.g., a teacher applying intervention techniques to comparison students, or comparison students overhearing intervention content). However, the ERCT standard's exception for personal/tutoring-style interventions applies here. The intervention was not delivered by the regular classroom teacher to a subset of students in situ; instead, dedicated external intervention instructors pulled small groups of 3 to 5 at-risk students out of the classroom for a separate 50-minute daily session "in addition to" and separate from core classroom instruction. This pull-out, small-group delivery model closely parallels one-to-one/small-group tutoring arrangements, where contamination risk from within-classroom spillover is much lower than for an in-classroom curricular change, because the comparison students do not experience the intervention instructor, materials, or session at all - they simply continue with business-as-usual classroom reading instruction. Because the intervention functionally resembles a personal/small-group tutoring program delivered outside the regular classroom rather than an in-class pedagogical change, the exception for tutoring-like interventions applies, and student-level randomisation is acceptable. Final: Criterion C is met via the tutoring/personal- teaching exception, despite randomisation being carried out at the individual-student level within schools.
    • E

      Exam-based Assessment

      • Primary outcome measures were well-established, widely used standardized instruments (WLPB-R, CTOPP/TOPP-S, DIBELS, TOWRE), even though one supplementary experimental spelling measure was also used.
      • "The Woodcock Language Proficiency Battery-Revised (WLPB-R; Woodcock, 1991...) is a well-standardized instrument with a normative sample of 6,359..." (p. 457)
      • Relevant Quotes: 1) "The Comprehensive Test of Phonological Processing (CTOPP; Wagner, Torgesen, & Rashotte, 1999) is a well-validated measure of wide-ranging phonological processing abilities... The Test of Phonological Processing-Spanish TOPP-S was developed to align with the English CTOPP." (p. 456) 2) "The Woodcock Language Proficiency Battery- Revised (WLPB-R; Woodcock, 1991; Woodcock & Munoz- Sandoval, 1995) is a well-standardized instrument with a normative sample of 6,359... Internal consistency values for the subtests used... have ranged from .77 to .95." (p. 457) 3) "The Dynamic Indicators of Basic Early Literacy Skills (DIBELS; Good & Kaminski, 2002)... were administered to assess oral reading fluency." (p. 458) 4) "The word reading efficiency subtest of the Test of Word Reading Efficiency (Torgesen, Wagner, & Rashotte, 1999) measures decontextualized reading fluency." (p. 458) 5) "An experimental spelling measure developed for this study consisted of 25 words of one or two syllables..." (p. 458) Detailed Analysis: The core outcome battery used to assess decoding, comprehension, oral language, phonological awareness, and fluency consists of nationally normed, widely used standardized instruments: the WLPB-R (a large, normed battery with over 6,000 normative cases), the CTOPP/TOPP-S, DIBELS/IDEL, and the TOWRE. These are not custom-built for this study; they are established assessments used across many studies of reading development, with published psychometric properties reported in the text. Only the spelling measure was purpose-built for this study, and it functioned as one supplementary outcome among many primarily standardized measures of reading and language. Because the study's central reading and language outcomes rely on widely recognised standardized tests rather than tests custom-designed to favor the intervention, Criterion E is satisfied. Final: Criterion E is met because standardized, widely used assessments (WLPB-R, CTOPP/TOPP-S, DIBELS, TOWRE) formed the core outcome battery.
    • T

      Term Duration

      • Because the stronger Year Duration criterion (Y) is met, the weaker Term Duration criterion is automatically considered met as well.
      • "...administered in both English and Spanish to all participants... prior to the intervention (October) and following its completion (May)." (p. 456)
      • Relevant Quotes: 1) "Eight different intervention instructors provided the Spanish or English intervention... from October through May." (p. 461) 2) "The same comprehensive battery of language/ literacy measures was administered in both English and Spanish to all participants by well-trained members of the research team prior to the intervention (October) and following its completion (May)." (p. 456) Detailed Analysis: Pretest occurred in October and posttest in May, an interval of roughly seven to eight months, which far exceeds the minimum one-term (3-4 month) requirement. As documented under Criterion Y below, this interval also satisfies the stronger year-long duration requirement, and per the ERCT standard, meeting Y automatically satisfies T. Final: Criterion T is met, both directly (the October-to-May interval exceeds one term) and by inheritance from the met Y criterion.
    • D

      Documented Control Group

      • The comparison group's demographics, baseline scores, and the amount/type of any supplemental instruction they received outside the study are documented in detail across four results tables and the narrative text.
      • "27 of 46 (59%) comparison group students received one or more types of reading instruction over their core instruction. The amount of time additional instruction was provided to these 27 students ranged from 7.9 to 93.2 hours..." (p. 465)
      • Relevant Quotes: 1) "All of the students were Hispanic, and 44% were female (42% in the Spanish study, 46% in the English study). The mean age at pretest among the 183 students who began the study was 6.6 years (SD = 0.4)..." (p. 455) 2) Tables 1-4 report, for the comparison ("C") group in each study/language: sample size (n), pretest mean and SD, posttest mean and SD, and gain scores for every outcome measure (pp. 462-463, 472-473, 476-477, 480-481). 3) "For the students in the Spanish study with additional instruction data, 27 of 46 (59%) comparison group students received one or more types of reading instruction over their core instruction. The amount of time additional instruction was provided to these 27 students ranged from 7.9 to 93.2 hours (M = 41.2 hours, SD = 27.2, Mdn = 31.2)..." (p. 465) 4) "Among the students in the English study with additional instruction data, 28 of 40 (70%) of those in the comparison group received one or more types of supplemental reading instruction over their core instruction..." (p. 465) 5) "As expected from the random assignment to groups, there were no significant group mean differences in performance on either of the intervention screening skills... in either language in either study (all ps > .05)." (p. 468) Detailed Analysis: The paper provides extensive documentation of the comparison ("business as usual") group: overall demographic composition, pretest equivalence checks confirming comparability with the intervention group, and detailed per-measure pretest/posttest statistics in Tables 1-4. Crucially, the authors also directly measured and reported how much supplemental reading instruction comparison students received from other sources outside the study, including exact percentages and hour ranges. This level of transparency about the control condition's composition and experience goes well beyond a bare mention of a control group. Final: Criterion D is met given the thorough demographic, baseline, and supplemental-instruction documentation of the comparison group.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was conducted at the individual- student level within participating schools, not at the level of entire schools.
      • "Randomization was carried out within schools without further stratification." (p. 455)
      • Relevant Quotes: 1) "Ten schools and 42 classrooms (Spanish study: six schools and 22 classrooms; English study: four schools and 20 classrooms) from three sites in Texas with large numbers of bilingual students were selected." (p. 454) 2) "Of these students, 94 were randomized into intervention (n = 42) and comparison (n = 52) groups... Randomization was carried out within schools without further stratification." (p. 455) 3) "For the English study... of these students, 96 were randomized into intervention (n = 46) and comparison (n = 50) groups." (p. 455) Detailed Analysis: Every participating school contributed both intervention and comparison students, since eligible at-risk students within each school were individually randomly assigned. No school, as a whole unit, was assigned entirely to either condition. This is explicitly the opposite of a school-level RCT design, where entire schools would be assigned wholesale to intervention or comparison. Final: Criterion S is not met because randomisation occurred among individual students within schools, not among whole schools.
    • I

      Independent Conduct

      • The same research team that designed the Proactive Reading/Lectura Proactiva curricula also delivered training, oversaw fidelity, collected the outcome data, and analyzed the results, with no independent external evaluator described.
      • "We used Proactive Reading (Mathes et al., 2005; Mathes, Torgesen, Wahl, Menchetti, & Grek, 1999)... In designing the reading intervention, Lectura Proactiva (Mathes, Linan-Thompson, Pollard-Durodola, Hagan, & Vaughn, 2001)..." (pp. 459, 460)
      • Relevant Quotes: 1) "We used Proactive Reading (Mathes et al., 2005; Mathes, Torgesen, Wahl, Menchetti, & Grek, 1999), a comprehensive, integrated intervention curriculum... that our team modified to include EL support lessons." (p. 459) 2) "In designing the reading intervention, Lectura Proactiva (Mathes, Linan-Thompson, Pollard-Durodola, Hagan, & Vaughn, 2001), we applied the same instructional design principles described for the EL program." (p. 460) 3) "The same comprehensive battery of language/ literacy measures was administered... by well- trained members of the research team." (p. 456) 4) "Two observers, in consultation with the research team leaders, conducted intervention validity checks..." (p. 464) Detailed Analysis: The authors (e.g., Mathes, Vaughn, Linan-Thompson, Pollard-Durodola, Hagan) are simultaneously the developers of the Proactive Reading and Lectura Proactiva curricula used as the interventions, the designers of the study, the trainers of intervention instructors, the assessors administering outcome measures ("members of the research team"), and the analysts of the results. There is no statement anywhere in the paper indicating that data collection or analysis was handled by an independent, third-party organisation blind to the study's hypotheses or unaffiliated with the intervention's developers. This is the classic pattern the Independent Conduct criterion is designed to flag. Final: Criterion I is not met because the intervention's developers were also responsible for implementing, assessing, and analyzing the study, with no independent evaluator involved.
    • Y

      Year Duration

      • Outcomes were measured in May after an intervention and assessment window running from October to May, covering roughly 7-8 months of an approximately 9-10 month academic year (i.e., well over 75%).
      • "Eight different intervention instructors provided the Spanish or English intervention... from October through May." (p. 461)
      • Relevant Quotes: 1) "The same comprehensive battery of language/ literacy measures was administered in both English and Spanish to all participants by well-trained members of the research team prior to the intervention (October) and following its completion (May)." (p. 456) 2) "Eight different intervention instructors provided the Spanish or English intervention (or both interventions) to treatment participants in small groups of 3 to 5 students for 50 minutes a day, 5 days a week, from October through May." (p. 461) 3) "Among the remaining students who completed the study, mean documented intervention times across the two studies ranged from 74.6 to 106.7 hours...or approximately 115 intervention sessions of 50 minutes each." (p. 464) Detailed Analysis: The intervention and outcome-tracking window ran from October (pretest, intervention start) through May (intervention end, posttest), an interval of approximately 7 to 8 months. A standard US academic year runs roughly September/August through May/June (about 9-10 months), so an October-to-May window covers on the order of 75-85% of the academic year - at or above the standard's 75% threshold for minor, technical shortfalls to still count as met. The approximately 90-100 hours of documented intervention delivered "115 sessions" over this span corroborates that this was a near-full-year undertaking rather than a brief, isolated intervention. Final: Criterion Y is met because the intervention and outcome measurement spanned October to May, comprising at least 75% of a full academic year.
    • B

      Balanced Control Group

      • The additional instructional time given to intervention students is the explicit treatment variable under investigation (comparing supplemental reading intervention versus business-as-usual core instruction), so the lack of an equivalent researcher-provided add-on for the comparison group does not violate the balance requirement.
      • "Both the Spanish and English interventions were designed to supplement and extend instructional time in reading and oral language. The experimental nature of the interventions justified this design, as opposed to providing these interventions instead of standard classroom instruction." (p. 464)
      • Relevant Quotes: 1) "Treatment students received approximately 115 sessions of supplemental reading daily for 50 minutes in groups of 3 to 5." (Abstract, p. 449) 2) "Both the Spanish and English interventions were designed to supplement and extend instructional time in reading and oral language. The experimental nature of the interventions justified this design, as opposed to providing these interventions instead of standard classroom instruction." (p. 464) 3) "While random assignment of participants to intervention and comparison groups controls for many experimental differences, information on the activities of these students outside the interventions provides an important context for interpreting group differences... Given that all students received a substantial amount of reading instruction from experienced teachers in schools recognized for their successful performance, differences between intervention and comparison at-risk students were probably more related to the content of the instruction received than simply to time on task." (pp. 464-465) 4) "In summary, the intervention students received less of this supplemental reading instruction than the comparison students in the Spanish study." (p. 465) "...the intervention and comparison students in the English study received equivalent amounts of additional instruction." (p. 465) Detailed Analysis: Step 1 (Determine Study Intent, applying the updated ERCT decision procedure for Criterion B): EXTRA_RESOURCES_PRESENT is true - the intervention group received an extra ~50 minutes/day (~90-100 hours total) of supplemental small-group reading instruction beyond core classroom reading, which the comparison group did not receive. The authors explicitly frame this additional time as the core treatment variable under study - the research question is whether adding this supplemental small-group intervention to standard classroom reading instruction improves outcomes relative to standard instruction alone ("business-as-usual"). This corresponds to RESOURCES_ARE_TREATMENT = true in the decision tree, which by itself is sufficient for the criterion to be met, regardless of whether the control group's resources were separately matched. Step 2/3 (Resource Integration Assessment): The additional daily small-group sessions are INTEGRAL to the treatment being tested, not a supplementary, separable add-on incidental to a different core manipulation. The authors' own framing ("designed to supplement and extend instructional time... as opposed to providing these interventions instead of standard classroom instruction") makes this explicit. Step 5 (Resource Balance Verification for integral resources): Because the added time is the treatment itself, the comparison group is expected to remain on standard core instruction ("business as usual"), which is exactly what the exception in the ERCT standard permits. Notably, the paper goes further than most studies by measuring and reporting how much supplemental instruction comparison students independently received from other sources, finding that in the Spanish study comparison students actually received somewhat more outside supplemental instruction on average than intervention students, and in the English study the two groups received equivalent amounts. This suggests no meaningful confound from an uncontrolled resource imbalance, further reinforcing (though not required for) the "met" determination. Final: Criterion B is met because the additional instructional time is explicitly the treatment variable being tested against a business-as-usual comparison condition, and the authors additionally document that comparison students were not resource-deprived relative to intervention students.
  • Level 3 Criteria

    • R

      Reproduced

      • The paper describes itself as replicating earlier Vaughn-team studies using a new, nonoverlapping sample, but this replication was conducted by the same research group rather than an independent team; an internet search for external citing/related work found no independent replication of this specific study by a different research team.
      • "However, the current study was an attempt to replicate previous results in a nonoverlapping sample of students so that sample-specific findings could be discerned from those that can be generalized or strengthened by occurring independently in two different samples of students." (p. 454)
      • Relevant Quotes (from this paper): 1) "The methods used in the current studies were the same as those used in two previously conducted intervention studies with first-grade EL learners described by Vaughn and colleagues (Vaughn et al., 2006, in press). However, the current study was an attempt to replicate previous results in a nonoverlapping sample of students..." (p. 454) 2) "Findings for students who were provided the intervention in Spanish were similar to those in a previous study (Vaughn et al., 2006)..." (p. 478) 3) "...although the performance levels of these students were slightly lower than those of students in a previous English intervention study (Vaughn et al., in press)." (p. 479) Internet Search for Independent Replication: A citation-index search (OpenAlex works citing this paper, doi 10.3102/00028312043003449) was conducted to look for independent replications by a different research team. Two directly related follow-up papers were identified, but both are by overlapping authorship with this paper (Vaughn, Cirino, Fletcher, Cardenas-Hagan, Francis, and colleagues), not an independent team, and are longitudinal follow-ups of the same original cohort rather than new replication attempts in a different context: Cirino et al. (2009), "One-Year Follow-Up Outcomes of Spanish and English Interventions for English Language Learners at Risk for Reading Problems," American Educational Research Journal, doi 10.3102/0002831208330214; and Vaughn et al. (2008), "Long-Term Follow-Up of Spanish and English Interventions for First-Grade English Language Learners at Risk for Reading Problems," Journal of Research on Educational Effectiveness, doi 10.1080/19345740802114749. No verbatim quotes are given from these two papers because full text was not accessible (publisher pages returned paywalled/403 errors); only bibliographic metadata could be confirmed, so no quotes are fabricated here. Beyond same-team follow-ups, the broader citing literature includes other EL reading-intervention studies by unrelated teams (e.g., Baker et al., 2012, "Effects of a paired bilingual reading program and an English-only program on the reading performance of English learners in Grades 1-3"; Fien et al., 2011 and 2015, on multitiered reading intervention for EL/at-risk first graders), but these test different curricula/programs with different designs and are not attempts to reproduce this specific study's intervention (Proactive Reading/Lectura Proactiva), sample, or methods. Detailed Analysis: Criterion R requires that the specific study (or its central experimental claim, in the same context and design) be replicated independently by a different research team in a peer-reviewed outlet. The paper explicitly positions itself as a replication of the authors' own prior Spanish study (Vaughn et al., 2006) and prior English study (Vaughn et al., in press), but both the original studies and this replication share overlapping authorship. The subsequent follow-up papers located via internet search are also by the same overlapping author team and track the same cohort forward in time rather than independently reproducing the study design with a new team. No independent replication by an unaffiliated research team was found. Final: Criterion R is not met because the only replication described (in the paper itself and via internet search) is by the same/overlapping research team, not an independent group.
    • A

      All-subject Exams

      • Outcomes were measured only in reading, phonological processing, and oral language; no other core subjects (e.g., mathematics, science) were assessed, and no vocational/specialised-education rationale is offered to justify this narrow focus.
      • "It included measures of letter knowledge, phonological awareness, rapid naming, language proficiency, reading decoding, reading comprehension, reading fluency, and spelling." (p. 456)
      • Relevant Quotes: 1) "The same comprehensive battery of language/ literacy measures was administered in both English and Spanish to all participants... It included measures of letter knowledge, phonological awareness, rapid naming, language proficiency, reading decoding, reading comprehension, reading fluency, and spelling." (p. 456) 2) The entire Results section (Tables 1-4) reports exclusively on letter naming, phonological processing, oral language, reading, and spelling outcomes; no mathematics, science, or other subject outcomes are reported anywhere in the paper. Detailed Analysis: Although Criterion E's prerequisite is satisfied (standardized instruments were used for the reading/ language outcomes that were measured), Criterion A additionally requires that all main subjects taught at that educational level be assessed, unless a specialised/vocational rationale is given (an exception intended for upper secondary/vocational contexts, not applicable to a Grade 1 reading intervention). This study is exclusively a reading/ literacy and oral-language intervention and makes no attempt to measure mathematics or other core subject areas, nor does it offer any exception-style rationale for this narrow scope. Final: Criterion A is not met because only reading and language outcomes were assessed, with no coverage of other main subjects such as mathematics.
    • G

      Graduation Tracking

      • Follow-up ended with the Grade 1 posttest in this paper; internet search located same-team follow-up papers extending tracking to Grade 2 and to Grade 4-5, but neither documents tracking through actual graduation/completion of the students' school stage.
      • "We hope to follow these students through third grade to determine the extent to which significant differences in reading between intervention and comparison students are maintained..." (p. 482)
      • Relevant Quotes (from this paper): 1) "We hope to follow these students through third grade to determine the extent to which significant differences in reading between intervention and comparison students are maintained and to assess relative growth and influence of oracy." (p. 482) 2) "We are currently investigating the extent to which the findings for intervention participants are maintained over time through second grade." (p. 483) 3) No data collection beyond the Grade 1 May posttest is reported anywhere in the Results or Discussion sections of this paper; both reported studies conclude with the single-year posttest. Follow-up Publications Identified via Internet Search: A citation-index search (OpenAlex works citing doi 10.3102/00028312043003449) located two subsequent papers by the same/overlapping author team that appear to track this cohort (or closely related cohorts from the companion English-intervention study) forward in time: 1) Cirino, Vaughn, Linan-Thompson, Cardenas-Hagan, Fletcher, & Francis (2009), "One-Year Follow-Up Outcomes of Spanish and English Interventions for English Language Learners at Risk for Reading Problems," American Educational Research Journal, doi 10.3102/0002831208330214. Per bibliographic database metadata (not a verbatim quote from the publisher's full text, which was inaccessible), this paper reports one-year follow-up data (i.e., Grade 2) for these EL learners, with the treatment group continuing to show advantages in decoding, spelling, fluency, and comprehension roughly a year after the intervention ended. 2) Vaughn, Cirino, Tolar, Fletcher, Cardenas-Hagan, Carlson, & Francis (2008), "Long-Term Follow-Up of Spanish and English Interventions for First-Grade English Language Learners at Risk for Reading Problems," Journal of Research on Educational Effectiveness, doi 10.1080/19345740802114749. Per bibliographic database metadata (again, full publisher text was not accessible, so no verbatim quote is given), this paper follows four samples of EL learners from two sequential Grade 1 cohorts 3-4 years postintervention, with assessments in Spring of Grade 4 or Grade 5, finding few statistically significant differences by that point although effect sizes generally still favored the earlier intervention group. No verbatim quotes are provided for these two follow-up papers because their full text could not be retrieved (publisher pages were paywalled/ returned 403 errors); only database-abstract metadata could be confirmed, and no quotes have been fabricated. Detailed Analysis: Even accounting for this newly located evidence, tracking of this cohort extends to roughly Grade 4-5 (about 3-4 years post-intervention), which is substantially longer than the single Grade-1 posttest reported in the original 2006 paper. However, per the ERCT standard's own worked example ("Students were tracked through to the end of their primary education, until Grade 6 graduation"), Graduation Tracking requires following students through to completion of their school stage (e.g., primary/elementary school), not merely for several additional years. Grade 4-5 falls one to two years short of a typical US elementary-school endpoint (Grade 5 or 6), and neither follow-up paper's available metadata claims tracking through actual graduation or completion of a school stage for this cohort. Final: Criterion G is not met. The original paper reports outcomes only through the end of Grade 1; same-team follow-up papers located via internet search extend tracking to Grade 2 (Cirino et al., 2009) and to approximately Grade 4-5 (Vaughn et al., 2008), but neither documents tracking through graduation/completion of the students' school stage.
    • P

      Pre-Registered

      • No pre-registration of the study protocol, hypotheses, or analysis plan is mentioned anywhere in the paper, and no evidence of registration was found via internet search; education-research pre-registration platforms (e.g., AEA RCT Registry, OSF Registries, SREE Registry) were established years after this study was conducted.
      • Relevant Quotes: 1) No sentence, footnote, or reference anywhere in the Method, Notes, or References sections mentions a trial registry, a registration number, or a pre-specified analysis plan published before data collection. 2) "Manuscript received September 9, 2005; Revision received January 20, 2006; Accepted February 27, 2006," with the intervention itself having been delivered years earlier (October-May of an unspecified prior school year), consistent with this being an older-style trial predating routine use of study pre-registration in education research. Internet Search: No trial-registry entry, registration ID, or pre-registered protocol referencing this study (Vaughn et al., 2006, this title, or its DOI) was located. This is unsurprising: major registries commonly used for education RCTs today (e.g., the AEA RCT Registry, OSF Registries/Preregistration, the SREE registry) were established starting around 2011-2013, several years after this study's data collection and 2006 publication, so pre-registration would not have been a realistic option for the authors at the time. Detailed Analysis: The ERCT standard requires explicit textual evidence of pre-registration (a registry platform, an ID, and a date preceding data collection) for this criterion to be met. This 2006 paper, typical of education RCTs from that period, contains no such statement, and no registry record was found online. Absent any quote referencing a registry or pre-specified protocol, the criterion cannot be marked as met. Final: Criterion P is not met because the paper contains no evidence of pre-registration and none could be located via internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.