The Effect of an Instructional Intervention on Middle School English Learners' Science and English Reading Achievement

Rafael Lara-Alecio, Fuhui Tong, Beverly J. Irby, Cindy Guerrero, Maggie Huerta, Yinan Fan

Published:
ERCT Check Date:
DOI: 10.1002/tea.21031
  • science
  • reading
  • L2 languages
  • K12
  • US
  • parent involvement
  • EdTech platform
0
  • C

    Although the four intermediate schools were initially randomly assigned, the comparison group's final analytic sample was augmented by teachers recruited non-randomly, leading the authors themselves to classify the overall design as quasi-experimental rather than a properly implemented RCT.

    "Because of the non-random addition of these teachers, this design of this overall project was considered to be quasi-experimental." (p. 994)

  • E

    In addition to district-developed benchmark tests, the study used the state-mandated standardized TAKS assessment and the nationally standardized DIBELS Oral Reading Fluency measure, both widely recognized standardized instruments.

    "The TAKS, a criterion-referenced assessment, measures student mastery of the content areas of state curriculum outlined in the TEKS." (p. 999)

  • T

    The intervention and outcome measurement spanned the full 2009-2010 school year, with final outcomes measured via TAKS and DIBELS in spring 2010, far exceeding one academic term.

    "Data were collected in the fall and spring of school year 2009-2010." (p. 1001)

  • D

    The comparison group's schools, teachers, and students are documented in detail, including demographics, teaching experience, class sizes, and the typical-practice instruction they received.

    "Table 1. School demographics for treatment and comparison group 2009-2010" (p. 995)

  • S

    Although the four schools were originally randomly assigned, the comparison group's sample was expanded with non-randomly recruited teachers, so the study does not represent a properly implemented school-level RCT.

    "Because of the non-random addition of these teachers, this design of this overall project was considered to be quasi-experimental." (p. 994)

  • I

    The same university research team that designed the intervention, trained the teachers, and monitored fidelity also conducted the evaluation and analysis, with no independent third-party evaluator.

    "This component consisted of ongoing training workshops for both teachers (biweekly) and paraprofessionals (monthly) provided by research coordinators for 3 hours per session." (p. 995)

  • Y

    The intervention and follow-up measurement covered the entire 2009-2010 academic year, satisfying the stronger year-duration requirement.

    "The larger research project from which the current study was derived, to our best knowledge, is the only longitudinal quasi-experimental design with science intervention of a full academic year among fifth grade ELLs and low-SES English proficient students." (p. 1006)

  • B

    Treatment classrooms received substantial additional, unmatched resources (professional development, EduSmart software, Science Saturdays, Family Involvement take-home packets) that the authors themselves acknowledge as a likely confound, without the study framing these extra resources as the explicit treatment variable being isolated.

    "Another limitation is that we compared conditions that differed in several enrichment components (e.g., family science support, Saturday science activities, and extra technology resources in the classroom) that may seem to confound the ability to compare the nature-of-the-instruction treatment." (p. 1004)

  • R

    No independent replication of this specific study or intervention (MSSELL) by a different research team is reported in the paper, and none was found through additional internet searches (Crossref, Semantic Scholar, OpenAlex citation records).

  • A

    The study measured only science and English reading outcomes, not the full range of core subjects such as mathematics or social studies.

    "Results are presented by construct measured, that is, science and reading achievement." (p. 1001)

  • G

    The study tracked students only through the end of fifth grade with no follow-up to graduation; the authors note plans to continue only into sixth grade, and no later publication tracking this cohort further was found via internet search.

    "Further, as we plan for the next step of the larger research project, we will take into consideration the rotation/classroom nature of science instruction in these schools as students complete sixth grade..." (p. 1004)

  • P

    No statement of pre-registration of the study protocol on a public registry prior to data collection is present anywhere in the paper, and no pre-registration record for this study was found via internet search.

Abstract

This study examined the effect of a quasi-experimental project on fifth grade English learners' achievement in state-mandated standards-based science and English reading assessment. A total of 166 treatment students and 80 comparison students from four randomized intermediate schools participated in the current project. The intervention consisted of on-going professional development and specific instructional science lessons with inquiry-based learning, direct and explicit vocabulary instruction, integration of reading and writing, and enrichment components including integration of technology, take-home science activities, and university scientists mentoring. Results suggested a significant and positive intervention effect in favor of the treatment students as reflected in higher performance in district-wide curriculum-based tests of science and reading and standardized tests of oral reading fluency.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Although the four intermediate schools were initially randomly assigned, the comparison group's final analytic sample was augmented by teachers recruited non-randomly, leading the authors themselves to classify the overall design as quasi-experimental rather than a properly implemented RCT.
      • "Because of the non-random addition of these teachers, this design of this overall project was considered to be quasi-experimental." (p. 994)
      • Relevant Quotes: 1) "The state law (Texas Education Code, 1995) prohibits random selection and assignment on the basis of individual students for program placement; therefore, in the larger research project, four intermediate schools with principals' approval from the district site were randomly assigned to conditions, resulting in two treatment (enhanced science practice) and two comparison (typical science practice) schools." (p. 994) 2) "When a school was assigned, teachers from that campus were then randomly selected to the assigned condition within that campus, and both ELLs and low SES non-ELLs in the same school received the same practice to allay contamination between experimental and comparison classrooms." (p. 994) 3) "Due to the low return response in the two comparison schools, a balanced design with four science teachers in their respective condition (i.e., two in treatment and two in comparison) could not be achieved. Therefore, in order to increase the sample size at student level, we recruited four more teachers within the same comparison schools that were already assigned based on the criteria that they had Spanish-speaking ELLs in their classrooms." (p. 994) 4) "Because of the non-random addition of these teachers, this design of this overall project was considered to be quasi-experimental." (p. 994) Detailed Analysis: On its face, the four intermediate schools were randomly assigned to condition, which would normally satisfy (and exceed) the class-level randomisation requirement. However, the actual analytic sample was not the product of that randomisation alone: because too few teachers in the two comparison schools volunteered, the researchers non-randomly recruited additional comparison-condition teachers/classrooms to inflate the student sample size. The authors explicitly and repeatedly describe the resulting design, in the abstract, methods, and discussion/limitation sections, as "quasi-experimental" precisely because of this non-random addition. There is no tutoring/personal teaching exception here (the intervention is a whole-classroom science curriculum, not one-to-one instruction). Because randomisation was not properly and cleanly implemented for the sample that was actually analysed, the criterion cannot be considered met. Criterion C is not met because the study's own authors classify the design as quasi-experimental due to non-random recruitment of additional comparison-group teachers/classrooms.
    • E

      Exam-based Assessment

      • In addition to district-developed benchmark tests, the study used the state-mandated standardized TAKS assessment and the nationally standardized DIBELS Oral Reading Fluency measure, both widely recognized standardized instruments.
      • "The TAKS, a criterion-referenced assessment, measures student mastery of the content areas of state curriculum outlined in the TEKS." (p. 999)
      • Relevant Quotes: 1) "State Standardized Test: Texas Assessment of Knowledge and Skills (TAKS). The TAKS, a criterion-referenced assessment, measures student mastery of the content areas of state curriculum outlined in the TEKS." (p. 999) 2) "The English form has a reported internal consistency ranging from 0.87 to 0.90, and predictive validity ranging from 0.56 to 0.79 with SAT and ACT (TEA, 2008)." (p. 1000) 3) "Dynamic Indicators of Basic Literacy Skills (DIBELS). DIBELS (Good and Kaminiski, 2002) includes a set of procedures and measures for assessing the acquisition of early literacy skills... ORF is reported to have a median alternate form reliability of 0.95..." (p. 1000) 4) "These benchmark tests were developed according to the scope and sequence of the fifth grade Texas Essential Knowledge and Skills (TEKS)... The information is limited about the reliability and validity of the benchmark tests at the district level..." (pp. 998-999) Detailed Analysis: The study used three tiers of assessment: (a) district-developed benchmark tests (custom-built and not standardized statewide), (b) the state-mandated TAKS, and (c) the nationally-used DIBELS Oral Reading Fluency subtest. Both TAKS and DIBELS are widely recognised, standardized instruments used far beyond this single study (TAKS statewide in Texas for Grades 3-12; DIBELS nationally), with documented psychometric properties reported independently of the authors' own study. Although the district benchmark tests are researcher/district-created and would not alone satisfy this criterion, the presence of TAKS and DIBELS as core outcome measures satisfies the requirement for a standardised, widely recognised exam-based assessment. Criterion E is met because the study's outcome measures include the state-standardized TAKS assessment and the nationally standardized DIBELS ORF subtest.
    • T

      Term Duration

      • The intervention and outcome measurement spanned the full 2009-2010 school year, with final outcomes measured via TAKS and DIBELS in spring 2010, far exceeding one academic term.
      • "Data were collected in the fall and spring of school year 2009-2010." (p. 1001)
      • Relevant Quotes: 1) "Data were collected in the fall and spring of school year 2009-2010. Science benchmark test was administered to each student every 6 weeks during the school year with a total of six tests." (p. 1001) 2) "TAKS data were collected during the spring of 2010." (p. 1001) 3) "In this study, Oral Reading Fluency (ORF) was administered to students in the beginning and end of fifth grade." (p. 1000) Detailed Analysis: The intervention began in the fall of the 2009-2010 school year and the primary outcome measures (TAKS, end-of-year DIBELS ORF) were collected in spring 2010, after a full academic year. This interval is far longer than the minimum one-term requirement. Criterion T is met because outcomes were measured after a full school year, well beyond one academic term.
    • D

      Documented Control Group

      • The comparison group's schools, teachers, and students are documented in detail, including demographics, teaching experience, class sizes, and the typical-practice instruction they received.
      • "Table 1. School demographics for treatment and comparison group 2009-2010" (p. 995)
      • Relevant Quotes: 1) "Table 1 demonstrates the demographics of these four schools." (p. 994), showing African American, Hispanic, White, Native American, Asian, Low SES, ELL percentages, and academic rating for each comparison school (C and D). (p. 995) 2) "Table 2 ... Comparison C 2 [teachers] 6 [rotations] 118 [students]; D 2 3 48." (p. 995) 3) "In typical practice in the comparison group, science is taught by certified or permitted bilingual/ESL education and science education teachers in English for the ELL students with no Spanish clarifications... science instruction in comparison classrooms varied from 80 to 90 minutes daily including one 5-E lesson cycle weekly... Teachers followed a locally developed science curriculum aligned to the TEKS." (p. 998) 4) "Although no training support was provided by the research team to the teachers in the comparison group, they attended workshops as a state requirement to fulfill a minimum of 30 professional development hours each year related to their content area. Typical practice also had computers, projectors, and document cameras (ELMOs) in the science classrooms." (p. 998) Detailed Analysis: The paper documents the comparison ("typical practice") condition thoroughly: school-level demographic tables, teacher/student counts by rotation, average teaching experience (8.4 years), and a detailed narrative of what typical-practice instruction looked like (curriculum, lesson length, classroom observations, existing technology, state-required professional development hours). This level of detail satisfies the documentation requirement. Criterion D is met because the comparison group's composition, baseline characteristics, and instructional conditions are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Although the four schools were originally randomly assigned, the comparison group's sample was expanded with non-randomly recruited teachers, so the study does not represent a properly implemented school-level RCT.
      • "Because of the non-random addition of these teachers, this design of this overall project was considered to be quasi-experimental." (p. 994)
      • Relevant Quotes: 1) "four intermediate schools with principals' approval from the district site were randomly assigned to conditions, resulting in two treatment ... and two comparison ... schools." (p. 994) 2) "we recruited four more teachers within the same comparison schools that were already assigned based on the criteria that they had Spanish-speaking ELLs in their classrooms. Because of the non-random addition of these teachers, this design of this overall project was considered to be quasi-experimental." (p. 994) Detailed Analysis: The S criterion requires a properly implemented school-level randomisation. While the four schools were the randomised units, the final comparison-group sample used in the analysis was augmented through non-random recruitment of additional teachers within those same schools. The authors' own explicit classification of the design as "quasi-experimental" because of this reflects that the stronger school-level RCT requirement is not properly met for the sample actually analysed. Criterion S is not met because the analysed comparison sample was not the product of clean school-level randomisation alone.
    • I

      Independent Conduct

      • The same university research team that designed the intervention, trained the teachers, and monitored fidelity also conducted the evaluation and analysis, with no independent third-party evaluator.
      • "This component consisted of ongoing training workshops for both teachers (biweekly) and paraprofessionals (monthly) provided by research coordinators for 3 hours per session." (p. 995)
      • Relevant Quotes: 1) "Our study was derived from a larger longitudinal (fifth to sixth grade), field-based research project targeting native both ELLs and low socioeconomic status (SES) non-ELLs..." (p. 994), authored and led by the Texas A&M/Sam Houston State University research team. 2) "This component consisted of ongoing training workshops for both teachers (biweekly) and paraprofessionals (monthly) provided by research coordinators for 3 hours per session." (p. 995) 3) "Observers were educators and had attended the professional workshops and had been trained by the principal investigators." (p. 1000) 4) Contract grant sponsor: National Science Foundation (NSF, grant # 0822343 to Texas A&M University, 0822153 to Sam Houston State University). (p. 987, footer) Detailed Analysis: The intervention was designed by the same research team (the paper's authors and their university-based research coordinators), who also delivered teacher training, conducted fidelity observations (with observers trained directly by the principal investigators), and performed the outcome analysis. There is no mention of an external, independent evaluation agency or blinded data collectors separate from the intervention design team. This concentration of design, delivery, monitoring, and analysis within one research group does not satisfy the independence requirement. Criterion I is not met because the same research team designed, implemented, monitored, and evaluated the intervention without independent third-party oversight.
    • Y

      Year Duration

      • The intervention and follow-up measurement covered the entire 2009-2010 academic year, satisfying the stronger year-duration requirement.
      • "The larger research project from which the current study was derived, to our best knowledge, is the only longitudinal quasi-experimental design with science intervention of a full academic year among fifth grade ELLs and low-SES English proficient students." (p. 1006)
      • Relevant Quotes: 1) "The larger research project from which the current study was derived, to our best knowledge, is the only longitudinal quasi-experimental design with science intervention of a full academic year among fifth grade ELLs and low-SES English proficient students." (p. 1006) 2) "Data were collected in the fall and spring of school year 2009-2010... TAKS data were collected during the spring of 2010. ... Data from DIBELS were collected at the beginning and end of school year." (p. 1001) Detailed Analysis: The intervention was implemented across the full 2009-2010 school year, with pre-measures in fall 2009 and outcome measures (benchmark tests throughout the year, TAKS and DIBELS ORF) collected at the end of the year in spring 2010. This spans the full academic year (well over the 75% threshold). Criterion Y is met because the study tracked students across an entire academic year from intervention start to final outcome measurement.
    • B

      Balanced Control Group

      • Treatment classrooms received substantial additional, unmatched resources (professional development, EduSmart software, Science Saturdays, Family Involvement take-home packets) that the authors themselves acknowledge as a likely confound, without the study framing these extra resources as the explicit treatment variable being isolated.
      • "Another limitation is that we compared conditions that differed in several enrichment components (e.g., family science support, Saturday science activities, and extra technology resources in the classroom) that may seem to confound the ability to compare the nature-of-the-instruction treatment." (p. 1004)
      • Relevant Quotes: 1) "This component consisted of ongoing training workshops for both teachers (biweekly) and paraprofessionals (monthly) provided by research coordinators for 3 hours per session." (p. 995), versus "Although no training support was provided by the research team to the teachers in the comparison group..." (p. 998) 2) "treatment classrooms were equipped with computers, a projector, a document camera, an interactive whiteboard, science-based educational software, such as EduSmart, internet resources, and a digital camera ... Science Saturdays with scientists. Over 150 treatment students traveled to one of the research universities... Family involvement in science (FIS). Take-home science materials were developed for students to work with their parents/family..." (p. 997) 3) "Typical practice also had computers, projectors, and document cameras (ELMOs) in the science classrooms. The major technological difference was that treatment classrooms had EduSmart, a science software aligned with the intervention." (p. 998) 4) "Another limitation is that we compared conditions that differed in several enrichment components (e.g., family science support, Saturday science activities, and extra technology resources in the classroom) that may seem to confound the ability to compare the nature-of-the-instruction treatment." (p. 1004) 5) "Our point, from an instructional perspective, is how to best allocate and utilize those minutes so as to provide quality science instruction at school. Further, it was the multiple components in our intervention that differed from the typical practice and served as the contrast, not the time..." (pp. 1005-1006) Detailed Analysis: Applying the decision tree: extra resources are clearly present (biweekly teacher/paraprofessional professional development, EduSmart software and extra hardware, Science Saturdays at a university, and Family Involvement in Science take-home packets), and these are not negligible in scope (3-hour biweekly PD sessions, multi-hour university visits). While the core instructional design (inquiry-based, literacy integrated lessons via the 5-E model) could be argued to be the intervention itself, the paper explicitly separates "enrichment components" (technology, Saturday science, family involvement) from the "well-controlled components," and does not frame these enrichment resources as the deliberate treatment variable under test; rather, the stated research question is about "nature-of-the-instruction," i.e., teaching approach, not extra resource provision. The comparison group received none of these extra resources, and the authors explicitly flag this as a limitation that "may seem to confound the ability to compare the nature-of-the-instruction treatment," i.e., an admission that the resource imbalance is not cleanly isolated from the instructional contrast being tested. Because the additional resources are substantial, non-negligible, not matched in the control group, and not clearly and exclusively framed as the explicit treatment variable (indeed the authors state the intended contrast was "not the time" or resources but instructional practice), this criterion is not met. Criterion B is not met because treatment classrooms received several substantial, unmatched extra resources that the authors themselves acknowledge may confound the intended instructional comparison.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study or intervention (MSSELL) by a different research team is reported in the paper, and none was found through additional internet searches (Crossref, Semantic Scholar, OpenAlex citation records).
      • Relevant Quotes: 1) No quote in the paper references a subsequent independent replication of this specific study; the paper instead cites related but distinct studies (Amaral et al., 2002; Lee et al., 2005, 2008; August et al., 2009; Santau et al., 2011) which are related science-inquiry/literacy interventions by other teams but not replications of this specific MSSELL project's design, sample, or measures. 2) "The larger research project from which the current study was derived, to our best knowledge, is the only longitudinal quasi-experimental design with science intervention of a full academic year among fifth grade ELLs..." (p. 1006), indicating the authors view their design as novel rather than a replication of prior work, and there is no forward reference to any subsequent replication. Detailed Analysis: Criterion R requires evidence that this specific study's intervention and design were independently replicated by a different research team in a peer-reviewed outlet. The paper cites related prior work on science-inquiry interventions for ELLs, but these are precursor/contextual studies, not replications of this particular MSSELL project. A search of Crossref, Semantic Scholar, and OpenAlex citation records for this paper (140 citing works) found subsequent studies that adopted similar literacy-integrated science instruction frameworks for English learners, but none that independently replicate this specific study's design, sample, and measures. No later independent replication of this specific study was found. Criterion R is not met because no independent replication of this specific study was found either in the paper or via internet search.
    • A

      All-subject Exams

      • The study measured only science and English reading outcomes, not the full range of core subjects such as mathematics or social studies.
      • "Results are presented by construct measured, that is, science and reading achievement." (p. 1001)
      • Relevant Quotes: 1) "Do students who are enrolled in the literacy-embedded science treatment condition classrooms perform better on the district science benchmark tests, and state standardized science test than do students in a comparison condition? ... on the district reading benchmark tests, and state standardized reading test... on the standardized English decoding measure..." (p. 994) 2) "In this study, we investigated the effectiveness of a quasi-experimental study on science intervention among fifth grade ELLs. We compared students' performance between treatment and comparison groups in the district benchmark tests in science and reading, state standardized tests in science and reading, and literacy development." (p. 1001) Detailed Analysis: The study's outcome measures are confined to science achievement, English reading/literacy achievement, and oral reading fluency. No measures of mathematics, social studies, or other core subjects are reported, and no rationale is given for why only these two domains are assessed (unlike the "specialized vocational" exception, this is a general fifth-grade curriculum where math and social studies are also accountability subjects per the paper's own opening discussion of Texas accountability). Therefore all-subject coverage is not demonstrated. Criterion A is not met because only science and reading/literacy outcomes were assessed, with no measurement of other core subjects such as mathematics or social studies.
    • G

      Graduation Tracking

      • The study tracked students only through the end of fifth grade with no follow-up to graduation; the authors note plans to continue only into sixth grade, and no later publication tracking this cohort further was found via internet search.
      • "Further, as we plan for the next step of the larger research project, we will take into consideration the rotation/classroom nature of science instruction in these schools as students complete sixth grade..." (p. 1004)
      • Relevant Quotes: 1) "A battery of assessments was given to student participants to evaluate the effect after 1 year of implementation of this project." (p. 998) 2) "Given the fact that the purpose of this current study was to evaluate the effect of first year of implementation, we did not present the results using multi-level structure." (p. 1003) 3) "Further, as we plan for the next step of the larger research project, we will take into consideration the rotation/classroom nature of science instruction in these schools as students complete sixth grade, and we will include the hierarchical modeling with cross-classification in the final analysis." (p. 1004) Detailed Analysis: This paper reports only the first-year (fifth grade) results of a larger longitudinal fifth-to-sixth-grade project. There is no data or discussion in this paper of outcomes tracked through graduation (e.g., end of middle school or high school); only a stated plan to continue into sixth grade is mentioned, which itself falls well short of tracking to graduation. A search of Crossref, Semantic Scholar, and OpenAlex for later publications by Lara-Alecio, Tong, Irby, Guerrero, or Huerta following this same fifth-grade cohort into sixth grade or beyond found no such follow-up paper; later works by overlapping authors (e.g., Irby et al., 2021; Tong et al., 2020; Guerrero et al., 2023) address different cohorts and grade levels, not a continuation of this MSSELL sample. Criterion G is not met because follow-up in this paper ends after one year (fifth grade), with no tracking to graduation, and no subsequent graduation-tracking publication for this cohort was located.
    • P

      Pre-Registered

      • No statement of pre-registration of the study protocol on a public registry prior to data collection is present anywhere in the paper, and no pre-registration record for this study was found via internet search.
      • Relevant Quotes: 1) No mention of a registered protocol, registry platform, or registration date appears anywhere in the Methods, Data Collection and Analysis, or acknowledgment sections of the paper. 2) "Contract grant sponsor: National Science Foundation (NSF, grant # 0822343 to Texas A&M University, 0822153 to Sam Houston State University)." (p. 987) is the only administrative disclosure provided, and it references funding, not pre-registration. Detailed Analysis: The paper contains no reference to any trial registry (e.g., ClinicalTrials.gov equivalent for education studies), no registration ID, and no statement of when a study protocol or analysis plan was published prior to data collection. This is unsurprising given the study was conducted in 2009-2010, before pre- registration of education RCTs became common practice. An internet search did not locate any pre-registration record (e.g., on AEA RCT Registry, OSF, or similar platforms) for this study or the larger MSSELL project. Absent any such evidence, this criterion cannot be considered met. Criterion P is not met because there is no mention of pre-registration anywhere in the paper, and no pre-registration record was found online.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.