Using the SIOP Model to Promote the Acquisition of Language and Science Concepts with English Learners

Jana Echevarria, Catherine Richards-Tutor, Rebecca Canges, and David Francis

Published:
ERCT Check Date:
DOI: 10.1080/15235882.2011.623600
  • science
  • L2 languages
  • K12
  • US
0
  • C

    Randomisation was conducted at the school level, which is a stronger unit than class-level and therefore satisfies the class-level RCT requirement.

    "In this study, a small, cluster-randomized trial with randomization at the school level was used to examine the impact of the SIOP Model." (p. 338)

  • E

    The science/language assessments were custom-built by the project researchers for this study rather than being a standardised, widely recognised exam.

    "The student measure used was designed by project researchers who have an expertise in measurement to quantify acquisition of the concepts and language of science of each of four units of study." (p. 340)

  • T

    The active intervention and outcome measurement window spanned only about eight weeks, well short of one academic term.

    "Additionally, the intervention in this study was only 8 weeks long." (p. 348)

  • D

    The control group's instructional content, pacing, teacher characteristics, and student language-subgroup composition are documented in detail.

    "Control teachers taught the same four topics of study using methods that they would normally use to teach these topics." (p. 338)

  • S

    Entire schools, not classes or students, were the unit of random assignment, directly satisfying the school-level RCT requirement.

    "Ten middle schools in one large urban district in Southern California were randomly assigned to either treatment (SIOP Model) or control (normal classroom science instruction)." (p. 338)

  • I

    The same research team that developed the SIOP Model also designed the lesson plans, trained and coached the treatment teachers, and collected the outcome data, with no independent evaluator described.

    "Jana Echevarria is Professor Emeritus at California State University, Long Beach and is a coauthor of the SIOP Model." (p. 334)

  • Y

    Because Criterion T (Term Duration) is not met, Criterion Y is automatically not met; independently, the roughly 8-week intervention window is far short of 75% of an academic year.

    "Additionally, the intervention in this study was only 8 weeks long." (p. 348)

  • B

    Both conditions received the same instructional time and covered the same four topics via a shared pacing guide; the only differential resource was teacher-facing SIOP training, which is integral to testing the instructional method itself rather than an unmatched student-facing resource.

    "Both treatment and control teachers were given a pacing guide for teaching the instructional units. This was done to ensure that students in each condition were receiving approximately the same amount of instructional time on each unit." (p. 342)

  • R

    Neither the paper nor a search of subsequent literature identifies an independent replication of this specific school-level cluster RCT of the SIOP Model on science outcomes.

  • A

    Because Criterion E is not met, Criterion A is automatically not met; additionally, only science content and language were assessed, with no other core subjects measured.

    "In our study we examined the efficacy of a model of instruction for English learners, the Sheltered Instruction Observation Protocol (SIOP) Model, in one content area, science."

  • G

    Because Criterion Y is not met, Criterion G is automatically not met; the paper also describes no follow-up beyond the immediate end-of-unit posttests within the 8-week intervention window.

  • P

    No pre-registration of the study protocol, hypotheses, or analysis plan is mentioned anywhere in the paper, nor was any registry entry found in a search.

Abstract

In this article we report findings from research through the Center for Research on the Educational Achievement and Teaching of English Language Learners (CREATE), a National Research and Development Center. In our study we examined the efficacy of a model of instruction for English learners, the Sheltered Instruction Observation Protocol (SIOP) Model, in one content area, science. Assessments measured the acquisition of academic language and science concepts among English learners, former English learners, and English Only students in middle school science classrooms. Results indicated that students in the SIOP group performed better than controls, although not to a significant degree. Reasons for these findings are explored. Due to differential attrition of schools in the control group, caution must be used in interpreting the study's findings.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was conducted at the school level, which is a stronger unit than class-level and therefore satisfies the class-level RCT requirement.
      • "In this study, a small, cluster-randomized trial with randomization at the school level was used to examine the impact of the SIOP Model." (p. 338)
      • Relevant Quotes: 1) "In this study, a small, cluster-randomized trial with randomization at the school level was used to examine the impact of the SIOP Model." (p. 338) 2) "Ten middle schools in one large urban district in Southern California were randomly assigned to either treatment (SIOP Model) or control (normal classroom science instruction)." (p. 338) 3) "The schools in each category (large and moderate) were randomly assigned to either treatment (SIOP Model) or control (normal classroom science instruction), ensuring that there was an equal distribution of type of school population in each condition." (p. 339) Detailed Analysis: The paper explicitly and repeatedly states that entire middle schools, not individual classes or students within a single class, were the unit of randomisation. This eliminates the contamination risk that the C criterion is designed to guard against, and it exceeds the class-level requirement. Final sentence: Criterion C is met because randomisation occurred at the school level, a stronger unit than the class level required by the standard.
    • E

      Exam-based Assessment

      • The science/language assessments were custom-built by the project researchers for this study rather than being a standardised, widely recognised exam.
      • "The student measure used was designed by project researchers who have an expertise in measurement to quantify acquisition of the concepts and language of science of each of four units of study." (p. 340)
      • Relevant Quotes: 1) "The student measure used was designed by project researchers who have an expertise in measurement to quantify acquisition of the concepts and language of science of each of four units of study." (p. 340) 2) "All science language and concept assessments were pilot tested the year prior to full implementation of the study (Year 1) and were modified based on results of the pilot test for Year 2 implementation." (p. 340) 3) "For each unit (Cell Division, Cell Structure and Function, Photosynthesis and Respiration, and Genetics), the assessment consisted of a reading passage on the science topic to be tested ... followed by three to eight ... multiple choice questions, three to eight ... short answer/fill in the blank questions, and one to two essay questions." (p. 340) Detailed Analysis: The outcome measures were researcher-designed unit tests created specifically for the four instructional topics used in this study, piloted and revised by the project team itself, rather than a widely recognised standardised exam (e.g., a state or national assessment). The essay portion was scored with a standardised rubric (IMAGE), but this scoring rubric does not make the underlying assessment instrument itself a standardised exam-the reading passages and questions were custom-built to match the study's own lesson content. This is precisely the kind of study-aligned, custom assessment the ERCT standard is designed to flag as unmet. Final sentence: Criterion E is not met because the assessments were custom-designed by the research team for this specific study rather than a recognised standardised exam.
    • T

      Term Duration

      • The active intervention and outcome measurement window spanned only about eight weeks, well short of one academic term.
      • "Additionally, the intervention in this study was only 8 weeks long." (p. 348)
      • Relevant Quotes: 1) "Additionally, the intervention in this study was only 8 weeks long. This is a very short period of time in which to expect teachers to develop a strong working knowledge of the SIOP Model, and an equally short time for a change in instruction to significantly impact the achievement of middle school students." (p. 348) 2) "Before teaching a unit, the pretest was administered to the students, and upon completion of the unit, students were administered the unit's posttest." (p. 342) 3) "In all of the schools, seventh-grade science included one semester of Biology." (p. 339) Detailed Analysis: The authors themselves explicitly state that "the intervention in this study was only 8 weeks long." Outcomes were measured immediately at the end of each of the four short instructional units, with the final combined posttest occurring at the close of this roughly two-month period. Although the surrounding Biology semester and the broader 2-year project (Year 1 pilot plus Year 2 study) were longer, the actual intervention-to-measurement interval that matters for this criterion is only about 8 weeks, which is shorter than the minimum one academic term (approximately 3-4 months) required. Final sentence: Criterion T is not met because the intervention and its outcome measurement spanned only about 8 weeks, short of the one-term minimum.
    • D

      Documented Control Group

      • The control group's instructional content, pacing, teacher characteristics, and student language-subgroup composition are documented in detail.
      • "Control teachers taught the same four topics of study using methods that they would normally use to teach these topics." (p. 338)
      • Relevant Quotes: 1) "Control teachers taught the same four topics of study using methods that they would normally use to teach these topics." (p. 338) 2) "With the attrition of two schools in the control after randomization, there was a total of three schools in the control group and five schools in the SIOP condition. The percent of English learners at SIOP and control schools was similar: SIOP schools ranged from 4.9% to 27.2% English learners, and control schools ranged from 4.7% to 39.9% English learners." (p. 339) 3) "There were eight seventh-grade teachers in the treatment group and four in the control with a total of 27 sections of science classes included in the SIOP condition and 15 sections in the control condition ... All of the control teachers were fully certified in science ... three of the four control teachers had completed their EL authorization at the time of the study." (p. 339) 4) Table 1 reports sample sizes for the Comparison group by language designation: "English Learners 112 30.1; Fluent English Proficient (3 years or less) 121 32.5; Fluent English Proficient (More than 3 years) 20 8.1; English Only 109 29.3; Total 372." (p. 339) Detailed Analysis: The paper documents the control condition in substantial detail: number of control schools and teachers, teacher certification status, number of class sections, percentage of English learners at each control school, and a full breakdown of the control sample by language-proficiency subgroup (Table 1). It also clearly states that control teachers taught the identical four topics using their normal methods, giving a clear baseline against which the SIOP condition can be compared. Final sentence: Criterion D is met because the control group's composition, characteristics, and instructional content are thoroughly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Entire schools, not classes or students, were the unit of random assignment, directly satisfying the school-level RCT requirement.
      • "Ten middle schools in one large urban district in Southern California were randomly assigned to either treatment (SIOP Model) or control (normal classroom science instruction)." (p. 338)
      • Relevant Quotes: 1) "In this study, a small, cluster-randomized trial with randomization at the school level was used to examine the impact of the SIOP Model. Ten middle schools in one large urban district in Southern California were randomly assigned to either treatment (SIOP Model) or control (normal classroom science instruction)." (p. 338) 2) "The schools in each category (large and moderate) were randomly assigned to either treatment (SIOP Model) or control (normal classroom science instruction), ensuring that there was an equal distribution of type of school population in each condition." (p. 339) 3) "However, before the onset of data collection, two schools in the control condition dropped out of the study so results must be interpreted as a quasiexperiment, with some caution." (p. 338) Detailed Analysis: The randomisation unit was explicitly the school: ten schools were block-randomised (stratified by EL concentration) to treatment or control. This directly meets the S criterion's requirement. It should be noted, however, that two control schools withdrew before data collection began, leaving three control and five treatment schools; the authors themselves flag this as breaking the clean randomised design and recommend the results be read as a quasi-experiment. This attrition affects confidence in group equivalence but does not change the fact that the assignment mechanism itself was school-level randomisation. Final sentence: Criterion S is met because assignment to condition was performed at the school level, notwithstanding the post-randomisation attrition of two control schools noted as a limitation by the authors.
    • I

      Independent Conduct

      • The same research team that developed the SIOP Model also designed the lesson plans, trained and coached the treatment teachers, and collected the outcome data, with no independent evaluator described.
      • "Jana Echevarria is Professor Emeritus at California State University, Long Beach and is a coauthor of the SIOP Model." (p. 334)
      • Relevant Quotes: 1) "Jana Echevarria is Professor Emeritus at California State University, Long Beach and is a coauthor of the SIOP Model. She is a Co-PI for CREATE, the National Research and Development Center for English language learners funded by IES." (p. 334) 2) "Teachers in the SIOP condition received training in the SIOP Model and then taught four science units using lesson plans and teaching methods that followed the SIOP Model." (p. 338) 3) "To ensure that teachers' delivery of the lesson plans followed the SIOP Model, coaching was provided to each treatment teacher by researchers who were experienced in implementing the model ... the coach observed and rated the lesson using the SIOP protocol." (p. 342) 4) "Researchers distributed and collected the assessments for each unit at both the treatment and control sites." (p. 342) Detailed Analysis: The lead author is herself a co-author/co-developer of the SIOP Model being tested, and she and the co-authors ("the research team") developed the study's lesson plans, trained the treatment teachers in the model, coached and rated their fidelity to the model, and personally distributed and collected the assessments used for both conditions. No external, third-party organisation independent of the model developers is described as running the data collection, lesson delivery oversight, or analysis. This is the opposite configuration from the Independent Conduct criterion's requirement that evaluation be free of involvement from the intervention's designers. Final sentence: Criterion I is not met because the intervention's designers/co-developers also designed, delivered oversight (coaching/fidelity rating), and collected data for the study, with no independent evaluator identified.
    • Y

      Year Duration

      • Because Criterion T (Term Duration) is not met, Criterion Y is automatically not met; independently, the roughly 8-week intervention window is far short of 75% of an academic year.
      • "Additionally, the intervention in this study was only 8 weeks long." (p. 348)
      • Relevant Quotes: 1) "Additionally, the intervention in this study was only 8 weeks long." (p. 348) 2) "In all of the schools, seventh-grade science included one semester of Biology." (p. 339) Detailed Analysis: Per the ERCT specification, if criterion T is not met, Y is automatically not met. Independently of that rule, the paper's own statement that the intervention lasted only about 8 weeks, with outcomes measured at the close of that period, falls far short of the 75%-of-an-academic-year (roughly 9-10 months) threshold required for Y. Final sentence: Criterion Y is not met, both because T is not met and because the roughly 8-week intervention window is far shorter than the required academic-year duration.
    • B

      Balanced Control Group

      • Both conditions received the same instructional time and covered the same four topics via a shared pacing guide; the only differential resource was teacher-facing SIOP training, which is integral to testing the instructional method itself rather than an unmatched student-facing resource.
      • "Both treatment and control teachers were given a pacing guide for teaching the instructional units. This was done to ensure that students in each condition were receiving approximately the same amount of instructional time on each unit." (p. 342)
      • Relevant Quotes: 1) "Both treatment and control teachers were given a pacing guide for teaching the instructional units. This was done to ensure that students in each condition were receiving approximately the same amount of instructional time on each unit." (p. 342) 2) "Control teachers taught the same four topics of study using methods that they would normally use to teach these topics." (p. 338) 3) "Since the purpose of the Year 2 study was to test the impact of the SIOP Model on student achievement, treatment teachers did not receive any additional training in science content or scientific inquiry, only in the SIOP Model. They were provided an intensive two-and-a-half-day training to introduce them to the SIOP Model and its components." (p. 341) Detailed Analysis: Applying the Criterion B decision tree: the intervention being tested is a change in teaching method (the SIOP Model) delivered by treatment teachers, not an addition of extra instructional time, materials, or budget provided directly to students. Both groups of students covered the identical four topics on a shared pacing guide, explicitly designed by the authors "to ensure that students in each condition were receiving approximately the same amount of instructional time on each unit." The one differential resource-the two-and-a-half-day SIOP training and subsequent coaching-was given only to treatment teachers, not students, and is an integral, defining part of the instructional method being tested (how teachers teach), not a separable supplementary resource that could have been balanced by giving control teachers the same training without changing their teaching method. Since no extra time or budget was extended to students in the treatment condition relative to control, the resources are considered balanced. Final sentence: Criterion B is met because student instructional time and topic coverage were explicitly equalised via a shared pacing guide, and the only differential input (teacher training in the SIOP Model) is integral to the instructional method under test rather than an unmatched student-facing resource.
  • Level 3 Criteria

    • R

      Reproduced

      • Neither the paper nor a search of subsequent literature identifies an independent replication of this specific school-level cluster RCT of the SIOP Model on science outcomes.
      • Relevant Quotes: 1) No quotes within the paper reference a prior or subsequent independent replication of this specific study. Detailed Analysis: An internet search for independent replications of this specific Echevarria, Richards-Tutor, Canges, and Francis (2011) cluster-randomized trial of the SIOP Model in middle school science found only descriptions and citations of the original study itself (e.g., in a What Works Clearinghouse intervention report on SIOP, which reviewed this same study rather than an independent replication of it), and other SIOP-related publications by overlapping author teams (e.g., Echevarria, Richards-Tutor, Chinn, & Ratleff, 2011, on fidelity, which analyzes the same study's data rather than an independent replication). A check of the paper's full citation record (Semantic Scholar, ~90 citing works through 2025) likewise surfaces no independent RCT replication of this specific school-level design; citing works are either narrative reviews, critiques (e.g., Krashen, 2013, "Does SIOP Research Support SIOP Claims"), or separate SIOP applications in different study designs and contexts. No peer-reviewed study by a different, independent research team replicating this specific science-focused, school-level cluster RCT design was found. Final sentence: Criterion R is not met because no independent replication of this specific study was found in the paper, in subsequent literature, or in its citation record.
    • A

      All-subject Exams

      • Because Criterion E is not met, Criterion A is automatically not met; additionally, only science content and language were assessed, with no other core subjects measured.
      • "In our study we examined the efficacy of a model of instruction for English learners, the Sheltered Instruction Observation Protocol (SIOP) Model, in one content area, science."
      • Relevant Quotes: 1) "In our study we examined the efficacy of a model of instruction for English learners, the Sheltered Instruction Observation Protocol (SIOP) Model, in one content area, science." (p. 334) 2) "In the current study we extend the extant research on the SIOP Model to examine its efficacy in one content area, science." (p. 338) Detailed Analysis: Per the ERCT specification, Criterion A automatically fails if Criterion E is not met, which is the case here. Independently, the study's own framing and its four assessed units (Cell Division, Cell Structure and Function, Photosynthesis and Respiration, and Genetics) confirm that only science content and its associated academic language were measured; no other core subjects (e.g., mathematics, social studies, general language arts outside the science context) were assessed. Final sentence: Criterion A is not met, both because criterion E is not met and because the study measured outcomes in science only.
    • G

      Graduation Tracking

      • Because Criterion Y is not met, Criterion G is automatically not met; the paper also describes no follow-up beyond the immediate end-of-unit posttests within the 8-week intervention window.
      • Relevant Quotes: 1) No quotes describing any follow-up data collection beyond the combined posttest administered "upon completion of the unit" (p. 342) at the end of the 8-week intervention were found. Detailed Analysis: Per the ERCT specification, if criterion Y (Year Duration) is not met, criterion G is automatically not met. This is confirmed by the paper's content: the last data collection point described is the posttest administered at the end of the fourth instructional unit, within the roughly 8-week intervention period, with no mention of any subsequent tracking of students, let alone tracking through to graduation. A check of the paper's citation record (~90 citing works through 2025) and of publications by the same author team found no follow-up study tracking this same cohort of students toward graduation; the one same-cohort follow-up identified (Echevarria, Richards-Tutor, Chinn, & Ratleff, 2011, on teacher fidelity) analyzes instructional fidelity, not long-term student outcomes. Final sentence: Criterion G is not met because Y is not met and no follow-up tracking beyond the immediate posttests, in this paper or in later publications, is described.
    • P

      Pre-Registered

      • No pre-registration of the study protocol, hypotheses, or analysis plan is mentioned anywhere in the paper, nor was any registry entry found in a search.
      • Relevant Quotes: 1) No quotes referencing a study registry, protocol registration platform, or pre-specified analysis plan were found anywhere in the Methods, Procedures, or Analysis Model sections of the paper. Detailed Analysis: The Methods and Analysis Model sections describe the study design, measures, and multilevel analysis approach in detail but contain no statement of prior registration on a public registry (e.g., a clinical-trial-style registry) nor any date for such registration. This 2011 study predates the widespread use of education-RCT registries (e.g., the AEA RCT Registry, launched 2013; OSF Registries), and a search for external evidence of pre-registration of this specific study likewise found no registry entry or protocol paper predating data collection. Final sentence: Criterion P is not met because there is no evidence, in the paper or externally, of a pre-registered protocol prior to data collection.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.