The impact of a systematic and explicit vocabulary intervention in Spanish with Spanish-speaking English learners in first grade

Johanna Cena, Doris Luft Baker, Edward J. Kame'enui, Scott K. Baker, Yonghan Park, Keith Smolkowski

Published:
ERCT Check Date:
DOI: 10.1007/s11145-012-9419-y
  • reading
  • language arts
  • L2 languages
  • K12
  • US
0
  • C

    Randomisation was carried out at the level of individual EL students within each of the two schools, not at the class or school level, and no one-to-one tutoring exception applies since students were taught in rotating small groups of 4-6.

    "ELs within each school were randomly assigned to either the VE-SETR (i.e., the treatment group), or the G-SETR (i.e., the comparison group)." (p. 1295)

  • E

    Alongside a custom researcher-developed vocabulary test, the study used three widely recognized, standardized, norm-referenced instruments (BVAT, TVIP, IDEL FLO) to measure outcomes.

    "The TVIP is an individually administered norm-referenced measure of receptive vocabulary and a screening test of verbal ability in Spanish (Dunn et al., 1986)." (p. 1298)

  • T

    The intervention and its outcome measurement spanned only about 8-10 weeks in total, far short of the one-term (3-4 month) minimum, a limitation the authors themselves acknowledge.

    "This study has two major limitations. First, the short duration of the intervention (8 weeks) may have accounted for the lack of statistically significant effects on more general outcome measures..." (p. 1306)

  • D

    The comparison (G-SETR) group's curriculum, training, group sizes, and pre/posttest characteristics are documented in detail in the text and in Tables 1 and 2.

    "The G-SETR vocabulary intervention was imbedded within the 90-min reading block using the Tesoros reading curriculum (Duran et al., 2008) and the systematic and explicit routines (i.e., the SETR)." (p. 1296)

  • S

    Only two schools participated, and randomisation occurred among individual EL students within each school rather than assigning whole schools to conditions.

    "ELs within each school were randomly assigned to either the VE-SETR (i.e., the treatment group), or the G-SETR (i.e., the comparison group)." (p. 1295)

  • I

    The same University of Oregon research team that designed the VE-SETR intervention also trained, coached, and observed teachers, with no independent or external evaluator mentioned.

    "Teachers also received coaching support from the researcher/trainer four times (once every 2 weeks) for approximately 30 min each time during the course of the 8-week study." (p. 1296)

  • Y

    Since criterion T (Term Duration) is not met, the stronger Year Duration criterion cannot be met either.

    "This study has two major limitations. First, the short duration of the intervention (8 weeks) may have accounted for the lack of statistically significant effects..." (p. 1306)

  • B

    Both groups received identical instructional time and the same 32 target words; the only extra resource (additional teacher training in the VE-SETR script) is integral to the instructional method being tested.

    "Teachers participating in the VE-SETR group received the same training as teachers in the G-SETR group with an additional 3-h training for the VE-SETR intervention." (p. 1296)

  • R

    No independent replication of this specific VE-SETR versus G-SETR study by an unrelated research team was found; two related studies located via citation search share overlapping authors with the original team.

  • A

    All outcome measures assess vocabulary, language proficiency, or oral reading fluency only; no other core subjects such as mathematics or science were assessed.

    "This study examined the impact of a 15-min daily explicit vocabulary intervention in Spanish on expressive and receptive vocabulary knowledge and oral reading fluency in Spanish, and on language proficiency in English." (Abstract)

  • G

    Since criterion Y (Year Duration) is not met, this criterion cannot be met either; the paper reports only immediate pre/post outcomes from the 8-week intervention, and no follow-up publication tracking this cohort to graduation was located.

    "The purpose of the larger research program was to examine the impact that Systematic and Explicit Teaching Routines (SETR), used in Spanish in first grade and in English in second grade, had on Spanish-speaking EL English and Spanish reading outcomes at the end of first, second, and third grades." (p. 1294)

  • P

    No statement of a pre-registered protocol, registry name, registration ID, or registration date is present anywhere in the paper, and no registry entry was found through internet search.

Abstract

This study examined the impact of a 15-min daily explicit vocabulary intervention in Spanish on expressive and receptive vocabulary knowledge and oral reading fluency in Spanish, and on language proficiency in English. Fifty Spanish-speaking English learners who received 90 min of Spanish reading instruction in an early transition model were randomly assigned to a treatment group (Vocabulary Enhanced Systematic and Explicit Teaching Routines [VE-SETR]) or a comparison group that received general vocabulary instruction using the standard reading curriculum with general strategies designed to increase the explicitness of instruction (General Systematic and Explicit Teaching Routines). Results indicated a statistically significant difference in depth of student Spanish vocabulary knowledge favoring the VE-SETR group. Differences on language proficiency in English, general vocabulary knowledge in Spanish, and oral reading fluency in Spanish were not statistically significant. Implications for future research are discussed.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was carried out at the level of individual EL students within each of the two schools, not at the class or school level, and no one-to-one tutoring exception applies since students were taught in rotating small groups of 4-6.
      • "ELs within each school were randomly assigned to either the VE-SETR (i.e., the treatment group), or the G-SETR (i.e., the comparison group)." (p. 1295)
      • Relevant Quotes: 1) "Fifty Spanish-speaking English learners who received 90 min of Spanish reading instruction in an early transition model were randomly assigned to a treatment group (Vocabulary Enhanced Systematic and Explicit Teaching Routines [VE-SETR]) or a comparison group..." (Abstract) 2) "ELs within each school were randomly assigned to either the VE-SETR (i.e., the treatment group), or the G-SETR (i.e., the comparison group)." (p. 1295) 3) "Students received reading instruction from nine teachers (i.e., in nine groups), four were certified teachers, and five were instructional assistants... Group size in the VE-SETR and the G-SETR group varied between four and six students." (p. 1295) 4) "To reduce teacher effects, every 2 weeks teachers and instructional assistants moved from one small group to another within their assigned condition." (p. 1295) Detailed Analysis: The unit of randomisation described in the Method section is the individual EL student: "ELs within each school were randomly assigned" to VE-SETR or G-SETR. This is explicitly student-level, not class-level or school-level, randomisation. Students were subsequently taught in small pull-out groups of four to six children led by one of nine rotating teachers/instructional assistants per condition, but this is a small-group instructional format, not the one-to-one personal tutoring described in the ERCT exception for criterion C. The paper never frames the intervention as individual tutoring, and does not provide the quote-level statement the exception procedure requires. Because randomisation occurred below the class level and the tutoring exception is not clearly applicable, contamination between conditions within the same schools cannot be ruled out. Criterion C is not met because randomisation was at the individual student level within schools rather than at the class or school level, and the small-group (4-6 student) delivery format does not satisfy the one-to-one tutoring exception.
    • E

      Exam-based Assessment

      • Alongside a custom researcher-developed vocabulary test, the study used three widely recognized, standardized, norm-referenced instruments (BVAT, TVIP, IDEL FLO) to measure outcomes.
      • "The TVIP is an individually administered norm-referenced measure of receptive vocabulary and a screening test of verbal ability in Spanish (Dunn et al., 1986)." (p. 1298)
      • Relevant Quotes: 1) "In this researcher-developed measure, students in the treatment and the comparison groups were tested before and after the study on all 32-target words (16 nouns and 16 verbs)." (p. 1297, DOK measure) 2) "This test measures a child's ability to use two languages to negotiate the meaning of academic content. It consists of three subtests from the Woodcock-Johnson Tests of Achievement-Revised... The norming sample included 5,602 subjects from over 100 different US communities." (BVAT, p. 1298) 3) "The TVIP is an individually administered norm-referenced measure of receptive vocabulary and a screening test of verbal ability in Spanish (Dunn et al., 1986). The test contains 125 translated items from the Peabody Picture Vocabulary Test-Revised... The test was normed with students from Mexico and Puerto Rico." (p. 1298) 4) "FLO is a standardized, timed, individually administered test of accuracy and fluency reading connected text in Spanish. It is part of the Indicadores Dinámicos del Exito en la Lectura, IDEL (Baker, Good III, Knutson, & Watson, 2006)." (p. 1299) Detailed Analysis: The study's primary hypothesis-testing outcome, the Depth of Knowledge (DOK) measure, is a custom researcher-developed instrument tailored to the 32 taught words, and by itself would not satisfy this criterion. However, the paper also administered and formally analyzed three additional, well-established standardized instruments: the Bilingual Verbal Ability Test (BVAT), the Test de Vocabulario en Imágenes Peabody (TVIP), and the IDEL Fluidez en la Lectura Oral (FLO). Each of these is described with norming samples, published reliability, and validity evidence, and none was created specifically for this study. Since the study reports outcomes on genuinely standardized, widely used exam-based assessments in addition to its custom measure, the requirement for a standard exam-based assessment is satisfied. Criterion E is met because the study used multiple widely recognized standardized instruments (BVAT, TVIP, IDEL FLO) to assess outcomes.
    • T

      Term Duration

      • The intervention and its outcome measurement spanned only about 8-10 weeks in total, far short of the one-term (3-4 month) minimum, a limitation the authors themselves acknowledge.
      • "This study has two major limitations. First, the short duration of the intervention (8 weeks) may have accounted for the lack of statistically significant effects on more general outcome measures..." (p. 1306)
      • Relevant Quotes: 1) "The VE-SETR intervention focused on teaching 32 vocabulary words during an 8-week period, approximately four words per week..." (p. 1295) 2) "FLO was given to all students approximately 2 weeks prior to the beginning of the intervention and approximately 2 weeks after the end of the intervention." (p. 1299) 3) Students were administered the TVIP "a week before the study began and approximately 1 week after the study was completed." (p. 1298-1299) 4) "This study has two major limitations. First, the short duration of the intervention (8 weeks) may have accounted for the lack of statistically significant effects on more general outcome measures such as the TVIP, the BVAT, and a negative effect on oral reading fluency." (p. 1306) Detailed Analysis: The core intervention lasted 8 weeks, and outcome measures were collected within about 1-2 weeks before and after this window, giving a total interval of roughly 8-12 weeks from intervention start to final measurement. This falls well short of the ERCT standard's minimum of one full academic term (approximately 3-4 months). The authors themselves explicitly identify the "short duration of the intervention (8 weeks)" as a study limitation. Criterion T is not met because the tracked interval from intervention start to outcome measurement was only about 8-12 weeks, shorter than one academic term.
    • D

      Documented Control Group

      • The comparison (G-SETR) group's curriculum, training, group sizes, and pre/posttest characteristics are documented in detail in the text and in Tables 1 and 2.
      • "The G-SETR vocabulary intervention was imbedded within the 90-min reading block using the Tesoros reading curriculum (Duran et al., 2008) and the systematic and explicit routines (i.e., the SETR)." (p. 1296)
      • Relevant Quotes: 1) "a comparison group that received general vocabulary instruction using the standard reading curriculum with general strategies designed to increase the explicitness of instruction (General Systematic and Explicit Teaching Routines)." (Abstract) 2) "The G-SETR vocabulary intervention was imbedded within the 90-min reading block using the Tesoros reading curriculum (Duran et al., 2008) and the systematic and explicit routines (i.e., the SETR). Teachers taught the reading program following the specified scope, sequence, and instructional guidance of the Tesoros reading program." (p. 1296) 3) Table 1, "Differences and similarities between the VE-SETR and the G-SETR vocabulary intervention," compares number of words, amount of instructional time, activities, training, and observations for both groups. (p. 1297) 4) Table 2 reports pretest and posttest means and standard deviations for the G-SETR group (N = 26) and VE-SETR group (N = 24) on every outcome measure. (p. 1301) Detailed Analysis: The paper provides a clear, detailed description of what the control (G-SETR) group received: the same core Tesoros reading curriculum, the same 32 target words, 15 min of daily vocabulary instruction within the same 90-min reading block, comparable teacher training, and the same fidelity-of-implementation observation protocol as the treatment group (Table 1). Table 2 further documents group sizes and baseline/ posttest scores on every measure, allowing direct comparison of the control group's characteristics against the treatment group. Criterion D is met because the control group's curriculum, instructional time, training, and baseline/posttest characteristics are documented in detail.
  • Level 2 Criteria

    • S

      School-level RCT

      • Only two schools participated, and randomisation occurred among individual EL students within each school rather than assigning whole schools to conditions.
      • "ELs within each school were randomly assigned to either the VE-SETR (i.e., the treatment group), or the G-SETR (i.e., the comparison group)." (p. 1295)
      • Relevant Quotes: 1) "ELs within each school were randomly assigned to either the VE-SETR (i.e., the treatment group), or the G-SETR (i.e., the comparison group)." (p. 1295) 2) "Both schools participated in the larger SETR study, and both schools used the Macmillan-McGraw Hill Tesoros (Duran et al., 2008) core reading program to teach reading in Spanish." (p. 1294-1295) Detailed Analysis: The study involved only two schools, and within each school individual EL students (not entire schools) were randomly allocated to the VE-SETR or G-SETR condition. Both conditions were present within both schools, which is the opposite of the school-level design this criterion requires (whole schools assigned entirely to one condition). Criterion S is not met because randomisation occurred at the individual student level within each of two schools, not at the school level.
    • I

      Independent Conduct

      • The same University of Oregon research team that designed the VE-SETR intervention also trained, coached, and observed teachers, with no independent or external evaluator mentioned.
      • "Teachers also received coaching support from the researcher/trainer four times (once every 2 weeks) for approximately 30 min each time during the course of the 8-week study." (p. 1296)
      • Relevant Quotes: 1) "Once words were selected, the first author reviewed the selected words with first grade teachers to determine the most appropriate vocabulary words to highlight in the lesson." (p. 1295) 2) "Teachers also received coaching support from the researcher/trainer four times (once every 2 weeks) for approximately 30 min each time during the course of the 8-week study." (p. 1296) 3) "This research was supported by Grant No. R324A090104A... funded by the US Department of Education Institute of Education Sciences to the Center on Teaching and Learning at the University of Oregon." (Acknowledgments, p. 1307) Detailed Analysis: The VE-SETR intervention was designed by the paper's first author in conjunction with the Center on Teaching and Learning research team, who also trained teachers, delivered ongoing coaching directly to VE-SETR teachers, and observed fidelity of implementation. There is no mention anywhere in the paper of an independent, third-party organization conducting data collection or analysis separate from the intervention design team; the same University of Oregon-based team appears responsible for design, training, coaching, and (implicitly) the reported analysis. Criterion I is not met because the intervention designers and the research/coaching team conducting the study are the same, with no documented independent evaluation.
    • Y

      Year Duration

      • Since criterion T (Term Duration) is not met, the stronger Year Duration criterion cannot be met either.
      • "This study has two major limitations. First, the short duration of the intervention (8 weeks) may have accounted for the lack of statistically significant effects..." (p. 1306)
      • Relevant Quotes: 1) "The VE-SETR intervention focused on teaching 32 vocabulary words during an 8-week period..." (p. 1295) 2) "This study has two major limitations. First, the short duration of the intervention (8 weeks) may have accounted for the lack of statistically significant effects on more general outcome measures..." (p. 1306) Detailed Analysis: The intervention and outcome tracking window spanned only about 8-12 weeks total, nowhere near the ~9-10 month academic year (or 75% thereof) required by this criterion. Per the ERCT standard's dependency rule, because the weaker Term Duration criterion (T) is not met, the stronger Year Duration criterion (Y) is automatically not met as well. Criterion Y is not met because the study duration (about 8-12 weeks) is far shorter than one academic term, let alone a full academic year.
    • B

      Balanced Control Group

      • Both groups received identical instructional time and the same 32 target words; the only extra resource (additional teacher training in the VE-SETR script) is integral to the instructional method being tested.
      • "Teachers participating in the VE-SETR group received the same training as teachers in the G-SETR group with an additional 3-h training for the VE-SETR intervention." (p. 1296)
      • Relevant Quotes: 1) Table 1: "Amount of time spent on vocabulary instruction — VE-SETR: Teachers provided 15 min of VE-SETR vocabulary instruction daily within the 90 min block... G-SETR: Teachers provided 15 min of G-SETR vocabulary instruction daily within the 90-min reading block..." (p. 1297) 2) "All 32 words in our study were taught in both conditions." (p. 1296) 3) "Teachers participating in the VE-SETR group received the same training as teachers in the G-SETR group with an additional 3-h training for the VE-SETR intervention." (p. 1296) 4) "Teachers also received coaching support from the researcher/trainer four times (once every 2 weeks) for approximately 30 min each time during the course of the 8-week study." (VE-SETR only, p. 1296) 5) "All teachers received three full days of training in the summer, two additional half-day trainings in October and December, and a 3-h training in February of the following year on how to use all the SETR..." (G-SETR/shared training, p. 1296) Detailed Analysis: Determining study intent: the intervention explicitly contrasts two ways of delivering the same amount of vocabulary instruction (a highly specified, scripted VE-SETR routine versus a teacher-discretion G-SETR routine), rather than testing whether extra student time or materials improve outcomes. Identifying resources: students in both conditions received the identical 15 minutes per day of vocabulary instruction within the same 90-min reading block, targeting the same 32 words, so there is no imbalance in student- facing instructional time or content. The only additional resource identified is teacher-facing: VE- SETR teachers received about 3 h of extra training plus roughly 2 h of ongoing coaching, on top of the substantial shared professional development (three summer days, two half-day sessions, and a 3-h session) that all teachers, including G-SETR teachers, received as part of the larger SETR project. Resource integration assessment: this extra teacher training is integral to implementing the specific scripted instructional method that constitutes the treatment variable itself, not a separable supplementary benefit to students; it is analogous to Example 2 in the ERCT standard's criterion-B guidance, where extra professional development tied to delivering the tested program does not violate balance. Applying the decision tree: extra resources are present (teacher training), but they are integral to the treatment being tested (RESOURCES_ARE_TREATMENT), and in addition student-facing time and content are matched (CONTROL_MATCHES_RESOURCES), so either branch of the decision tree independently supports "met." Since student time and materials were matched and the sole extra resource is intrinsic to the intervention's design, the control condition can be considered to have received a comparable "business as usual" level of core inputs. Criterion B is met because student instructional time and content were balanced between groups, and the additional teacher training given to the VE-SETR group is integral to the scripted instructional method under investigation rather than a separable, unmatched resource.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific VE-SETR versus G-SETR study by an unrelated research team was found; two related studies located via citation search share overlapping authors with the original team.
      • Relevant Quotes: 1) "This study was part of a 4-year, national longitudinal program of research conducted in 37 schools in Oregon, Washington, and Texas." (p. 1294) 2) No sentence in the paper claims or references an independent replication of this specific VE-SETR/ G-SETR contrast. Detailed Analysis: The paper does not reference any prior or concurrent independent replication of its findings. An internet citation search (OpenAlex/Crossref, 32 works citing this paper as of the check date) identified two related studies, neither of which is an independent replication because both share authors with the original team: (a) "Does Supplemental Instruction Support the Transition From Spanish to English Reading Instruction for First-Grade English Learners at Risk of Reading Difficulties?" (D. L. Baker, Burns, Kame'enui, Smolkowski, & S. K. Baker, 2016, Learning Disability Quarterly, DOI 10.1177/0731948715616757). It shares four of the six authors of the present study (D. L. Baker, Kame'enui, Smolkowski, S. K. Baker) and tests a different, broader supplemental-instruction design for at-risk readers rather than reproducing this specific VE-SETR vs. G-SETR contrast with the same population and measures. (b) "Exploring the Effects of a Spanish Vocabulary Intervention to Teach Words in Depth to Second-Grade Students in Chile" (D. L. Baker, Granada Azcárraga, Pomes Correa, Lepe-Martinez, & Smolkowski, 2019, Reading & Writing Quarterly, DOI 10.1080/10573569.2018.1523763). This is a similarly designed depth-of-vocabulary intervention conducted in a different country (Chile) and grade (second grade), but it shares two authors (D. L. Baker and Smolkowski) with the present study and so does not constitute an independent replication by a different research team. No independent replication by an unrelated research team of this specific study was located. Criterion R is not met because no independent replication of this specific study by a different, unrelated research team was found.
    • A

      All-subject Exams

      • All outcome measures assess vocabulary, language proficiency, or oral reading fluency only; no other core subjects such as mathematics or science were assessed.
      • "This study examined the impact of a 15-min daily explicit vocabulary intervention in Spanish on expressive and receptive vocabulary knowledge and oral reading fluency in Spanish, and on language proficiency in English." (Abstract)
      • Relevant Quotes: 1) "This study examined the impact of a 15-min daily explicit vocabulary intervention in Spanish on expressive and receptive vocabulary knowledge and oral reading fluency in Spanish, and on language proficiency in English." (Abstract) 2) The Measures section (pp. 1297-1299) lists only the Depth of Knowledge (DOK) vocabulary assessment, the Bilingual Verbal Ability Test (BVAT), the Test de Vocabulario en Imágenes Peabody (TVIP), and IDEL Fluidez en la Lectura Oral (FLO); no assessment of mathematics, science, or any other core subject is described anywhere in the paper. Detailed Analysis: Every outcome measure reported in this study targets vocabulary, language proficiency, or oral reading fluency in Spanish and/or English. The paper does not assess, mention, or provide any rationale for omitting other core academic subjects such as mathematics or science. No exception (e.g., a specialized vocational focus) is claimed or applicable here. Criterion A is not met because only language/ vocabulary/reading-fluency outcomes were assessed, with no coverage of other main subjects.
    • G

      Graduation Tracking

      • Since criterion Y (Year Duration) is not met, this criterion cannot be met either; the paper reports only immediate pre/post outcomes from the 8-week intervention, and no follow-up publication tracking this cohort to graduation was located.
      • "The purpose of the larger research program was to examine the impact that Systematic and Explicit Teaching Routines (SETR), used in Spanish in first grade and in English in second grade, had on Spanish-speaking EL English and Spanish reading outcomes at the end of first, second, and third grades." (p. 1294)
      • Relevant Quotes: 1) "This study was part of a 4-year, national longitudinal program of research conducted in 37 schools in Oregon, Washington, and Texas." (p. 1294) 2) "The purpose of the larger research program was to examine the impact that Systematic and Explicit Teaching Routines (SETR), used in Spanish in first grade and in English in second grade, had on Spanish-speaking EL English and Spanish reading outcomes at the end of first, second, and third grades." (p. 1294) 3) No results or follow-up data beyond the immediate posttest of the 8-week first-grade intervention are reported anywhere in this paper. Detailed Analysis: Although the paper mentions a larger 4-year longitudinal SETR research program that followed cohorts through third grade, this specific paper only reports pre- and post-test outcomes from the 8-week first-grade vocabulary intervention; no long-term or graduation-level follow-up data are presented. An internet search of papers citing this study and of the authors' subsequent publications (e.g., Baker et al., 2016, Learning Disability Quarterly; Baker et al., 2019, Reading & Writing Quarterly) found no follow-up publication that tracks this specific 50-student first-grade VE-SETR/G-SETR cohort through to graduation or any later stage. Per the ERCT standard's dependency rule, since criterion Y (Year Duration) is not met, criterion G (Graduation Tracking) is not met either. Criterion G is not met because no tracking beyond the immediate 8-week intervention posttest is reported in this paper or in any located follow-up publication, and criterion Y was not met.
    • P

      Pre-Registered

      • No statement of a pre-registered protocol, registry name, registration ID, or registration date is present anywhere in the paper, and no registry entry was found through internet search.
      • Relevant Quotes: No quotes were found in this paper referencing a study registry (e.g., ClinicalTrials.gov), a pre-registration identifier, or a pre-registration date. The Method, Data analysis, and Acknowledgments sections describe funding (Grant No. R324A090104A) and procedures but make no mention of prospective registration. Detailed Analysis: Absent any textual evidence of a pre-registered protocol or registry entry filed before data collection began, this criterion cannot be considered met. An internet search for a trial or study registry entry associated with Grant No. R324A090104A, the "SETR" project, or this specific VE-SETR/G-SETR study did not locate any pre-registration record; education RCTs funded around this period (the study predates its 2012 online publication) were not routinely pre-registered on a public registry. Criterion P is not met because no pre-registration statement, registry link, or registration date is provided in the paper or located through internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.