Abstract
This study compared traditional and multimedia-enhanced read-aloud vocabulary instruction and investigated whether the effects differed for English-language learners (ELLs) and non-English-language learners (non-ELLs). Results indicate that although there was no added benefit of multimedia-enhanced instruction for non-ELLs, there was a positive effect for ELLs on a researcher-designed measure and on a measure of general vocabulary knowledge. Furthermore, for children in the multimedia-enhanced condition, the gap between non-ELLs and ELLs in knowledge of instructional words was closed, and the gap in general vocabulary knowledge was narrowed. The multimedia support did not negatively impact non-ELLs, indicating the potential of multimedia-enhanced vocabulary instruction for ELLs in inclusive settings.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Children, not intact classes, were the unit of random assignment, with students from the same grade being individually split across the two conditions and teachers.
- "Then, across pairs of teachers, children were randomly assigned to the nonmultimedia or the multimedia condition."
Relevant Quotes:
1) "For assignment to condition, teachers were paired across grades (e.g., the 2 first grade teachers were paired together). One teacher was randomly assigned to the nonmultimedia condition, and the other was assigned to the multimedia condition. Then, across pairs of teachers, children were randomly assigned to the nonmultimedia or the multimedia condition." (p. 309)
2) "For instance, of the 25 children in first grade, 13 were randomly assigned to the nonmultimedia condition, and 12 were randomly assigned to the multimedia condition." (p. 309)
3) "Therefore, due to random assignment to condition, some children received intervention from their regular homeroom teacher, and some received intervention from the other teacher at their grade level. Random assignment within grades rather than across the entire sample was necessary to limit disruption of the students' schedules." (p. 309)
Detailed Analysis:
Criterion C requires randomisation at the class (or stronger, school) level, not at the level of individual students within a class, to avoid contamination. Although teachers were first randomly assigned to a condition, the paper then explicitly states that individual children were randomly assigned to condition across the paired teachers, not as intact pre-existing classes. This means that within a single grade level, children originally in the same homeroom class were split between the two conditions, with "some children" moving to "the other teacher at their grade level." The explicit unit of randomisation for children, as described by the authors themselves, was the individual student, not the classroom.
This is not a one-to-one tutoring intervention (it is a small-group, classroom-based read-aloud and video intervention delivered to reconstituted groups of children), so the tutoring exception to the class-level requirement does not apply. Because children from the same original classroom cohort were individually reassigned to different condition groups and moved between teachers, there is a plausible risk of contamination (e.g., informal contact, shared recess/lunch, siblings) between conditions within the same grade and school, and the paper's own description characterises the unit of assignment as the individual child rather than the intact class.
Criterion C is not met because the paper explicitly describes random assignment of individual children to condition (crossing original classroom boundaries), not random assignment of intact classes, and no tutoring exception applies.
-
E
Exam-based Assessment
- The study used the PPVT-III, a widely recognised standardised vocabulary test, as one of its primary outcome measures.
- "The Peabody Picture Vocabulary Test–Third Edition (PPVT–III; Dunn & Dunn, 1997) was used to assess children's general vocabulary knowledge. This is a commonly used norm-referenced measure of receptive vocabulary..."
Relevant Quotes:
1) "General vocabulary knowledge. The Peabody Picture Vocabulary Test–Third Edition (PPVT–III; Dunn & Dunn, 1997) was used to assess children's general vocabulary knowledge. This is a commonly used norm-referenced measure of receptive vocabulary in which children choose one of four pictures that corresponds to the target word given orally by the test administrator." (p. 309)
2) "A researcher-designed measure was used to assess children's knowledge of words targeted in the intervention. The target vocabulary assessment (TVA) was based on the measure used in Beck and McKeown (2007)." (p. 309)
3) "Children's knowledge of science concepts taught in the intervention was assessed via a researcher-designed science assessment (SCI)." (p. 309)
Detailed Analysis:
The study used three outcome measures: a researcher-designed target vocabulary assessment (TVA), a researcher-designed science-content assessment (SCI), and the Peabody Picture Vocabulary Test-Third Edition (PPVT-III). The TVA and SCI are explicitly researcher-created instruments tailored to the specific intervention content, which would not on their own satisfy the standardised-exam requirement. However, the PPVT-III is a well-established, commercially published, norm-referenced standardised test of receptive vocabulary that is widely used in research and practice well beyond this study, and it served as one of the study's primary outcome measures (general vocabulary knowledge), with results reported and analysed in the same way as the other measures.
Because a widely recognised, standardised assessment (PPVT-III) was used as a primary outcome alongside the custom measures, the study satisfies the requirement of employing a standard, validated exam-based assessment, even though it also used supplementary researcher-made instruments.
Criterion E is met because the study used the PPVT-III, a widely recognised, standardised, norm-referenced vocabulary test, as one of its primary outcome measures.
-
T
Term Duration
- Outcomes were assessed at the end of the 12-week intervention, which approximates one academic term (roughly 3 months).
- "In both of these conditions, teachers implemented a scripted intervention lesson 45 min per day for 3 days a week over the course of 12 weeks." (p. 307)
Relevant Quotes:
1) "In both of these conditions, teachers implemented a scripted intervention lesson 45 min per day for 3 days a week over the course of 12 weeks. The length and duration of the intervention were held constant across conditions." (p. 307)
2) "The intervention in each of the conditions was similar in many ways. In both conditions, the intervention consisted of four 3-week cycles, one cycle for each habitat." (p. 308)
3) "Both prior to and following the intervention, children were assessed on knowledge of words targeted in the intervention, general vocabulary knowledge, and knowledge of science concepts taught in the intervention." (p. 309)
Detailed Analysis:
The intervention itself ran for 12 weeks (four 3-week cycles), and posttest outcome assessment occurred following the completion of this 12-week period. Twelve weeks is approximately three months, which falls within the ERCT standard's definition of a term ("typically defined as a semester or equivalent (approximately 3-4 months)"). While no exact calendar dates are given, the described duration from intervention start to the posttest measurement is close enough to a full academic term to satisfy this weaker, Level 1 duration criterion.
Final sentence: Criterion T is met because outcomes were measured after a 12-week intervention period, which approximates one academic term.
-
D
Documented Control Group
- The control (nonmultimedia) group's demographics, baseline scores, and instructional content are documented in detail, including formal tests confirming baseline equivalence with the treatment group.
- "The ANOVAs showed that there was no difference between conditions on pretest TVA... PPVT... or SCI... Furthermore, the chi-square tests revealed that there was no difference between conditions on language background... or SES..."
Relevant Quotes:
1) "Table 1. Demographic Characteristics of the Children in the Sample" listing grade, gender, race, socioeconomic status, and primary language broken down by "Nonmultimedia group" and "Multimedia group." (p. 310, Table 1)
2) "The ANOVAs showed that there was no difference between conditions on pretest TVA, F(1, 84) = 0.10, p = .756; PPVT, F(1, 84) = 0.60, p = .440; or SCI, F(1, 84) = 0.52, p = .471. Furthermore, the chi-square tests revealed that there was no difference between conditions on language background... or SES..." (p. 310)
3) "In the nonmultimedia condition, teachers read each book on 3 days. They also implemented scripted curricula that accompanied the read-aloud books for their condition." (p. 308)
4) Table 2 reports pretest and posttest means and standard deviations for the "Nonmultimedia group" separately for the total sample, non-ELLs, and ELLs on TVA, PPVT, and SCI. (p. 311, Table 2)
Detailed Analysis:
The nonmultimedia (control) condition is documented in considerable detail. Table 1 reports its demographic composition (grade level distribution, gender, race, socioeconomic status, and primary language) alongside the multimedia group. Table 2 reports its baseline (pretest) and posttest scores on all three outcome measures. The methods section explicitly describes what the control group received: the same scripted curriculum content and instructional time as the multimedia group, but delivered via three days of book reading per book cycle instead of two days of reading plus a video day. Preliminary chi-square and ANOVA tests explicitly confirm that the control group did not differ from the treatment group at baseline on demographics or pretest scores.
This level of detail — demographic breakdown, baseline equivalence testing, and a clear description of exactly what instructional content the control group received — satisfies the requirement for a well-documented control group.
Criterion D is met because the paper documents the control group's demographics, baseline scores, and exact instructional content in detail, including formal baseline-equivalence tests.
-
Level 2 Criteria
-
S
School-level RCT
- The study was conducted in a single school with randomisation at the teacher/child level, not across multiple schools.
- "The study was set in a small, semiurban public school in the northeast."
Relevant Quotes:
1) "The study was set in a small, semiurban public school in the northeast." (p. 307)
2) "For assignment to condition, teachers were paired across grades... One teacher was randomly assigned to the nonmultimedia condition, and the other was assigned to the multimedia condition. Then, across pairs of teachers, children were randomly assigned to the nonmultimedia or the multimedia condition." (p. 309)
Detailed Analysis:
The study took place in a single school, with randomisation occurring at the level of teacher pairs and individual children within that one school, not across multiple schools. There is no mention of any other school participating, and no school-level random assignment took place; instead, teachers and then individual children within the one school were assigned to condition. This falls far short of the S criterion, which requires randomisation among multiple schools (or equivalent institutional units).
Criterion S is not met because the study was conducted in a single school with randomisation of teachers and individual children, not of schools.
-
I
Independent Conduct
- The same researchers who designed the curriculum also trained teachers and led the study; only isolated scoring tasks were performed by blinded research assistants, not a fully independent evaluation team.
- "Teachers were trained on implementing the scripted curriculum during a 1-day in-school session with the researchers."
Relevant Quotes:
1) "Teachers were trained on implementing the scripted curriculum during a 1-day in-school session with the researchers." (p. 309)
2) "Finally, to document teacher fidelity to the intervention, a research assistant (RA) unaware of the study design completed fidelity checklists at two randomly selected times for each teacher." (p. 309)
3) "RAs scored this assessment by giving 1 point for every independent, correct statement about the habitat... Each child's assessment was scored separately by two RAs whose interrater reliability was 86%." (p. 309)
Detailed Analysis:
The intervention curriculum, videos, and assessment instruments (TVA, SCI) were designed by the study's own authors (Silverman & Hines), who also trained the teachers, observed classrooms for implementation feedback, and directed the overall research design and analysis. While a research assistant "unaware of the study design" scored fidelity checklists and other RAs scored the open-ended science assessment, this reflects blinding of specific scoring tasks, not an independent third-party organisation conducting the trial's design, data collection, and analysis as a whole. There is no statement that data collection or the overall evaluation was conducted by an external, independent evaluation team distinct from the intervention's designers; the same author-researchers who designed the curriculum appear to have led training, observation, and analysis throughout.
Criterion I is not met because the same researchers who designed the intervention also trained teachers, oversaw implementation, and led the overall study, with only isolated scoring tasks performed by blinded research assistants rather than a genuinely independent evaluation team.
-
Y
Year Duration
- The study spanned only about 12 weeks, far short of 75% of an academic year, and Criterion T is also not met.
- "In both of these conditions, teachers implemented a scripted intervention lesson 45 min per day for 3 days a week over the course of 12 weeks."
Relevant Quotes:
1) "In both of these conditions, teachers implemented a scripted intervention lesson 45 min per day for 3 days a week over the course of 12 weeks." (p. 307)
2) "This study was conducted in a limited number of classrooms and for a short duration." (p. 312)
Detailed Analysis:
As established under Criterion T, the entire intervention and its outcome assessment spanned only about 12 weeks, with no evidence of tracking extending toward a full academic year (approximately 75% of 9-10 months). Because Criterion T is not met, Criterion Y cannot be met per the standard's rule that Y requires T. Independently, the actual duration (12 weeks) is far short of 75% of an academic year in any case.
Criterion Y is not met because Criterion T is not met, and independently the study's 12-week duration is far short of the year-long (75% of an academic year) requirement.
-
B
Balanced Control Group
- Both conditions received identical instructional time and the same core curriculum, differing only in the medium used for the final lesson of each cycle, so no resource imbalance exists.
- "In both of these conditions, teachers implemented a scripted intervention lesson 45 min per day for 3 days a week over the course of 12 weeks. The length and duration of the intervention were held constant across conditions."
Relevant Quotes:
1) "In both of these conditions, teachers implemented a scripted intervention lesson 45 min per day for 3 days a week over the course of 12 weeks. The length and duration of the intervention were held constant across conditions." (p. 307)
2) "In the nonmultimedia condition, teachers read each book on 3 days. They also implemented scripted curricula that accompanied the read-aloud books for their condition. In the multimedia condition, teachers read each book on 2 days. Then, for 3 days at the end of the cycle, teachers showed children different clips from a video that related to the habitat for that cycle." (p. 308)
3) "Teachers in both conditions taught the same scripted curriculum for the first two lessons for each book." (p. 308)
Detailed Analysis:
Applying the Criterion B decision procedure: extra time or budget was not added to the multimedia condition relative to the control — both conditions received exactly the same total amount of instructional time (45 min/day, 3 days/week, 12 weeks) and the same scripted curriculum for the first two lessons of every book. The only difference is the content/format of the third lesson each cycle: the control group had a third day of book re-reading and review, while the multimedia group had three days of video-based review lessons instead. No additional minutes, sessions, staff, or materials budget were given to either group; the video days simply substitute for book-review days within the same fixed schedule.
Because there is no extra time or resource given to the treatment group beyond what the control group received (EXTRA_RESOURCES_PRESENT is false — the manipulation is a substitution of content/medium within an identical time/resource envelope, not an addition), the control condition is trivially balanced with the treatment condition in terms of time and resources.
Criterion B is met because both conditions received identical amounts of instructional time and the same scripted curriculum, differing only in the medium (book review vs. video) used for the final lesson of each cycle, with no additional resources given to either group.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study by a different research team was found.
Relevant Quotes:
1) "In this study, we investigated the effects of a research-based, read-aloud vocabulary intervention that was enhanced with multimedia support for vocabulary learning." (p. 311)
2) "Further research on the effect of multimedia on vocabulary instruction should proceed in a few directions. First, research on the effect of multimedia should be implemented on a larger scale so that researchers can consider teacher and grade-level effects." (p. 312, Future Directions)
Detailed Analysis:
The paper itself calls for further, larger-scale research on this specific intervention design, indicating that at the time of publication no independent replication existed. An internet search for subsequent independent replications of this specific study — a comparison of multimedia-enhanced versus non-multimedia read-aloud vocabulary instruction using this same habitat-themed curriculum and book/video set with pre-K–second grade ELL and non-ELL children — did not identify any peer-reviewed study by a different research team that reproduced this specific experimental comparison and design. The search located many papers that cite Silverman and Hines (2009) as background literature on multimedia and vocabulary instruction (e.g., subsequent studies on media type and vocabulary learning for linguistically diverse students, and reviews of multimedia vocabulary instruction for ELLs), but none of these constitute an independent replication of this specific study's design, curriculum, and comparison; they instead test different interventions, populations, or measures.
Criterion R is not met because no independent replication of this specific study, by a different research team, was found in the peer-reviewed literature.
-
A
All-subject Exams
- Outcomes were limited to vocabulary and science-content knowledge, mostly via researcher-made measures, not standardised exams spanning all main subjects.
- "Both prior to and following the intervention, children were assessed on knowledge of words targeted in the intervention, general vocabulary knowledge, and knowledge of science concepts taught in the intervention."
Relevant Quotes:
1) "Both prior to and following the intervention, children were assessed on knowledge of words targeted in the intervention, general vocabulary knowledge, and knowledge of science concepts taught in the intervention." (p. 309)
2) "Knowledge of target words... General vocabulary knowledge... Science concepts knowledge." (p. 309, Assessments subsection headers)
Detailed Analysis:
The study measured target vocabulary knowledge (TVA), general vocabulary knowledge (PPVT), and science-content knowledge related to the habitats taught (SCI). It did not assess other core academic subjects such as mathematics or broader reading comprehension/language arts achievement beyond vocabulary. Only the PPVT among these measures is a widely recognised standardised exam (satisfying Criterion E); the TVA and SCI are researcher-made instruments and do not themselves count as standardised exam-based assessments of "all subjects." No exception rationale (e.g., a specialised upper-secondary/vocational focus) is offered or applicable here, since this is a general vocabulary intervention for young children, not a specialised course.
Criterion A is not met because the study assessed only vocabulary and science-content knowledge using mostly researcher-made instruments, not standardised exams across all main subjects.
-
G
Graduation Tracking
- Criterion Y is not met, outcomes were only measured immediately before and after the short intervention, and no follow-up publication tracking this cohort toward graduation was found.
- "This study was conducted in a limited number of classrooms and for a short duration."
Relevant Quotes:
1) "This study was conducted in a limited number of classrooms and for a short duration." (p. 312, Limitations)
2) "Both prior to and following the intervention, children were assessed on knowledge of words targeted in the intervention, general vocabulary knowledge, and knowledge of science concepts taught in the intervention." (p. 309)
Detailed Analysis:
As established under Criterion Y, the study's outcome tracking spanned only about 12 weeks and did not reach even the term-length benchmark, let alone 75% of an academic year; per the standard's rule that Criterion G requires Criterion Y, this alone means Criterion G cannot be met. The study only assessed children immediately before and after the 12-week intervention. There is no description of any longer-term follow-up, let alone tracking of participants through graduation from their educational stage. The authors explicitly note the short duration as a limitation and propose future, larger-scale research rather than reporting any completed long-term follow-up. An internet search for later publications by Silverman and/or Hines that might report longer-term or graduation follow-up on this same pre-K through second-grade cohort did not identify any such follow-up publication tracking this specific sample.
Criterion G is not met because Criterion Y is not met, outcomes were only measured immediately before and after the 12-week intervention, and no follow-up publication tracking this cohort toward graduation was identified via internet search.
-
P
Pre-Registered
- No mention of a pre-registered protocol or registry reference was found anywhere in the paper, and none was located via internet search.
Relevant Quotes:
No quote referencing a pre-registration platform, registry ID, or pre-registration date was found anywhere in the paper, including the Method, Results, and Discussion sections.
Detailed Analysis:
The paper (received February 2008, published 2009) contains no statement that the study protocol, hypotheses, or analysis plan were registered on any public registry (e.g., a clinical-trials-style registry or the Open Science Framework) before data collection began. There is no registry link, ID, or registration date given anywhere in the text. An internet search for a pre-registration record of this study (e.g., on OSF, ClinicalTrials.gov, or the AEA/AsPredicted-style registries) did not locate any such record, which is unsurprising given that formal pre-registration of education RCT protocols was not a widespread practice at the time this study was conducted (2008 data collection, 2009 publication).
Criterion P is not met because no evidence of pre-registration of the study protocol was found in the paper or via internet search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.