Abstract
This study examined the added value of a vocabulary plus phonological awareness (vocab+) intervention against a phonological awareness (PA only) intervention only. The vocabulary intervention built networks among words through attention to morphological and semantic relationships. This supplementary classroom instruction augmented existing literacy curriculum in a 71 primarily Spanish-speaking English Learners (EL) in grade one, many "at-risk" of reading difficulty. Vocab+ lessons drew from expository text, and words were revisited throughout the program. The PA only group received a previously validated phonological awareness and decoding instruction with no vocabulary instruction. The treatment group (vocab+) spent 30% of the intervention on PA and decoding. Students demonstrated expected gains in vocabulary, while maintaining gains in phonological decoding equivalent to those of the PA only group. This study demonstrates initial justification for dedicating limited instructional minutes to vocabulary building in early literacy interventions while still dedicating a small portion of time to phonological awareness and decoding.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Intact classrooms, not individual students within a classroom, were the unit of randomisation.
- "First, intact classrooms were randomly assigned to treatment (three Vocab+ classrooms) or treatment-control (two PAD classrooms) conditions." (p. 17)
Relevant Quotes:
1) "First, intact classrooms were randomly assigned to treatment (three Vocab+ classrooms) or treatment-control (two PAD classrooms) conditions." (p. 17)
2) "Second, students were rank-ordered within classroom on Nonsense Word Fluency pretest scores, a measure of phonological decoding (see Measures)." (p. 17)
3) "Therefore, there were three classrooms, each with two Vocab+ -SR groups and two Vocab+-MA groups and two PAD classrooms." (p. 18)
Detailed Analysis:
The paper explicitly states that whole, intact classrooms were randomly allocated to either the Vocab+ treatment condition or the PAD treatment-control condition. Only after this class-level assignment were students within the Vocab+ classrooms rank-ordered and grouped, and those small groups (not individual students across conditions) were randomly assigned to the SR or MA variant. Crucially, the PAD-vs-Vocab+ contrast, which is the primary comparison of interest for this criterion, was randomised at the classroom level, preventing within-classroom contamination between the two main arms.
Criterion C is met because the primary treatment-versus-control contrast was randomised at the level of intact classrooms.
-
E
Exam-based Assessment
- The primary vocabulary outcome measure was a custom, study-specific test built directly from the taught words, not a standardized exam.
- "To measure knowledge of taught vocabulary words, the target vocabulary test (VOC) was administered in a group format at pretest and at posttest. Words tested were all taught to both MA and SR conditions." (p. 18)
Relevant Quotes:
1) "Target vocabulary knowledge (experimental). To measure knowledge of taught vocabulary words, the target vocabulary test (VOC) was administered in a group format at pretest and at posttest. Words tested were all taught to both MA and SR conditions." (p. 18)
2) "Both receptive and expressive measures [PPVT, EOWVT] were obtained only as baseline information to validate the target vocabulary knowledge test, as neither is sensitive to change over a short time period." (p. 18)
3) "Phonological decoding. The Dynamic Indicators of Basic Early Literacy Skills (DIBELS ...) Nonsense Word Fluency (NWF; first grade) subtest measures the alphabetic principle..." (p. 18)
Detailed Analysis:
The paper explicitly labels its primary vocabulary outcome measure as "experimental," and describes it as a test built directly from the specific words taught in the intervention, with items authored by the researchers (silly/true sentences per target word). This is precisely the pattern the ERCT standard's E criterion problem statement warns against: a custom instrument tightly aligned to the intervention content, which can inflate apparent effectiveness. The standardized PPVT and EOWVT were used only at baseline/pretest to validate the custom VOC measure, not as outcome measures of change. The Nonsense Word Fluency (NWF) subtest of DIBELS is a standardized, widely-used measure and was used as a genuine pre-post outcome, which would partially satisfy E on its own. However, since the central research question and primary contrast of interest (added value of vocabulary instruction) was assessed with the custom VOC test rather than a standardized exam, the overall exam-based assessment requirement is not satisfied.
Criterion E is not met because the paper's central vocabulary outcome relies on a custom, researcher-made test tied directly to intervention content, rather than a standardized exam.
-
T
Term Duration
- The intervention and its outcome measurement spanned only about eight weeks, far short of a full academic term.
- "... four days a week for eight weeks, for a total of 435 possible instructional minutes." (p. 19)
Relevant Quotes:
1) "All groups received an average of 15 minutes of supplementary instruction per day, four days a week for eight weeks, for a total of 435 possible instructional minutes." (p. 18-19)
2) "On average, students received 26.3 days of instruction (SD = 2.26) across an 8-week period, and therefore an average of 394.5 minutes." (p. 19)
3) "Pre and post-test student performance was measured on target vocabulary knowledge, PA, and phonological decoding." (p. 18)
Detailed Analysis:
The intervention ran for eight weeks in total, and outcome measures (VOC, NWF) were collected as a pretest/posttest pair bracketing this eight-week period, with no indication of a longer follow-up window. Eight weeks is roughly half of a typical 3-4 month academic term, and there is no mention of measurement occurring at any later date beyond the immediate posttest.
Criterion T is not met because the interval between intervention start and outcome measurement was only about eight weeks, well short of one academic term.
-
D
Documented Control Group
- The control (PAD) group's size, demographics, baseline scores, and instructional content are described in detail.
- "The PAD condition was based on a previously validated intervention from Project La Patera, a 100% PA and decoding intervention..." (p. 19)
Relevant Quotes:
1) "The PAD condition was based on a previously validated intervention from Project La Patera, a 100% PA and decoding intervention that explicitly teaches identification, production, and manipulation of sounds in word- and sentence-level decoding (Gerber et al., 2004; Leafstedt, Richards, & Gerber, 2004)." (p. 19)
2) "Table 1. Participant Information ... PAD: n = 18, % male = 50, ELL status = 94, NWF Pretest M = 35.06, SD = 17.91." (p. 17)
3) "After group assignment, one treatment-control classroom (PAD only) dropped out during the pre-test phase citing logistics and lack of teacher interest. Class composition and pre-intervention achievement characteristics of students in this classroom did not significantly differ, practically or statistically, from those of remaining students." (p. 18)
Detailed Analysis:
The paper provides a detailed description of the control (PAD) condition: its size (18 students, one classroom), gender and ELL-status composition, baseline NWF performance (Table 1), and the specific previously validated curriculum (Project La Patera) it received. The paper also transparently reports and justifies the attrition of a second PAD classroom. This level of detail allows the control group's comparability and treatment to be properly assessed.
Criterion D is met because the control group's composition, baseline characteristics, and instructional content are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Only a single school participated, so randomisation could not occur at the school level.
- "The participating school was a Title I ... southern California elementary school..." (p. 17)
Relevant Quotes:
1) "The participating school was a Title I (i.e., high poverty, low achieving) southern California elementary school, and had been designated as Program Improvement Year 1 under the Federal No Child Left Behind Act of 2001." (p. 17)
2) "The sampling frame comprised 97 first-grade students from five intact classrooms." (p. 17)
Detailed Analysis:
The study was conducted within a single elementary school, with randomisation occurring among the five intact classrooms inside that one school. There is no school-level randomisation across multiple schools; the "school" is a fixed, non-randomised unit here, and the study explicitly frames randomisation as occurring among classrooms, not schools.
Criterion S is not met because randomisation occurred at the classroom level within a single school, not across multiple randomised schools.
-
I
Independent Conduct
- The same research/development team designed, delivered, and assessed the intervention, with no independent evaluator.
- "The first author and trained research assistants administered assessments at school sites." (p. 18)
Relevant Quotes:
1) "The first author and trained research assistants administered assessments at school sites." (p. 18)
2) "Students were instructed in small groups by trained undergraduate and graduate researchers." (p. 18)
3) "In all three conditions instructors were trained in the Core Intervention Model (Gerber et al 2004; Richards & Leafstedt, 2010)." (p. 19)
Detailed Analysis:
The intervention (a refinement of the authors' own Core Intervention Model and prior Project La Patera work) was designed by the paper's authors, delivered by researchers trained and supervised by the study team, and the outcome assessments were administered by the first author and research assistants under her direction. There is no mention of an external, third-party organisation conducting data collection, analysis, or oversight independent of the intervention designers.
Criterion I is not met because the same research team designed the intervention, trained and supervised the instructors, and administered the outcome assessments themselves.
-
Y
Year Duration
- Since the term-duration criterion (T) is not met, the stronger year-duration criterion cannot be met either.
- "... four days a week for eight weeks, for a total of 435 possible instructional minutes." (p. 19)
Relevant Quotes:
1) "All groups received an average of 15 minutes of supplementary instruction per day, four days a week for eight weeks..." (p. 18-19)
2) "On average, students received 26.3 days of instruction (SD = 2.26) across an 8-week period, and therefore an average of 394.5 minutes." (p. 19)
Detailed Analysis:
The entire study, from intervention start to final outcome measurement, spanned only eight weeks. This is far below even the weaker one-term requirement, let alone the 75%-of-an-academic-year threshold (about 9-10 months) required for criterion Y. Per the ERCT rule that Y cannot be met if T is not met, and given the actual duration is a small fraction of a school year, this criterion clearly fails.
Criterion Y is not met because the study duration (eight weeks) is a small fraction of an academic year and the weaker term-duration criterion (T) is also not met.
-
B
Balanced Control Group
- All conditions received the identical total amount of instructional time; only the content emphasis within that shared time differed.
- "However, the proportion of PAD varied between the PAD condition (100% PAD) and the Vocab+ conditions (70% vocabulary instruction, 30% PAD)." (p. 19)
Relevant Quotes:
1) "All groups received an average of 15 minutes of supplementary instruction per day, four days a week for eight weeks, for a total of 435 possible instructional minutes." (p. 18-19)
2) "All instructional conditions shared the following features: frequency, duration and number of lessons; group size; expository passage read-aloud and instructional model (CIM)." (p. 19)
3) "However, the proportion of PAD varied between the PAD condition (100% PAD) and the Vocab+ conditions (70% vocabulary instruction, 30% PAD)." (p. 19)
Detailed Analysis:
Applying the Criterion B decision procedure: the treatment (Vocab+) group did not receive any extra time, budget, or materials relative to the treatment-control (PAD) group. All groups received the same frequency (4 days/week), duration (8 weeks), session length (~15 min/day), group size, shared read-aloud passages, and instructional model (CIM). The only difference is how those identical minutes were allocated across content (100% PAD vs. 70% vocabulary/30% PAD). Since no additional resources were present in the intervention arm (EXTRA_RESOURCES_PRESENT = false), the decision procedure resolves at step 1 and the balance requirement is trivially satisfied. This conclusion is unaffected by the updated Criterion B definition, since the trigger condition (extra time or budget in the intervention arm) is simply absent here.
Criterion B is met because all conditions received identical total instructional time and structure, differing only in the content allocated within that shared time.
-
Level 3 Criteria
-
R
Reproduced
- No evidence of independent replication of this specific study was found in the paper or via external search.
Relevant Quotes:
No quotes in the paper reference a replication of this specific study by an independent team.
Detailed Analysis:
The paper does not mention any prior or subsequent independent replication of this exact vocab+ vs. PAD comparison. An OpenAlex citation search identified 18 works citing this paper (Filippini, Gerber, & Leafstedt, 2012) as of 2026; all 18 were reviewed by title and authorship. They are meta-analyses, systematic reviews, or unrelated vocabulary and morphology intervention studies by different author teams (e.g., Silverman, Johnson, Keane, & Khanna, 2020; Rogde, Hagen, Melby-Lervag, & Lervag, 2019; Colenbrander et al., 2024; Bengochea & Sembiante, 2023, 2024; Li, Snow, Ely, Frijters, Geva, & Chen, 2023). None of these works reproduce this specific vocab+ vs. PAD experimental design with a new sample in a new context; each cites this study only as one data point within a broader synthesis. No independent replication of this specific study was located in any available source.
Criterion R is not met because no independent replication of this specific study was found in the paper or in external literature.
-
A
All-subject Exams
- Only reading-related constructs (vocabulary, phonological decoding) were assessed; no other core subjects were measured.
- "Target vocabulary knowledge, phonological decoding, word reading, reading fluency, and comprehension were measured..." (p. 18)
Relevant Quotes:
1) "Target vocabulary knowledge, phonological decoding, word reading, reading fluency, and comprehension were measured using a combination of standardized and experimental measures." (p. 18)
2) "Breadth of vocabulary knowledge. The Peabody Picture Vocabulary Test-Revised (PPVT...) was administered individually as an indicator of receptive vocabulary knowledge in English." (p. 18)
Detailed Analysis:
All outcome measures in this study (VOC, NWF, PPVT, EOWVT) fall within the single domain of language and early literacy. No other core subjects such as mathematics or science were assessed, and no rationale is given for restricting outcomes to literacy only (this is a general first-grade literacy intervention, not a specialised vocational or upper-secondary program that would justify a narrower exception). Since only reading/language outcomes were measured, this fails the all-subject requirement regardless of the mixed verdict on criterion E.
Criterion A is not met because only reading- and language-related outcomes were assessed, with no coverage of other core subjects.
-
G
Graduation Tracking
- Criterion Y (Year Duration) is not met, which automatically fails this criterion; additionally, no follow-up or graduation-tracking publication by the same authors was located.
- "Findings in this study should be interpreted with caution, and used primarily as guidance toward future research due to lack of statistically reliable findings." (p. 23)
Relevant Quotes:
1) "Pre and post-test student performance was measured on target vocabulary knowledge, PA, and phonological decoding." (p. 18)
2) "Given the large effect sizes but low reliability of findings in this study, future research with larger sample sizes is clearly needed..." (p. 23)
Detailed Analysis:
The paper reports only a single pretest/posttest pair bracketing the eight-week intervention. There is no mention of any longer-term follow-up of these first-grade students, let alone tracking through to the end of their primary education or graduation. Per the ERCT rule that Criterion G cannot be met if Criterion Y (Year Duration) is not met, and Y is not met here, Criterion G fails on that basis alone. An OpenAlex author search for Alexis L. Filippini and Jill M. Leafstedt was also conducted to check for later follow-up publications. One later co-authored paper was found, "Longitudinal Prediction of 1st and 2nd Grade English Oral Reading Fluency in English Language Learners" (Solari, Aceves, Higareda, Richards-Tutor, Filippini, & Gerber, 2013), but it tracks a related sample only through 2nd grade, far short of graduation, and does not report outcomes for this study's specific vocab+/PAD cohort. No companion publication tracking this cohort through graduation was found.
Criterion G is not met because outcomes were measured only immediately after the eight-week intervention, the weaker Year Duration criterion (Y) is also not met, and no graduation-tracking follow-up publication was found.
-
P
Pre-Registered
- The paper contains no statement or reference to a pre-registered study protocol, and no registry record was found via external search.
Relevant Quotes:
No quotes in the paper reference pre-registration, a trial registry, or a published protocol.
Detailed Analysis:
The methods, results, and discussion sections make no mention of the study's hypotheses, design, or analysis plan having been registered on any public registry (e.g., ClinicalTrials.gov, AEA RCT Registry, OSF) prior to data collection. Given the paper explicitly describes itself as "exploratory," it is consistent with the study not having a pre-registered protocol. No pre-registration record for this study was located via external database search (OSF Registries, ClinicalTrials.gov-style registries), which is unsurprising given the study's small-scale, exploratory nature and the era (data collection era predates widespread pre-registration norms in education research).
Criterion P is not met because no pre-registration statement, registry link, or date is present anywhere in the paper or in external sources.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.