Abstract
The introduction of second language (L2) education in kindergartens is ubiquitous in many places globally; nevertheless, research in these settings is scarce compared with that on older learners. L2 vocabulary development is especially germane to these very young learners, rendering this a research-worthy topic. The present study examined the effects of researcher-designed explicit vocabulary instruction compared with implicit instruction on English-as-a-second- language participants' (N=157) gains in not only the target vocabulary items, but also general vocabulary as well as phonological awareness. Statistically significant differences were found in all vocabulary tasks and the phonemic awareness task with small to large effect sizes. These showed that, in addition to the target vocabulary, the participants receiving explicit vocabulary instruction also had greater gains in receptive and expressive general vocabulary and phonemic awareness. The article culminates in delineating the children's differential achievements, followed by a brief discussion of the limitations and implications.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was conducted at the kindergarten (school) level rather than within a single class, which is a stronger design than the class-level requirement and therefore automatically satisfies it.
- "The current study adopted a cluster randomized trial, in which we randomly assigned schools, but not individual participants, into the experimental and the control groups." (p. 675)
Relevant Quotes:
1) "The current study adopted a cluster randomized trial, in which we randomly assigned schools, but not individual participants, into the experimental and the control groups." (p. 675)
2) "Two of the kindergartens were assigned randomly to the control group (n=61) and three of them to the experimental group (n=96)." (p. 675)
Detailed Analysis:
The paper explicitly describes a cluster randomised trial in which whole kindergartens (schools), not individual students or single classes within a school, were assigned to condition. Five kindergartens in total were randomised, two to control and three to the experimental arm. The ERCT standard states that a stronger school-level RCT automatically satisfies the weaker class-level (C) criterion.
Final summary: Criterion C is met because randomisation occurred at the whole-kindergarten (school) level, which is stronger than and therefore satisfies the class-level requirement.
-
E
Exam-based Assessment
- The primary target-vocabulary outcome measures were researcher-developed tasks created specifically for this study, rather than standard, widely recognised exams.
- "Two measures developed by the researchers were used to assess the learning of 16 target vocabulary items specified in the EVI, namely receptive target word definition and expressive target word definition tasks." (p. 678)
Relevant Quotes:
1) "Both standardized and researcher-developed tasks were employed in the study. The latter was recommended by the National Reading Panel (2000) to assess children's target vocabulary development, as such measures would be more sensitive to content-specific word gains from specific vocabulary intervention." (p. 677)
2) "Two measures developed by the researchers were used to assess the learning of 16 target vocabulary items specified in the EVI, namely receptive target word definition and expressive target word definition tasks." (p. 678)
3) "English receptive vocabulary was measured using the Peabody Picture Vocabulary Test III (PPVT-III) developed by Dunn and Dunn (1997)." (p. 678)
4) "The children's English expressive vocabulary was measured with a picture-naming task (Learning Disabilities Association of Alberta, 2009)." (p. 678)
5) "Two measures of phonological awareness were included: English syllable deletion and English initial phoneme identification (Learning Disabilities Association of Alberta, 2009)." (p. 679)
Detailed Analysis:
The ERCT Standard requires that outcome assessment rely on standard, widely recognised tests rather than instruments designed specifically for the study, in order to avoid bias from a test overly aligned with the intervention. The authors themselves acknowledge using "both standardized and researcher-developed tasks," and explicitly state that the two target-vocabulary tasks (receptive and expressive target word definition) were measures "developed by the researchers" solely to assess the 16 target words taught in the EVI programme. These are, by the authors' own description, custom instruments tailored to the intervention content.
The remaining measures are somewhat more standardised: the PPVT-III is a widely used, commercially published standardized vocabulary test, while the picture-naming and phonological awareness tasks are drawn from a published "Reading Readiness Screening Tool" (Learning Disabilities Association of Alberta, 2009), which is a regional screening instrument rather than a nationally or internationally recognised standardised exam comparable to a curriculum-based achievement test.
Because the primary target-vocabulary outcomes (the measures most tied to the intervention's specific content) were custom-built by the researchers for this study, and the remaining instruments are, at best, a screening tool and a vocabulary test rather than standardised academic exams, the overall assessment battery does not meet the bar of standard, widely recognised exam-based assessment required by the ERCT standard.
Final summary: Criterion E is not met because the study's target-vocabulary outcome measures were custom, researcher-developed tasks built specifically for this intervention, rather than standard, widely recognised exams.
-
T
Term Duration
- The intervention and its posttest measurement together spanned only about 8 weeks, well short of a full academic term.
- "it was an 8-week intervention program with two sessions per week, each lasting about 30 min." (p. 676)
Relevant Quotes:
1) "With regard to the structure, it was an 8-week intervention program with two sessions per week, each lasting about 30 min." (p. 676)
2) "The data collection was undertaken by trained experimenters, at pretest and posttest, in a quiet room in the attending school. The assessments lasted around 30 min." (p. 679)
Detailed Analysis:
The ERCT Standard requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Here, the entire EVI/IVI programme ran for only 8 weeks (two 30-minute sessions per week), and the posttest was administered directly after the programme concluded, with no indication of any delayed follow-up assessment. Eight weeks is approximately two months, clearly shorter than the minimum one-term threshold, and no quote in the paper suggests a longer interval between intervention start and outcome measurement.
Final summary: Criterion T is not met because the intervention and its posttest measurement spanned only about 8 weeks, short of the required one academic term.
-
D
Documented Control Group
- The control group (kindergartens receiving Implicit Vocabulary Instruction) is clearly documented in terms of size, demographics, materials used, and who delivered the instruction.
- "Control group instruction was provided by a trained research assistant. She conducted shared reading activities using the same four picture books as used in the EVI." (p. 677)
Relevant Quotes:
1) "Two of the kindergartens were assigned randomly to the control group (n=61) and three of them to the experimental group (n=96)." (p. 675)
2) "The IVI followed the same structure as the EVI; four storybooks were used and each one was presented for four sessions in 2 weeks. No target word was specified for each session, but the participants were exposed to the same words through shared reading in each session." (p. 677)
3) "Control group instruction was provided by a trained research assistant. She conducted shared reading activities using the same four picture books as used in the EVI." (p. 677)
4) "Preliminary analyses using t-tests showed that there were no significant differences between the means of the two groups at pretest, ps>.22, indicating that both groups were similar on all measures prior to receiving the intervention." (p. 680)
Detailed Analysis:
The paper documents the control group's size (n=61 across two kindergartens), the content and structure of the instruction it received (implicit, incidental exposure to the same storybooks and target words, without explicit teaching), who delivered it (a trained research assistant), and confirms baseline equivalence with the experimental group via non-significant pretest differences. This level of detail satisfies the requirement for a well-documented control group.
Final summary: Criterion D is met because the control group's size, composition, instructional content, and baseline comparability with the experimental group are all clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation was conducted at the level of whole kindergartens, satisfying the school-level RCT requirement.
- "The current study adopted a cluster randomized trial, in which we randomly assigned schools, but not individual participants, into the experimental and the control groups." (p. 675)
Relevant Quotes:
1) "The current study adopted a cluster randomized trial, in which we randomly assigned schools, but not individual participants, into the experimental and the control groups." (p. 675)
2) "A total of 157 children ... were recruited from five Hong Kong kindergartens for the present study. Two of the kindergartens were assigned randomly to the control group (n=61) and three of them to the experimental group (n=96)." (p. 675)
Detailed Analysis:
The ERCT S criterion requires randomisation at the school (educational institution) level rather than merely at the class level. Here, "school" corresponds to the kindergarten as a whole institution. All five participating kindergartens were randomised as complete units into either the control or experimental condition, with no mixing of conditions within a single kindergarten. This directly matches the school-level randomisation requirement.
Final summary: Criterion S is met because entire kindergartens (schools), rather than classes or students within a school, were the unit of randomisation.
-
I
Independent Conduct
- The researchers who designed the intervention also designed the control condition, trained all instructors, and oversaw fidelity checks, with no independent or third-party evaluator described.
- "All teachers received 6 hours of training conducted by the researchers." (p. 677)
Relevant Quotes:
1) "The experimental group received the Explicit Vocabulary Instruction (EVI) developed by the researchers and the control group received Implicit Vocabulary Instruction (IVI)." (p. 675)
2) "All teachers received 6 hours of training conducted by the researchers. The training covered the rationale of the intervention, knowledge and skills related to implementing the English learning activities." (p. 677)
3) "On-site school support was provided to the teachers and the research assistant in both the experimental and control groups and lesson observations were conducted once a week by other research assistants to ensure program integrity." (p. 677)
4) "Control group instruction was provided by a trained research assistant." (p. 677)
Detailed Analysis:
Criterion I requires that the study be conducted independently of the team that designed the intervention, typically via a third-party evaluator with no involvement in the intervention's design. In this study, the same research team designed both the EVI and IVI programmes, trained the participating teachers and the research assistant who delivered the control condition, provided ongoing on-site support, and used their own research assistants (rather than an external body) to monitor fidelity of implementation. While some fidelity ratings were done by an assistant "blind to the research hypotheses," this blinding is an internal procedural safeguard, not independent third-party conduct of the trial itself. There is no mention of an external evaluation agency, government body, or other independent organisation running the study or analysing the data.
Final summary: Criterion I is not met because the intervention's designers (the researchers) also trained instructors, supervised delivery, and monitored fidelity themselves, with no independent third-party conducting the evaluation.
-
Y
Year Duration
- Because criterion T (Term Duration) is not met, and the actual intervention/measurement period of about 8 weeks is far short of a full academic year, criterion Y is not met.
- "it was an 8-week intervention program with two sessions per week, each lasting about 30 min." (p. 676)
Relevant Quotes:
1) "it was an 8-week intervention program with two sessions per week, each lasting about 30 min." (p. 676)
Detailed Analysis:
The Year Duration criterion requires that outcomes be measured at least 75% of one academic year (roughly 9-10 months) after the intervention begins. The entire study, from intervention start to posttest, spanned only about 8 weeks, with no indication of any longer-term follow-up. Per the ERCT specification, if the weaker Term Duration criterion (T) is not met, the stronger Year Duration criterion (Y) cannot be met either. Both the explicit cascade rule and the actual reported duration point to the same conclusion.
Final summary: Criterion Y is not met because the study spanned only about 8 weeks, far short of the required academic-year duration, and criterion T is also not met.
-
B
Balanced Control Group
- Both the experimental and control conditions received identical instructional time and identical storybook materials, differing only in the explicitness of the vocabulary teaching, so no extra resources were present to begin with.
- "The amount of instructional time and the storybooks used were the same in the two conditions." (p. 675)
Relevant Quotes:
1) "The amount of instructional time and the storybooks used were the same in the two conditions." (p. 675)
2) "The IVI followed the same structure as the EVI; four storybooks were used and each one was presented for four sessions in 2 weeks." (p. 677)
3) "Both the EVI and the IVI were supplemental programs for the participants (i.e., the participating children attended their business-as-usual English lessons during the study period)." (p. 675)
Detailed Analysis:
Applying the ERCT decision procedure for criterion B: the intervention (EVI) did not add extra instructional time, budget, or materials relative to the control (IVI) -- both groups received the identical number of sessions (an 8-week programme, four storybooks, sessions of the same length), and both continued their normal, business-as-usual English lessons alongside the supplemental programme. Since EXTRA_RESOURCES_PRESENT is false (no additional time, budget, or materials were given to either arm relative to the other), the decision tree resolves at its first branch and the criterion is met without needing to assess whether any resource was integral to the treatment. The sole difference between arms was the explicitness of the vocabulary teaching embedded within otherwise identical instructional time and materials, which is precisely the manipulation the study intends to isolate.
Final summary: Criterion B is met because instructional time and materials (storybooks, session count and length) were identical between the experimental and control groups, isolating explicitness of instruction as the sole manipulated variable.
-
Level 3 Criteria
-
R
Reproduced
- No evidence of an independent replication of this specific study by a different research team was found in the paper or through internet searches.
Relevant Quotes:
1) "This corroborates the findings from previous studies (e.g., Collins, 2010; Sun, 2017; Yeung et al., 2016) and the meta-analysis by Marulis and Neuman (2010) that the explicit instruction model is more effective than the incidental model in teaching target vocabulary." (p. 683)
Detailed Analysis:
Criterion R requires that this specific study be independently replicated by a different research team in a different context, published in a peer-reviewed journal. The discussion section cites broadly related prior work on explicit versus implicit vocabulary instruction (including an earlier paper by the same first author, Yeung et al., 2016) as converging evidence, but none of these are replications of this particular cluster-randomised EVI/IVI design with Hong Kong kindergarteners; they are earlier, related studies that motivated the present design, not independent reproductions of it.
To check for later independent replication, internet searches were conducted for citing works, other EVI/IVI or explicit-versus-implicit L2 vocabulary cluster-RCTs in kindergarten settings, and any papers referencing this study's DOI (10.1007/s11145-019-09982-3) as a replication target. These searches surfaced only conceptually related but methodologically distinct vocabulary-instruction studies (e.g., work on embodied instruction, other L2 vocabulary intervention designs, and unrelated meta- analyses), none of which report an independent replication of this specific EVI-versus-IVI cluster-randomised trial by a different research team. No such replication paper could be identified.
Final summary: Criterion R is not met because no independent replication of this specific study was found, only citations to earlier, related but distinct studies, confirmed by additional internet searches that found no later replication.
-
A
All-subject Exams
- Because criterion E (Exam-based Assessment) is not met, and only vocabulary and phonological awareness were assessed rather than all main subjects, criterion A is not met.
Relevant Quotes:
1) "Two researcher-developed program-specific tasks were used to measure the gains in target vocabulary. Two general vocabulary measures were also included ... Two measures of phonological awareness were used to examine the program effects on phonological awareness." (pp. 677-678)
Detailed Analysis:
Per the ERCT specification, criterion A (All-subject Exams) requires criterion E to be met as a prerequisite; if E is not met, A cannot be met either. In addition, the outcome measures in this study cover only L2 vocabulary and phonological awareness, not the full range of core subjects taught at the kindergarten level, and no rationale for a specialised, single-subject focus (as permitted for upper secondary/vocational contexts) is provided or applicable here.
Final summary: Criterion A is not met both because criterion E is not met and because only vocabulary and phonological awareness, not all main subjects, were assessed.
-
G
Graduation Tracking
- Because criterion Y (Year Duration) is not met, and no subsequent paper tracking this specific cohort to graduation could be found, criterion G is not met.
Relevant Quotes:
1) "Third, further study is needed to examine the impact of long-term explicit instruction and its sustained effect on vocabulary development in children learning L2s." (p. 685)
Detailed Analysis:
Per the ERCT specification, if criterion Y (Year Duration) is not met, criterion G (Graduation Tracking) is also not met. Consistent with this, the paper reports only a single pretest/posttest design over an 8-week programme, with the authors themselves noting in the limitations that "further study is needed to examine the impact of long-term explicit instruction," explicitly acknowledging the absence of any longer-term or graduation-tracking follow-up.
Internet searches were conducted for later papers by the same author team (Yeung, Ng, Qiao, Tsang) that might track the same cohort of 157 kindergarteners from these five Hong Kong kindergartens through to graduation. These searches identified other work by the first author on Hong Kong ESL vocabulary and reading development (e.g., an earlier, and thus not a follow-up, longitudinal study, Liu, Yeung, Lin, & Wong, 2017, "English expressive vocabulary growth and its unique role in predicting English word reading," which studies a different cohort of 141 children observed from K2 to K3 and predates the present study), but no publication reporting continued tracking of this study's specific EVI/IVI cohort through to kindergarten graduation or beyond was identified.
Final summary: Criterion G is not met because there was no long-term or graduation follow-up reported in the paper, criterion Y is also not met, and internet searches found no follow-up publication tracking this specific cohort to graduation.
-
P
Pre-Registered
- The paper contains no statement or reference indicating that the study protocol was pre-registered before data collection began, and no registry entry could be found online.
Relevant Quotes:
(No quotes found: the paper contains no mention of a trial registry, registration number, or pre-registration date anywhere in the Method, Results, or Acknowledgements sections.)
Detailed Analysis:
Criterion P requires a documented pre-registration of the study's hypotheses, methods, and planned analyses on a public registry before data collection began. A thorough review of the Method section, Procedures, Acknowledgements, and reference list reveals no mention of any registry platform (e.g., ClinicalTrials.gov, ISRCTN, AEA RCT Registry, OSF Registries) or a pre-registration date. The Acknowledgements section only references funding support ("Quality Education Fund, HKSAR").
Internet searches for a pre-registration record associated with this study (by title, authors, DOI, and the "Quality Education Fund" grant reference QEF2012/0316) found no matching entry in any public trial or study registry.
Final summary: Criterion P is not met because no pre-registration statement, registry link, or registration date is provided anywhere in the paper, and no registry record was found through internet searches.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.