Level 1 Criteria
-
C Class-level RCT
- Allocation was made at the kindergarten level with all children in a kindergarten sharing one condition, which satisfies the class-level randomisation requirement.
- "Randomization was conducted at the kindergarten level, with each kindergarten assigned to either experimental or control condition. All children within a given kindergarten received the same allocation." (p. 8)
- Relevant Quotes: 1) "Children were randomly allocated to two groups: experimental and control groups. Randomization was conducted at the kindergarten level, with each kindergarten assigned to either experimental or control condition. All children within a given kindergarten received the same allocation." (p. 8) 2) "The experimental group comprised 250 children (116 boys, 134 girls) in 50 Arab kindergarten classrooms (5 children each). The control group comprised 153 children (73 boys, 80 girls) in 13 kindergarten classrooms (5-16 children each) who received business-as-usual instruction, namely instruction following the standard preschool language education curriculum." (p. 8) 3) "A total of 403 Palestinian-Arabic-speaking kindergarten children in Israel (189 boys and 214 girls) ranging in age from 57 to 70 months (M = 64.4, SD = 3.1) participated in the study." (p. 7) Detailed Analysis: Criterion C requires that randomisation be performed on intact groups (classes or larger units) rather than on individual students inside a single classroom, so that contamination between treated and untreated children sharing a teacher and a room is avoided. The paper is explicit on this point: the unit of allocation was the kindergarten, and the authors state that every child within a given kindergarten received the same allocation. This is a cluster design at a level at least as strong as the class, which the ERCT standard accepts as satisfying the class-level requirement. The description also supplies the elements the standard asks to check for: the unit of randomisation (kindergarten), the total sample (403 children), and the number of clusters per arm (50 experimental and 13 control kindergarten classrooms). The paper does not describe the mechanism of the random draw (e.g. random number generator, stratification), which is a reporting weakness, but the unit of randomisation itself is stated unambiguously and is the element criterion C tests. Crucially, no reading of the Method supports allocation of individual children within a shared classroom, so the contamination risk criterion C targets is avoided regardless of how the cluster counts are interpreted. Criterion C is met because allocation was performed at the kindergarten level with all children in a kindergarten sharing the same condition, which is at or above the class level.
-
E Exam-based Assessment
- The paper states outright that every outcome measure was custom-built for the study, so no standardised exam-based assessment was used.
- "All testing measures were developed for the study and they systematically targeted SpA structures (identical in SpA and StA) and StA structures (cognate and unique)..." (p. 12)
- Relevant Quotes: 1) "All testing measures were developed for the study and they systematically targeted SpA structures (identical in SpA and StA) and StA structures (cognate and unique) in order to test the contribution of the intervention to the development of skills in SpA and StA. Lexical items for the pre and post tasks were also derived from Al-Fanus corpus (Haj and Saiegh-Haddad, in preparation)." (p. 12) 2) "SpA and StA receptive vocabulary (N = 30, Cronbach Alpha = .70) Children were asked to choose the picture (out of four pictures presented on a computer screen) that matched an oral presentation of a concrete noun." (p. 12) 3) "SpA and StA morphological awareness (N = 24, Cronbach Alpha = .84). To measure morphological awareness, a morphological analogy task was used that required children to produce a morphologically complex word (inflected or derived) in a simple word: complex word pair using the morphological analogy/relatedness illustrated in the previous pair of words." (p. 13) 4) "Letter name knowledge (N items = 29; Cronbach Alpha = .95) The 29 Arabic letters in their basic form were presented randomly on a computer screen and children were required to give the name of each letter." (p. 13) 5) "Familiarity was measured based on the subjective assessment of 10 kindergarten teachers." (p. 12) 6) "All participants were administered the Raven's Colored Progressive Matrices (RCPM; Raven, 1965) to assess nonverbal intelligence and scored within the normal range (8-13, M = 9.5, SD = 1.5)." (p. 7) Detailed Analysis: Criterion E requires that educational outcomes be measured with standardised, widely recognised exam-based assessments rather than instruments created by the researchers for the purposes of the trial. The paper answers this question directly and negatively in the first line of its measures section: "All testing measures were developed for the study." Every outcome instrument reported - receptive and expressive vocabulary, pseudoword repetition for lexico-phonological representations, syllable and phoneme awareness, morphological analogy, and letter name and letter sound knowledge - is an author-constructed task built from the authors' own Al-Fanus lexical corpus, with item familiarity calibrated by a panel of ten kindergarten teachers rather than by any national or standardised norming procedure. Only Cronbach alphas and inter-rater ICCs are reported; there is no reference to national curriculum tests, ministry examinations, or published normed batteries for the outcome measures. Two instruments in the paper derive from published paradigms - Raven's Coloured Progressive Matrices and computerised EF tasks based on the Simon task, the Automated Working Memory Assessment, and the Dimensional Change Card Sort. However, Raven's was used only as a screening/covariate measure of nonverbal intelligence and not as an educational outcome, and the EF tasks measure cognitive rather than curricular attainment; neither is a standardised exam-based assessment of educational achievement in the sense required by criterion E. The instruments were also tightly aligned with the trained content (items drawn from the same corpus used to build the intervention activities), which is precisely the alignment bias the criterion is designed to detect. Criterion E is not met because the paper explicitly states that all outcome measures were developed by the researchers for this study, with no standardised exam used to assess educational attainment.
-
T Term Duration
- The intervention ran for 15 weeks (about 3.5 months) with outcomes measured at its end, meeting the one-term interval requirement.
- "ADMIN consisted of 45 intervention sessions (3 sessions each week) and was implemented by the kindergarten teachers over a period of 15 weeks..." (p. 9)
- Relevant Quotes: 1) "ADMIN consisted of 45 intervention sessions (3 sessions each week) and was implemented by the kindergarten teachers over a period of 15 weeks, in accordance with earlier research (Dallasheh-Khatib et al., 2014; Levin et al., 2008)." (p. 9) 2) "Each intervention session lasted for 20 min embedded within the kindergarten's existing daily schedule and delivered in small groups of five children." (p. 9) 3) "The pretest measures were administered at the start of the school year (September-October) and the posttest measures were administered at the end of the intervention (May-June of the same year)." (p. 14) 4) "The intervention program was structured into six instructional units... Syllable-level phonological awareness (Weeks 1-3)... Letter knowledge (Weeks 4-7)... Phoneme-level phonological awareness (Week 8)... Morphological awareness - inflection (Weeks 9-11)... Morphological awareness - derivation (Weeks 12-13)... Vocabulary and semantic fields (Weeks 14-15)." (pp. 10-11) 5) "First, the study focuses on immediate post-intervention outcomes. Longitudinal research is needed to determine whether the observed improvements in kindergarten extend to literacy performance in primary school." (p. 25) Detailed Analysis: Criterion T asks whether the primary outcome was measured at least one full academic term (approximately 3-4 months) after the intervention began. The paper gives an explicit intervention length of 15 weeks, delivered in six sequential units spanning weeks 1 through 15, and states that the posttest was administered at the end of the intervention. Fifteen weeks is approximately 3.5 months, which falls within the standard's own definition of a term as roughly 3-4 months. The wider anchoring dates reinforce this. The pretest was taken at the start of the school year in September-October and the posttest in May-June of the same academic year, so the whole measurement window spans roughly eight months of the school year, and the 15-week intervention block sits inside it. On either reading - the 15-week intervention-start to posttest interval, or the wider pretest-to-posttest window - the interval from the start of the intervention to the primary outcome measurement is at least one academic term. The standard also explicitly permits short interventions provided outcome tracking reaches a term from the start, and the tracking here does. Criterion T is met because outcomes were measured at the end of a 15-week (~3.5 month) intervention block running within a September-to-June school year, an interval of at least one academic term from intervention start.
-
D Documented Control Group
- The control group's size, demographics, baseline equivalence, and business-as-usual condition are all documented in detail, including fidelity observations.
- "The control group comprised 153 children (73 boys, 80 girls) in 13 kindergarten classrooms (5-16 children each) who received business-as-usual instruction, namely instruction following the standard preschool language education curriculum." (p. 8)
- Relevant Quotes: 1) "The control group comprised 153 children (73 boys, 80 girls) in 13 kindergarten classrooms (5-16 children each) who received business-as-usual instruction, namely instruction following the standard preschool language education curriculum." (p. 8) 2) "There was no statistically significant group difference in gender (x2(1) = .07, p = .798). Group difference in age was not significant [Experimental group: M = 64.3, SD = 3.2; Control group: M = 64.6, SD = 2.8: t (401) = .93, p = .35], or on Raven's matrices scores [Experimental group: M = 9.5, SD = 1.5; Control group: M = 9.5, SD = 1.6: t (401) = .04, p = .969]." (p. 8) 3) "All children were monolingual speakers of Palestinian Arabic and attended kindergartens serving mid-low socioeconomic status (SES) populations, as classified by the Ministry of Education's standardized assessment protocols employing a 10-point scale and evaluating 16 variables (e.g., family economic indicators, parental education, and household composition)." (pp. 7-8) 4) "The control group received business-as-usual instructions which mandates that teachers train children in emergent literacy skills. The domains may include phonological awareness, morphological awareness, letter knowledge. Yet, the curriculum does not detail the activities that teachers can use to train these skills." (p. 9) 5) "Based on data collected in the current study (See the supplementary information on the fidelity of implementations), while both the experimental and control group reported similar content components (e.g., phonological awareness, letter knowledge, morphological awareness, vocabulary), the control group, and unlike the experimental group, did not report use of EF components or diglossia-centered instruction." (p. 9) 6) "Table 1 presents the mean, SD, t-values and Cohen's d effect sizes of children's performance on the SpA metalinguistic and lexical skill tasks by time of testing (Pretest, Posttest)." (p. 15) Detailed Analysis: Criterion D requires that the control group be documented well enough that a reader can judge comparability: size, composition, baseline performance, and what the control condition actually consisted of. The paper supplies all four. Size and composition are given (153 children, 73 boys and 80 girls, in 13 kindergarten classrooms of 5-16 children each), as are the shared eligibility characteristics (monolingual Palestinian Arabic speakers, mid-low SES classified on the Ministry of Education's 16-variable scale, no hearing/vision impairment or diagnosed developmental disability). Baseline equivalence is reported statistically for gender, age, and Raven's nonverbal IQ, all non-significant, and pretest means and standard deviations for the control arm are tabulated alongside the experimental arm for every outcome task. The condition the control group experienced is also characterised rather than merely named: it is the standard preschool language education curriculum, which does mandate emergent literacy work (phonological awareness, morphological awareness, letter knowledge) but does not prescribe activities or address diglossia. The authors go further and report observational fidelity data for control teachers as well as experimental teachers, confirming similar content components but no EF or diglossia-centred instruction. Criterion D is met because the control group's size, demographics, baseline scores, and business-as-usual instructional condition are all explicitly documented, including observed fidelity data.