Investigating the role of Arabic diglossia-centered multi-domain intervention (ADMIN) in enhancing Palestinian-Arabic speaking kindergarten children's emergent literacy

Lina Haj, Rachel Schiff, Ola Ghawi-Dakwar, Elinor Saiegh-Haddad

Published:
ERCT Check Date:
DOI: 10.1007/s11145-026-10766-9
  • reading
  • language arts
  • kindergarten
  • Asia
0
  • C

    Allocation was made at the kindergarten level with all children in a kindergarten sharing one condition, which satisfies the class-level randomisation requirement.

    "Randomization was conducted at the kindergarten level, with each kindergarten assigned to either experimental or control condition. All children within a given kindergarten received the same allocation." (p. 8)

  • E

    The paper states outright that every outcome measure was custom-built for the study, so no standardised exam-based assessment was used.

    "All testing measures were developed for the study and they systematically targeted SpA structures (identical in SpA and StA) and StA structures (cognate and unique)..." (p. 12)

  • T

    The intervention ran for 15 weeks (about 3.5 months) with outcomes measured at its end, meeting the one-term interval requirement.

    "ADMIN consisted of 45 intervention sessions (3 sessions each week) and was implemented by the kindergarten teachers over a period of 15 weeks..." (p. 9)

  • D

    The control group's size, demographics, baseline equivalence, and business-as-usual condition are all documented in detail, including fidelity observations.

    "The control group comprised 153 children (73 boys, 80 girls) in 13 kindergarten classrooms (5-16 children each) who received business-as-usual instruction, namely instruction following the standard preschool language education curriculum." (p. 8)

  • S

    Whole kindergartens - the preschool institutions implementing the programme - were the units of random assignment, though the paper's interchangeable use of "kindergarten" and "kindergarten classroom" leaves the exact number of institutions unclear.

    "Randomization was conducted at the kindergarten level, with each kindergarten assigned to either experimental or control condition. All children within a given kindergarten received the same allocation." (p. 8)

  • I

    The authors designed the intervention, built the measures, trained the testers and fidelity coders, and analysed the data, with no independent evaluator reported.

    "This study is part of a PhD project that was conducted at Bar-Ilan University under the supervision of Prof. Elinor Saiegh-Haddad and Prof. Rachel Schiff." (p. 25)

  • Y

    The intervention ran 15 weeks with outcomes taken immediately at its end and no longer-term follow-up, well below 75% of an academic year.

    "First, the study focuses on immediate post-intervention outcomes. Longitudinal research is needed to determine whether the observed improvements in kindergarten extend to literacy performance in primary school." (p. 25)

  • B

    Children's instructional time and content domains were comparable across arms, and the extra 30-hour teacher training and weekly coaching are integral to the ADMIN package being tested.

    "Level of instruction, classroom management and length of session were reportedly similar with no significant differences between the two groups." (p. 9)

  • R

    ADMIN is presented as a first-of-its-kind programme and internet searching identified no independent replication by a different research team.

    "Research testing early literacy interventions in Arabic is scarce and none has directly integrated diglossia as an intervention component and tested the effect of integrating diglossia into the intervention content and procedures on promoting SpA and StA skills separately." (p. 2)

  • A

    Criterion E was not met and outcomes covered only language and emergent literacy, with no assessment of other main curriculum domains such as numeracy.

    "The study specifically tested the role of ADMIN in promoting children's lexical skills... and metalinguistic awareness... in SpA and StA, as well as their EF skills..." (p. 20)

  • G

    Criterion Y failed, only immediate post-intervention outcomes were collected, and no published follow-up tracks the cohort to graduation.

    "First, the study focuses on immediate post-intervention outcomes. Longitudinal research is needed to determine whether the observed improvements in kindergarten extend to literacy performance in primary school." (p. 25)

  • P

    The paper reports ethics and funding approvals but contains no pre-registration reference, and no registry entry for this trial could be found online.

Abstract

Arabic-speaking children grow up in diglossia where different language varieties serve complementary functions; They use Spoken Arabic (SpA) for everyday communication but acquire literacy in Standard Arabic (StA), the only variety with a conventional orthography. Research shows that while the linguistic distance between SpA and StA challenges StA language and literacy acquisition with an advantage for linguistic structures that are identical in the two varieties, skills in SpA and StA are interdependent. Furthermore, Executive Functions (EF) predict StA comprehension reflecting the need to manage the competition between SpA and StA. The study investigated the effectiveness of Arabic Diglossia-centered Multi-domain Intervention (ADMIN) in promoting Palestinian-Arabic-speaking kindergarten children's language and literacy skills in SpA and in StA, as well as their EF skills. ADMIN is diglossia-centered; It trains SpA-identical structures before gradually transitioning to StA structures while fostering diglossic awareness. ADMIN is also multi-domain; It trains language, literacy and EFs. 403 children (Mean age = 64.4, SD = 3.1 months) were randomly assigned to an experimental group (n = 250) or a control (n = 153) group. The experimental group received three 20-min ADMIN sessions a week for 15 weeks, whereas the control group received business-as-usual instruction. Multi-level modeling for nested data showed greater gains in the experimental group in SpA and StA language (receptive and expressive vocabulary, lexico-phonological representations) and metalinguistic skills (phonological and morphological awareness), and in EF and letter knowledge, compared to the control group. The findings have implications for early literacy education in Arabic diglossia and beyond.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Allocation was made at the kindergarten level with all children in a kindergarten sharing one condition, which satisfies the class-level randomisation requirement.
      • "Randomization was conducted at the kindergarten level, with each kindergarten assigned to either experimental or control condition. All children within a given kindergarten received the same allocation." (p. 8)
      • Relevant Quotes: 1) "Children were randomly allocated to two groups: experimental and control groups. Randomization was conducted at the kindergarten level, with each kindergarten assigned to either experimental or control condition. All children within a given kindergarten received the same allocation." (p. 8) 2) "The experimental group comprised 250 children (116 boys, 134 girls) in 50 Arab kindergarten classrooms (5 children each). The control group comprised 153 children (73 boys, 80 girls) in 13 kindergarten classrooms (5-16 children each) who received business-as-usual instruction, namely instruction following the standard preschool language education curriculum." (p. 8) 3) "A total of 403 Palestinian-Arabic-speaking kindergarten children in Israel (189 boys and 214 girls) ranging in age from 57 to 70 months (M = 64.4, SD = 3.1) participated in the study." (p. 7) Detailed Analysis: Criterion C requires that randomisation be performed on intact groups (classes or larger units) rather than on individual students inside a single classroom, so that contamination between treated and untreated children sharing a teacher and a room is avoided. The paper is explicit on this point: the unit of allocation was the kindergarten, and the authors state that every child within a given kindergarten received the same allocation. This is a cluster design at a level at least as strong as the class, which the ERCT standard accepts as satisfying the class-level requirement. The description also supplies the elements the standard asks to check for: the unit of randomisation (kindergarten), the total sample (403 children), and the number of clusters per arm (50 experimental and 13 control kindergarten classrooms). The paper does not describe the mechanism of the random draw (e.g. random number generator, stratification), which is a reporting weakness, but the unit of randomisation itself is stated unambiguously and is the element criterion C tests. Crucially, no reading of the Method supports allocation of individual children within a shared classroom, so the contamination risk criterion C targets is avoided regardless of how the cluster counts are interpreted. Criterion C is met because allocation was performed at the kindergarten level with all children in a kindergarten sharing the same condition, which is at or above the class level.
    • E

      Exam-based Assessment

      • The paper states outright that every outcome measure was custom-built for the study, so no standardised exam-based assessment was used.
      • "All testing measures were developed for the study and they systematically targeted SpA structures (identical in SpA and StA) and StA structures (cognate and unique)..." (p. 12)
      • Relevant Quotes: 1) "All testing measures were developed for the study and they systematically targeted SpA structures (identical in SpA and StA) and StA structures (cognate and unique) in order to test the contribution of the intervention to the development of skills in SpA and StA. Lexical items for the pre and post tasks were also derived from Al-Fanus corpus (Haj and Saiegh-Haddad, in preparation)." (p. 12) 2) "SpA and StA receptive vocabulary (N = 30, Cronbach Alpha = .70) Children were asked to choose the picture (out of four pictures presented on a computer screen) that matched an oral presentation of a concrete noun." (p. 12) 3) "SpA and StA morphological awareness (N = 24, Cronbach Alpha = .84). To measure morphological awareness, a morphological analogy task was used that required children to produce a morphologically complex word (inflected or derived) in a simple word: complex word pair using the morphological analogy/relatedness illustrated in the previous pair of words." (p. 13) 4) "Letter name knowledge (N items = 29; Cronbach Alpha = .95) The 29 Arabic letters in their basic form were presented randomly on a computer screen and children were required to give the name of each letter." (p. 13) 5) "Familiarity was measured based on the subjective assessment of 10 kindergarten teachers." (p. 12) 6) "All participants were administered the Raven's Colored Progressive Matrices (RCPM; Raven, 1965) to assess nonverbal intelligence and scored within the normal range (8-13, M = 9.5, SD = 1.5)." (p. 7) Detailed Analysis: Criterion E requires that educational outcomes be measured with standardised, widely recognised exam-based assessments rather than instruments created by the researchers for the purposes of the trial. The paper answers this question directly and negatively in the first line of its measures section: "All testing measures were developed for the study." Every outcome instrument reported - receptive and expressive vocabulary, pseudoword repetition for lexico-phonological representations, syllable and phoneme awareness, morphological analogy, and letter name and letter sound knowledge - is an author-constructed task built from the authors' own Al-Fanus lexical corpus, with item familiarity calibrated by a panel of ten kindergarten teachers rather than by any national or standardised norming procedure. Only Cronbach alphas and inter-rater ICCs are reported; there is no reference to national curriculum tests, ministry examinations, or published normed batteries for the outcome measures. Two instruments in the paper derive from published paradigms - Raven's Coloured Progressive Matrices and computerised EF tasks based on the Simon task, the Automated Working Memory Assessment, and the Dimensional Change Card Sort. However, Raven's was used only as a screening/covariate measure of nonverbal intelligence and not as an educational outcome, and the EF tasks measure cognitive rather than curricular attainment; neither is a standardised exam-based assessment of educational achievement in the sense required by criterion E. The instruments were also tightly aligned with the trained content (items drawn from the same corpus used to build the intervention activities), which is precisely the alignment bias the criterion is designed to detect. Criterion E is not met because the paper explicitly states that all outcome measures were developed by the researchers for this study, with no standardised exam used to assess educational attainment.
    • T

      Term Duration

      • The intervention ran for 15 weeks (about 3.5 months) with outcomes measured at its end, meeting the one-term interval requirement.
      • "ADMIN consisted of 45 intervention sessions (3 sessions each week) and was implemented by the kindergarten teachers over a period of 15 weeks..." (p. 9)
      • Relevant Quotes: 1) "ADMIN consisted of 45 intervention sessions (3 sessions each week) and was implemented by the kindergarten teachers over a period of 15 weeks, in accordance with earlier research (Dallasheh-Khatib et al., 2014; Levin et al., 2008)." (p. 9) 2) "Each intervention session lasted for 20 min embedded within the kindergarten's existing daily schedule and delivered in small groups of five children." (p. 9) 3) "The pretest measures were administered at the start of the school year (September-October) and the posttest measures were administered at the end of the intervention (May-June of the same year)." (p. 14) 4) "The intervention program was structured into six instructional units... Syllable-level phonological awareness (Weeks 1-3)... Letter knowledge (Weeks 4-7)... Phoneme-level phonological awareness (Week 8)... Morphological awareness - inflection (Weeks 9-11)... Morphological awareness - derivation (Weeks 12-13)... Vocabulary and semantic fields (Weeks 14-15)." (pp. 10-11) 5) "First, the study focuses on immediate post-intervention outcomes. Longitudinal research is needed to determine whether the observed improvements in kindergarten extend to literacy performance in primary school." (p. 25) Detailed Analysis: Criterion T asks whether the primary outcome was measured at least one full academic term (approximately 3-4 months) after the intervention began. The paper gives an explicit intervention length of 15 weeks, delivered in six sequential units spanning weeks 1 through 15, and states that the posttest was administered at the end of the intervention. Fifteen weeks is approximately 3.5 months, which falls within the standard's own definition of a term as roughly 3-4 months. The wider anchoring dates reinforce this. The pretest was taken at the start of the school year in September-October and the posttest in May-June of the same academic year, so the whole measurement window spans roughly eight months of the school year, and the 15-week intervention block sits inside it. On either reading - the 15-week intervention-start to posttest interval, or the wider pretest-to-posttest window - the interval from the start of the intervention to the primary outcome measurement is at least one academic term. The standard also explicitly permits short interventions provided outcome tracking reaches a term from the start, and the tracking here does. Criterion T is met because outcomes were measured at the end of a 15-week (~3.5 month) intervention block running within a September-to-June school year, an interval of at least one academic term from intervention start.
    • D

      Documented Control Group

      • The control group's size, demographics, baseline equivalence, and business-as-usual condition are all documented in detail, including fidelity observations.
      • "The control group comprised 153 children (73 boys, 80 girls) in 13 kindergarten classrooms (5-16 children each) who received business-as-usual instruction, namely instruction following the standard preschool language education curriculum." (p. 8)
      • Relevant Quotes: 1) "The control group comprised 153 children (73 boys, 80 girls) in 13 kindergarten classrooms (5-16 children each) who received business-as-usual instruction, namely instruction following the standard preschool language education curriculum." (p. 8) 2) "There was no statistically significant group difference in gender (x2(1) = .07, p = .798). Group difference in age was not significant [Experimental group: M = 64.3, SD = 3.2; Control group: M = 64.6, SD = 2.8: t (401) = .93, p = .35], or on Raven's matrices scores [Experimental group: M = 9.5, SD = 1.5; Control group: M = 9.5, SD = 1.6: t (401) = .04, p = .969]." (p. 8) 3) "All children were monolingual speakers of Palestinian Arabic and attended kindergartens serving mid-low socioeconomic status (SES) populations, as classified by the Ministry of Education's standardized assessment protocols employing a 10-point scale and evaluating 16 variables (e.g., family economic indicators, parental education, and household composition)." (pp. 7-8) 4) "The control group received business-as-usual instructions which mandates that teachers train children in emergent literacy skills. The domains may include phonological awareness, morphological awareness, letter knowledge. Yet, the curriculum does not detail the activities that teachers can use to train these skills." (p. 9) 5) "Based on data collected in the current study (See the supplementary information on the fidelity of implementations), while both the experimental and control group reported similar content components (e.g., phonological awareness, letter knowledge, morphological awareness, vocabulary), the control group, and unlike the experimental group, did not report use of EF components or diglossia-centered instruction." (p. 9) 6) "Table 1 presents the mean, SD, t-values and Cohen's d effect sizes of children's performance on the SpA metalinguistic and lexical skill tasks by time of testing (Pretest, Posttest)." (p. 15) Detailed Analysis: Criterion D requires that the control group be documented well enough that a reader can judge comparability: size, composition, baseline performance, and what the control condition actually consisted of. The paper supplies all four. Size and composition are given (153 children, 73 boys and 80 girls, in 13 kindergarten classrooms of 5-16 children each), as are the shared eligibility characteristics (monolingual Palestinian Arabic speakers, mid-low SES classified on the Ministry of Education's 16-variable scale, no hearing/vision impairment or diagnosed developmental disability). Baseline equivalence is reported statistically for gender, age, and Raven's nonverbal IQ, all non-significant, and pretest means and standard deviations for the control arm are tabulated alongside the experimental arm for every outcome task. The condition the control group experienced is also characterised rather than merely named: it is the standard preschool language education curriculum, which does mandate emergent literacy work (phonological awareness, morphological awareness, letter knowledge) but does not prescribe activities or address diglossia. The authors go further and report observational fidelity data for control teachers as well as experimental teachers, confirming similar content components but no EF or diglossia-centred instruction. Criterion D is met because the control group's size, demographics, baseline scores, and business-as-usual instructional condition are all explicitly documented, including observed fidelity data.
  • Level 2 Criteria

    • S

      School-level RCT

      • Whole kindergartens - the preschool institutions implementing the programme - were the units of random assignment, though the paper's interchangeable use of "kindergarten" and "kindergarten classroom" leaves the exact number of institutions unclear.
      • "Randomization was conducted at the kindergarten level, with each kindergarten assigned to either experimental or control condition. All children within a given kindergarten received the same allocation." (p. 8)
      • Relevant Quotes: 1) "Children were randomly allocated to two groups: experimental and control groups. Randomization was conducted at the kindergarten level, with each kindergarten assigned to either experimental or control condition. All children within a given kindergarten received the same allocation." (p. 8) 2) "The experimental group comprised 250 children (116 boys, 134 girls) in 50 Arab kindergarten classrooms (5 children each). The control group comprised 153 children (73 boys, 80 girls) in 13 kindergarten classrooms (5-16 children each)..." (p. 8) 3) "Each intervention session lasted for 20 min embedded within the kindergarten's existing daily schedule and delivered in small groups of five children." (p. 9) 4) "Teachers also received weekly guidance from a designated research assistant who arrived to each kindergarten classroom once a week and accompanied the teacher during the implementation of the intervention session providing modeling and feedback." (pp. 9-10) 5) "Tasks in pretest and posttest were individually administered in a quiet room within the kindergarten center by the same research assistants proficient in testing children and who spoke the same Arabic variety spoken by the children." (p. 14) Detailed Analysis: Criterion S requires randomisation among the educational institutions or units that implement the intervention. The ERCT standard states explicitly that 'school' here means the educational institution or implementing unit and gives preschool centers as a qualifying example. In the Israeli system a kindergarten (gan) is a standalone early-years institution with its own teacher and premises, and the paper refers to the "kindergarten center" as the physical site where testing took place. The intervention was delivered by the kindergarten teachers within each kindergarten's own daily schedule, so the kindergarten is the implementing unit here. This verification specifically interrogated whether the reported "50 Arab kindergarten classrooms (5 children each)" describes 50 institutions or merely 50 small delivery groups. The concern is real, because the intervention was "delivered in small groups of five children," exactly matching the five-children-per-classroom figure, and because a 50-versus-13 split of 63 units is a striking imbalance for a simple random allocation. The most consistent reading of the Method is that each participating experimental kindergarten contributed a single group of five children (the delivery group size dictated how many children per site could be served), whereas control kindergartens, being under no such delivery constraint, contributed 5 to 16 children each; this reconciles 250 = 50 x 5 with 153 across 13 sites. The statement that a research assistant "arrived to each kindergarten classroom once a week" and that testing occurred "within the kindergarten center" is consistent with distinct physical sites rather than parallel groups inside one building. Decisively for this criterion, the level of allocation is stated unambiguously and identically under either reading: the kindergarten was assigned as a whole and "All children within a given kindergarten received the same allocation," so no kindergarten contained both conditions and allocation was never at the individual-child or within-institution level. What is genuinely under-reported is the number of allocated institutions and the randomisation mechanism, since the paper uses "kindergarten" and "kindergarten classroom" interchangeably and never states a cluster count in institutional terms; this is a material reporting weakness that should be flagged, but it concerns how many units were randomised rather than at what level. Criterion S is met because randomisation was conducted at the kindergarten level, the kindergarten being the preschool institution implementing the intervention, although the paper's conflation of "kindergarten" with "kindergarten classroom" leaves the exact number of allocated institutions unverifiable.
    • I

      Independent Conduct

      • The authors designed the intervention, built the measures, trained the testers and fidelity coders, and analysed the data, with no independent evaluator reported.
      • "This study is part of a PhD project that was conducted at Bar-Ilan University under the supervision of Prof. Elinor Saiegh-Haddad and Prof. Rachel Schiff." (p. 25)
      • Relevant Quotes: 1) "This study is part of a PhD project that was conducted at Bar-Ilan University under the supervision of Prof. Elinor Saiegh-Haddad and Prof. Rachel Schiff." (Acknowledgments, p. 25) 2) "Funding Open access funding provided by Bar-Ilan University. Funding was provided by the Ministry of Education, Grant No: 10208 (2018-2020), Elinor Saiegh-Haddad" (p. 25) 3) "Items for the intervention activities derived from a lexical corpus totaling 19,836-word tokens compiled based on 70 Ministry-of-Education recommended storybooks for children (aka Al-Fanus Library) (Haj and Saiegh-Haddad, in preparation)." (p. 9) 4) "The teachers implementing ADMIN received a 30-h training course to familiarize them with the theoretical framework and the specific content and procedures of the intervention. Teachers also received weekly guidance from a designated research assistant who arrived to each kindergarten classroom once a week and accompanied the teacher during the implementation of the intervention session providing modeling and feedback. This designated research assistant also collected fidelity data (as explained later)." (pp. 9-10) 5) "Tasks in pretest and posttest were individually administered in a quiet room within the kindergarten center by the same research assistants proficient in testing children and who spoke the same Arabic variety spoken by the children. The research assistants underwent a three-day training workshop by the authors." (p. 14) 6) "In a previous study, we tested the effectiveness of metalinguistic awareness training in kindergarten children focusing on phonological and morphological awareness and showed that an intervention that takes linguistic distance, both phonological and morphological, into account can better help children bridge the gap in metalinguistic awareness between SpA and StA (Saiegh-Haddad, 2023)." (p. 5) Detailed Analysis: Criterion I requires that the trial be conducted independently of the people who designed the intervention, or that some explicit third-party oversight of implementation, measurement, and analysis be documented. Neither condition holds here. The ADMIN programme is the authors' own creation, built on their own prior theoretical work (MAWRID, the diglossia-centred approach of Saiegh-Haddad 2023) and on their own unpublished Al-Fanus lexical corpus. The trial is a PhD project supervised by two of the co-authors at the same university, and the grant is held by the senior author. Independence is absent at every operational stage as well. The outcome instruments were developed by the authors; the research assistants who administered pretests and posttests were trained by the authors; and the research assistants who collected the implementation fidelity data were the same people embedded weekly in the experimental classrooms to coach the teachers, which means fidelity ratings were produced by staff who were both non-blind and invested in the intervention's delivery. The paper contains no disclosure statement, no conflict of interest declaration, and no mention of an external evaluator, an independent statistician, or blinded assessors. Nothing in the paper resembles the documented separations shown in the standard's examples of acceptable exceptions. Criterion I is not met because the intervention's designers also designed the measures, trained the testers and coaches, and ran the analysis, with no external or third-party evaluation reported.
    • Y

      Year Duration

      • The intervention ran 15 weeks with outcomes taken immediately at its end and no longer-term follow-up, well below 75% of an academic year.
      • "First, the study focuses on immediate post-intervention outcomes. Longitudinal research is needed to determine whether the observed improvements in kindergarten extend to literacy performance in primary school." (p. 25)
      • Relevant Quotes: 1) "ADMIN consisted of 45 intervention sessions (3 sessions each week) and was implemented by the kindergarten teachers over a period of 15 weeks, in accordance with earlier research (Dallasheh-Khatib et al., 2014; Levin et al., 2008)." (p. 9) 2) "The experimental group received three 20-min ADMIN sessions a week for 15 weeks, whereas the control group received business-as-usual instruction." (Abstract, p. 1) 3) "The pretest measures were administered at the start of the school year (September-October) and the posttest measures were administered at the end of the intervention (May-June of the same year)." (p. 14) 4) "Several limitations are in turn. First, the study focuses on immediate post-intervention outcomes. Longitudinal research is needed to determine whether the observed improvements in kindergarten extend to literacy performance in primary school." (p. 25) Detailed Analysis: Criterion Y requires that outcomes be measured at least 75% of an academic year (roughly 7 to 7.5 months of a 9-10 month year) after the intervention begins. The paper's own account of the intervention is a 15-week block of 45 sessions, with the posttest administered "at the end of the intervention." Fifteen weeks is approximately 3.5 months, which is well under 75% of an academic year, and no follow-up measurement was taken after that endpoint. The pretest-to-posttest window is longer - September-October to May-June, about eight months - but that window is anchored on the pretest, not on the start of the intervention, and the paper gives no date for when the 15-week block began. The sequence of six instructional blocks (weeks 1-3, 4-7, 8, 9-11, 12-13, 14-15) ending at the posttest indicates the programme concluded immediately before the May-June testing, which places its start in roughly February. On the criterion's own terms - the interval from intervention start to outcome measurement - the evidence supports about 3.5 months rather than a year. The authors themselves characterise their results as "immediate post-intervention outcomes" and call for longitudinal research, confirming that no year-long tracking was performed. Criterion Y is not met because outcomes were measured immediately at the end of a 15-week intervention, far short of 75% of an academic year from intervention start.
    • B

      Balanced Control Group

      • Children's instructional time and content domains were comparable across arms, and the extra 30-hour teacher training and weekly coaching are integral to the ADMIN package being tested.
      • "Level of instruction, classroom management and length of session were reportedly similar with no significant differences between the two groups." (p. 9)
      • Relevant Quotes: 1) "Each intervention session lasted for 20 min embedded within the kindergarten's existing daily schedule and delivered in small groups of five children." (p. 9) 2) "The teachers implementing ADMIN received a 30-h training course to familiarize them with the theoretical framework and the specific content and procedures of the intervention. Teachers also received weekly guidance from a designated research assistant who arrived to each kindergarten classroom once a week and accompanied the teacher during the implementation of the intervention session providing modeling and feedback." (pp. 9-10) 3) "The control group received business-as-usual instructions which mandates that teachers train children in emergent literacy skills. The domains may include phonological awareness, morphological awareness, letter knowledge. Yet, the curriculum does not detail the activities that teachers can use to train these skills. The curriculum also acknowledges diglossia. Yet, it does not explain how it might be addressed in literacy instruction for children." (p. 9) 4) "Based on data collected in the current study... while both the experimental and control group reported similar content components (e.g., phonological awareness, letter knowledge, morphological awareness, vocabulary), the control group, and unlike the experimental group, did not report use of EF components or diglossia-centered instruction. Level of instruction, classroom management and length of session were reportedly similar with no significant differences between the two groups." (p. 9) 5) "15 sessions (33%; one session per week) were observed by trained research assistants for fidelity testing of the teachers in the experimental and the control groups." (p. 11) 6) "The study investigated the effectiveness of Arabic Diglossia-centered Multi-domain Intervention (ADMIN) in promoting Palestinian-Arabic-speaking kindergarten children's language and literacy skills in SpA and in StA, as well as their EF skills." (Abstract, p. 1) Detailed Analysis: Working through the decision tree: the first question is whether the intervention adds time or budget relative to the control. Instructional time for children was not increased - the 20-minute sessions were "embedded within the kindergarten's existing daily schedule," and the fidelity observations confirm that "level of instruction, classroom management and length of session were reportedly similar with no significant differences between the two groups." Both arms covered the same content domains (phonological awareness, morphological awareness, letter knowledge, vocabulary). The contrast is therefore one of instructional approach - diglossia-centred sequencing and embedded EF strategies versus the undifferentiated business-as-usual curriculum - rather than of quantity of education delivered to children. There is nonetheless a real resource asymmetry on the teacher side: experimental teachers received a 30-hour training course, detailed written session protocols, and weekly in-class modelling and feedback from a research assistant, none of which the control teachers received. This must be named explicitly. The question the standard poses is whether that support is integral to the intervention package being tested or a separable add-on that could have been balanced. ADMIN is defined as a specific content sequence and a specific set of EF strategies that kindergarten teachers must be taught to deliver; the training and coaching are the delivery mechanism for that sequence, not an independent extra. This matches the standard's own worked example in which teacher professional development supplied only to the intervention arm was judged integral to the intervention package rather than a confounding add-on. Delivery format also differed (small groups of five in the experimental arm), which the paper does not match in the control condition and which is a genuine limitation on the isolation of the specific ADMIN content; but it is likewise part of how the programme is defined and delivered. Criterion B is met because children's instructional time and content domains were comparable across arms, and the additional teacher training and weekly coaching supplied to the experimental arm are integral components of the ADMIN package being tested rather than separable extra resources.
  • Level 3 Criteria

    • R

      Reproduced

      • ADMIN is presented as a first-of-its-kind programme and internet searching identified no independent replication by a different research team.
      • "Research testing early literacy interventions in Arabic is scarce and none has directly integrated diglossia as an intervention component and tested the effect of integrating diglossia into the intervention content and procedures on promoting SpA and StA skills separately." (p. 2)
      • Relevant Quotes: 1) "Research testing early literacy interventions in Arabic is scarce and none has directly integrated diglossia as an intervention component and tested the effect of integrating diglossia into the intervention content and procedures on promoting SpA and StA skills separately. Moreover, no intervention has combined language, emergent literacy and executive functions..." (p. 2) 2) "While the unique challenges that diglossia poses for Arabic-speaking children in acquiring literacy have received increasing attention, research on effective interventions in this context remains limited (Saiegh-Haddad, 2023)." (p. 6) 3) "Moreover, Saiegh-Haddad (2023) tested the effectiveness of a diglossia-centered intervention that first trained metalinguistic awareness in SpA and then in StA and demonstrated significant improvements in the experimental compared to the control group in morphological analogies and syllable blending in SpA and in StA." (p. 6) 4) "Third, the study was conducted among speakers of PA residing in Israel... generalization to other Arabic-speaking contexts requires consideration of cultural, regional, and sociolinguistic variability that may moderate intervention effects." (p. 25) Detailed Analysis: Criterion R requires that this specific study have been independently replicated by a different research team in a different context and published in a peer-reviewed journal. The paper positions ADMIN as novel on its own account, stating that no prior intervention has integrated diglossia as an intervention component or combined language, emergent literacy and EF training. A programme presented as the first of its kind cannot yet have been replicated. Internet searching was carried out for this verification across Springer Nature Link, publisher sites, and general web sources for any independent trial of ADMIN or of a diglossia-centred multi-domain kindergarten programme by another author team. No such replication was found. The closest related works all belong to the same team or test different programmes: Saiegh-Haddad (2023), "Embracing diglossia in early literacy education in Arabic: A pilot intervention study with kindergarten children" (Oxford Review of Education, 49(1), 48-68), is prior work by the senior author of this paper and is described as the foundation ADMIN builds on, and it was funded under the same Ministry of Education grant (No. 10208, 2018-2020); same-team antecedents do not satisfy the criterion. The other cited Arabic kindergarten intervention studies - Levin, Saiegh- Haddad, Hende and Ziv (2008), "Early literacy in Arabic: An intervention study among Israeli Palestinian kindergartners" (Applied Psycholinguistics, 29(3), 413-436), and Dallasheh- Khatib, Ibrahim and Karni (2014) - tested different programmes (phonological awareness and letter knowledge, and morphological training respectively) and are cited only for session-count precedent and general context, not as replications of ADMIN; the 2008 study also shares an author with the present paper. This paper was accepted on 14 January 2026 and published in 2026, so no independent replication would yet be expected; none was identified, and the authors themselves flag generalisation beyond Palestinian-Arabic speakers in Israel as an open question. Criterion R is not met because ADMIN is presented as a novel programme and no independent replication by another research team could be found in the peer-reviewed literature.
    • A

      All-subject Exams

      • Criterion E was not met and outcomes covered only language and emergent literacy, with no assessment of other main curriculum domains such as numeracy.
      • "The study specifically tested the role of ADMIN in promoting children's lexical skills... and metalinguistic awareness... in SpA and StA, as well as their EF skills..." (p. 20)
      • Relevant Quotes: 1) "All testing measures were developed for the study..." (p. 12) 2) "The study specifically tested the role of ADMIN in promoting children's lexical skills (receptive and expressive vocabulary, lexico-phonological representations) and metalinguistic awareness (syllable awareness; phoneme awareness, morphological awareness) in SpA and StA, as well as their EF skills (non-verbal memory and shifting skills) in the experimental compared to the control group." (p. 20) 3) "In addition to metalinguistic awareness, we also tested the effectiveness of the intervention in enhancing children's letter knowledge, a proximal emergent literacy skill." (p. 23) 4) "1. What is the role of ADMIN in promoting children's SpA language (receptive and expressive vocabulary; lexico-phonological representations) and metalinguistic awareness (syllable awareness; phoneme awareness; morphological awareness) in the experimental compared to the control group?" (p. 7) Detailed Analysis: Criterion A requires assessment of all main subjects taught at the relevant educational level, using standardised exam-based assessments, and the standard states directly that if criterion E is not met then criterion A cannot be met. Criterion E fails here because the paper states that all outcome measures were developed for the study, so criterion A fails on that prerequisite alone. The substantive coverage also falls short independently. Every outcome reported falls within a single broad domain - Arabic language and emergent literacy (vocabulary, lexico-phonological representations, syllable and phoneme awareness, morphological awareness, letter name and letter sound knowledge) - supplemented by three non-curricular executive function tasks. Early numeracy, which is a core domain of the kindergarten curriculum alongside language and literacy, was not measured at all, nor was any other curricular area. The paper offers no rationale for restricting outcomes to the trained domain, and the exception for highly specialised upper-secondary or vocational interventions does not apply to a general kindergarten programme. Criterion A is not met because criterion E failed and, independently, only language and emergent literacy outcomes were assessed with no measurement of other main kindergarten curriculum domains such as numeracy.
    • G

      Graduation Tracking

      • Criterion Y failed, only immediate post-intervention outcomes were collected, and no published follow-up tracks the cohort to graduation.
      • "First, the study focuses on immediate post-intervention outcomes. Longitudinal research is needed to determine whether the observed improvements in kindergarten extend to literacy performance in primary school." (p. 25)
      • Relevant Quotes: 1) "The pretest measures were administered at the start of the school year (September-October) and the posttest measures were administered at the end of the intervention (May-June of the same year)." (p. 14) 2) "Several limitations are in turn. First, the study focuses on immediate post-intervention outcomes. Longitudinal research is needed to determine whether the observed improvements in kindergarten extend to literacy performance in primary school." (p. 25) 3) "Our results demonstrate that our intervention has helped children contract this distance and develop greater awareness of StA morphological structure in kindergarten, thereby providing them with a stronger foundation for later literacy acquisition in StA." (p. 23) Detailed Analysis: Criterion G requires that participants be tracked through to graduation from the relevant educational stage, and the standard specifies that if criterion Y is not met then criterion G cannot be met. Criterion Y fails here, so criterion G fails on that prerequisite. The substance confirms it. Measurement stopped at the posttest administered at the end of the intervention, and the authors state explicitly in their limitations that the study "focuses on immediate post-intervention outcomes" and that longitudinal research would be needed to see whether gains carry into primary school. There is no follow-up wave, no description of tracking procedures, and no linkage to later school records. Internet searching was carried out for follow-up publications by the same authors tracking this cohort. The only later-stage follow-up located is a conference presentation, not a peer-reviewed paper: "A longitudinal study of the contribution of a diglossia-centered Arabic literacy intervention in kindergarten on reading skills in the 3rd grade" by Elinor Saiegh-Haddad, Rachel Schiff, Lina Haj and Ola Ghawi-Dakwar, presented at the 37th Annual Symposium on Arabic Linguistics (Long Island University, 23 February 2024). Only the title and author list are published in the symposium programme; no abstract text is available, so no verbatim quotes from it can be given, and it is not possible to confirm whether it follows this ADMIN cohort or the earlier Saiegh-Haddad (2023) pilot cohort. In any case, third-grade reading outcomes are not graduation from primary school, and a conference talk is not a peer-reviewed publication of graduation tracking. No published paper tracking any of these cohorts to graduation was found. Criterion G is not met because criterion Y failed, the authors state that only immediate post-intervention outcomes were collected, and no published follow-up tracking participants to graduation could be identified.
    • P

      Pre-Registered

      • The paper reports ethics and funding approvals but contains no pre-registration reference, and no registry entry for this trial could be found online.
      • Relevant Quotes: 1) "Authorization for data collection was obtained from the Ministry of Education's chief scientist (approval No. 10208, 2018-2020), and from Bar-Ilan University's IRB. Written parental consent was obtained for all participating children." (p. 8) 2) "Funding Open access funding provided by Bar-Ilan University. Funding was provided by the Ministry of Education, Grant No: 10208 (2018-2020), Elinor Saiegh-Haddad" (p. 25) 3) "Supplementary Information The online version contains supplementary material available at https://doi.org/10.1007/s11145-026-10766-9." (p. 25) 4) "Acknowledgments This study is part of a PhD project that was conducted at Bar-Ilan University under the supervision of Prof. Elinor Saiegh-Haddad and Prof. Rachel Schiff." (p. 25) Detailed Analysis: Criterion P requires a pre-registered study protocol - a registry entry naming hypotheses, methods and planned analyses, dated before data collection began, with a citable link or ID. The paper contains no such reference anywhere. There is no mention of ClinicalTrials.gov, ISRCTN, AEA RCT Registry, OSF, AsPredicted, or any other registry; no registration number; and no registration date. Internet searching was carried out for this verification against trial and study registries and the publisher record for any registration associated with this trial, these authors, or Ministry of Education grant No. 10208. No pre-registration entry was found, and the article record carries no trial registration statement. What the paper does provide is ethics and authorisation documentation - Ministry of Education chief scientist approval No. 10208 covering 2018-2020 and Bar-Ilan University IRB approval - together with a grant number. These are governance and funding approvals, not public pre-registration of hypotheses and analysis plans, and they do not serve the selective-reporting protection that criterion P tests. The supplementary information referenced concerns fidelity of implementation and session structure rather than a registered protocol. Note also that the analysis strategy is described only within the results narrative, with no reference to a pre-specified analysis plan or pre-specified primary outcome among the many outcome measures reported. Criterion P is not met because neither the paper nor any searched registry contains a pre-registration reference, registry ID, or registration date.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.