Teaching young multilingual learners: impacts of a professional learning programme on teachers' practices and students' language and literacy skills

Leslie Babinski, Steven Amendum, Madeline Carrig, Steven Knotek, Marta Sanchez

Published:
ERCT Check Date:
DOI: 10.1080/01434632.2025.2472880
  • reading
  • L2 languages
  • K12
  • US
  • digital assessment
0
  • C

    Randomisation was conducted at the school level, which is stronger than and therefore satisfies the class-level RCT requirement.

    "Schools were blocked by district and then assigned to treatment or control group within district using a permuted blocks randomisation procedure."

  • E

    The study used MAP Growth Reading K-2, a nationally normed standardised assessment, satisfying the exam-based assessment requirement.

    "MAP Growth Reading K-2 ... is a nationally normed achievement test that is designed to measure the general academic literacy achievement and growth of K-2 students ..."

  • T

    Outcomes were measured across a full academic year, which exceeds and therefore satisfies the one-term duration requirement.

    "MAP Growth Reading K-2 ... was used to assess students' growth in language and literacy at three time points during the school year, including beginning-of-year baseline, mid-year, and end-of-year."

  • D

    The paper reports pre/post outcome means by condition but does not provide a description of demographic characteristics broken down by condition, nor an explicit statement of what the control group received or did instead of the PL.

    "Descriptive statistics, including mean and standard deviation, for pre and post intervention teacher outcomes are provided in Table 3." (p. 2452)

  • S

    Randomisation was explicitly conducted at the school level, satisfying the school-level RCT requirement.

    "Schools were blocked by district and then assigned to treatment or control group within district using a permuted blocks randomisation procedure."

  • I

    The study authors who designed and led the trial are co-founders of the company commercialising the intervention, and no independent evaluation team is described, so independent conduct is not met.

    "The BELLA professional learning program is available from FigStar Learning (figstar.org), cofounded by Drs. Babinski and Amendum. If the program is commercially successful in the future, Drs. Babinski and Amendum may benefit financially."

  • Y

    The PL programme and outcome tracking spanned a full academic year, satisfying the year duration requirement.

    "Using a randomised control trial, we evaluated the impact of a year-long professional learning (PL) programme for teachers ..."

  • B

    The additional PL time is the explicit treatment variable being tested, so the control group's lack of equivalent time is by design rather than an unbalanced confound.

    "In total, teachers in the intervention group spent about 35 hours engaged in PL activities across the school year."

  • R

    All related supporting RCTs are by the same author team; no independent replication by a different research group was found in the paper or via internet search.

  • A

    Only language and literacy outcomes were measured, with no assessment of other core subjects such as mathematics or science.

    "Test results yield an RIT score, which is an overall language and literacy score based on four domains or goal areas: Foundational Skills, Language and Writing, Literature and Information, and Vocabulary Use and Functions."

  • G

    Outcome tracking ended at the close of the single intervention school year, with no evidence of follow-up toward graduation found in the paper or through internet search.

  • P

    No reference to a pre-registered protocol, registry, or registration date is found anywhere in the paper or via internet search.

Abstract

Using a randomised control trial, we evaluated the impact of a year-long professional learning (PL) programme for teachers designed to support their multilingual learners' language and literacy skills. The PL included evidence-based instructional practices for literacy, a semi-structured model for collaboration, and a strengths-based approach in instruction. Participants included 39 kindergarten and first-grade teachers and 106 multilingual learners (MLs) from 13 schools in two districts in the Southeastern United States. Using a mixed modelling approach, we found significantly higher literacy growth on the Measures of Academic Progress (MAP) for MLs whose teachers participated in the PL. The PL also had a significant impact on teachers' collaboration planning and processes with their students' English as a Second Language (ESL) teachers. Effect sizes for the impact on teachers' use of evidence-based instructional strategies and collaboration frequency were large but not statistically significant. The findings from this study show positive impacts of a teacher professional learning programme on teachers' practices and young MLs' literacy growth.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was conducted at the school level, which is stronger than and therefore satisfies the class-level RCT requirement.
      • "Schools were blocked by district and then assigned to treatment or control group within district using a permuted blocks randomisation procedure."
      • Relevant Quotes: 1) "Randomisation procedures for condition were conducted after all teacher consent forms were collected. Schools were blocked by district and then assigned to treatment or control group within district using a permuted blocks randomisation procedure." (Method, Sampling and randomisation procedures) 2) "Thirty-nine classroom teachers (18 kindergarten and 21 first grade) from 13 schools enrolled in the RCT." (Participants and settings) 3) "ESL teachers participated alongside either kindergarten or first-grade teachers from the same school." (Description of the professional learning programme, PL process) Detailed Analysis: All three quotes were checked against the PDF and are verbatim. The paper explicitly states that the unit of randomisation was the school ("Schools were ... assigned to treatment or control group"), which is consistent with the design in which ESL and classroom teachers from the same school worked together as a school-based team. This design choice avoids within-school contamination, since an entire school's participating teachers are placed in the same condition. The ERCT Standard treats a stronger school-level randomisation as automatically satisfying the weaker class-level criterion. Final: Criterion C is met because randomisation occurred at the school level, which is stronger than and therefore satisfies the class-level requirement.
    • E

      Exam-based Assessment

      • The study used MAP Growth Reading K-2, a nationally normed standardised assessment, satisfying the exam-based assessment requirement.
      • "MAP Growth Reading K-2 ... is a nationally normed achievement test that is designed to measure the general academic literacy achievement and growth of K-2 students ..."
      • Relevant Quotes: 1) "MAP Growth Reading K-2 (Northwest Evaluation Association 2019) was used to assess students' growth in language and literacy at three time points during the school year ... MAP Growth Reading K-2 is a nationally normed achievement test that is designed to measure the general academic literacy achievement and growth of K-2 students ..." (Measures, Student language and literacy skills) 2) "Test-retest correlations for MAP Growth Reading K-2 RIT scores range from .82 to .86 and reliability coefficients range from .95 to .97 for kindergarten through second grade students (Northwest Evaluation Association 2011)." (Measures) Detailed Analysis: Both quotes were checked against the PDF and are verbatim. The primary student outcome measure, MAP Growth Reading K-2, is described as a nationally normed achievement test with documented reliability statistics, not an instrument custom-built for this study. This satisfies the requirement for a standardised, widely recognised exam-based assessment. Final: Criterion E is met because the study used MAP Growth Reading K-2, a nationally normed, standardised achievement test.
    • T

      Term Duration

      • Outcomes were measured across a full academic year, which exceeds and therefore satisfies the one-term duration requirement.
      • "MAP Growth Reading K-2 ... was used to assess students' growth in language and literacy at three time points during the school year, including beginning-of-year baseline, mid-year, and end-of-year."
      • Relevant Quotes: 1) "Using a randomised control trial, we evaluated the impact of a year-long professional learning (PL) programme ..." (Abstract) 2) "MAP Growth Reading K-2 ... was used to assess students' growth in language and literacy at three time points during the school year, including beginning-of-year baseline, mid-year, and end-of-year." (Measures) Detailed Analysis: Both quotes were checked against the PDF and are verbatim. The intervention and its outcome tracking spanned an entire academic year, from a beginning-of-year baseline through an end-of-year assessment. Since the ERCT Standard specifies that a year-long duration automatically satisfies the weaker term-duration requirement, and this interval far exceeds one academic term, this criterion is met. Final: Criterion T is met because outcomes were tracked across a full academic year, well beyond the minimum one-term requirement.
    • D

      Documented Control Group

      • The paper reports pre/post outcome means by condition but does not provide a description of demographic characteristics broken down by condition, nor an explicit statement of what the control group received or did instead of the PL.
      • "Descriptive statistics, including mean and standard deviation, for pre and post intervention teacher outcomes are provided in Table 3." (p. 2452)
      • Relevant Quotes: 1) "See Table 2 for teacher and student demographic information." (p. 2450) 2) "Descriptive statistics, including mean and standard deviation, for pre and post intervention teacher outcomes are provided in Table 3." (p. 2452) 3) Table 3 header: "Teacher Outcome Variable ... Control Mean (SD) ... Intervention Mean (SD)" with rows including "Instructional Strategies - Pre 1.95 (1.64)" for the control group. (p. 2452) 4) "Condition = Intervention" is coded as a dummy/effects-coded predictor throughout the analytic models (e.g. p. 2453-2456), but no narrative text describes what activity, business- as-usual instruction, or alternative professional development (if any) the control-condition teachers and students experienced during the school year. Detailed Analysis: Criterion D requires clear documentation of the control group's demographic characteristics, baseline performance, and the conditions/ treatment it received. This paper provides control-group baseline and post scores for the teacher-outcome variables in Table 3 (partially satisfying the "baseline performance" requirement), but Table 2, which presents teacher and student demographic characteristics (gender, experience, degree, race/ethnicity, home language, country of birth, etc.), reports only aggregate figures for the full sample (n = 39 teachers, n = 106 students) and is not broken down by condition, and no counts of teachers or students per condition are reported anywhere in the paper. There is therefore no way to verify from the text whether the control and intervention groups were demographically comparable, or even how large each group was. Moreover, nowhere in the Method section is there an explicit statement describing what, if anything, control-condition teachers did instead of the BELLA PL programme (e.g. "business as usual" instruction, a different/lighter PL offering, or no PL at all). The absence of this description means readers cannot confirm that the control condition was a stable, well-specified comparison condition as opposed to an unknown or variable set of practices. Because the demographic breakdown by condition, the size of the control group, and an explicit description of the control condition's treatment are all missing, the documentation of the control group falls short of what criterion D requires, despite the partial information available via Table 3. Criterion D is not met because the paper lacks a condition-level demographic and size breakdown and an explicit description of the control group's treatment/business-as-usual condition.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was explicitly conducted at the school level, satisfying the school-level RCT requirement.
      • "Schools were blocked by district and then assigned to treatment or control group within district using a permuted blocks randomisation procedure."
      • Relevant Quotes: 1) "Schools were blocked by district and then assigned to treatment or control group within district using a permuted blocks randomisation procedure." (Sampling and randomisation procedures) 2) "Thirty-nine classroom teachers (18 kindergarten and 21 first grade) from 13 schools enrolled in the RCT." (Participants and settings) Detailed Analysis: Both quotes were checked against the PDF and are verbatim. The quote explicitly identifies the school as the unit assigned to condition, not individual classes or teachers within a school. With 13 schools and 39 teachers overall, this is a genuine school-level cluster randomised design, matching the ERCT Standard's definition of 'school' as the implementing institution. Final: Criterion S is met because randomisation was conducted at the school level.
    • I

      Independent Conduct

      • The study authors who designed and led the trial are co-founders of the company commercialising the intervention, and no independent evaluation team is described, so independent conduct is not met.
      • "The BELLA professional learning program is available from FigStar Learning (figstar.org), cofounded by Drs. Babinski and Amendum. If the program is commercially successful in the future, Drs. Babinski and Amendum may benefit financially."
      • Relevant Quotes: 1) "The BELLA professional learning program is available from FigStar Learning (figstar.org), cofounded by Drs. Babinski and Amendum. If the program is commercially successful in the future, Drs. Babinski and Amendum may benefit financially." (Disclosure statement) 2) "Trained observers, blind to condition, scored the presence or absence of between three and seven indicators on the Classroom Observation Tool ..." (Measures, Classroom observations: instructional strategies) 3) No statement anywhere in the paper (Method, Data-analytic strategy, Disclosure statement, or Funding sections) indicates that data collection, implementation coaching, or analysis was carried out by an evaluator or team independent of the intervention's designers; the author team (Babinski, Amendum, Carrig, Knotek, Sanchez) both designed the PL programme and conducted the trial, its implementation coaching, and its statistical analysis. Detailed Analysis: Quotes 1 and 2 were checked against the PDF and are verbatim. The disclosure statement reveals that two of the study authors, who also designed and led this trial, are co-founders of the company (FigStar Learning) that commercialises the "BELLA" PL programme evaluated here and stand to benefit financially from its future commercial success. While classroom observers were blinded to condition for the instructional-strategy and cultural-wealth ratings, this addresses measurement bias for those specific outcomes only; it is not equivalent to independent conduct of the overall study. Unlike the ERCT Standard's exception examples (e.g. a study where the paper explicitly states the provider "did not participate in data collection, analysis, or the formulation of conclusions," or where an independent government body is stated to have "designed and led" the trial), this paper contains no comparable statement distancing data collection, analysis, or conclusions from the conflicted authors. The same research team that designed the PL programme also recruited schools, delivered the coaching, ran the ANCOVA/mixed models, and wrote the conclusions, with no mention of an external, independent evaluation organisation overseeing analysis or interpretation. Final: Criterion I is not met because the study was designed, implemented, and analysed by the same team that developed and commercially benefits from the intervention, with no independent evaluator or analyst described, despite a disclosed financial conflict of interest.
    • Y

      Year Duration

      • The PL programme and outcome tracking spanned a full academic year, satisfying the year duration requirement.
      • "Using a randomised control trial, we evaluated the impact of a year-long professional learning (PL) programme for teachers ..."
      • Relevant Quotes: 1) "Using a randomised control trial, we evaluated the impact of a year-long professional learning (PL) programme for teachers ..." (Abstract) 2) "... coherence across a sustained yearlong implementation, and collective participation with colleagues (c.f., Desimone 2009; Desimone and Garet 2015)." (Conceptual framework & prior research on professional learning) 3) "MAP Growth Reading K-2 ... was used to assess students' growth in language and literacy at three time points during the school year, including beginning-of-year baseline, mid-year, and end-of-year." (Measures) Detailed Analysis: All quotes were checked against the PDF and are verbatim. The PL programme and its associated data collection explicitly span a full academic year, from a beginning-of-year baseline to an end-of-year assessment, with a mid-year check-in as well. This comfortably meets the requirement of tracking outcomes across at least 75% of an academic year. Final: Criterion Y is met because the intervention and outcome tracking spanned a full academic year.
    • B

      Balanced Control Group

      • The additional PL time is the explicit treatment variable being tested, so the control group's lack of equivalent time is by design rather than an unbalanced confound.
      • "In total, teachers in the intervention group spent about 35 hours engaged in PL activities across the school year."
      • Relevant Quotes: 1) "Teachers received four interactive days of PL (7 hours each day) and four follow-up coaching sessions (40 minutes each) and participated in an average of 11 weekly collaboration meetings (30 minutes each). In total, teachers in the intervention group spent about 35 hours engaged in PL activities across the school year." (Description of the professional learning programme) 2) "(1) What impact does the PL programme have on kindergarten and first-grade classroom teachers' use of evidence-based instructional strategies, collaboration, and use of a strengths-based approach in instruction as compared to control classroom teachers?" (Research questions) 3) No quote describes control-condition teachers receiving an equivalent 35 hours of alternative professional learning or coaching; the comparison throughout is intervention PL versus a "Condition = Control" reference group receiving no such programme. Detailed Analysis: Quotes 1 and 2 were checked against the PDF and are verbatim. Applying the updated Criterion B decision procedure: extra teacher time/resources are present (about 35 hours of PL across the year); those resources are integral to and are precisely the treatment variable under study, as the research questions frame the whole comparison as PL teachers "as compared to control classroom teachers" who receive no such programme. Per the decision tree, when the additional resource is itself the treatment being tested, the control group may remain "business as usual" without a matched resource, and the criterion is met by design rather than by resource balancing. Final: Criterion B is met because the additional professional learning time given to the intervention group is itself the treatment variable under study, making the untreated control group the appropriate comparison.
  • Level 3 Criteria

    • R

      Reproduced

      • All related supporting RCTs are by the same author team; no independent replication by a different research group was found in the paper or via internet search.
      • Relevant Quotes: 1) "Babinski, L. M., S. J. Amendum, S. E. Knotek, M. Sánchez, and P. Malone. 2018. 'Improving Young English Learners' Language and Literacy Skills Through Teacher Professional Development: A Randomized Controlled Trial.' American Educational Research Journal 55 (1): 117-143." (References) 2) "Babinski, L., S. J. Amendum, M. Carrig, S. E. Knotek, J. C. Mann, and M. Sánchez. 2024a. 'Professional Learning for ESL Teachers: A Randomized Controlled Trial to Examine the Impact on Instruction, Collaboration, and Cultural Wealth.' Education Sciences 14:690." (References) 3) "Results from this study are consistent with previous studies that found a significant impact of the PL programme on teachers' use of evidence-based instructional strategies (Babinski et al. 2018) and ESL teachers' use of literacy strategies (Babinski et al. 2024a, 2024b)." (Discussion, Teacher use of instructional strategies) Internet search findings: A CrossRef/citation search identified five papers citing this study (Nguyen 2026 on Vietnamese EFL teacher perceptions; Pennington et al. 2026 on multilingual literacy pedagogies with families; Neugebauer et al. 2026 on aligning pre/in-service teacher PD; Gutiérrez et al. 2026 on pandemic-era kindergarten oral language; and Vaughn et al. 2025 on multilingual learner agency). None of these describe an independent replication of this specific BELLA/PL programme by a different research team; they are topically related but distinct studies, some appearing to be conceptual or narrative pieces rather than RCT replications. Detailed Analysis: All three in-paper quotes were checked and are verbatim. The related prior and companion RCTs cited by the authors as consistent with the current findings are all authored by the same core research team (Babinski, Amendum, and colleagues), so they do not constitute independent replication. The internet search for external citing/related work likewise did not surface a peer-reviewed independent replication of this specific PL programme's design and results by a different research team. Final: Criterion R is not met because no independent replication by a different research team was identified, either in the paper's own references or through internet search of the surrounding literature.
    • A

      All-subject Exams

      • Only language and literacy outcomes were measured, with no assessment of other core subjects such as mathematics or science.
      • "Test results yield an RIT score, which is an overall language and literacy score based on four domains or goal areas: Foundational Skills, Language and Writing, Literature and Information, and Vocabulary Use and Functions."
      • Relevant Quotes: 1) "MAP Growth Reading K-2 ... was used to assess students' growth in language and literacy at three time points during the school year ..." (Measures) 2) "Test results yield an RIT score, which is an overall language and literacy score based on four domains or goal areas: Foundational Skills, Language and Writing, Literature and Information, and Vocabulary Use and Functions." (Measures) 3) No mention anywhere in the Measures or Results sections of assessing mathematics, science, or any other core subject beyond language/literacy. Detailed Analysis: Quotes 1 and 2 were checked against the PDF and are verbatim. The study's only academic outcome measure is MAP Growth Reading K-2, which, although broken into four literacy-related goal areas, covers a single subject domain (language and literacy). No standardised assessment of mathematics, science, or other core subjects is reported, and the paper offers no rationale for a specialised, single-subject focus of the kind allowed under the vocational/upper- secondary exception, since this is a general K-1 elementary programme. Final: Criterion A is not met because only language and literacy outcomes were assessed, with no coverage of other core subjects.
    • G

      Graduation Tracking

      • Outcome tracking ended at the close of the single intervention school year, with no evidence of follow-up toward graduation found in the paper or through internet search.
      • Relevant Quotes: 1) "MAP Growth Reading K-2 ... was used to assess students' growth in language and literacy at three time points during the school year, including beginning-of-year baseline, mid-year, and end-of-year." (Measures) 2) "There are several limitations that should be considered when interpreting the findings from this study. First, the sample size was relatively small." (Limitations) 3) No statement anywhere in the Results, Discussion, or Limitations sections describes any follow-up beyond the end of the single school year studied, nor reference to a planned or completed follow-up study tracking this cohort to graduation. Internet search findings: A search of citing papers and related grantee-submission/ERIC records for this grant (IES R305A180336) and this author team (e.g. ERIC ED676418, the MDPI 2024 companion paper, and the 2024 "Learning Professional" piece on teacher collaboration) did not surface any subsequent publication tracking this specific 106-student, 13-school cohort of kindergarten/first-grade MLs through later grades or graduation. The 2024 and 2025 companion papers by the same team examine instruction, collaboration, and cultural wealth outcomes within the same or an adjacent single-year trial window, not longer-term follow-up. Detailed Analysis: Outcome tracking in this study stopped at the end-of-year MAP assessment within the single school year of the intervention. The Limitations section discusses sample size and the value of larger future trials but does not mention any longer-term or graduation-tracking follow-up of the kindergarten and first-grade students in this cohort, and no companion follow-up publication tracking this specific cohort was found among the references or via internet search. Final: Criterion G is not met because no tracking beyond the single intervention school year is reported or referenced, and no follow-up publication was found through internet search.
    • P

      Pre-Registered

      • No reference to a pre-registered protocol, registry, or registration date is found anywhere in the paper or via internet search.
      • Relevant Quotes: 1) No quote referencing a trial registry (e.g. ClinicalTrials.gov, AEA RCT Registry, OSF Registries, IES RCT Registry), a registration ID, or a pre-registration date appears anywhere in the Method, Data-analytic strategy, Disclosure statement, or Funding sections of the paper. 2) "This work was supported by Institute of Education Sciences [grant number R305A180336]." (Funding) Internet search findings: A search of the IES grant/ERIC records for grant R305A180336 (Babinski, Duke University) and the associated ERIC record for this paper (ED676418) did not identify any trial registry entry or pre-registration ID for this study. Detailed Analysis: Quote 2 was checked against the PDF and is verbatim. The paper describes the IES grant that funded the study but contains no statement anywhere about a publicly pre-registered protocol, hypotheses, or analysis plan filed before data collection began. Internet search of the grant and article records did not locate a registry entry either. In the absence of any quoted reference to a registry platform or registration date, and no registry entry found via internet search, this criterion cannot be confirmed as met. Final: Criterion P is not met because no pre-registration reference or date is provided anywhere in the paper, and no registry entry was found through internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.