Level 1 Criteria
-
C Class-level RCT
- Randomisation was conducted at the school level, which is stronger than and therefore satisfies the class-level RCT requirement.
- "Schools were blocked by district and then assigned to treatment or control group within district using a permuted blocks randomisation procedure."
- Relevant Quotes: 1) "Randomisation procedures for condition were conducted after all teacher consent forms were collected. Schools were blocked by district and then assigned to treatment or control group within district using a permuted blocks randomisation procedure." (Method, Sampling and randomisation procedures) 2) "Thirty-nine classroom teachers (18 kindergarten and 21 first grade) from 13 schools enrolled in the RCT." (Participants and settings) 3) "ESL teachers participated alongside either kindergarten or first-grade teachers from the same school." (Description of the professional learning programme, PL process) Detailed Analysis: All three quotes were checked against the PDF and are verbatim. The paper explicitly states that the unit of randomisation was the school ("Schools were ... assigned to treatment or control group"), which is consistent with the design in which ESL and classroom teachers from the same school worked together as a school-based team. This design choice avoids within-school contamination, since an entire school's participating teachers are placed in the same condition. The ERCT Standard treats a stronger school-level randomisation as automatically satisfying the weaker class-level criterion. Final: Criterion C is met because randomisation occurred at the school level, which is stronger than and therefore satisfies the class-level requirement.
-
E Exam-based Assessment
- The study used MAP Growth Reading K-2, a nationally normed standardised assessment, satisfying the exam-based assessment requirement.
- "MAP Growth Reading K-2 ... is a nationally normed achievement test that is designed to measure the general academic literacy achievement and growth of K-2 students ..."
- Relevant Quotes: 1) "MAP Growth Reading K-2 (Northwest Evaluation Association 2019) was used to assess students' growth in language and literacy at three time points during the school year ... MAP Growth Reading K-2 is a nationally normed achievement test that is designed to measure the general academic literacy achievement and growth of K-2 students ..." (Measures, Student language and literacy skills) 2) "Test-retest correlations for MAP Growth Reading K-2 RIT scores range from .82 to .86 and reliability coefficients range from .95 to .97 for kindergarten through second grade students (Northwest Evaluation Association 2011)." (Measures) Detailed Analysis: Both quotes were checked against the PDF and are verbatim. The primary student outcome measure, MAP Growth Reading K-2, is described as a nationally normed achievement test with documented reliability statistics, not an instrument custom-built for this study. This satisfies the requirement for a standardised, widely recognised exam-based assessment. Final: Criterion E is met because the study used MAP Growth Reading K-2, a nationally normed, standardised achievement test.
-
T Term Duration
- Outcomes were measured across a full academic year, which exceeds and therefore satisfies the one-term duration requirement.
- "MAP Growth Reading K-2 ... was used to assess students' growth in language and literacy at three time points during the school year, including beginning-of-year baseline, mid-year, and end-of-year."
- Relevant Quotes: 1) "Using a randomised control trial, we evaluated the impact of a year-long professional learning (PL) programme ..." (Abstract) 2) "MAP Growth Reading K-2 ... was used to assess students' growth in language and literacy at three time points during the school year, including beginning-of-year baseline, mid-year, and end-of-year." (Measures) Detailed Analysis: Both quotes were checked against the PDF and are verbatim. The intervention and its outcome tracking spanned an entire academic year, from a beginning-of-year baseline through an end-of-year assessment. Since the ERCT Standard specifies that a year-long duration automatically satisfies the weaker term-duration requirement, and this interval far exceeds one academic term, this criterion is met. Final: Criterion T is met because outcomes were tracked across a full academic year, well beyond the minimum one-term requirement.
-
D Documented Control Group
- The paper reports pre/post outcome means by condition but does not provide a description of demographic characteristics broken down by condition, nor an explicit statement of what the control group received or did instead of the PL.
- "Descriptive statistics, including mean and standard deviation, for pre and post intervention teacher outcomes are provided in Table 3." (p. 2452)
- Relevant Quotes: 1) "See Table 2 for teacher and student demographic information." (p. 2450) 2) "Descriptive statistics, including mean and standard deviation, for pre and post intervention teacher outcomes are provided in Table 3." (p. 2452) 3) Table 3 header: "Teacher Outcome Variable ... Control Mean (SD) ... Intervention Mean (SD)" with rows including "Instructional Strategies - Pre 1.95 (1.64)" for the control group. (p. 2452) 4) "Condition = Intervention" is coded as a dummy/effects-coded predictor throughout the analytic models (e.g. p. 2453-2456), but no narrative text describes what activity, business- as-usual instruction, or alternative professional development (if any) the control-condition teachers and students experienced during the school year. Detailed Analysis: Criterion D requires clear documentation of the control group's demographic characteristics, baseline performance, and the conditions/ treatment it received. This paper provides control-group baseline and post scores for the teacher-outcome variables in Table 3 (partially satisfying the "baseline performance" requirement), but Table 2, which presents teacher and student demographic characteristics (gender, experience, degree, race/ethnicity, home language, country of birth, etc.), reports only aggregate figures for the full sample (n = 39 teachers, n = 106 students) and is not broken down by condition, and no counts of teachers or students per condition are reported anywhere in the paper. There is therefore no way to verify from the text whether the control and intervention groups were demographically comparable, or even how large each group was. Moreover, nowhere in the Method section is there an explicit statement describing what, if anything, control-condition teachers did instead of the BELLA PL programme (e.g. "business as usual" instruction, a different/lighter PL offering, or no PL at all). The absence of this description means readers cannot confirm that the control condition was a stable, well-specified comparison condition as opposed to an unknown or variable set of practices. Because the demographic breakdown by condition, the size of the control group, and an explicit description of the control condition's treatment are all missing, the documentation of the control group falls short of what criterion D requires, despite the partial information available via Table 3. Criterion D is not met because the paper lacks a condition-level demographic and size breakdown and an explicit description of the control group's treatment/business-as-usual condition.