Level 1 Criteria
-
C Class-level RCT
- Although the four intermediate schools were initially randomly assigned, the comparison group's final analytic sample was augmented by teachers recruited non-randomly, leading the authors themselves to classify the overall design as quasi-experimental rather than a properly implemented RCT.
- "Because of the non-random addition of these teachers, this design of this overall project was considered to be quasi-experimental." (p. 994)
- Relevant Quotes: 1) "The state law (Texas Education Code, 1995) prohibits random selection and assignment on the basis of individual students for program placement; therefore, in the larger research project, four intermediate schools with principals' approval from the district site were randomly assigned to conditions, resulting in two treatment (enhanced science practice) and two comparison (typical science practice) schools." (p. 994) 2) "When a school was assigned, teachers from that campus were then randomly selected to the assigned condition within that campus, and both ELLs and low SES non-ELLs in the same school received the same practice to allay contamination between experimental and comparison classrooms." (p. 994) 3) "Due to the low return response in the two comparison schools, a balanced design with four science teachers in their respective condition (i.e., two in treatment and two in comparison) could not be achieved. Therefore, in order to increase the sample size at student level, we recruited four more teachers within the same comparison schools that were already assigned based on the criteria that they had Spanish-speaking ELLs in their classrooms." (p. 994) 4) "Because of the non-random addition of these teachers, this design of this overall project was considered to be quasi-experimental." (p. 994) Detailed Analysis: On its face, the four intermediate schools were randomly assigned to condition, which would normally satisfy (and exceed) the class-level randomisation requirement. However, the actual analytic sample was not the product of that randomisation alone: because too few teachers in the two comparison schools volunteered, the researchers non-randomly recruited additional comparison-condition teachers/classrooms to inflate the student sample size. The authors explicitly and repeatedly describe the resulting design, in the abstract, methods, and discussion/limitation sections, as "quasi-experimental" precisely because of this non-random addition. There is no tutoring/personal teaching exception here (the intervention is a whole-classroom science curriculum, not one-to-one instruction). Because randomisation was not properly and cleanly implemented for the sample that was actually analysed, the criterion cannot be considered met. Criterion C is not met because the study's own authors classify the design as quasi-experimental due to non-random recruitment of additional comparison-group teachers/classrooms.
-
E Exam-based Assessment
- In addition to district-developed benchmark tests, the study used the state-mandated standardized TAKS assessment and the nationally standardized DIBELS Oral Reading Fluency measure, both widely recognized standardized instruments.
- "The TAKS, a criterion-referenced assessment, measures student mastery of the content areas of state curriculum outlined in the TEKS." (p. 999)
- Relevant Quotes: 1) "State Standardized Test: Texas Assessment of Knowledge and Skills (TAKS). The TAKS, a criterion-referenced assessment, measures student mastery of the content areas of state curriculum outlined in the TEKS." (p. 999) 2) "The English form has a reported internal consistency ranging from 0.87 to 0.90, and predictive validity ranging from 0.56 to 0.79 with SAT and ACT (TEA, 2008)." (p. 1000) 3) "Dynamic Indicators of Basic Literacy Skills (DIBELS). DIBELS (Good and Kaminiski, 2002) includes a set of procedures and measures for assessing the acquisition of early literacy skills... ORF is reported to have a median alternate form reliability of 0.95..." (p. 1000) 4) "These benchmark tests were developed according to the scope and sequence of the fifth grade Texas Essential Knowledge and Skills (TEKS)... The information is limited about the reliability and validity of the benchmark tests at the district level..." (pp. 998-999) Detailed Analysis: The study used three tiers of assessment: (a) district-developed benchmark tests (custom-built and not standardized statewide), (b) the state-mandated TAKS, and (c) the nationally-used DIBELS Oral Reading Fluency subtest. Both TAKS and DIBELS are widely recognised, standardized instruments used far beyond this single study (TAKS statewide in Texas for Grades 3-12; DIBELS nationally), with documented psychometric properties reported independently of the authors' own study. Although the district benchmark tests are researcher/district-created and would not alone satisfy this criterion, the presence of TAKS and DIBELS as core outcome measures satisfies the requirement for a standardised, widely recognised exam-based assessment. Criterion E is met because the study's outcome measures include the state-standardized TAKS assessment and the nationally standardized DIBELS ORF subtest.
-
T Term Duration
- The intervention and outcome measurement spanned the full 2009-2010 school year, with final outcomes measured via TAKS and DIBELS in spring 2010, far exceeding one academic term.
- "Data were collected in the fall and spring of school year 2009-2010." (p. 1001)
- Relevant Quotes: 1) "Data were collected in the fall and spring of school year 2009-2010. Science benchmark test was administered to each student every 6 weeks during the school year with a total of six tests." (p. 1001) 2) "TAKS data were collected during the spring of 2010." (p. 1001) 3) "In this study, Oral Reading Fluency (ORF) was administered to students in the beginning and end of fifth grade." (p. 1000) Detailed Analysis: The intervention began in the fall of the 2009-2010 school year and the primary outcome measures (TAKS, end-of-year DIBELS ORF) were collected in spring 2010, after a full academic year. This interval is far longer than the minimum one-term requirement. Criterion T is met because outcomes were measured after a full school year, well beyond one academic term.
-
D Documented Control Group
- The comparison group's schools, teachers, and students are documented in detail, including demographics, teaching experience, class sizes, and the typical-practice instruction they received.
- "Table 1. School demographics for treatment and comparison group 2009-2010" (p. 995)
- Relevant Quotes: 1) "Table 1 demonstrates the demographics of these four schools." (p. 994), showing African American, Hispanic, White, Native American, Asian, Low SES, ELL percentages, and academic rating for each comparison school (C and D). (p. 995) 2) "Table 2 ... Comparison C 2 [teachers] 6 [rotations] 118 [students]; D 2 3 48." (p. 995) 3) "In typical practice in the comparison group, science is taught by certified or permitted bilingual/ESL education and science education teachers in English for the ELL students with no Spanish clarifications... science instruction in comparison classrooms varied from 80 to 90 minutes daily including one 5-E lesson cycle weekly... Teachers followed a locally developed science curriculum aligned to the TEKS." (p. 998) 4) "Although no training support was provided by the research team to the teachers in the comparison group, they attended workshops as a state requirement to fulfill a minimum of 30 professional development hours each year related to their content area. Typical practice also had computers, projectors, and document cameras (ELMOs) in the science classrooms." (p. 998) Detailed Analysis: The paper documents the comparison ("typical practice") condition thoroughly: school-level demographic tables, teacher/student counts by rotation, average teaching experience (8.4 years), and a detailed narrative of what typical-practice instruction looked like (curriculum, lesson length, classroom observations, existing technology, state-required professional development hours). This level of detail satisfies the documentation requirement. Criterion D is met because the comparison group's composition, baseline characteristics, and instructional conditions are clearly documented.