Level 1 Criteria
-
C Class-level RCT
- Randomisation was performed at the individual student level within a single secondary school, not at the class or school level, and no tutoring exception applies.
- "The students were assigned to either the experimental or the control group upon a random basis." (p. 55)
- Relevant Quotes: 1) "The study adopts a pre-post control group design where forty students were randomly assigned to either a control or an experimental group." (Abstract, p. 50) 2) "The sample of the study consisted of forty male students in the first-year from a secondary school in Taif City, Saudi Arabia (n=40). The students were assigned to either the experimental or the control group upon a random basis." (p. 55) Detailed Analysis: The paper explicitly states that individual students (n=40) from a single secondary school were randomly assigned to either the experimental or control condition. There is no mention of randomising by class or by school; instead, forty students within the same first-year cohort were split into two groups of twenty on a random basis. This raises the risk of contamination between the experimental and control students, as they attend the same school and could interact. The intervention (teaching a reading strategy) is not a personal one-to-one tutoring intervention, so the "personal teaching" exception described in the ERCT standard does not apply here. Because randomisation occurred at the student level within one school rather than at the class or school level, and no valid exception applies, criterion C is not met.
-
E Exam-based Assessment
- The study used a custom-designed Reading Comprehension Skills Test and Attitude Scale created and validated by the researcher, not a standardised, widely recognised exam.
- "To achieve that end, the reading texts as well as the Reading Comprehension Skills Test were submitted to the judgment of a group of eleven jury members who agreed on their validity and suitability." (p. 55)
- Relevant Quotes: 1) "A quantitative study with a quasi-experimental design was implemented through applying two different instruments: Reading Comprehension Skills Test and Attitude Scale towards learning EFL." (Abstract, p. 50) 2) "The study was pretested with a comparison group of 15 male students to determine the reliability of the reading comprehension study. Then, after analyzing the data obtained, three items were omitted from the entire study, and four more items were updated." (p. 55) 3) "To achieve that end, the reading texts as well as the Reading Comprehension Skills Test were submitted to the judgment of a group of eleven jury members who agreed on their validity and suitability. Four jury members suggested some modifications and accordingly modifications have been made." (p. 55) 4) "To assess the empirical validity of the reading comprehension test, Pearson Correlation between the reading comprehension tests (authentic and inauthentic) and TOEFL were measured as (0.68) and (0.63) respectively." (p. 55) Detailed Analysis: The reading comprehension and attitude instruments used in this study were developed specifically by the researcher for this study, piloted with a comparison group, revised based on item analysis, and reviewed by a jury of eleven experts. Although the researcher correlated the custom test with TOEFL scores to argue for its validity, this correlational check does not make the instrument itself a standardised, widely recognised exam; it remains a researcher-designed measure tailored to the intervention's specific reading texts. Because the assessment instruments were custom-built for this study rather than being standard, widely recognised examinations, criterion E is not met.
-
T Term Duration
- Outcomes were measured immediately after a short, twelve-session intervention lasting only a few weeks, far short of a full academic term.
- "The experimental groups received instruction using metacognitive Think-Aloud strategy for twelve sessions with forty minutes each." (p. 55)
- Relevant Quotes: 1) "Finally, the experimental groups received instruction using metacognitive Think-Aloud strategy for twelve sessions with forty minutes each. The teaching sessions were 12 sessions including 12 reading texts." (p. 55) 2) "The analysis of data using t-test showed that the experimental group achieved significantly higher scores than the control group on the post-performance of the test of reading comprehension skills..." (p. 56) Detailed Analysis: The intervention consisted of only twelve 40-minute sessions (approximately eight hours of instruction in total), and outcomes were measured via a post-test administered immediately following completion of these sessions. There is no indication that measurement occurred a full academic term (roughly 3-4 months) after the intervention began; instead, the entire intervention-to-measurement window appears to span only a few weeks. Because the interval from intervention start to outcome measurement is far shorter than one academic term, criterion T is not met.
-
D Documented Control Group
- The control group's size, baseline test scores, and the "business as usual" traditional treatment it received are clearly documented and compared with the experimental group.
- "Students of the experimental group were instructed by using metacognitive Think-Aloud strategy, whereas, the control group received traditional treatment such as skimming and scanning techniques." (Abstract, p. 50)
- Relevant Quotes: 1) "Students of the experimental group were instructed by using metacognitive Think-Aloud strategy, whereas, the control group received traditional treatment such as skimming and scanning techniques." (Abstract, p. 50) 2) "Table 1. Means, standard deviations, t-value of means and significance of differences of the two groups in the pre-performance on reading comprehension skills test... Cont. 20 31.82 8.83... 38 0.04." (p. 55) 3) "Table 2. Means, standard deviations, t-value of means and significance of differences of the two groups in the pre-performance on the Attitude Scale... Cont. 20 30.01 7.41 38 0.48." (p. 55) 4) "Pre-test data on the reading comprehension skills test showed group equivalence as the t-value (0.04) was insignificant at p ≤ .05 level as in table (1)." (p. 55) Detailed Analysis: The paper documents the control group's size (n=20), provides its pre-test means and standard deviations for both the reading comprehension test and the attitude scale, and confirms baseline equivalence with the experimental group via non-significant t-tests. It also explicitly states what treatment the control group received during the study (traditional skimming and scanning instruction) rather than leaving it undefined. This level of detail is sufficient to assess whether the control group was comparable at baseline and to understand what activities it engaged in. Because the control group's baseline characteristics, size, and treatment are clearly documented, criterion D is met.