Level 1 Criteria
-
C Class-level RCT
- Two entire intact classes, not individual students within one class, were randomly assigned to the control and experimental conditions.
- "Two intact classes (n = 50) of grade 9 learners of English as a foreign language were randomly assigned to control and experimental conditions." (p. 74)
- Relevant Quotes: 1) "The study employed a pre-test - post-test control group design whereby two intact classes were randomly assigned to control and experimental conditions." (p. 69) 2) "Two intact classes (n = 50) of grade 9 learners of English as a foreign language were randomly assigned to control and experimental conditions." (p. 74) 3) "A sample of 50 EFL learners enrolled in two sections of grade 9 was randomly assigned to control and experimental conditions." (p. 75) Detailed Analysis: The unit of randomisation here is the intact class/section, not the individual student. The authors explicitly state that two whole sections (classes) of grade 9 were assigned as units to the control and experimental conditions, rather than splitting students within a single classroom into two groups. This design avoids the specific contamination problem the criterion is meant to guard against (students in the same room receiving different conditions), because each class as a whole received only one condition. This is not a school-level RCT (only one school and two sections were involved, so the stronger S criterion is not met), but it clearly satisfies the weaker class-level requirement. Criterion C is met because randomisation occurred at the level of two whole intact classes rather than within a single class.
-
E Exam-based Assessment
- The reading comprehension tests were adapted from a Paideia lesson-plan website rather than a standardised, validated exam, and no standardised exam was used for the anxiety outcome either.
- "Specifically, we used part one of the test given by Paideia Active Learning" ... on Shakespeare's 'The Seven Ages of Man' as a pre-test and a part of the "Mending Wall" reading questions retrieved from Paideia Active Learning as well ... as the post tests." (p. 74)
- Relevant Quotes: 1) "Two reading achievement tests were used as pre-test and post-test measures of reading comprehension. Specifically, we used part one of the test given by Paideia Active Learning" (https://www.paideia.org/lesson-plans/seven-ages- of-man/) on Shakespeare's 'The Seven Ages of Man' as a pre-test and a part of the "Mending Wall" reading questions retrieved from Paideia Active Learning as well (https://www.paideia.org/lesson- plans/mending-wall/) as the post tests (See Appendix C)." (p. 74) 2) "Both tests consisted of response and comprehension question items and assessed the main idea, literal, analysis, synthesis, inference, application, and creative comprehension." (p. 74) 3) "In addition, the Reading Anxiety Scale (Young, 1999) was used to assess the readers' levels of reading anxiety prior to and after the treatment. This scale consists of 15 Likert-type items, 7 of which are negatively stated to avoid a response set." (p. 74) 4) Appendix C reproduces both tests, which consist of open-ended discussion-style comprehension questions (e.g., "What images comes to mind when you read about the different 'ages of man'...?") taken directly from downloadable lesson-plan question sets on the paideia.org website, not from any national, state, or otherwise psychometrically validated standardised examination. (p. 85-89) Detailed Analysis: The ERCT E criterion requires a widely recognised, standardised exam-based assessment rather than a study-specific or informally sourced instrument. Here the reading achievement pre-test and post-test are lesson-plan comprehension question sets downloaded from the Paideia Active Learning website, tied to two specific poems ("The Seven Ages of Man" and "Mending Wall"). These are teacher resource materials designed to accompany lesson plans, not a standardised, psychometrically validated exam with established norms, reliability, or validity evidence. The paper reports no information on the instrument's standardisation, norming population, or psychometric properties beyond noting the general comprehension categories assessed. The Reading Anxiety Scale (Young, 1999) is a published affective scale, but it measures anxiety, not an academic exam outcome, and the paper itself reports only "moderately high" internal consistency (alpha = .611) for this specific sample, which is not evidence of a widely validated standardised instrument's use here either. Because the core reading comprehension outcome was measured with a non-standardised, lesson-plan- derived instrument, Criterion E is not satisfied.
-
T Term Duration
- The intervention and outcome measurement together spanned only six weeks, far short of a full academic term.
- "All the participants received the treatment for a period of 6 weeks at the rate of 5 hours per week." (p. 76)
- Relevant Quotes: 1) "All the participants received the treatment for a period of 6 weeks at the rate of 5 hours per week." (p. 76) 2) "Each seminar was conducted during 60 minutes for the seminar itself and 20 minutes for group evaluation at the end." (p. 77) 3) "The ANCOVA results of the analysis of the experimental versus control group post-test reading achievement scores, after having controlled for pre-test scores existing differences, were found to be statistically significant..." (p. 77) — the post-test was administered at the end of the same 6-week treatment period described above, with no later follow-up measurement reported. Detailed Analysis: The entire intervention, from pre-test through treatment to post-test, took place over six weeks at five hours per week. There is no indication that outcomes were measured any later than immediately following the conclusion of the six-week treatment period. A full academic term is typically defined as roughly 3-4 months; six weeks is well under half of that minimum, and no evidence in the paper suggests a longer interval between intervention start and outcome measurement. Because the study duration falls well short of the one-term minimum required, and because the study is also far short of a full year (so the stronger Y criterion cannot be met either), Criterion T is not met.
-
D Documented Control Group
- The control group's regular-instruction procedures, activities, and shared demographic/baseline characteristics with the experimental group are described in reasonable detail.
- "Instruction in the control group consisted of regular reading comprehension practice based on the pedagogical implications of the interactive theory of reading which requires focus on the different steps of the reading comprehension process." (p. 76)
- Relevant Quotes: 1) "Meanwhile, participants in the control group read the same texts according the procedures of regular reading comprehension instruction." (p. 76) 2) "Instruction in the control group consisted of regular reading comprehension practice based on the pedagogical implications of the interactive theory of reading which requires focus on the different steps of the reading comprehension process. Specifically, instruction in the control group consisted or pre-reading, during reading, post reading stages in which a range of activities were used in order to activate readers' background knowledge, build vocabulary, check comprehension, and reflect on what is read. Examples of the activities used in the control group include brainstorming based on titles and illustrations; vocabulary learning strategies such as structural analysis, guessing meaning from context, and dictionary; literal and higher-order comprehension checks; and reflection." (p. 76) 3) "As such, the experimental and the control group included EFL learners from economically underprivileged families and were all native speakers of Arabic. Thirty (n = 37) participants were Lebanese with limited English language proficiency and 13 (n = 13) participants are Syrian refugees with comparable limited-English proficiency ... The age of the participants ranged from 14-16 years." (p. 75) 4) "Participants' pre-test scores on the dependent variables under investigation were as covariates in order to mathematically adjust for potential pre-existing difference between the control and the experimental group." (p. 74) Detailed Analysis: The paper describes, in specific procedural detail, what the control group actually experienced (pre-reading/during-reading/post-reading activities such as brainstorming, vocabulary strategies, and comprehension checks), rather than a vague "business as usual" label. It also documents the shared demographic profile (nationality mix, socio-economic background, native language, age range) of the overall sample from which both intact classes were drawn, and it uses pre-test scores as covariates specifically to characterise and adjust for baseline differences between the control and experimental groups. While the paper does not give a separate numeric breakdown of the control group's size or demographics apart from the experimental group, the combination of a detailed activity description and baseline-score covariate adjustment provides adequate documentation to assess comparability. Criterion D is met because the control condition's instructional activities and the sample's baseline characteristics are clearly and specifically documented.