Level 1 Criteria
-
C Class-level RCT
- Participants were randomly assigned as individual students drawn from mixed majors into six conditions, not as intact classes or schools, and this is not a one-to-one tutoring intervention.
- "They were randomly assigned into six different groups. Each group consisted of 25-30 students majoring in chemistry, communication engineering, biology, material engineering, medicine, etc." (p. 5)
- Relevant Quotes: 1) "One hundred and forty-four intermediate-level Chinese EFL learners participated in this study and were divided randomly into six groups." (Abstract) 2) "They were randomly assigned into six different groups. Each group consisted of 25-30 students majoring in chemistry, communication engineering, biology, material engineering, medicine, etc." (p. 5) 3) "This experiment was a 2 x 3 between-group design. Question and text types were independent variables... We randomly grouped the final 144 participants into one of the six conditions." (p. 6) 4) "During their English classroom learning sessions, their teachers gave them instructions on how to finish the reading tasks. After that, each participant received a set of corresponding materials..." (p. 7) Detailed Analysis: The unit of randomisation here is the individual student. Participants were recruited as a convenience sample spanning many different majors (chemistry, communication engineering, biology, material engineering, medicine, etc.) and were individually allocated to one of six experimental conditions. There is no description of whole intact classes or schools being assigned as a block to a single condition; rather, students appear to have been mixed across conditions even while attending "classroom learning sessions," raising the possibility that students from the same class ended up in different conditions (contamination risk). The intervention (embedded adjunct questions inserted into a reading passage) is a one-time group-administered experimental manipulation, not a personal one-to-one tutoring intervention, so the tutoring exception described in the ERCT standard does not apply. Criterion C is not met because randomisation occurred at the individual-student level rather than the class or school level, and no tutoring exception applies.
-
E Exam-based Assessment
- The outcome measures (ten custom multiple-choice questions per text and a written-recall/pausal-unit scoring scheme) were researcher-designed for this study, not a widely recognised standardised exam.
- "Both texts had ten multiple-choice questions written in English." (p. 7)
- Relevant Quotes: 1) "The materials used for the pre-test were taken from TEM4 (Test for English Majors-level 4), a large-scale standardized test in China that evaluates the language proficiency level of English majors in China. The pre-test consisted of four reading passages with a total of 30 multiple-choice questions on reading comprehension." (p. 6) 2) "Both texts had ten multiple-choice questions written in English." (p. 7) 3) "To measure participants' scores in the written recall task, we calculated the number of pausal units participants could recall from the reading texts... There were 33 pausal units in text 1 and 29 in text 2." (p. 7) 4) "The development of the embedded questions for text 2 was under discussion among a group of experienced language teachers in the university... However, the inserted questions for text 1 were adopted from previous researchers [5]." (p. 6) Detailed Analysis: TEM4 is used only as a screening/pre-test instrument to confirm participants' baseline reading proficiency, not as the outcome (post-intervention) assessment. The actual outcome measures used to evaluate the effect of the adjunct questions are (a) ten multiple-choice questions per text, and (b) a written-recall task scored by counting recalled "pausal units," both of which were developed or assembled specifically by the research team for this study (the embedded adjunct questions themselves were also either adapted from a prior study or developed in-house by university teachers, but these are the independent variable, not the outcome test). Neither the ten-item MC test nor the pausal-unit written-recall scoring is a widely recognised, standardised examination. Criterion E is not met because the primary outcome instruments are custom, researcher-assembled measures rather than standardised exams.
-
T Term Duration
- The entire experiment, from reading the passage to completing all assessment tasks, was administered in a single session and completed within 45 minutes, far short of one academic term.
- "Upon completion, we noted that all participants finished within 45 min." (p. 7)
- Relevant Quotes: 1) "The materials were arranged in the following order: (1) reading passage; (2) topic familiarity questionnaire; (3) written recall task; (4) multiple-choice questions." (p. 7) 2) "Participants were told not to read back during the experiment and were given enough time to complete all the tasks." (p. 7) 3) "Upon completion, we noted that all participants finished within 45 min." (p. 7) 4) "In the eighth week of the fall semester of 2023, a reading test was given to about 200 recruited participants..." (p. 7) (this refers to the separate screening/pre-test, not the main intervention.) Detailed Analysis: The intervention (reading a single passage with or without embedded adjunct questions) and all outcome measurements (written recall and multiple-choice test) occurred back-to-back within one 45-minute sitting. There is no indication of any delay between the "intervention" (reading with embedded questions) and the measurement of outcomes; they occur in the same session. This is vastly shorter than the minimum one-term (roughly 3-4 month) interval required by the standard. Criterion T is not met because intervention and outcome measurement occurred in the same 45-minute session.
-
D Documented Control Group
- The control ("no question") group's size, composition, and procedure are clearly documented, and baseline reading competence was confirmed to be equivalent across all six randomly assigned groups.
- "Three groups dealt with expository text, with 21 participants for no question condition... The number of participants who dealt with narrative text was 26... in the same condition order as the expository text." (p. 6)
- Relevant Quotes: 1) "After taking a reading comprehension test, only about 153 participants who showed no significant differences in their reading competence joined the formal study (F(5, 138) = 0.827; p = 0.533)." (p. 5) 2) "Three groups dealt with expository text, with 21 participants for no question condition, 22 for the what questions condition, and 22 for the why questions condition. The number of participants who dealt with narrative text was 26, 28, and 25, respectively, in the same condition order as the expository text." (p. 6) 3) "no embedded question groups (Groups 1 and 4), embedded what question groups (Groups 2 and 5), and embedded why question groups (Groups 3 and 6)." (p. 7) 4) Table 2 reports Mean, SD and N for the "Control" condition separately for both text types (multiple choice and written recall scores). Detailed Analysis: The paper documents the exact size of the no-question (control) condition for each text (n = 21 for the expository text, n = 26 for the narrative text) and describes precisely what this group did: they read the same passage as the other groups but without any embedded questions, then completed the same questionnaire, written-recall task, and multiple-choice test. Because all 144 participants were drawn from the same screened, demographically homogeneous pool (average age 19.5, 80% male/20% female, all intermediate CET4-level EFL learners) and were randomly assigned to conditions, and because a pre-test confirmed no baseline differences in reading competence across the six groups (F(5, 138) = 0.827, p = 0.533), the control group's characteristics and conditions are adequately documented for comparison. Criterion D is met because the control condition's size, procedure, and baseline equivalence are clearly documented.