Level 1 Criteria
-
C Class-level RCT
- Randomisation was carried out at the individual student level within a single school rather than by class or school, and the intervention is a self-study mobile system, not one-to-one tutoring, so no exception applies.
- "A total of 598 EFL learners were recruited from a senior secondary school using the convenience sampling technique. They were randomly assigned into the experimental group (n = 278) or the control group (n = 320)." (p. 4)
- Relevant Quotes: 1) "A quasi-experimental design, involving an experimental group (n = 278) and a control group (n = 320), was adopted to examine its effectiveness on students' grammar learning." (p. 1, Abstract) 2) "A total of 598 EFL learners were recruited from a senior secondary school using the convenience sampling technique. They were randomly assigned into the experimental group (n = 278) or the control group (n = 320)." (p. 4) 3) "Although the participants were assigned into the conditions randomly, whether the two groups were equivalent in the use of SRL strategies had not been examined." (p. 10) Detailed Analysis: Criterion C requires randomisation of entire classes (or schools), unless the intervention is one-to-one personal teaching such as tutoring. The paper states that 598 individual students from one senior secondary school were "randomly assigned" to the experimental or control condition. There is no mention of classes or schools being the unit of randomisation; the unit is clearly the individual student. The intervention is a self-directed mobile learning system used by students as a supplement to their regular classroom instruction, not a personal tutoring or one-to-one teaching intervention, so the tutoring exception does not apply. Because treated and control students attend the same school (and plausibly the same classes), contamination between conditions cannot be ruled out. Note also that the authors themselves label the design "quasi-experimental" even though they describe random assignment of individuals. Criterion C is not met because randomisation was performed at the student level within a single school without a valid tutoring exception.
-
E Exam-based Assessment
- Outcomes were measured with 50-item multiple-choice grammar tests custom-designed by two EFL teachers for this study, not a widely recognised standardised exam.
- "Both the pre- and post- grammar tests comprised of 50 multiple choices questions, designed by two experienced EFL teachers, to examine the grammar structures that were introduced in the system, with each question worth two points." (p. 6)
- Relevant Quotes: 1) "Each of the pre- and posttests was a cumulative review of the grammatical structures that were taught in class as well as in the system." (p. 6) 2) "Both the pre- and post- grammar tests comprised of 50 multiple choices questions, designed by two experienced EFL teachers, to examine the grammar structures that were introduced in the system, with each question worth two points." (p. 6) 3) "The internal consistency of the tests was acceptable for both the pre- and post-tests (above.80)." (p. 6) Detailed Analysis: Criterion E requires that outcomes be measured with a standard, widely recognised standardised exam rather than an instrument specially designed for the study. Here the pre- and post-tests were explicitly "designed by two experienced EFL teachers" to examine the grammar structures introduced in the system under study. Although the content was aligned with the nationally sanctioned Senior High School English Curriculum, the tests themselves are researcher/teacher-made instruments constructed for this study and closely aligned to the intervention content, which is exactly the situation the criterion is designed to guard against. No named national or state standardised examination (e.g., the College Entrance Examination) was used as the outcome measure. Criterion E is not met because the outcome measure was a custom-made grammar test rather than a recognised standardised exam.
-
T Term Duration
- The intervention ran for 16 weeks (a full semester) from the pre-test at the beginning of the semester to the post-test after the treatment, which covers at least one academic term.
- "After that, the 16-week treatment started and the participants from the experimental group used the system to study English grammar for at least 15 min every day." (p. 7)
- Relevant Quotes: 1) "At the beginning of the semester, participants in both groups were invited to complete the pre-grammar test in classroom environments." (p. 6) 2) "After that, the 16-week treatment started and the participants from the experimental group used the system to study English grammar for at least 15 min every day." (p. 7) 3) "To measure the instructional effects of the system on students' performance in grammar texts, participants in both groups completed the post- grammar test after the treatment." (p. 7) 4) "The experimental group ... used the system to learn grammar, complete grammar exercise, and engage in for a semester, while the control group only use the system to submit weekly assignments." (p. 1, Abstract) Detailed Analysis: Criterion T requires the interval from intervention start to outcome measurement to be at least one full academic term (approximately 3-4 months). The intervention started shortly after the pre-test at the beginning of the semester, lasted 16 weeks (described as "a semester"), and the post-test was administered after the treatment ended. Sixteen weeks is approximately four months, which meets the definition of a full term/semester between intervention start and outcome measurement. Criterion T is met because outcomes were measured after a 16-week, semester-long intervention period, which is at least one full academic term after the intervention began.
-
D Documented Control Group
- The control group's size, gender proportion, year level, English-learning history, baseline test scores, and exact conditions (system access limited to weekly assignments) are clearly documented.
- "No statistical differences were found between the two groups concerning their average age, gender proportion, length of English learning (see Table 1)." (p. 4)
- Relevant Quotes: 1) "No statistical differences were found between the two groups concerning their average age, gender proportion, length of English learning (see Table 1)." (p. 4) 2) "TABLE 1 | Participant demographic information. ... Control group 320 62.2% [male] 10 [year level] 6.14 (0.78) [average length of English learning]" (p. 4) 3) "Participants in the control group, however, had limited access to the functions of the system. They could only login the system to complete grammar exercises as assignments and submit their answers. Differing from the exercises provided for the experimental group, which were selected and recommended based on students' learning history and performance, the exercises used in the control group were chosen based on the year level without considering their learning progress." (p. 7) 4) "It can be seen from Figure 8 that the experimental group scored higher than the control group in the post-test, while the two groups obtained similar scores in the pre-test." (p. 9) Detailed Analysis: Criterion D requires detailed documentation of the control group, including demographics, baseline performance, and the treatment it received. The paper reports the control group's size (n = 320), gender proportion (62.2% male), year level (Grade 10), and average length of English learning (6.14 years, SD 0.78) in Table 1, confirms baseline comparability on demographics, and reports baseline (pre-test) grammar scores for both groups. It also describes precisely what the control group did during the study: they attended the same English courses and could only log into the system to complete and submit weekly year-level grammar assignments, without the personalised recommendation, feedback, and e-portfolio functions. This is sufficient to assess comparability and the control condition. Criterion D is met because the control group's composition, baseline characteristics, and conditions are clearly documented.