Level 1 Criteria
-
C Class-level RCT
- Randomisation was done at the individual student level within one university, not at the class or school level, and the intervention was group classroom instruction, not one-to-one tutoring.
- "The participants were randomly divided into five groups: control group (CG; n = 23), syllabic perception instruction (SPe; n = 21), syllabic production instruction (SPr; n = 22), phonemic perception instruction (PPe; n = 24), and phonemic production instruction (PPr; n = 25)." (p. 5)
- Relevant Quotes: 1) "Recruitment efforts resulted in a total of 119 participants. However, four members either chose to drop out or were unable to attend one or more of the training sessions. These individuals' data were excluded from the final analyses, resulting in a final sample of N = 115." (p. 5) 2) "The participants were randomly divided into five groups: control group (CG; n = 23), syllabic perception instruction (SPe; n = 21), syllabic production instruction (SPr; n = 22), phonemic perception instruction (PPe; n = 24), and phonemic production instruction (PPr; n = 25)." (p. 5) 3) "Each experimental group received two treatment sessions of 30 minutes in duration from the lead researcher, with all groups following the same pattern of: explicit instruction lecture (10 minutes), teacher-led activities (10 minutes), pair work with a classmate (5 minutes) and finally worksheet completion (5 minutes)." (p. 7) Detailed Analysis: The unit of randomisation was the individual student: participants recruited at a single small university in rural Japan were "randomly divided into five groups." There is no statement that intact classes or schools were assigned to conditions. The ERCT C criterion requires randomisation of entire classes (or schools) to prevent contamination, unless the intervention is one-to-one tutoring or personal teaching. Here the treatment was delivered as group classroom instruction (lectures, teacher-led activities, pair work, class-wide drills and performances in front of the class), so the tutoring exception does not apply. Students from the same institution were assigned to different conditions, creating exactly the contamination risk the criterion is designed to avoid. Criterion C is not met because randomisation was conducted at the individual student level for a group-based classroom intervention.
-
E Exam-based Assessment
- Outcomes were measured with a custom researcher-designed elicitation instrument scored on a 9-point Likert scale by raters, not with any widely recognised standardised exam.
- "A PowerPoint slideshow of 30 slides was designed to test 10 English words three times each: in a free-response style question, a direct translation task from Japanese to English, and finally a read-aloud word list." (pp. 5-6)
- Relevant Quotes: 1) "A PowerPoint slideshow of 30 slides was designed to test 10 English words three times each: in a free-response style question, a direct translation task from Japanese to English, and finally a read-aloud word list." (pp. 5-6) 2) "Prior to the start of the experiment, a total of three piloting sessions were conducted by the lead researcher with the help of eight bilingual L1 Japanese individuals of similar age and ability as the target population." (p. 6) 3) "Each utterance of the target words was rated on a 9-point Likert scale, where only the ends of the scale were defined." (p. 10) 4) "The first author (a bilingual speaker of American English/Japanese, with a background in applied linguistics and over 20 years teaching experience in Japan) was the main expert coder." (p. 10) Detailed Analysis: The outcome measure was a bespoke three-part elicitation instrument (free response, translation, word list) created and piloted by the researchers specifically for this study, with pronunciation accuracy judged subjectively by expert raters on a 9-point Likert scale. This is precisely the situation the E criterion warns against: a custom test aligned to the taught content (the same 10 target words were used in the treatment sessions and the assessment). No national, state-wide, or otherwise widely recognised standardised exam (e.g., TOEIC, TOEFL, Eiken) was used to measure outcomes. Criterion E is not met because the study relied on a custom-made, researcher-designed assessment rather than a standardised exam.
-
T Term Duration
- The entire study spanned only four weeks from pretest to delayed posttest, far short of the one academic term (roughly 3-4 months) required.
- "Although the duration of the study was relatively short (a total of 4 weeks from pretest to delayed posttest)..." (p. 5)
- Relevant Quotes: 1) "Participants (N = 115) received two weeks of instruction on either segmental or suprasegmental features of English, using either a perception- or a production-based method, with progress assessed in a pre/post/delayed posttest study design." (p. 1, Abstract) 2) "Although the duration of the study was relatively short (a total of 4 weeks from pretest to delayed posttest), and normal classwork at the university does not consist of any PI, the establishment of the control group was necessary..." (p. 5) 3) "The same three-part testing instrument was used for the collection at all three time points: pretest, posttest (immediately after the second treatment), and delayed posttest (exactly two weeks after the second treatment)." (p. 7) 4) "As the treatments only lasted two weeks, it is likely that the duration was not long enough for learners to fully form phonetic representations and they were still relying heavily on memory." (p. 16) Detailed Analysis: The T criterion requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. Here the intervention consisted of two 30-minute sessions over two weeks, and the final (delayed) posttest was administered exactly two weeks after the second treatment, giving a total interval of only about four weeks from pretest to final measurement. The authors themselves acknowledge the short duration as a limitation. Four weeks is well below the minimum term-long follow-up window. Criterion T is not met because the interval from intervention start to final outcome measurement was only about four weeks, far less than one academic term.
-
D Documented Control Group
- The control group is clearly documented, including its size, baseline scores with confidence intervals, and confirmation that it received only normal coursework with no PI.
- "The control group received no treatment for the duration of the study, apart from their normal coursework at the university, and were only asked to convene for the pre- and posttests." (p. 5)
- Relevant Quotes: 1) "The participants were randomly divided into five groups: control group (CG; n = 23)..." (p. 5) 2) "The control group received no treatment for the duration of the study, apart from their normal coursework at the university, and were only asked to convene for the pre- and posttests." (p. 5) 3) "...normal classwork at the university does not consist of any PI, the establishment of the control group was necessary to determine if any gains demonstrated by the experimental groups could be attributed solely to the treatments, or if there were other factors (e.g., test practice effects) that needed to be considered." (p. 5) 4) "CG 3.92 (.34) [3.77, 4.06] 4.05 (.46) [3.86, 4.25] 4.07 (.45) [3.87, 4.26]" (Table 3, p. 12) 5) "This study was conducted in an EFL setting at a small university in rural Japan... the students' oral proficiency level could be described as low-intermediate." (p. 5) Detailed Analysis: The D criterion requires the control group to be documented in terms of composition, baseline performance, and treatment received. The paper reports the control group's size (n = 23), explicitly states what it received (no treatment other than normal university coursework, which contains no pronunciation instruction), and provides its baseline pretest mean, standard deviation, and confidence interval in Table 3, alongside posttest and delayed posttest scores and gain scores (Table 4). The population from which all groups (including controls) were drawn is described (Japanese tertiary EFL students, low-intermediate oral proficiency, aged 18-20 per the plain language summary). Baseline comparability can be checked directly from the reported pretest scores. While demographic breakdown per group is limited, the documentation matches the level accepted under this criterion for a well-described no-treatment baseline group. Criterion D is met because the control group's size, baseline performance, and conditions are clearly documented.