Abstract
In English as a Foreign Language (EFL) education, addressing intelligibility presents unique challenges, necessitating exploration of interventions aimed at raising phonological awareness. This practical study investigated the effectiveness of instruction on syllable perception among Japanese university students in English pronunciation, with the goal of improving English intelligibility. The study employed different instructional methods to two groups of university freshmen. One group, the Phonological Instruction (PI) group, received instruction that explicitly used linguistic terms such as phoneme and syllable, along with phonemic transcriptions as their representations. The other group, referred to as the non-PI group, was taught without the use of such terminology or phonemic symbols. A total of 38 Japanese EFL students participated in the study. Both groups received 20 minutes of instruction per week over seven weeks in their first semester. They counted syllables in nonwords before and after instruction. A generalized linear mixed model was conducted to examine the effects of the two types of instruction. Despite the absence of significant effects observed between pre- and post-interventions regardless of the instruction type, this study may be considered an innovative endeavor to address challenges in Japanese university English education, particularly in the domain of pronunciation instruction.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was carried out at the level of two intact classes (an economics seminar and a science seminar), not among individual students within a single class.
- "The conditions of the educational intervention were randomly assigned to the two classes: the economics class received instruction using the non-PI method (non-PI group), while the science class received instruction with the PI method (PI group)." (p. 244)
Relevant Quotes:
1) "The conditions of the educational intervention were randomly assigned to the two classes: the economics class received instruction using the non-PI method (non-PI group), while the science class received instruction with the PI method (PI group)." (p. 244)
2) "This study involved 43 first-year students newly enrolled at a Japanese national university during the first semester. They belonged to two separate freshman start-up seminars: one with 18 economics students and the other with 25 students majoring in symbiotic science." (p. 244)
Detailed Analysis:
The paper explicitly states that the intervention condition (non-PI vs. PI) was assigned to whole intact classes (freshman seminars), not to individual students within one shared class. This is precisely the unit of randomisation the ERCT "C" criterion requires, since it prevents within-class contamination between conditions. The design only involves two classes in total (one class per condition), which is a methodologically weak instantiation of class-level randomisation (no replication of classes within a condition), but the criterion itself concerns the unit of randomisation rather than the number of clusters per arm.
Final sentence: Criterion C is met because the paper clearly documents that whole classes, not individual students within a class, were assigned to each instructional condition.
-
E
Exam-based Assessment
- The outcome measure was a researcher-adapted nonword syllable-counting task, not a widely recognised standardised exam; the one genuinely standardised instrument used (TOEIC Speaking) served only as a baseline descriptive, not as the pre/post outcome.
- "The present study, based on the procedures outlined by He and Sakuma (2022), employed a similar intervention approach distributed across seven spaced sessions, replacing the original repetition task with a syllable counting task." (p. 245)
Relevant Quotes:
1) "The task of identifying the number of syllables in a nonword was evaluated based on whether the syllable count was correct or not. Scoring was binary, with no provision for partial scores." (p. 251)
2) "The CNRep comprises 40 nonwords, encompassing two, three, four, and five syllables, with ten in each category." (p. 245)
3) "The present study, based on the procedures outlined by He and Sakuma (2022), employed a similar intervention approach distributed across seven spaced sessions, replacing the original repetition task with a syllable counting task." (p. 245)
4) "A total of 13 students of the non-PI group and 20 students of the PI group followed the recommendation and completed the TOEIC speaking test to assess their general speaking proficiency... While the score of the non-PI group seems relatively high, the difference of the scores was not significant, t(31) = 1.77, p = .086)." (p. 244)
Detailed Analysis:
The primary dependent variable analysed via the GLMM (Table 5) is performance on a syllable-counting task using CNRep nonwords, itself an adaptation of the original CNRep repetition task customised specifically for this study's purposes. While CNRep has been used across several prior studies of phonological awareness, it is a laboratory phonological-memory research instrument rather than a widely recognised, standardised educational exam (e.g., a national curriculum test). The one genuinely standardised, externally administered instrument in the study, the official TOEIC Speaking Test, was used only once, as a background/baseline descriptive comparison between the two intact groups (Table 1), and was explicitly not significant between groups; it was not used as the pre-/post-intervention outcome measuring instructional effect. Because the criterion requires that the outcome measuring the intervention's effect itself be a standardised exam, and here that measure is a custom-adapted nonword task, this criterion is not satisfied.
Final sentence: Criterion E is not met because the pre/post outcome measure was a study-specific adaptation of a phonological research task rather than a standardised exam-based assessment.
-
T
Term Duration
- The intervention and its outcome measurement together spanned only about seven weeks, well short of one full academic term, with the post-test given immediately at the end of the final session.
- "Both groups received 20 minutes of instruction per week over seven weeks in their first semester." (p. 237)
Relevant Quotes:
1) "Both groups received 20 minutes of instruction per week over seven weeks in their first semester." (p. 237)
2) "The two participating groups of students each received a distinct type of intervention: either the non-PI or the PI method. Both groups engaged in 20-minute sessions per week over the course of seven weeks (see 3.3.2 for detail)." (p. 247)
3) "After the intervention, during the last four minutes of the wrap-up session, which involved a reflection on the highlights from the previous seven session, the post-test was administered." (p. 247)
Detailed Analysis:
The intervention consisted of seven weekly 20-minute sessions, and the outcome (post-test) was collected immediately at the end of the seventh week, in the wrap-up session, with no additional delayed follow-up measurement. Seven weeks is substantially shorter than a full academic term (typically 3-4 months, i.e., roughly 12-16 weeks), and there is no term-long tracking of outcomes past the intervention's end.
Final sentence: Criterion T is not met because the interval from intervention start to outcome measurement was only about seven weeks, short of one academic term.
-
D
Documented Control Group
- Although the design compares two active instructional arms rather than an untreated control, both arms are thoroughly documented in terms of size, exclusions, prior English background, and baseline TOEIC scores.
- "Therefore, the results of 17 non-PI group students and 21 PI group students were analyzed in the present study." (p. 244)
Relevant Quotes:
1) "One economics student and three science students were excluded from the analysis because they were absent from either introductory or wrap-up session. In addition, one Chinese student from the science class was excluded... Therefore, the results of 17 non-PI group students and 21 PI group students were analyzed in the present study." (p. 244)
2) "None of the students had lived overseas before, and all had received six years of English education in secondary school." (p. 244)
3) Table 1, "Descriptive Statistics of the TOEIC Speaking Test": Non-PI n = 13, M = 81.54, SD = 12.14; PI n = 20, M = 69.50, SD = 22.35 (p. 245)
Detailed Analysis:
This study does not include a genuine "business as usual" no-treatment control; instead, the non-PI arm and PI arm serve as mutual comparators, each receiving active instruction that differs only in whether phonological terminology and transcription are used (see criterion B). Nevertheless, the paper documents each arm's composition in reasonable detail: sample sizes before and after exclusions, reasons for exclusion, shared prior English- education background (six years of secondary English, no overseas residence), and baseline TOEIC Speaking scores by group with descriptive statistics and a significance test. This level of detail allows a reader to assess whether the comparison arm was comparable at baseline, which is the underlying intent of criterion D, even though the comparator is an alternative active treatment rather than an untreated control group.
Final sentence: Criterion D is met because the comparison group's size, exclusions, background, and baseline scores are clearly documented, even though the study lacks a true untreated control condition.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred between two classes at a single university, not among multiple schools.
- "The conditions of the educational intervention were randomly assigned to the two classes." (p. 244)
Relevant Quotes:
1) "The conditions of the educational intervention were randomly assigned to the two classes: the economics class received instruction using the non-PI method (non-PI group), while the science class received instruction with the PI method (PI group)." (p. 244)
2) "This study involved 43 first-year students newly enrolled at a Japanese national university during the first semester." (p. 244)
Detailed Analysis:
The entire study was conducted within a single institution (one Japanese national university), with only two intact classes (seminars) serving as the units of assignment. There is no indication that multiple schools or institutions were involved or randomised. This does not meet the stronger school-level requirement, which calls for randomisation among separate educational institutions or sites.
Final sentence: Criterion S is not met because randomisation was confined to two classes within one university, not across multiple schools.
-
I
Independent Conduct
- The same authors designed the teaching materials, delivered all instruction, and conducted the data analysis, with no independent evaluation team involved.
- "The students were informed by their seminar instructors that participation in the eight-session instruction provided by the first author was mandatory." (p. 244)
Relevant Quotes:
1) "The students were informed by their seminar instructors that participation in the eight-session instruction provided by the first author was mandatory and would contribute to their course credits." (p. 244)
2) "SIP, a compact material designed by the first author summarizing phonological elements for the PI method, introduces key concepts through five diagrams and two tables..." (p. 245)
3) "Before the instruction began, most students attended individual interviews with the first author to discuss their English learning history..." (p. 247)
Detailed Analysis:
The first author both designed the SIP teaching material and personally delivered all instructional sessions to both groups, and the two named authors (rather than an independent evaluation team) conducted the statistical analysis reported in section 3.4 and 4.1. The only external parties mentioned are native-speaker judges who commented briefly on Haiku presentations in the final session, and the IIBC, which administered the TOEIC test; neither functioned as an independent evaluator of the intervention's overall design, delivery, or main analysis. There is no statement of third-party oversight of data collection or analysis for the primary syllable-counting outcome.
Final sentence: Criterion I is not met because the intervention was designed, delivered, and analysed by the same research team without independent conduct.
-
Y
Year Duration
- Since criterion T (Term Duration) is not met, criterion Y is automatically not met; the study duration of seven weeks is far shorter than an academic year regardless.
- "Both groups received 20 minutes of instruction per week over seven weeks in their first semester." (p. 237)
Relevant Quotes:
1) "Both groups received 20 minutes of instruction per week over seven weeks in their first semester." (p. 237)
2) "After the intervention, during the last four minutes of the wrap-up session... the post-test was administered." (p. 247)
Detailed Analysis:
Per the criteria-specific instruction that Y cannot be met if T is not met, and given that the entire study (from intervention start to final outcome measurement) spanned only about seven weeks with no extended follow-up, the Year Duration requirement of tracking outcomes across roughly 75% of an academic year is clearly not satisfied.
Final sentence: Criterion Y is not met because the study duration was only about seven weeks, and criterion T was also not met.
-
B
Balanced Control Group
- Both instructional arms received an identical number and length of sessions covering matched topics; only the instructional style (terminology-based vs. intuitive) differed, with no extra time or resources for either arm.
- "The primary difference between these two interventions lies in their instructional approach, but they share the same content distribution." (p. 247)
Relevant Quotes:
1) "The primary difference between these two interventions lies in their instructional approach, but they share the same content distribution." (p. 247)
2) "Both groups engaged in 20-minute sessions per week over the course of seven weeks (see 3.3.2 for detail)." (p. 247)
3) Table 2, "Comparison of the Non-PI and the PI methods": both columns list matched session topics (English Rhythm, Stress Unit Sound, Orthographic Transparency/Schwa, Phonics/Syllable Structure, Intonation, Mora Intervention, Evaluation) across the same seven sessions. (pp. 248-249)
Detailed Analysis:
Applying the criterion B decision procedure: no extra time or budget is present in an unbalanced way here, since both arms received exactly the same weekly dosage (20 minutes, seven weeks) and the same sequence of topics, as shown session-by-session in Table 2. The only systematic difference is whether phonological terminology and phonemic transcription were used to present that shared content (non-PI vs. PI), not a difference in the quantity of instructional time, materials cost, or teacher attention. Because EXTRA_RESOURCES_PRESENT is false (no supplementary resource was withheld from either arm), the "no extra resources present" branch of the decision tree applies directly and the criterion is trivially satisfied, without needing to invoke the integral-resource exception.
Final sentence: Criterion B is met because both groups received matched instructional time and topic coverage, differing only in presentation style rather than in the quantity of resources provided.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific comparative study by another research team was found in the paper or via a dedicated internet search.
Relevant Quotes:
1) "To the best of our knowledge, no prior research has directly compared these methods, making this a unique attempt to identify the more effective approach." (p. 243)
2) The reference list cites related but distinct prior studies by the same or overlapping authors (e.g., He & Sakuma, 2022; Ishikawa, 2009; Takayama, 2010, 2021, 2023, 2024) that examine non-PI or PI approaches individually, but none replicates this specific head-to-head comparison design. (pp. 256-259)
Detailed Analysis:
The authors themselves state that no prior study has directly compared the non-PI and PI methods, underscoring that this is a novel design rather than a replication of an existing study. An internet search (July 2026) for subsequent studies by other author teams replicating this specific non-PI vs. PI comparative design in Japanese EFL university students, using the paper's title, author names, and the SIP/CNRep materials as search terms, did not surface any independent replication; the only related hits were the source paper itself and unrelated phonology literature. Since the study was only published in 2025, insufficient time may have elapsed for an independent replication to appear, but as of this review no such replication exists.
Final sentence: Criterion R is not met because no independent replication of this specific study was found in the paper or in an internet search.
-
A
All-subject Exams
- Criterion E (a standardised exam-based outcome) is not met, and in any case the study measured only phonological syllable perception, not performance across main subjects.
Relevant Quotes:
1) "They counted syllables in nonwords before and after instruction." (p. 237)
2) "The task of identifying the number of syllables in a nonword was evaluated based on whether the syllable count was correct or not." (p. 251)
Detailed Analysis:
Per the criteria-specific instruction, if criterion E is not met then criterion A cannot be met either. Beyond that gating rule, the study's sole outcome measure is a phonological syllable-counting task focused narrowly on English pronunciation perception; no other academic subjects (e.g., general English proficiency beyond pronunciation, mathematics, science) were assessed via standardised exams.
Final sentence: Criterion A is not met because criterion E was not met and only a single, narrow outcome domain was assessed.
-
G
Graduation Tracking
- Criterion Y (Year Duration) is not met, the study ended data collection immediately after the seven-week intervention, and no follow-up publication tracking these students toward graduation was found via internet search.
- "After the intervention, during the last four minutes of the wrap-up session... the post-test was administered." (p. 247)
Relevant Quotes:
1) "After the intervention, during the last four minutes of the wrap-up session, which involved a reflection on the highlights from the previous seven session, the post-test was administered." (p. 247)
2) No mention anywhere in the Results, Discussion, or Conclusion sections of any follow-up testing beyond the immediate post-test, nor of tracking participants toward graduation. (pp. 253-255)
Detailed Analysis:
Per the criteria-specific instruction, if criterion Y is not met then criterion G cannot be met. Independently, the paper's own narrative confirms outcomes were collected only once, immediately following the final intervention session, with the Conclusion instead proposing future studies with longer or more intensive designs rather than reporting any completed longer-term or graduation tracking. An internet search (July 2026) for subsequent papers by Min He or Shuichi Takaki that might report longer-term or graduation tracking of this same freshman cohort did not identify any such follow-up publication; no evidence of graduation tracking was found in any source.
Final sentence: Criterion G is not met because there was no follow-up tracking beyond the immediate post-test, criterion Y was also not met, and no follow-up publication was found.
-
P
Pre-Registered
- No statement of pre-registration, registry platform, or registration date is present anywhere in the paper, and no registry entry was found via internet search.
Relevant Quotes:
1) "As per university regulations, this experiment is exempt from ethical review for studies involving human subjects." (p. 244)
2) No mention of any trial registry (e.g., UMIN-CTR, ISRCTN, OSF pre-registration) or a pre-specified analysis plan filed before data collection appears anywhere in the Method, Results, or Acknowledgement sections. (pp. 244-255)
Detailed Analysis:
The paper discusses an exemption from ethical review but makes no reference to a public pre-registration of the study's hypotheses, design, or planned analyses prior to data collection. In fact, the statistical model itself was determined post hoc by comparing candidate models via AIC after the maximal model failed to converge (p. 253), indicating an exploratory rather than pre-registered analytic approach. An internet search (July 2026) for a pre-registration record for this study (by title, authors, and institution) did not locate any entry in common registries.
Final sentence: Criterion P is not met because no pre-registration reference or registry link is provided anywhere in the paper, and none was found through an internet search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.