The Effects of Suprasegmental Phonological Training on English Reading Comprehension: Evidence from Chinese EFL Learners

Gang Cui, Yuemin Wang, and Xiaoyun Zhong

Published:
ERCT Check Date:
DOI: 10.1007/s10936-020-09743-2
  • reading
  • L2 languages
  • higher education
  • China
0
  • C

    Randomisation was conducted at the individual student level within a single university cohort taught by the same teacher, not at the class or school level, so the criterion is not met.

    "They were randomly divided into two groups with 80 in the experimental group and 80 in the control group." (p. 323)

  • E

    Authentic IELTS reading tests, an internationally recognised standardised exam, were used for both pretest and posttest, satisfying the exam-based assessment criterion.

    "Authentic IELTS reading tests (6 passages, with an average length of 1181 words) were used for both pretest and posttest, as IELTS is an international standardized test..." (p. 324)

  • T

    Outcomes were measured immediately after a 12-week (approximately one academic term) training programme that began right after the pretest, satisfying the minimum term-duration requirement.

    "the experimental group was given an extra 12-week phonological training which lasted for 2 hours each week..." (p. 324)

  • D

    The control group's size, teaching conditions, and baseline reading performance are documented and shown to be statistically equivalent to the experimental group at pretest.

    "A pretest was given to ensure that the two groups didn't have significant difference in the initial level of reading comprehension (df=158, t=-1.638, p>0.05)." (p. 323)

  • S

    Randomisation occurred among individual students within a single university, not among schools, so the school-level criterion is not met.

    "They were randomly divided into two groups with 80 in the experimental group and 80 in the control group." (p. 323)

  • I

    The study appears to have been designed, implemented, and analysed by the same research team, with no statement of independent or third-party evaluation.

    "provided by an American teacher who graduated from University of Oregon with an MA degree in English education and a TESOL certificate and four years' teaching experience in oral English." (p. 324)

  • Y

    The intervention and follow-up period lasted only 12 weeks, far short of the 75%-of-an-academic-year requirement.

    "the experimental group was given an extra 12-week phonological training which lasted for 2 hours each week..." (p. 324)

  • B

    The experimental group received 24 extra hours of instruction from a native English speaker that the control group did not receive, and the authors themselves flag this as an unaddressed confounding factor rather than framing the extra time as the deliberate treatment variable.

    "the experimental group received 24-h more instruction provided by a native speaker of English. Although the instruction focused on suprasegmentals, more L2 exposure could be a confounding factor for the improvement in the experimental group." (p. 329)

  • R

    No independent replication of this specific study is mentioned in the paper, and an internet search of citing literature found no study presenting itself as a replication of this specific study.

  • A

    Only English reading comprehension was assessed; no other core academic subjects were measured, and no rationale for this narrow focus is provided.

    "Authentic IELTS reading tests (6 passages, with an average length of 1181 words) were used for both pretest and posttest..." (p. 324)

  • G

    There is no follow-up beyond the immediate posttest after the 12-week training, no follow-up publication by the same authors was found, and the Year Duration criterion is not met, so graduation tracking cannot be considered met.

  • P

    The paper contains no statement or reference indicating that the study protocol was pre-registered before data collection, and no pre-registration record was found via internet search.

Abstract

This study aims to investigate the effect of suprasegmental phonological training on connected-text reading comprehension of Chinese university students with different English reading proficiency levels. A sample of 160 freshmen was recruited and randomly divided into experimental and control groups, and the experimental group was given a 12-week training on stress, intonation and rhythm in English. Comparison and analysis of the subjects' reading comprehension performance, involving overall accuracy and speed as well as literal and inferential comprehension, reveal that: (1) suprasegmental phonological training exerts positive effects on the subjects' overall reading comprehension, especially on reading time and literal comprehension; (2) lower-proficiency readers improve more remarkably than higher-proficiency readers in terms of overall accuracy and literal comprehension, while the effect of the training on reading time is significant regardless of the subjects' reading proficiency. The results indicate that with explicit instruction and intensive exposure to suprasegmental knowledge, students' automaticity in lower level processing, such as parsing and understanding propositional messages, can be increased. From a perspective of interaction among different cognitive and psychological processes of reading comprehension, this study can shed light on developing students' reading comprehension in EFL contexts.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was conducted at the individual student level within a single university cohort taught by the same teacher, not at the class or school level, so the criterion is not met.
      • "They were randomly divided into two groups with 80 in the experimental group and 80 in the control group." (p. 323)
      • Relevant Quotes: 1) "The subjects of this study consist of 160 freshmen (non-English majors whose native language are all Chinese) from a university in southeast China." (p. 323) 2) "They were randomly divided into two groups with 80 in the experimental group and 80 in the control group." (p. 323) 3) "Both groups were taught by the same English teacher, but only the experimental group received the phonological training." (p. 323) Detailed Analysis: The paper describes randomisation as occurring among individual students drawn from a single cohort of freshmen at one university, not among intact classes or schools. The quote "randomly divided into two groups" refers to the 160 individual subjects, and the fact that "both groups were taught by the same English teacher" confirms that students from the same regular class(es) were split between the experimental and control conditions rather than whole classes being assigned as units. There is no description of randomisation occurring above the individual-student level, and no measures are described to prevent contamination between students in the two conditions who share the same regular English classroom. The intervention consists of group phonological training sessions (2 hours/week for 12 weeks), not one-to-one tutoring, so the tutoring exception does not apply. Criterion C is not met because randomisation was conducted at the individual student level within a shared classroom context, not at the class or school level, and no tutoring exception applies.
    • E

      Exam-based Assessment

      • Authentic IELTS reading tests, an internationally recognised standardised exam, were used for both pretest and posttest, satisfying the exam-based assessment criterion.
      • "Authentic IELTS reading tests (6 passages, with an average length of 1181 words) were used for both pretest and posttest, as IELTS is an international standardized test..." (p. 324)
      • Relevant Quotes: 1) "Authentic IELTS reading tests (6 passages, with an average length of 1181 words) were used for both pretest and posttest, as IELTS is an international standardized test to measure EFL learners' general performance in reading comprehension with great validity and reliability." (p. 324) 2) "The scoring was based on the keys released by IELTS, one point for each correct answer, zero for wrong or partial answers." (p. 324) Detailed Analysis: The study used authentic IELTS reading test passages and items for both pretest and posttest, with official IELTS answer keys used for scoring. IELTS is a widely recognised, internationally standardised assessment, not an instrument custom-built for this study. The authors explicitly justify their choice by citing the test's established validity and reliability. This directly satisfies the requirement for a standard, widely recognised exam-based assessment. Criterion E is met because the study used authentic, unmodified IELTS reading tests, a widely recognised standardised assessment.
    • T

      Term Duration

      • Outcomes were measured immediately after a 12-week (approximately one academic term) training programme that began right after the pretest, satisfying the minimum term-duration requirement.
      • "the experimental group was given an extra 12-week phonological training which lasted for 2 hours each week..." (p. 324)
      • Relevant Quotes: 1) "The training program began right after the pretest. Both groups had their English classes as usual, while the experimental group was given an extra 12-week phonological training which lasted for 2 hours each week..." (p. 324) 2) "After finishing the phonological training, both groups took the posttest." (p. 325) Detailed Analysis: The intervention began immediately after the pretest and ran for 12 consecutive weeks, with the posttest (primary outcome measurement) administered immediately after the training concluded. Twelve weeks corresponds to approximately three months, which falls within the standard's definition of a term ("a semester or equivalent, approximately 3-4 months"), and is consistent with a typical academic term/quarter length. The interval between intervention start and outcome measurement therefore spans essentially the full 12-week programme, approximating a full academic term, rather than a brief few-week intervention followed by an immediate test. Criterion T is met because the interval from intervention start to outcome measurement spans approximately 12 weeks, which approximates the minimum one-term duration required by the standard.
    • D

      Documented Control Group

      • The control group's size, teaching conditions, and baseline reading performance are documented and shown to be statistically equivalent to the experimental group at pretest.
      • "A pretest was given to ensure that the two groups didn't have significant difference in the initial level of reading comprehension (df=158, t=-1.638, p>0.05)." (p. 323)
      • Relevant Quotes: 1) "They were randomly divided into two groups with 80 in the experimental group and 80 in the control group." (p. 323) 2) "A pretest was given to ensure that the two groups didn't have significant difference in the initial level of reading comprehension (df=158, t=-1.638, p>0.05). Both groups were taught by the same English teacher, but only the experimental group received the phonological training." (p. 323) 3) Table 3 reports control-group pretest means and standard deviations for total score (M=24.96, SD=3.56), reading time (M=61.14, SD=5.11), literal comprehension (M=68.85%, SD=15.81), and inferential comprehension (M=59.31%, SD=9.00). (p. 325) Detailed Analysis: The paper documents the control group's size (N=80), confirms it received identical regular English instruction from the same teacher as the experimental group, and explicitly states that only the experimental group received the additional phonological training, i.e. the control group's condition was "business as usual." Quantitative baseline (pretest) data for the control group are reported separately from the experimental group across all four outcome measures in Table 3, and a formal statistical test confirms no significant baseline difference between groups. While the paper does not break down demographic characteristics (age, gender) separately by group, the size, instructional conditions, and quantitative baseline performance of the control group are documented in enough detail to assess comparability. Criterion D is met because the control group's size, teaching conditions, and baseline performance are documented and shown to be statistically equivalent to the experimental group.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred among individual students within a single university, not among schools, so the school-level criterion is not met.
      • "They were randomly divided into two groups with 80 in the experimental group and 80 in the control group." (p. 323)
      • Relevant Quotes: 1) "The subjects of this study consist of 160 freshmen ...from a university in southeast China." (p. 323) 2) "They were randomly divided into two groups with 80 in the experimental group and 80 in the control group." (p. 323) Detailed Analysis: The study was conducted at a single university, with individual freshmen students randomly assigned to experimental or control conditions. There is no mention of multiple schools or institutions being randomised, nor of randomisation occurring above the individual-student level. This falls well short of the school-level randomisation required by criterion S. Criterion S is not met because randomisation occurred among individual students at a single university, not among schools.
    • I

      Independent Conduct

      • The study appears to have been designed, implemented, and analysed by the same research team, with no statement of independent or third-party evaluation.
      • "provided by an American teacher who graduated from University of Oregon with an MA degree in English education and a TESOL certificate and four years' teaching experience in oral English." (p. 324)
      • Relevant Quotes: 1) "The training program began right after the pretest. Both groups had their English classes as usual, while the experimental group was given an extra 12-week phonological training which lasted for 2 hours each week, provided by an American teacher who graduated from University of Oregon with an MA degree in English education and a TESOL certificate and four years' teaching experience in oral English." (p. 324) 2) No statement in the Methods, Acknowledgments, or Compliance sections describes an external or independent team conducting data collection, scoring, or analysis separate from the study's authors. Detailed Analysis: The paper describes the phonological training itself as being delivered by an external instructor (an American teacher), but this concerns delivery of the intervention only, not independent conduct of the evaluation. There is no quote indicating that data collection, scoring, or statistical analysis were carried out by a third party independent of the authors who designed the study. The "Compliance with Ethical Standards" section addresses only conflict of interest and informed consent, not evaluator independence. Absent any statement of independent oversight of data collection or analysis, the criterion is not satisfied. Criterion I is not met because there is no evidence that data collection or analysis was conducted independently of the research team that designed the study.
    • Y

      Year Duration

      • The intervention and follow-up period lasted only 12 weeks, far short of the 75%-of-an-academic-year requirement.
      • "the experimental group was given an extra 12-week phonological training which lasted for 2 hours each week..." (p. 324)
      • Relevant Quotes: 1) "The training program began right after the pretest...the experimental group was given an extra 12-week phonological training which lasted for 2 hours each week." (p. 324) 2) "After finishing the phonological training, both groups took the posttest." (p. 325) Detailed Analysis: The entire study, from intervention start to final outcome measurement, spans only 12 weeks. This is far short of 75% of an academic year (approximately 9-10 months), which the standard requires for the stronger Year Duration criterion. There is no evidence of any longer-term follow-up beyond the immediate posttest. Criterion Y is not met because the tracking period covers only about 12 weeks, well below 75% of an academic year.
    • B

      Balanced Control Group

      • The experimental group received 24 extra hours of instruction from a native English speaker that the control group did not receive, and the authors themselves flag this as an unaddressed confounding factor rather than framing the extra time as the deliberate treatment variable.
      • "the experimental group received 24-h more instruction provided by a native speaker of English. Although the instruction focused on suprasegmentals, more L2 exposure could be a confounding factor for the improvement in the experimental group." (p. 329)
      • Relevant Quotes: 1) "Both groups had their English classes as usual, while the experimental group was given an extra 12-week phonological training which lasted for 2 hours each week, provided by an American teacher..." (p. 324) 2) "...only the experimental group received the phonological training." (p. 323) 3) "Meanwhile, the experimental group received 24-h more instruction provided by a native speaker of English. Although the instruction focused on suprasegmentals, more L2 exposure could be a confounding factor for the improvement in the experimental group." (p. 329) Detailed Analysis: The experimental group received an additional 24 hours of in-person instruction (12 weeks x 2 hours) from a native English-speaking teacher, on top of the regular English classes both groups attended. The control group received no compensatory time, alternative activity, or placebo exposure to match this extra instructional time; they simply continued their normal English classes. Applying the criterion B decision logic: extra resources are clearly present (24 extra hours of native-speaker instruction) and this is not a negligible amount of time. The study's stated purpose is to test the effect of suprasegmental phonological training as a specific instructional content, not to test the effect of additional instructional time or native-speaker exposure per se, so the "resources are the explicit treatment variable" exception is not clearly established by the authors' framing of intent. Crucially, the control group's resources were not matched, and the authors themselves explicitly acknowledge in the limitations section that the extra 24 hours of native-speaker exposure "could be a confounding factor for the improvement in the experimental group," directly undermining any claim that the additional resource was cleanly isolated as the intended treatment variable rather than an unaddressed confound. This self-identified imbalance is precisely the scenario criterion B is designed to flag. Criterion B is not met because the experimental group received substantially more instructional time and native-speaker exposure than the control group, with no balancing resource provided to the control group, and the authors explicitly acknowledge this imbalance as an unresolved confound.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study is mentioned in the paper, and an internet search of citing literature found no study presenting itself as a replication of this specific study.
      • Relevant Quotes: No quotes in the paper reference a prior or subsequent independent replication of this specific suprasegmental phonological training intervention with Chinese EFL university students. Detailed Analysis: The paper cites related prior work on phonological training and reading (e.g. Sun et al. 2013; Li and Chen 2016) as background literature, but these studies examine different populations (young children) and different training foci (segmental phonological awareness, morphological awareness), and are not presented as replications of this study's design or findings. An internet search (Semantic Scholar citation records for DOI 10.1007/s10936-020-09743-2) found citing works such as a 2023 Heliyon study on English liaison acquisition by Chinese EFL learners and a 2024 study on suprasegmental phonology instruction in an Algerian school; both examine different interventions, populations and contexts, and neither presents itself as an independent replication of this specific study's design or findings. No evidence was found that this specific study has been independently reproduced by a different research team in a comparable context. Criterion R is not met because no independent replication of this study is referenced in the paper or found via internet search.
    • A

      All-subject Exams

      • Only English reading comprehension was assessed; no other core academic subjects were measured, and no rationale for this narrow focus is provided.
      • "Authentic IELTS reading tests (6 passages, with an average length of 1181 words) were used for both pretest and posttest..." (p. 324)
      • Relevant Quotes: 1) "Authentic IELTS reading tests (6 passages, with an average length of 1181 words) were used for both pretest and posttest, as IELTS is an international standardized test to measure EFL learners' general performance in reading comprehension..." (p. 324) 2) The Results and Discussion sections report only total score, reading time, literal comprehension, and inferential comprehension, all sub-measures of the single IELTS reading test; no other subject-area outcomes are reported. Detailed Analysis: While criterion E is met (IELTS is a standardised exam), the study exclusively measures English reading comprehension. No other core subjects (e.g. mathematics, science) taken by these university freshmen were assessed, and the paper does not provide an explicit rationale for why only reading was measured as an acceptable narrow-scope exception (such as a vocational or highly specialised programme justification). As an EFL reading-specific intervention, the study's outcome measure is inherently limited to one subject, and the all-subject exception is not invoked or justified in the text. Criterion A is not met because only reading comprehension was assessed, with no other core subjects measured and no explicit exception rationale provided.
    • G

      Graduation Tracking

      • There is no follow-up beyond the immediate posttest after the 12-week training, no follow-up publication by the same authors was found, and the Year Duration criterion is not met, so graduation tracking cannot be considered met.
      • Relevant Quotes: 1) "After finishing the phonological training, both groups took the posttest." (p. 325) 2) The Discussion and Conclusion sections propose only future research directions (e.g. longer training duration, larger sample), with no mention of any completed or planned follow-up of these subjects toward graduation. Detailed Analysis: Outcome measurement ended with the posttest immediately following the 12-week training programme. There is no evidence of any further tracking of the subjects, let alone tracking through to their graduation from university. An internet search (Semantic Scholar author and citation records) found no subsequent publications by Gang Cui, Yuemin Wang, or Xiaoyun Zhong tracking this same cohort of 160 freshmen toward graduation. As the weaker Year Duration criterion (Y) is also not met, the stronger Graduation Tracking criterion cannot be met either, per the standard's cascading requirement. Criterion G is not met because no follow-up beyond the immediate posttest is reported, no follow-up publication by the same authors was found via internet search, and criterion Y is not met.
    • P

      Pre-Registered

      • The paper contains no statement or reference indicating that the study protocol was pre-registered before data collection, and no pre-registration record was found via internet search.
      • Relevant Quotes: No quotes anywhere in the paper (including the Methodology, Compliance with Ethical Standards, or Acknowledgments sections) mention a study registry, registration ID, or pre-registration date. Detailed Analysis: The "Compliance with Ethical Standards" section addresses only conflict of interest and informed consent. No registry platform (e.g. ClinicalTrials.gov, AsPredicted, OSF) or pre-registration date is referenced anywhere in the text. An internet search of the paper's external identifiers (DOI 10.1007/s10936-020-09743-2, PubMed ID 33151474) and of common pre-registration registries found no corresponding pre-registration record for this study. Criterion P is not met because there is no evidence of a pre-registered protocol in the paper or in available registries.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.