Abstract
The purpose of this study is to explore the impact of teaching subskills, namely micro- and macro-skills, with a speaking-listening model on the improvement of listening competence. The research included 112 Chinese tertiary students with intermediate English proficiency who were recruited from around the country. Before attending a listening class, the experimental group engaged in oral practice of the subskills, while the control one engaged in conventional listening-oriented preparation before attending a listening class. A randomized controlled trial (RCT), as well as a questionnaire, were used to assess the listening skills. Following the results of the test analysis, we concluded that practicing listening subskills, first verbally and subsequently audibly, had a substantial impact on the development of listening competence. This efficiency was particularly evident when it came to growing discourse and pragmatical listening skills, rather than developing grammatical and sociolinguistic competence. The results of the questionnaire indicated that there was minimal difference between the two groups in terms of listening strategic competence. Our findings were confirmed by coding the interview data, which revealed that tertiary students' self-agency and class participation had increased. The findings indicate that teaching tertiary students listening with speaking before listening in a computer-mediated communication (CMC) setting has an uneven influence on their development of listening skills.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was performed at the individual student level within one university course, not at the class or school level, and the intervention is whole-class instruction, not one-to-one tutoring, so no exception applies.
- "Intermediate-level EFL learners enrolled in the Listening/Speaking courses at a midwestern university in China. They were randomly assigned to two pedagogical conditions (Speaking-preceding-listening, Listening-oriented) for the RCT." (p. 5)
Relevant Quotes:
1) "Intermediate-level EFL learners enrolled in the Listening/Speaking courses at a midwestern university in China. They were randomly assigned to two pedagogical conditions (Speaking-preceding-listening, Listening-oriented) for the RCT." (p. 5)
2) "The research included 112 Chinese tertiary students with intermediate English proficiency who were recruited from around the country." (p. 1, Abstract)
3) "Table 2 summarizes language background information in the final sample (N = 112; Mage = 20.38; 35 female)." (p. 5)
4) "The classes (both EG and CG) were held once a week for 2 h each in classrooms equipped with computer technology for CMC." (p. 7)
Detailed Analysis:
Criterion C requires that entire classes (or, stronger, whole schools) be the unit of randomisation, unless the intervention is one-to-one tutoring/personal teaching. The paper states that individual learners enrolled in the Listening/Speaking courses "were randomly assigned to two pedagogical conditions." The unit of randomisation is thus the individual student, who was allocated into one of two instructed groups (EG n = 56, CG n = 56) at a single university. No pre-existing intact classes or schools were randomised, and no description of class-level or cluster-level allocation is given anywhere in the paper. The intervention is group classroom instruction delivered by a teacher in a CMC-equipped classroom, not personal tutoring, so the tutoring exception does not apply.
Criterion C is not met because randomisation was carried out at the individual student level rather than at the class or school level, and no valid exception applies.
-
E
Exam-based Assessment
- Outcomes were measured with a researcher-assembled listening test compiled from a Pearson practice extract and textbook exercises, piloted by the authors themselves, not a recognised standardised exam.
- "The test material was delivered at normal speed (120-150/min) and consisted of three tasks: (1) the responsive listening task, extracted from Pearson English International Certificate; (2) the extensive listening task; (3) the selective listening tasks consisting of two listening Cloze tests, one dialog and one long speech, borrowed from Tactics for Listening at Intermediate level (Richards, 2004...)." (p. 6)
Relevant Quotes:
1) "Two listening comprehension tests were administrated to both groups at the outset and the end of the study." (p. 6)
2) "The test material was delivered at normal speed (120-150/min) and consisted of three tasks: (1) the responsive listening task, extracted from Pearson English International Certificate; (2) the extensive listening task; (3) the selective listening tasks consisting of two listening Cloze tests, one dialog and one long speech, borrowed from Tactics for Listening at Intermediate level (Richards, 2004; see Supplementary Appendix 1)." (p. 6)
3) "The test was tallied with integer points; 1-13 questions accounted for two points each, 14~17 gap-fillings for 3 x 4 = 12 points, and 18~21 are six points in total. The reliability analysis of a pilot test showed that Cronbach's alpha = 0.78." (p. 6)
4) "To investigate learners' language competence underlying their performance, the question items were categorized according to Buck's framework of describing listening ability (Buck, 2001)." (p. 6)
Detailed Analysis:
Criterion E requires that outcomes be measured with a standard, widely recognised standardised exam, not an instrument specially assembled for the study. Here the outcome measure is a composite test constructed by the researchers: one task was "extracted from" practice material of the Pearson English International Certificate, other tasks were "borrowed from" the commercial textbook Tactics for Listening, and the whole instrument was scored with a scheme devised by the authors and validated only via their own pilot (Cronbach's alpha = 0.78) and confirmatory factor analysis. Although some items originate from recognised sources, the students did not sit an actual administered standardised examination (e.g., the full PTE, CET, IELTS, or TOEFL); they took a bespoke researcher- compiled test. This is exactly the kind of study-specific instrument the criterion is designed to exclude, since it may be aligned with the taught content.
Criterion E is not met because the outcome was a custom, researcher-assembled listening test rather than an administered, widely recognised standardised exam.
-
T
Term Duration
- The intervention ran across a full 12-week teaching semester with the post-test administered at the end of the semester, which corresponds to one academic term ("a semester or equivalent") from intervention start to measurement.
- "The randomized controlled trial lasted for 12 weeks." (p. 7)
Relevant Quotes:
1) "The classes (both EG and CG) were held once a week for 2 h each in classrooms equipped with computer technology for CMC... The randomized controlled trial lasted for 12 weeks." (p. 7)
2) "In week 1, the teacher introduced the teaching syllabus and research aim and asked students to sign a consent form. Besides, both groups took a mock test before they completed the pretest. From Week 2 to Week 11, EG received instructions with the speaking-listening model while CG with the conventional listening-oriented instruction model. In week 12, both groups took a post-listening test and a questionnaire." (p. 7)
3) "To ensure the internal consistency between the two tests, a test-retest model or RCT was adopted in that the two-time points for testing were over 2 months." (p. 6)
4) "After the instruction intervention over a semester, the mean score differences on the post-test between the two groups were larger (CG, M = 22.9; EG, M = 26.2)." (p. 10)
5) "Specifically, participants were excluded from the final statistical analysis if their oral class attendance was lower than 70%... the achievements they made over the semester were maximally due to our designated treatments." (p. 5)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (a semester or equivalent, roughly 3-4 months) after the intervention begins. The trial spanned 12 weeks: instruction under the two models ran from Week 2 to Week 11 and the post-listening test was taken in Week 12, so outcome measurement occurred at the end of the full teaching period, roughly 10-11 weeks (about 2.5-3 months) after the intervention began. The paper repeatedly frames this period as the teaching semester ("the instruction intervention over a semester", "the achievements they made over the semester"), i.e., the intervention filled the course's semester and outcomes were collected at semester end. This matches the standard's definition of a term as "a semester or equivalent", although the 12-week window sits at the short end of the approximately 3-4 month range, which should be noted.
Criterion T is met because outcomes were measured at the end of a semester-long, 12-week instructional period, i.e., approximately one full academic term after the intervention began.
-
D
Documented Control Group
- The control group's size, demographics, baseline scores and the exact instruction it received are documented in detail in Table 2, Table 6 and the description of the listening-oriented model.
- "Learners of the CG were instructed with the listening-oriented method illustrated in Figure 2. As for the knowledge of listening subskills, they received the explicit instruction of online resources on the computer, but without oral practice." (p. 7)
Relevant Quotes:
1) "Table 2 summarizes language background information in the final sample (N = 112; Mage = 20.38; 35 female). All participants reported Chinese as their only native language and English as their second language. They had learned English for about 11 (M = 11.95) years and generally were at the intermediate proficiency level." (p. 5)
2) "TABLE 2 | Participant background information. Speaking-preceding-listening / Listening-oriented: N 56 / 56; Age 20.42 / 20.34; Years of education in English 11.8 / 12.1; Latest comprehensive English examination 71.8 (7.48) / 69.7 (8.67)." (p. 5)
3) "Mann-Whitney U test of the latest semester's comprehensive English examination revealed no significant group differences (p = 0.06 > 0.05)." (p. 5)
4) "Learners of the CG were instructed with the listening-oriented method illustrated in Figure 2. As for the knowledge of listening subskills, they received the explicit instruction of online resources on the computer, but without oral practice. For the textbook, they followed the conventional model: pre-listening, while-listening, and post-listening activities." (p. 7)
5) "Table 7 presents no significant difference between the CG and EG in the pre-listening test, t (110) = 0.34, p = 0.74 > 0.05, confirming that the listening proficiency was not statistically different between the two groups before treatments." (p. 10)
6) "Participants in both the experimental group (EG) and the control group (CG) were instructed with the same textbook: New Horizon College English - Viewing, Listening and Speaking 3." (p. 5)
Detailed Analysis:
Criterion D requires detailed documentation of the control group: who they are, their baseline characteristics, and what they received. The paper reports the control group's size (n = 56), age, gender composition of the sample, years of English education, prior comprehensive English examination scores with a baseline equivalence test, and pre-test listening scores with descriptive statistics (Table 6: M = 21.1, SD = 9.41) plus a formal comparability test. The condition the control group experienced is also described concretely: same textbook, same weekly 2-hour class, conventional pre-/while-/post-listening instruction without oral practice. This is sufficient to judge comparability of the groups and what "business as usual" consisted of.
Criterion D is met because the control group's composition, baseline performance and instructional condition are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- The trial randomised individual students within a single university; no schools or institutional units were randomised.
- "Intermediate-level EFL learners enrolled in the Listening/Speaking courses at a midwestern university in China. They were randomly assigned to two pedagogical conditions (Speaking-preceding-listening, Listening-oriented) for the RCT." (p. 5)
Relevant Quotes:
1) "Intermediate-level EFL learners enrolled in the Listening/Speaking courses at a midwestern university in China. They were randomly assigned to two pedagogical conditions (Speaking-preceding-listening, Listening-oriented) for the RCT." (p. 5)
2) "The classes (both EG and CG) were held once a week for 2 h each in classrooms equipped with computer technology for CMC." (p. 7)
Detailed Analysis:
Criterion S requires randomisation among schools or equivalent implementing institutions (sites, centres, etc.). This study took place at one university ("a midwestern university in China") and allocated individual students to the two conditions. There is no mention of multiple schools, sites or institutions, let alone their random assignment. With a single institution and student-level allocation, the school-level requirement cannot be satisfied.
Criterion S is not met because randomisation occurred at the student level within one university, not among schools or institutional units.
-
I
Independent Conduct
- The two authors designed the instructional model and also ran and analysed the trial through the course teacher, with no external or third-party evaluation team documented.
- "Both authors listed have made a substantial, direct, and intellectual contribution to the work, and approved it for publication." (p. 14, Author Contributions)
Relevant Quotes:
1) "Both authors listed have made a substantial, direct, and intellectual contribution to the work, and approved it for publication." (p. 14, Author Contributions)
2) "As for the instruction context, the blended model consisting of the online teaching and large classroom teaching with the same local language teacher was adopted in this study." (p. 5)
3) "The interviews were mainly about learners' perceptions of the relationship between speaking subskills practice and their listening improvement, which were carried out by the teacher at the end of the semester after the class for the experimental group only." (p. 7)
4) "The researcher experimented with these samples for two reasons." (p. 5)
5) "Conflict of Interest: The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest." (p. 15)
Detailed Analysis:
Criterion I requires the study to be conducted independently of those who designed the intervention, or at least with documented third-party oversight of data collection and analysis. In this paper the authors conceived the speaking-listening instructional model, designed the tests and questionnaire, and carried out the experiment ("The researcher experimented with these samples..."). Teaching, testing and interviewing were done by the course teacher within the study team, and the analysis was performed by the authors. No external evaluation agency, independent data collectors, or blinded assessors are mentioned anywhere. The conflict-of-interest declaration addresses only commercial relationships, not evaluator independence.
Criterion I is not met because the intervention's designers themselves implemented and evaluated the study without any documented independent oversight.
-
Y
Year Duration
- The whole trial lasted only 12 weeks, far less than 75% of an academic year (about 9-10 months).
- "The randomized controlled trial lasted for 12 weeks." (p. 7)
Relevant Quotes:
1) "The randomized controlled trial lasted for 12 weeks." (p. 7)
2) "In week 1, the teacher introduced the teaching syllabus and research aim... From Week 2 to Week 11, EG received instructions with the speaking-listening model while CG with the conventional listening-oriented instruction model. In week 12, both groups took a post-listening test and a questionnaire." (p. 7)
3) "To ensure the internal consistency between the two tests, a test-retest model or RCT was adopted in that the two-time points for testing were over 2 months." (p. 6)
Detailed Analysis:
Criterion Y requires outcomes to be measured at least 75% of a full academic year (roughly 9-10 months) after the intervention begins. Here the entire trial, from introduction to post-test, spanned 12 weeks within a single semester, and the interval between pre-test and post-test was described as "over 2 months." Twelve weeks is roughly three months, i.e., about 30% of a 9-10 month academic year, well below the 75% threshold. No longer follow-up measurement is reported.
Criterion Y is not met because tracking lasted only 12 weeks, far short of 75% of an academic year.
-
B
Balanced Control Group
- Both groups received the same weekly 2-hour class, the same textbook and the same subskill content over 12 weeks, with the experimental group's CMC speaking tools replacing (not adding to) class time as an integral part of the speaking-listening model being tested.
- "The textbook material in each unit was selectively covered for EG since the speaking session took half of the class time." (p. 8-9)
Relevant Quotes:
1) "The classes (both EG and CG) were held once a week for 2 h each in classrooms equipped with computer technology for CMC. EG had access to the designated website (see text footnote 3) and voice chat room like WeChat, while CG employed CMC through watching and listening." (p. 7)
2) "Participants in both the experimental group (EG) and the control group (CG) were instructed with the same textbook: New Horizon College English - Viewing, Listening and Speaking 3. Additionally, for the subskills, both groups were introduced to the same online video materials from YouTube." (p. 5)
3) "Besides, for homework, all participants were required to find their authentic listening resources based on what they had learned in class." (p. 6)
4) "The overall syllabus for both CG and EG is shown in Table 5. A total of seven units of the textbook plus micro- and macro-skills were covered over the semester." (p. 8)
5) "The textbook material in each unit was selectively covered for EG since the speaking session took half of the class time." (p. 8-9)
6) "Learners of the CG were instructed with the listening-oriented method illustrated in Figure 2. As for the knowledge of listening subskills, they received the explicit instruction of online resources on the computer, but without oral practice." (p. 7)
7) "Especially, learners of EG were facilitated with CMC in speaking sessions during the treatment, a speaking practice web with Automatic speech recognition (ASR), and a voice chat room like WeChat group (a social software application)." (p. 6)
Detailed Analysis:
Criterion B asks whether the control condition received comparable time and resources so that the effect of the specific intervention is isolated. Following the decision tree: (1) Did the intervention add extra time or budget? Instructional time was identical - both groups had one 2-hour class per week for 12 weeks, the same textbook, the same syllabus (Table 5), the same online subskill videos, and the same homework requirement. The experimental group did not receive additional class time; instead, half of its existing class time was reallocated from listening work to oral practice ("the speaking session took half of the class time," with textbook material "selectively covered" to compensate). (2) The only extra resources for EG were the free speaking-practice website with ASR and the WeChat-style voice chatroom. These tools are the delivery mechanism of the speaking-listening model itself - the study explicitly tests oral CMC practice of subskills versus listening-oriented preparation - so they are integral to the treatment being tested rather than a separable add-on, and both groups' classes were equally equipped with CMC computers. The control condition was an active one: it received the same explicit subskill instruction and spent the equivalent time on listening activities. Any negligible out-of-class use of the practice website is part of the tested instructional contrast.
Criterion B is met because both conditions received equal instructional time, materials and subskill content, and the CMC speaking tools given to the experimental group are an integral component of the speaking-listening model under test rather than an unmatched extra resource.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this speaking-listening subskills RCT by another team was reported in the paper or found in a search of citing literature.
Relevant Quotes:
1) "Due to transient and invisible attributes of speaking and listening, research on the noticing effect of speaking to listening and its learning outcomes is relatively limited." (p. 4)
2) "This research lends support to the critical role of oral output in listening development in SLA, which is well recognized theoretically but has not been experimentally evident to date." (p. 12)
3) "Future research may examine the format's effect on L2 outcomes in different L1 contexts and age groups." (p. 13)
Detailed Analysis:
Criterion R requires that the study be independently replicated by a different research team, in a different context, in a peer-reviewed journal. The paper itself positions this trial as novel, noting that the speaking- preceding-listening effect "has not been experimentally evident to date" and calling for future research in other contexts. Related earlier studies it cites (Izumi, 2002; Izumi and Izumi, 2004; Linebaugh and Roche, 2015; Zalbidea, 2021) are prior investigations of output-input models on different constructs, not replications of this study.
A search of the literature citing this article (via Semantic Scholar, DOI 10.3389/fpsyg.2022.836013) returned nine citing works published 2023-2026, including a systematic review of technology-enhanced L2 listening development (Zhang, Zou and Cheng, 2023), a systematic review of RCTs in English language education (Sijali, Poudel and Dahal, 2026), and several unrelated instructional studies (e.g., on visual aids for speaking, Google Forms listening practice, think-pair-share). None of these is an independent replication of this specific 12-week speaking-listening subskills RCT by a different research team; they are reviews or studies of different interventions. No peer-reviewed replication of this study was found.
Criterion R is not met because no independent replication of this study has been published or referenced.
-
A
All-subject Exams
- Only English listening competence was assessed, with a non-standardised custom test (criterion E fails), so all-subject standardised assessment is clearly not satisfied.
- "Two listening comprehension tests were administrated to both groups at the outset and the end of the study." (p. 6)
Relevant Quotes:
1) "Two listening comprehension tests were administrated to both groups at the outset and the end of the study." (p. 6)
2) "A randomized controlled trial (RCT), as well as a questionnaire, were used to assess the listening skills." (p. 1, Abstract)
3) "To investigate learners' language competence underlying their performance, the question items were categorized according to Buck's framework of describing listening ability (Buck, 2001)." (p. 6)
Detailed Analysis:
Criterion A requires standardised exam-based assessment of all main subjects taught at the educational level, and it presupposes criterion E. Criterion E is not met here (the listening test was researcher-assembled), so criterion A automatically fails. In addition, the only academic outcome measured is English listening comprehension (subdivided into four linguistic competence categories); no other subjects of the university curriculum were assessed. While the intervention is a specialised L2 course, the paper offers no standardised multi-subject assessment nor an explicit rationale framed against other subjects.
Criterion A is not met because criterion E fails and the study assessed only English listening competence rather than all main subjects.
-
G
Graduation Tracking
- Measurement ended with the Week 12 post-test, no follow-up to graduation is reported or found in later publications, and the prerequisite criterion Y is not met.
- "In week 12, both groups took a post-listening test and a questionnaire." (p. 7)
Relevant Quotes:
1) "In week 12, both groups took a post-listening test and a questionnaire. The interviews were conducted only with the experimental group." (p. 7)
2) "Additionally, the grammatical skills of the experimental group in word recognition still need further investigation due to possible memory decay in the post-treatment period." (p. 13)
3) "Therefore, further research is needed to strengthen our understanding of the impact of the novel format on the development of learners' listening competence." (p. 13)
Detailed Analysis:
Criterion G requires tracking participants until graduation from their educational stage. Data collection ended with the post-test, questionnaire and interviews in Week 12 of the same semester. The authors themselves flag post-treatment effects as an open question and call for further research, indicating no longer-term follow-up.
A search of works citing this article and of publications by the same authors (Jinman Zhao, Chang In Lee) found no follow-up study tracking the same 112-student cohort through to graduation; the nine citing works identified are reviews or unrelated instructional studies by other authors, not longitudinal follow-ups of this cohort. Moreover, per the standard's dependency rule, criterion G cannot be met when criterion Y is not met, and Y fails here.
Criterion G is not met because tracking stopped at the end of the 12-week trial with no graduation follow-up found.
-
P
Pre-Registered
- The paper contains no mention of any pre-registered protocol, registry platform, or registration date, and no registry entry for this trial was found.
Relevant Quotes:
1) "The studies involving human participants were reviewed and approved by the Shanxi Agricultural University. The patients/participants provided their written informed consent to participate in this study." (p. 14, Ethics Statement)
2) "The original contributions presented in this study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding author." (p. 14, Data Availability Statement)
Detailed Analysis:
Criterion P requires the full study protocol to be registered on a public registry before data collection begins, with verifiable timing. The paper reports ethics approval and informed consent but contains no reference to any trial registry (e.g., ClinicalTrials.gov, OSF, AEA registry), no registration ID, and no registration date. Since the paper gives no registry name or ID to check, no further registry lookup was possible; no independent evidence of pre-registration for this trial was found.
Criterion P is not met because no pre-registration of the study protocol is mentioned anywhere in the paper or found elsewhere.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.