Abstract
There has been little research investigating how mode of input affects incidental vocabulary learning, and no study examining how it affects the learning of multiword items. The aim of this study was to investigate incidental learning of L2 collocations in three different modes: reading, listening, and reading while listening. One hundred thirty-eight second-year college students learning EFL in Taiwan were randomly assigned to three experimental groups (reading, listening, reading while listening) and a no treatment control group. The experimental groups encountered 17 target collocations in the same graded reader. Learning was measured using two tests that involved matching the component words and recalling their meanings. The results indicated that the reading while listening condition was most effective while the reading and listening conditions contributed to similarly sized gains. The findings suggest that listening may play a more important role in learning collocations than single-word items.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Individual students, not intact classes, were randomly assigned to the reading, listening, reading-while- listening, and control conditions, and no tutoring exception applies.
- "The participants were randomly assigned to one of four treatments: reading, listening, reading while listening, and a no treatment control." (p. 41)
Relevant Quotes:
1) "The participants of the present study were 138 second-year college students learning English as a foreign language in Taiwan." (p. 41)
2) "The participants were randomly assigned to one of four treatments: reading, listening, reading while listening, and a no treatment control. The data for any participants who missed a treatment or testing session was excluded from the study. This left 112 students who completed the full intervention. Twenty-five participants completed the listening treatment, 28 completed the reading treatment, 26 completed the simultaneous reading and listening treatment, and 33 were in the no treatment control." (p. 41)
3) "The treatment was conducted in six 50-minute classes over a 3-week period. All classes were taught by the same teacher." (p. 41)
Detailed Analysis:
The paper explicitly states that individual participants (not intact classes or schools) were the unit of randomisation: "the participants were randomly assigned to one of four treatments." There is no quote anywhere in the methodology describing pre-existing classes being assigned wholesale to a single condition; instead, students drawn from a pool of 138 second-year college students were allocated to one of four groups. This raises the standard contamination concern that the ERCT C criterion is designed to guard against, since students who ended up in different conditions could plausibly have been drawn from the same classes or peer groups and could discuss the reading material with one another.
The exception for personal tutoring/one-to-one teaching does not apply here: this is a graded-reader-based literacy/listening intervention delivered to groups of students in class sessions, not one-to-one tutoring, so the weaker student-level allocation is not excused by the exception clause.
Because there is no quoted evidence of class- or school- level randomisation, and no valid exception applies, criterion C is not met.
-
E
Exam-based Assessment
- The two dependent measures (Test A matching/recall, Test B meaning recall) were custom-built by the authors specifically for this study's 17 target collocations, not standardised exams.
- "Two tests were created to measure knowledge of the 17 target collocations." (p. 43)
Relevant Quotes:
1) "Two tests were created to measure knowledge of the 17 target collocations. The first test (Test A) was made up of two components: a matching component and a meaning recall component." (p. 43)
2) "The second test (Test B) used a meaning recall format. ... In Test B, the 18 collocations were presented to the participants and they had to translate the items into Chinese." (p. 43)
3) "Test A was piloted with native and nonnative speakers who did not take part in the study and found to work correctly." (p. 43)
Detailed Analysis:
The outcome measures (Test A and Test B) were purpose- built by the researchers to assess knowledge of the exact 17 (18, including the motivational filler item) target collocations selected for this study. They are not a standardised, widely recognised exam such as a national curriculum test or a validated external vocabulary test; rather, they are researcher-designed instruments piloted informally with a small convenience sample of native and nonnative speakers. This is precisely the kind of custom, intervention-aligned assessment the ERCT E criterion warns against, since it is tailored exactly to the treatment content rather than to a general, externally validated construct.
Because the assessments used are custom-made rather than standardised, criterion E is not met.
-
T
Term Duration
- Outcomes were measured only about six to seven weeks after the intervention began, far short of a full academic term.
- "Week 1 The pretest (Test A) was administered ... Week 4 ... the immediate posttests ... Week 8 The delayed posttests (Tests A and B) were administered to all participants." (Table 1, p. 42)
Relevant Quotes:
1) "The treatment was conducted in six 50-minute classes over a 3-week period." (p. 41)
2) Table 1: "Week 1 The pretest (Test A) was administered to all participants"; "Week 2 Experimental groups completed the first four chapters of the treatment"; "Week 4 Experimental groups completed the final three chapters of the treatment and all participants completed the immediate posttests (Tests A and B)"; "Week 8 The delayed posttests (Tests A and B) were administered to all participants." (p. 42)
3) "Knowledge of the target items was measured in a pretest 1 week before the treatment, in an immediate posttest at the conclusion of the treatment, and in a delayed posttest 4 weeks after the treatment." (p. 41)
Detailed Analysis:
The intervention itself started in week 2 and ended in week 4, and the final (delayed) measurement occurred in week 8 of the study timeline, i.e. roughly 6 weeks after the intervention began and about 4 weeks after the intervention concluded. This interval is far shorter than the minimum of one academic term (roughly 3-4 months) required by criterion T. There is no quote suggesting a longer follow-up window, and the Year Duration criterion is also not met (see Y), so the weaker Year-implies-Term exception does not apply either.
Because the interval from intervention start to final measurement is only a few weeks, criterion T is not met.
-
D
Documented Control Group
- The control group's size, composition, baseline scores, and exact activities (no reading/listening, only the same tests) are clearly documented.
- "The control group did not read or listen to audio versions of the book, but completed the same dependent measures as the experimental groups." (p. 41)
Relevant Quotes:
1) "33 were in the no treatment control." (p. 41)
2) "The control group did not read or listen to audio versions of the book, but completed the same dependent measures as the experimental groups. The inclusion of the control group ensured that any learning that occurred could be attributed to the treatments alone." (p. 41)
3) "The participants of the present study were 138 second-year college students learning English as a foreign language in Taiwan. They had been learning English for an average of 9 years and their English proficiency was estimated to range from low- to high- intermediate." (p. 41)
4) Table 3: "Pretest A ... C (n = 33) .61 (.56)" (p. 45)
Detailed Analysis:
The paper documents the control group's size (n = 33), its exact treatment (no exposure to the reading/listening material, but administration of identical pretest, immediate, and delayed posttests), and its baseline performance (Pretest A mean = .61, SD = .56 in Table 3). The broader participant description (average 9 years of English study, low- to high-intermediate proficiency) also applies to the randomly-assigned control subsample. Together this constitutes a reasonably detailed description of who the control group was, their baseline characteristics, and what (lack of) treatment they received.
Because clear documentation of the control group's size, baseline performance, and conditions is provided, criterion D is met.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred among individual students within a single college cohort, not among schools.
- "The participants were randomly assigned to one of four treatments: reading, listening, reading while listening, and a no treatment control." (p. 41)
Relevant Quotes:
1) "The participants of the present study were 138 second-year college students learning English as a foreign language in Taiwan." (p. 41)
2) "The participants were randomly assigned to one of four treatments: reading, listening, reading while listening, and a no treatment control." (p. 41)
Detailed Analysis:
The study was carried out within a single institution (one cohort of second-year EFL college students), with randomisation performed at the individual student level. There is no mention of multiple schools, let alone schools being randomised as whole units to different conditions. Since the weaker class-level criterion (C) is already not met, the stronger school-level criterion cannot be met either.
Because there is no evidence of school-level randomisation, criterion S is not met.
-
I
Independent Conduct
- The same two authors designed the intervention, built the target-item list and both dependent measures, and there is no statement of an independent third-party team conducting the study.
- "Two tests were created to measure knowledge of the 17 target collocations." (p. 43)
Relevant Quotes:
1) "A total of 17 two-word collocations were selected as target items. Three criteria were used to select the target items." (p. 42)
2) "Two tests were created to measure knowledge of the 17 target collocations." (p. 43)
3) "The treatment was conducted in six 50-minute classes over a 3-week period. All classes were taught by the same teacher." (p. 41)
4) "We would like [to] express thanks to Emeritus Professor John Read of the University of Auckland for providing some useful suggestions while the authors were developing the tests." (p. 35, footnote)
Detailed Analysis:
There is no statement anywhere in the paper of an external or independent research/evaluation team designing, delivering, or scoring the intervention. The target collocation selection, the pilot testing of the instruments, and the design of Test A and Test B are all described as done by "the authors" themselves (Webb and Chang), consistent with their long-standing individual research programme on incidental vocabulary learning. The acknowledgment thanking a colleague for "useful suggestions" on test development does not amount to an independent evaluation body; it is informal academic consultation, not third-party conduct of the trial. No quote indicates that data collection or analysis was handled by anyone independent of the authors who designed the study.
Because the same authors designed, implemented, and presumably analysed the study without a documented independent third party, criterion I is not met.
-
Y
Year Duration
- Since Term Duration (T) is not met, Year Duration cannot be met either; the whole study spanned only about 8 weeks.
- "Week 8 The delayed posttests (Tests A and B) were administered to all participants." (Table 1, p. 42)
Relevant Quotes:
1) Table 1: total study timeline runs from "Week 1" (the pretest) to "Week 8" (delayed posttest). (p. 42)
2) "Knowledge of the target items was measured in a pretest 1 week before the treatment, in an immediate posttest at the conclusion of the treatment, and in a delayed posttest 4 weeks after the treatment." (p. 41)
Detailed Analysis:
Per the criteria-specific instruction, if criterion T (Term Duration) is not met, criterion Y is automatically not met. Independently, the entire study, from pretest to final delayed posttest, spans only about 8 weeks, which is nowhere near 75% of an academic year (~9-10 months).
Criterion Y is not met.
-
B
Balanced Control Group
- The extra reading/listening exposure is itself the explicit treatment variable under investigation, so the no-treatment control lacking that exposure is a valid "business as usual" baseline rather than an unbalanced confound.
- "The control group did not read or listen to audio versions of the book, but completed the same dependent measures as the experimental groups. The inclusion of the control group ensured that any learning that occurred could be attributed to the treatments alone." (p. 41)
Relevant Quotes:
1) "The aim of this study was to investigate incidental learning of L2 collocations in three different modes: reading, listening, and reading while listening." (p. 35, abstract)
2) "The experimental groups encountered 17 target collocations in the same graded reader in one of three input modes: reading, listening, and reading while listening. ... The control group did not read or listen to audio versions of the book, but completed the same dependent measures as the experimental groups. The inclusion of the control group ensured that any learning that occurred could be attributed to the treatments alone." (p. 41)
3) "The treatment was conducted in six 50-minute classes over a 3-week period." (p. 41)
Detailed Analysis:
Applying the criterion B decision tree: extra time/ exposure is present (the three experimental groups spend six 50-minute classes reading and/or listening to a graded reader, while the control group does not). The key question is whether this extra exposure is itself the treatment variable being tested (RESOURCES_ARE_TREATMENT) rather than an incidental, separable add-on.
Here the entire research question is "does exposure to L2 input, and in which mode, produce incidental collocation learning?" The no-treatment control's sole purpose, as explicitly stated by the authors, is to establish a baseline against which any gains from exposure can be attributed to the treatments themselves, i.e., the presence or absence of meaning-focused input is precisely what is being manipulated and measured. This is analogous to studies where additional resources (e.g., a digital learning tool, extra tutoring) are the explicit treatment variable and the control is legitimately left at a "business as usual" (here, no-input) baseline. The control group did complete identical pretests and posttests, so testing conditions themselves were held constant across groups; only exposure to the graded reader differed, and that exposure is the object of study, not an incidental confound.
Because the additional reading/listening exposure is explicitly and transparently the treatment variable being investigated, and the control group is a legitimate no- input baseline used only to isolate incidental learning effects, criterion B is met.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study (same design, materials, and research questions, conducted by a different team in a different context) was found in the paper or in a subsequent internet-based literature search.
Relevant Quotes:
There is no quote in the paper referencing a prior or contemporaneous independent replication of this specific study, as the paper itself presents itself as the first study of its kind: "It is the first study to compare the effects of three modes of L2 input on learning collocations." (p. 47)
Detailed Analysis:
An internet search for subsequent research citing or conceptually replicating this paper (Webb & Chang, 2022, SSLA) identified several related but non-replicating studies by other research teams: Dang, Lu, and Webb (2022), "Incidental Learning of Collocations in an Academic Lecture Through Different Input Modes," Language Learning; Yuan, X., and Tang, J. (2025), "Incidental Learning of Collocations Under Different Input Modes and the Mediating Role of Perceptual Learning Style," published in a peer-reviewed journal; and other work on reading-while-listening and multimodal collocation learning (e.g., Pu, on young EFL learners' incidental collocation learning through multimodal input). These studies examine related research questions (mode-of-input effects on collocation learning) but use different target collocations, different source texts (an academic lecture rather than the graded reader "A Kiss before Dying"), different participant populations (e.g., young learners or different EFL contexts), and different specific designs; none of them present themselves as an independent replication of this specific study's design, materials, and claims. A separate conceptual multisite replication effort identified in the search (Peters, Puimege, and Szudarski, 2023) targets a different Webb, Newton, and Chang (2013) paper on repetition and incidental learning of multiword units, not the present 2022 study. Per the ERCT standard, conceptually related studies on the same general topic by other teams do not constitute an independent reproduction of this particular study.
Because no independent replication of this specific study was identified, criterion R is not met.
-
A
All-subject Exams
- The study measured knowledge of only 17 collocations drawn from one graded reader, not performance across all main subjects, and criterion E (a prerequisite) is not met.
- "Two tests were created to measure knowledge of the 17 target collocations." (p. 43)
Relevant Quotes:
1) "A total of 17 two-word collocations were selected as target items." (p. 42)
2) "Two tests were created to measure knowledge of the 17 target collocations." (p. 43)
Detailed Analysis:
Per the criteria-specific instruction, criterion A requires criterion E to be met as a prerequisite; since E is not met (the assessments are custom-built, not standardised), A cannot be met either. Independently, the outcome measures are narrowly focused on vocabulary/ collocation knowledge derived from a single graded reader and do not assess broader academic subjects at all; this is an L2 vocabulary-acquisition study, not a multi-subject curriculum evaluation, so the "all main subjects" coverage required by criterion A is not present.
Criterion A is not met.
-
G
Graduation Tracking
- Since Year Duration (Y) is not met, Graduation Tracking cannot be met; the study's longest follow-up was a delayed posttest four weeks after the intervention ended, and no follow-up publication tracking this cohort was found.
- "The delayed posttests (Tests A and B) were administered to all participants." (Table 1, p. 42)
Relevant Quotes:
1) "Knowledge of the target items was measured in a pretest 1 week before the treatment, in an immediate posttest at the conclusion of the treatment, and in a delayed posttest 4 weeks after the treatment." (p. 41)
2) Table 1 shows the study timeline ending at "Week 8" with the delayed posttests, with no further follow-up described. (p. 42)
Detailed Analysis:
Per the criteria-specific instruction, since criterion Y (Year Duration) is not met, criterion G is automatically not met. Independently, there is no mention anywhere in the paper of any follow-up beyond the 4-week delayed posttest, let alone tracking of participants until graduation from their degree programme. An internet search for subsequent papers by Webb and/or Chang tracking this same cohort of 138 Taiwanese college students did not identify any follow-up publication continuing measurement of this cohort; the searches instead surfaced only separate, later studies by Webb, Chang, and other authors using different samples and materials (e.g., Webb and Chang's own earlier 2012 and 2015 studies, which precede rather than follow this 2022 paper). No evidence of graduation tracking was found.
Criterion G is not met.
-
P
Pre-Registered
- There is no mention anywhere in the paper of a pre-registered protocol, registry platform, or registration date, and no internet search evidence of a registration was found.
Relevant Quotes:
No quotes referencing pre-registration, a study registry, or a registration ID/date were found anywhere in the paper, including the methodology, results, and discussion sections.
Detailed Analysis:
The paper contains no statement of the study protocol having been pre-registered on any platform (e.g., OSF, AsPredicted, a clinical-trials-style registry) prior to data collection, and provides no registration link or date. An internet search for a pre-registration record associated with this study (Webb & Chang, mode of input, incidental collocation learning) did not locate any OSF, AsPredicted, or similar registry entry linked to this paper. Given the complete absence of any such reference in the paper or from external search, there is no basis to conclude the protocol was pre-registered.
Because no evidence of pre-registration is present, criterion P is not met.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.