Learning Vocabulary Through Reading, Listening, and Viewing: Which Mode of Input Is Most Effective?

Yanxue Feng, Stuart Webb

Published:
ERCT Check Date:
DOI: 10.1017/S0272263119000494
  • reading
  • L2 languages
  • higher education
  • China
0
  • C

    The paper calls its own design "quasi-experimental" and only describes random assignment of students to classes by the university, not random assignment of the four classes to treatment conditions.

    "They were in four classes that were randomly assigned by the university, with 21 second-year students being in one class and 55 third-year students divided into the other three."

  • E

    Outcome measures were custom checklist and multiple-choice tests built around 43 documentary-specific target words, not standardised exams.

    "Checklist and multiple-choice tests were designed to measure knowledge of target words."

  • T

    The treatment was a single session, with outcomes measured immediately and one week later, far short of one academic term.

    "After seven days, the participants completed the treatment in their assigned groups in separate classrooms."

  • D

    The control group's size, baseline vocabulary level, and lack of treatment are clearly documented.

    "The Control Group took the posttest but did not complete a treatment."

  • S

    Randomisation occurred among four classes at a single university, not among multiple schools.

    "The participants were students majoring in English Translation at a university in China."

  • I

    The same researchers designed the materials, ran the treatments, and analysed the data, with no independent evaluator mentioned.

  • Y

    The study tracked outcomes for only about two weeks, and criterion T is also not met, so Y automatically fails.

    "After another seven days, all participants took the same checklist and multiple-choice test again as a delayed posttest."

  • B

    Exposure to the input mode is itself the explicit treatment variable being tested, so the no-treatment control condition is met by design.

    "The Control Group took the posttest but did not complete a treatment."

  • R

    No independent replication of this specific study was found; the authors state it is the first study of its kind, and a later similarly designed study by a different team does not present itself as a replication.

    "This study fills research gaps in two ways. It is the first study to compare incidental vocabulary learning through written, audio, and audiovisual input." (p. 515)

  • A

    Only vocabulary knowledge was assessed, and criterion E is not met, so A automatically fails.

  • G

    Tracking ended one week after treatment, criterion Y is not met, and no follow-up publications tracking these participants toward graduation were found.

    "After another seven days, all participants took the same checklist and multiple-choice test again as a delayed posttest."

  • P

    No pre-registration of the study protocol is mentioned anywhere in the paper, and no internet search identified a registry entry for this study.

Abstract

This study used a pretest-posttest-delayed posttest design at one-week intervals to determine the extent to which written, audio, and audiovisual L2 input contributed to incidental vocabulary learning. Seventy-six university students learning EFL in China were randomly assigned to four groups. Each group was presented with the input from the same television documentary in different modes: reading the printed transcript, listening to the documentary, viewing the documentary, and a nontreatment control condition. Checklist and multiple-choice tests were designed to measure knowledge of target words. The results showed that L2 incidental vocabulary learning occurred through reading, listening, and viewing, and that the gain was retained in all modes of input one week after encountering the input. However, no significant differences were found between the three modes on the posttests indicating that each mode of input yielded similar amounts of vocabulary gain and retention. A significant relationship was found between prior vocabulary knowledge and vocabulary learning, but not between frequency of occurrence and vocabulary learning. The study provides further support for the use of L2 television programs for language learning.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • The paper calls its own design "quasi-experimental" and only describes random assignment of students to classes by the university, not random assignment of the four classes to treatment conditions.
      • "They were in four classes that were randomly assigned by the university, with 21 second-year students being in one class and 55 third-year students divided into the other three."
      • Relevant Quotes: 1) "The research was a quasi-experimental study in an EFL context with 76 participants ranging in age from 19 to 21." (p. 506) 2) "They were in four classes that were randomly assigned by the university, with 21 second-year students being in one class and 55 third-year students divided into the other three." (p. 506) 3) "Twenty-one participants were assigned to a reading group; fifteen participants were assigned to a listening group; twenty-one participants were in a viewing group; and nineteen were assigned to a control group. Data collection took place during their class time and their preassigned classes were used as the experimental and control groups." (p. 506) Detailed Analysis: Criterion C requires a clearly described randomisation process at the class level (or stronger), not merely that pre-existing classes happened to be used as the study groups. The only randomisation described in the paper concerns how individual students were sorted into the university's four ordinary course sections ("randomly assigned by the university"), which is a routine administrative process for forming classes and is unrelated to this study's treatment assignment. No quote anywhere in the Method section states that the four pre-existing classes were themselves randomly allocated by the researchers to the four conditions (reading, listening, viewing, control); instead, the "preassigned classes were used as the experimental and control groups," which describes a convenience mapping of intact groups onto conditions rather than randomisation of the treatment assignment. This is consistent with the authors' own explicit description of the design as "quasi-experimental" rather than a randomised controlled trial. Because Criterion C requires that the assignment to conditions itself be randomised at the class level or higher, and the paper provides no evidence of this beyond an unrelated administrative process, and the tutoring exception does not apply (this is a classroom-based comparison of input modes, not one-to-one tutoring), the criterion is not met. Criterion C is not met because the paper only documents random assignment of students to pre-existing classes for ordinary course administration, not random assignment of classes to the study's treatment conditions, and the authors themselves describe the design as quasi-experimental.
    • E

      Exam-based Assessment

      • Outcome measures were custom checklist and multiple-choice tests built around 43 documentary-specific target words, not standardised exams.
      • "Checklist and multiple-choice tests were designed to measure knowledge of target words."
      • Relevant Quotes: 1) "Checklist and multiple-choice tests were designed to measure knowledge of target words." (Abstract) 2) "This paper-and-pencil test required participants to respond yes or no to indicate whether they knew the provided words. Adapted from the yes/no EFL vocabulary test designed by Meara (1992), this test consisted of 60 test items: the 43 target words, 10 words that were expected to be known, and 7 nonwords." (p. 509) 3) "This test was a prompted recognition four-choice test with the key and three distractors in the participants' L1 (Mandarin, see example in Table 3)." (p. 510) 4) "Forty-three words encountered in the transcript of the television documentary were selected as target items." (p. 509) Detailed Analysis: Both outcome instruments (the checklist test and the multiple-choice test) were purpose-built by the researchers around the 43 target words that appear in the specific documentary used as the treatment material. While the checklist format was adapted from an existing yes/no vocabulary test template (Meara, 1992), the actual test content (word items, distractors) was custom-created for this study and is not a standardised, widely recognised exam such as a national test. The Vocabulary Levels Test (VLT) used elsewhere in the study is a standardised instrument, but it served only as a baseline covariate of prior vocabulary knowledge, not as the outcome measure of learning from the treatment. Criterion E is not met because the primary outcome measures are researcher-designed tests tailored to the study's specific target words rather than standardised assessments.
    • T

      Term Duration

      • The treatment was a single session, with outcomes measured immediately and one week later, far short of one academic term.
      • "After seven days, the participants completed the treatment in their assigned groups in separate classrooms."
      • Relevant Quotes: 1) "In the first week of the study, all participants completed the VLT followed by a pretest consisting of a checklist test and then a multiple-choice test." (p. 510) 2) "After seven days, the participants completed the treatment in their assigned groups in separate classrooms." (p. 510) 3) "A posttest consisting of the same checklist and multiple-choice tests was completed by the participants immediately after the treatment." (p. 510) 4) "After another seven days, all participants took the same checklist and multiple-choice test again as a delayed posttest." (p. 510-511) Detailed Analysis: The entire study, from the pretest through the single treatment session (reading/listening/viewing one documentary) to the delayed posttest, spans only about two weeks. Outcomes were measured immediately after the one-time treatment and again one week later. This tracking window is far shorter than the one full academic term (roughly 3-4 months) required by the T criterion. Criterion T is not met because the interval from intervention start to final measurement was only about one week, well short of a full term.
    • D

      Documented Control Group

      • The control group's size, baseline vocabulary level, and lack of treatment are clearly documented.
      • "The Control Group took the posttest but did not complete a treatment."
      • Relevant Quotes: 1) "...and nineteen were assigned to a control group." (p. 506) 2) Table 1 reports the Control group's Vocabulary Levels Test scores: "N 19, M 122.32, SD 18.63." (p. 507) 3) "The Control Group took the posttest but did not complete a treatment." (p. 510) 4) Tables 4 and 5 report the Control group's checklist and multiple-choice pretest, posttest, and delayed posttest means and standard deviations (N=19). (p. 511) Detailed Analysis: The paper documents the control group's sample size (19 participants), its baseline vocabulary knowledge (VLT mean and SD, shown alongside the other three groups and confirmed statistically equivalent via one-way ANOVA, F(3,72)=2.717, p=.051), and explicitly states that this group received no treatment and only completed the pretest, immediate posttest, and delayed posttest. This level of detail allows readers to assess the comparability of the control group to the experimental groups at baseline. Criterion D is met because the control group's size, baseline vocabulary characteristics, and no-treatment condition are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred among four classes at a single university, not among multiple schools.
      • "The participants were students majoring in English Translation at a university in China."
      • Relevant Quotes: 1) "The participants were students majoring in English Translation at a university in China." (p. 506) 2) "They were in four classes that were randomly assigned by the university..." (p. 506) Detailed Analysis: The study was conducted with four classes drawn from a single university department (English Translation majors at one university in China). There is no indication that multiple schools or institutions were involved, and no school-level (or higher) randomisation is described. The unit of randomisation was the class within one institution, which is weaker than the school-level requirement. Criterion S is not met because the study involved only one educational institution with class-level, not school-level, randomisation.
    • I

      Independent Conduct

      • The same researchers designed the materials, ran the treatments, and analysed the data, with no independent evaluator mentioned.
      • Relevant Quotes: 1) "This study used a pretest-posttest-delayed posttest design at one-week intervals to determine the extent to which written, audio, and audiovisual L2 input contributed to incidental vocabulary learning." (Abstract) 2) The Method, Materials, Instruments, Procedure, and Results sections describe the same two authors (Feng & Webb) as having designed the materials, selected the target words, and administered the study, with no mention of a third-party or independent evaluation team. (pp. 506-511) Detailed Analysis: No statement anywhere in the paper indicates that data collection or analysis was performed by an independent, external team unconnected to the researchers who designed the study. The tests, target word selection, and treatments were all designed and, by all indications, administered and analysed by the same research team (Feng and Webb). There is no acknowledgment of independent evaluators, external data collectors, or blinded administrators. Criterion I is not met because there is no evidence of independent conduct separate from the researchers who designed the study.
    • Y

      Year Duration

      • The study tracked outcomes for only about two weeks, and criterion T is also not met, so Y automatically fails.
      • "After another seven days, all participants took the same checklist and multiple-choice test again as a delayed posttest."
      • Relevant Quotes: 1) "After seven days, the participants completed the treatment..." (p. 510) 2) "After another seven days, all participants took the same checklist and multiple-choice test again as a delayed posttest." (p. 510-511) Detailed Analysis: Per the ERCT standard, Y requires tracking of at least 75% of an academic year, and automatically fails if the weaker T (Term Duration) criterion is not met. Since the entire study spanned only about two weeks from pretest to delayed posttest, it falls far short of even one academic term, let alone a year. Criterion Y is not met because the total study duration was approximately two weeks, and the prerequisite T criterion was also not met.
    • B

      Balanced Control Group

      • Exposure to the input mode is itself the explicit treatment variable being tested, so the no-treatment control condition is met by design.
      • "The Control Group took the posttest but did not complete a treatment."
      • Relevant Quotes: 1) "The aim of this study was to compare vocabulary learning through reading, listening, and viewing." (p. 505) 2) "Each group was presented with the input from the same television documentary in different modes: reading the printed transcript, listening to the documentary, viewing the documentary, and a nontreatment control condition." (Abstract) 3) "The Control Group took the posttest but did not complete a treatment." (p. 510) Detailed Analysis: Applying the Criterion B decision tree: the "additional resource" here is exposure to the L2 input itself (reading/listening/viewing the documentary). This exposure is explicitly the primary treatment variable under investigation -- the study's central research question is whether, and how much, each mode of input contributes to incidental vocabulary learning relative to no input at all. The Control Group's role is to establish a no-treatment baseline (to control for test-retest effects and any vocabulary gain from simply taking the tests twice), not to serve as a "business as usual" comparison for a resource imbalance that needs to be neutralised. Because the differential exposure is the explicit, integral treatment variable being tested and is clearly framed as such, the control condition receiving no input does not create an unaddressed confound. Criterion B is met because the difference in input exposure between treatment and control groups is the explicit treatment variable under study, not an extraneous resource imbalance.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study was found; the authors state it is the first study of its kind, and a later similarly designed study by a different team does not present itself as a replication.
      • "This study fills research gaps in two ways. It is the first study to compare incidental vocabulary learning through written, audio, and audiovisual input." (p. 515)
      • Relevant Quotes: 1) "This study fills research gaps in two ways. It is the first study to compare incidental vocabulary learning through written, audio, and audiovisual input." (p. 515) 2) "Only one study has compared audiovisual input and reading-while-listening but did not include written-only and audio-only input (Neuman & Koskinen, 1992)." (p. 515) 3) From Aldukhayel (2022), "Comparing L2 Incidental Vocabulary Learning Through Viewing, Listening, and Reading," Journal of Language Teaching and Research: "Using a pretest-posttest-delayed posttest design, this study recruited 95 university EFL students who were randomly assigned to four groups. The same TV documentary was presented to each group in four different modes: viewing the documentary, listening to the documentary, reading the printed transcript, and a control condition." Feng and Webb (2020) is cited only in Aldukhayel's reference list; the article contains no explicit statement that it replicates Feng and Webb's study or directly compares its results to theirs. Detailed Analysis: An internet search (conducted 2026-07-28) for independent replications identified the related peer-reviewed study by Aldukhayel (Qassim University, Saudi Arabia, 2022), which mirrors the reading/ listening/viewing/control design with a different EFL university population (Saudi Arabic-L1 learners) and a different documentary and target-word set. However, this study does not present itself as an explicit replication of Feng and Webb's specific study (it cites Feng and Webb only as background literature and does not compare its numerical results against theirs), so it does not constitute a documented independent reproduction under the ERCT standard. No other study was found that explicitly reproduces this specific design (same materials/target words, same comparison of reading/listening/viewing/control conditions) by an independent team with a direct comparison of results to this paper. Criterion R is not met because no independent replication meeting the standard's requirements was identified.
    • A

      All-subject Exams

      • Only vocabulary knowledge was assessed, and criterion E is not met, so A automatically fails.
      • Relevant Quotes: 1) "Checklist and multiple-choice tests were designed to measure knowledge of target words." (Abstract) Detailed Analysis: The study measured only vocabulary knowledge (via custom checklist and multiple-choice tests), not performance across all main subjects. Per the ERCT standard, Criterion A also requires that Criterion E (standardised exam-based assessment) be met as a prerequisite; since E was not met, A cannot be met either. Criterion A is not met because only a single, non-standardised vocabulary measure was used and the prerequisite Criterion E was not satisfied.
    • G

      Graduation Tracking

      • Tracking ended one week after treatment, criterion Y is not met, and no follow-up publications tracking these participants toward graduation were found.
      • "After another seven days, all participants took the same checklist and multiple-choice test again as a delayed posttest."
      • Relevant Quotes: 1) "After another seven days, all participants took the same checklist and multiple-choice test again as a delayed posttest. The participants were given sufficient time for everyone to finish the tests. This was followed by a 10-minute debriefing session..." (pp. 510-512) 2) No mention of any further follow-up appears anywhere in the Discussion, Future Directions, or Conclusion sections. Detailed Analysis: Measurement stopped one week after the treatment with the delayed posttest, followed only by a debriefing session. There is no indication of any further tracking of participants, let alone tracking until they complete their degree or educational stage. An internet search for follow-up papers by Feng and/or Webb that continued tracking this same cohort of 76 EFL students toward graduation did not identify any such publication. Per the ERCT standard, Criterion G also automatically fails because the prerequisite Criterion Y (Year Duration) is not met. Criterion G is not met because no follow-up beyond one week after treatment is reported or found in subsequent publications, and the prerequisite Criterion Y was also not satisfied.
    • P

      Pre-Registered

      • No pre-registration of the study protocol is mentioned anywhere in the paper, and no internet search identified a registry entry for this study.
      • Relevant Quotes: 1) No statement referencing a study registry, registration ID, or date of pre-registration appears in the Method, Procedure, or any other section of the paper. 2) The paper does note: "The experiment in this article earned an Open Materials badge for transparent practices. The materials are available at www.iris-database.org..." (p. 499), which concerns sharing of materials, not pre-registration of hypotheses and analysis plans. Detailed Analysis: While the paper received an Open Materials badge indicating the research materials were made publicly available, this is distinct from pre-registering the study's hypotheses, methods, and planned analyses before data collection. No registry platform, registration ID, or pre-registration date is mentioned anywhere in the text. An internet search for a pre-registration record for this study (e.g., on OSF or a similar registry) did not locate any matching entry. Criterion P is not met because there is no evidence of a pre-registered protocol in the paper or in external registries.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.