The impact of input, input repetition, and task repetition on L2 lexical use and fluency in speaking

Phuong-Thao Duong, Maribel Montero Perez, Long Quoc Nguyen, Piet Desmet, Elke Peters

Published:
ERCT Check Date:
DOI: 10.14746/ssllt.29727
  • L2 languages
  • higher education
  • Asia
0
  • C

    Randomisation was performed at the individual student level within three universities, not at the class or school level, and the intervention (video input before an oral task) is not a personal tutoring intervention that would qualify for the exception.

    "The participants were randomly allocated to one of three groups: an input group (N = 29), an input repetition group (N = 32), and a no-input group (N = 29)." (p. 107)

  • E

    Outcomes were measured with a custom picture-prompted oral narrative task analysed by research tools (HD-D, TAALES, Praat), not with any widely recognised standardised exam.

    "The present study employed the same picture-prompted narrative task used in Duong et al. (2021b)." (p. 107)

  • T

    The whole study spanned only three sessions over a few days, with the repeat oral task administered two days after the immediate task, far short of one academic term.

    "The three groups repeated the same oral task after two days." (Abstract, p. 102)

  • D

    The no-input baseline group is clearly documented, with its size, condition, demographics, and baseline vocabulary and working memory scores reported and compared across groups.

    "Participants in the no-input group completed only the oral task and did not watch any videos." (p. 109)

  • S

    Randomisation was at the individual student level; no schools or institutions were randomly assigned to conditions.

    "Ninety university students learning English as a foreign language were randomly assigned to one of three groups..." (Abstract, p. 102)

  • I

    The same research team designed the intervention, delivered the sessions, transcribed, coded, and analysed all data, with no external or third-party evaluation reported.

    "Each speech was audio-recorded and then transcribed manually by the first author then verified by the third author." (p. 110)

  • Y

    With outcomes measured two days after the intervention began, the study falls vastly short of the required 75% of an academic year; criterion T is also unmet, so Y cannot be met.

    "The three groups repeated the same oral task after two days." (Abstract, p. 102)

  • B

    The extra input exposure given to the treatment groups is itself the treatment variable being tested against a deliberate no-input baseline, and the added video time (about 6-18 minutes) is minimal and integral to the design.

    "Each group received a different treatment: one group was asked to view L2 captioned videos before performing an oral task ... and one group performed an oral task without watching the videos (i.e., no-input group/baseline group)." (p. 106)

  • R

    No independent replication of this specific study by a different research team is reported in the paper or found in an external search; the authors themselves call for replication studies.

    "As the participants of both studies are Vietnamese EFL learners and did not follow a task-based language program, replication studies with different learners' backgrounds are needed." (p. 116)

  • A

    Since criterion E is not met and only L2 speaking outcomes were measured, with no assessment of other core subjects, the all-subject exams criterion fails.

    "The six dependent variables included two measures of lexical use (i.e., lexical diversity and lexical sophistication) and four measures of fluency..." (p. 106)

  • G

    Measurement ended two days after the intervention with no follow-up, so participants were not tracked to graduation; criterion Y is also unmet, which rules G out.

    "Second, because the repeat task was administered only two days after the initial task, readers should be cautious in generalizing the findings to scenarios where tasks are repeated more than once." (p. 118)

  • P

    The paper contains no mention of any pre-registration, registry, or published protocol for the study.

Abstract

The present study investigates the impact of meaningful input on L2 learners' vocabulary use and their fluency in oral performance (immediate and repeat tasks), as well as whether the effects are mediated by learners' prior vocabulary knowledge and working memory. Ninety university students learning English as a foreign language were randomly assigned to one of three groups: input (N = 29), input repetition (N = 32), and no-input (i.e., baseline group) (N = 29). The input group watched L2 videos prior to performing an immediate oral task, whereas the input repetition group watched the same videos not only before but also after the immediate oral task. The no-input group only performed the oral tasks without watching the videos. The three groups repeated the same oral task after two days. Results did not show a significant effect of task repetition, input, and input repetition on learners' lexical use and fluency. However, the fluency and lexical complexity in learners' L2 speech can be predicted by their receptive vocabulary knowledge and working memory capacity to some extent.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was performed at the individual student level within three universities, not at the class or school level, and the intervention (video input before an oral task) is not a personal tutoring intervention that would qualify for the exception.
      • "The participants were randomly allocated to one of three groups: an input group (N = 29), an input repetition group (N = 32), and a no-input group (N = 29)." (p. 107)
      • Relevant Quotes: 1) "Ninety university students learning English as a foreign language were randomly assigned to one of three groups: input (N = 29), input repetition (N = 32), and no-input (i.e., baseline group) (N = 29)." (Abstract, p. 102) 2) "The participants were randomly allocated to one of three groups: an input group (N = 29), an input repetition group (N = 32), and a no-input group (N = 29)." (p. 107) 3) "The study included 90 Vietnamese undergraduates (50 females and 40 males, 19-21 years old) who studied English as a foreign language in three different universities." (p. 106) 4) "We collected data in three one-to-one sessions." (p. 109) Detailed Analysis: The unit of randomisation was the individual student: 90 undergraduates were randomly allocated to one of three conditions. There is no mention of classes or schools being randomised. The ERCT C criterion requires randomisation of entire classes (or schools) unless the intervention is personal teaching such as one-to-one tutoring. Although data collection occurred in one-to-one sessions with the researchers, the intervention itself is exposure to captioned videos plus an oral narrative task in a lab-style experiment, not a personal tutoring or one-to-one teaching programme, so the tutoring exception does not apply. Because participants came from three different universities and were individually assigned, contamination between conditions cannot be ruled out by design. Criterion C is not met because randomisation was carried out at the individual student level and the intervention does not qualify for the personal-tutoring exception.
    • E

      Exam-based Assessment

      • Outcomes were measured with a custom picture-prompted oral narrative task analysed by research tools (HD-D, TAALES, Praat), not with any widely recognised standardised exam.
      • "The present study employed the same picture-prompted narrative task used in Duong et al. (2021b)." (p. 107)
      • Relevant Quotes: 1) "The present study employed the same picture-prompted narrative task used in Duong et al. (2021b). In this task, learners had to manipulate and reorganize the information processed from the videos to create their own story in our study." (p. 107) 2) "The six dependent variables included two measures of lexical use (i.e., lexical diversity and lexical sophistication) and four measures of fluency (i.e., frequency of pauses, between clauses and within clauses, frequency of total repairs, articulation rate)." (p. 106) 3) "Each speech was audio-recorded and then transcribed manually by the first author then verified by the third author." (p. 110) 4) "Receptive vocabulary knowledge was assessed using the English-Vietnamese Vocabulary Size Test by Nguyen and Nation (2011)..." (pp. 107-108) Detailed Analysis: The primary outcomes (lexical diversity, lexical sophistication, and four fluency measures) were derived from recordings of a researcher-designed picture-prompted narrative task, analysed with research software (TAALES, TAALED, Praat). This is a custom, study-specific assessment, not a standardised, widely recognised exam. The Vocabulary Size Test and Productive Levels Test are research instruments used only as covariates (prior knowledge), not as outcome exams, and TOEIC scores are reported only to describe participants' proficiency, not as an outcome measure. Criterion E is not met because outcomes were assessed with a custom oral narrative task rather than a standardised exam.
    • T

      Term Duration

      • The whole study spanned only three sessions over a few days, with the repeat oral task administered two days after the immediate task, far short of one academic term.
      • "The three groups repeated the same oral task after two days." (Abstract, p. 102)
      • Relevant Quotes: 1) "The three groups repeated the same oral task after two days." (Abstract, p. 102) 2) "We collected data in three one-to-one sessions. In the first session, the input and input repetition group were asked to watch the videos twice on a computer screen (without taking notes) and then perform the immediate oral task." (p. 109) 3) "Two days later, in the second session, all three groups did not watch the videos but were asked to complete the same oral task in three minutes." (p. 109) 4) "Two prior vocabulary knowledge tests were completed in the third session with a 15-minute break in between." (p. 110) Detailed Analysis: The ERCT T criterion requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Here the intervention consisted of a single session of video viewing and an immediate oral task, with the final outcome measurement (the repeat oral task) taken only two days later. The entire data collection covered three short sessions within days, orders of magnitude shorter than an academic term, and no longer-term follow-up is reported. Criterion T is not met because the interval from intervention start to outcome measurement was only two days.
    • D

      Documented Control Group

      • The no-input baseline group is clearly documented, with its size, condition, demographics, and baseline vocabulary and working memory scores reported and compared across groups.
      • "Participants in the no-input group completed only the oral task and did not watch any videos." (p. 109)
      • Relevant Quotes: 1) "The participants were randomly allocated to one of three groups: an input group (N = 29), an input repetition group (N = 32), and a no-input group (N = 29)." (p. 107) 2) "Participants in the no-input group completed only the oral task and did not watch any videos." (p. 109) 3) "The study included 90 Vietnamese undergraduates (50 females and 40 males, 19-21 years old) who studied English as a foreign language in three different universities... Their English proficiency ranged from 340 to 750 points (M = 501.24, SD = 93.16) on the Test of English for International Communication (TOEIC), corresponding to A2-B1 level." (pp. 106-107) 4) "Table 1 reveals that participants in all three groups performed similarly in the receptive vocabulary test and productive levels test. An ANOVA showed that the differences were not significant: receptive test (F(2, 87) = .195 , p = .823) and productive test (F(2, 87) = 2.784, p = .067)." (p. 112) 5) "Participants were shown to have good and comparable scores on the working memory tests (see Table 2)." (p. 112) Detailed Analysis: The control (no-input/baseline) group is explicitly defined: its size (N = 29), what it did (performed the oral tasks only, without any video input), and its baseline characteristics. Tables 1 and 2 report baseline receptive and productive vocabulary scores and three working memory scores per group with statistical comparisons showing group equivalence, and overall demographics (age, gender, nationality, proficiency range) are described. This provides adequate documentation to judge comparability of the control group. Criterion D is met because the baseline group's size, condition, and baseline performance are clearly documented and compared with the treatment groups.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was at the individual student level; no schools or institutions were randomly assigned to conditions.
      • "Ninety university students learning English as a foreign language were randomly assigned to one of three groups..." (Abstract, p. 102)
      • Relevant Quotes: 1) "Ninety university students learning English as a foreign language were randomly assigned to one of three groups: input (N = 29), input repetition (N = 32), and no-input (i.e., baseline group) (N = 29)." (Abstract, p. 102) 2) "The study included 90 Vietnamese undergraduates ... who studied English as a foreign language in three different universities." (p. 106) Detailed Analysis: The S criterion requires random assignment of whole schools or comparable institutional units. In this study, individual students recruited from three universities were randomised to conditions; the universities themselves were not units of randomisation, and no school-level assignment procedure is described anywhere in the paper. Criterion S is not met because randomisation occurred at the individual student level, not the school level.
    • I

      Independent Conduct

      • The same research team designed the intervention, delivered the sessions, transcribed, coded, and analysed all data, with no external or third-party evaluation reported.
      • "Each speech was audio-recorded and then transcribed manually by the first author then verified by the third author." (p. 110)
      • Relevant Quotes: 1) "Each speech was audio-recorded and then transcribed manually by the first author then verified by the third author." (p. 110) 2) "The first and third authors performed independent coding. 20% of the transcripts were coded first. High interrater reliability was established on the five variables..." (p. 110) 3) "We used short English videos (total time = 6 minutes) with captions (i.e., subtitles in English), which were first introduced in Duong et al. (2021a)." (p. 107) 4) "Following Duong et al. (2021b), this study employed a between-subjects design..." (p. 106) Detailed Analysis: The intervention materials (videos, oral task, tests) were developed or selected by the authors themselves, building on their own earlier studies (Duong et al., 2021a, 2021b), and the authors personally collected, transcribed, coded, and analysed the data (first and third authors did the transcription and coding). There is no mention of an external evaluation team, independent enumerators, or third-party oversight anywhere in the paper. The "independent coding" by the first and third authors refers to interrater reliability within the author team, not independence from the intervention designers. Criterion I is not met because the intervention designers themselves conducted, coded, and analysed the study without independent oversight.
    • Y

      Year Duration

      • With outcomes measured two days after the intervention began, the study falls vastly short of the required 75% of an academic year; criterion T is also unmet, so Y cannot be met.
      • "The three groups repeated the same oral task after two days." (Abstract, p. 102)
      • Relevant Quotes: 1) "The three groups repeated the same oral task after two days." (Abstract, p. 102) 2) "Two days later, in the second session, all three groups did not watch the videos but were asked to complete the same oral task in three minutes." (p. 109) Detailed Analysis: The Y criterion requires outcome tracking covering at least 75% of a full academic year (~9-10 months) from intervention start. This study's entire timeline, from first video viewing to the final repeat oral task, spans two days, plus a third short session for covariate tests. Since the weaker T (term duration) criterion is already not met, Y is automatically not met as well. Criterion Y is not met because the whole study lasted only a few days, nowhere near an academic year.
    • B

      Balanced Control Group

      • The extra input exposure given to the treatment groups is itself the treatment variable being tested against a deliberate no-input baseline, and the added video time (about 6-18 minutes) is minimal and integral to the design.
      • "Each group received a different treatment: one group was asked to view L2 captioned videos before performing an oral task ... and one group performed an oral task without watching the videos (i.e., no-input group/baseline group)." (p. 106)
      • Relevant Quotes: 1) "Following Duong et al. (2021b), this study employed a between-subjects design which involves two independent variables: input (the between-subjects variable) at three levels (no-input vs. input vs. input repetition) and task repetition (the within-subjects variable)." (p. 106) 2) "Each group received a different treatment: one group was asked to view L2 captioned videos before performing an oral task (i.e., input group), another group was asked to view the same videos before and after performing the oral task (i.e. input repetition group), and one group performed an oral task without watching the videos (i.e., no-input group/baseline group)." (p. 106) 3) "We used short English videos (total time = 6 minutes) with captions (i.e., subtitles in English)..." (p. 107) 4) "RQ1: To what extent does input exposure affect L2 learners' lexical use and fluency in an immediate and a repeat oral task?" (p. 106) Detailed Analysis: Applying the criterion B decision procedure: extra resources are present (the input and input repetition groups watch 6-minute captioned videos once or twice, i.e. roughly 6-18 minutes total, that the no-input group does not receive), so the "no extra resources" met-by-default branch does not apply. However, this extra exposure IS the explicit treatment variable under test - the study's central research question (RQ1) asks "to what extent does input exposure affect L2 learners' lexical use and fluency," with the no-input condition expressly labelled the baseline/control group by design. Per the decision tree, when the additional resource is the primary treatment variable being tested, a business-as-usual (here, no-input) control is acceptable and the criterion is met, without needing the control group to be given a matching activity. All other inputs - the oral task itself, planning time, and testing procedures - are identical across the three groups, reinforcing that the video input is the sole manipulated variable. Criterion B is met because the additional video input is the explicit treatment variable being tested against a designed no-input baseline, with all other inputs identical across groups.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study by a different research team is reported in the paper or found in an external search; the authors themselves call for replication studies.
      • "As the participants of both studies are Vietnamese EFL learners and did not follow a task-based language program, replication studies with different learners' backgrounds are needed." (p. 116)
      • Relevant Quotes: 1) "As the participants of both studies are Vietnamese EFL learners and did not follow a task-based language program, replication studies with different learners' backgrounds are needed." (p. 116) 2) "While Nguyen and Boers's study is the only one that has suggested that the input-output-cycle could foster vocabulary acquisition, we argue that this technique might help learners perform better in lexical use and fluency in the repeat oral performance... However, this hypothesis has yet to be verified empirically." (p. 104) 3) "To the best of our knowledge, no study has looked into whether task repetition combined with meaningful input exposure can aid the development of lexical use and fluency." (p. 104) Detailed Analysis: The paper positions itself as the first study of its kind, and the authors explicitly call for replication with other learner populations. Related prior studies cited (e.g., Duong et al., 2021b; Duong & Le, 2022) are by overlapping author teams and are precursors, not replications of this study. A search of the citing literature (via Semantic Scholar, checked 2026-07-27) for this article (DOI 10.14746/ssllt.29727) found only four citing works: Alanazi (2026) on listening tasks, Vanbuel and Bugge (2026) on citizenship testing, Abdi Tabari et al. (2025) "Task repetition and L2 oral performance: A meta-analysis," and Al-Abri et al. (2024) on MALL-based lexical fluency. Abdi Tabari et al. (2025) is a meta-analytic synthesis that may include this study as one data point rather than an independent replication of its specific design (three Vietnamese-university input/input-repetition conditions with a two-day repeat task); none of the four citing works constitute an independent, peer-reviewed replication of this study by a different research team in a different context. Criterion R is not met because no independent published replication of this study exists, either referenced in the paper or located through a citation search.
    • A

      All-subject Exams

      • Since criterion E is not met and only L2 speaking outcomes were measured, with no assessment of other core subjects, the all-subject exams criterion fails.
      • "The six dependent variables included two measures of lexical use (i.e., lexical diversity and lexical sophistication) and four measures of fluency..." (p. 106)
      • Relevant Quotes: 1) "The six dependent variables included two measures of lexical use (i.e., lexical diversity and lexical sophistication) and four measures of fluency (i.e., frequency of pauses, between clauses and within clauses, frequency of total repairs, articulation rate)." (p. 106) Detailed Analysis: Criterion A requires standardised exam-based assessment across all main subjects. As a prerequisite, criterion E must be met, but it is not: the outcomes come from a custom oral narrative task. Moreover, only English speaking performance (lexical use and fluency) was measured; no other subjects of the university curriculum were assessed, and no justification for a specialised-intervention exception involving standardised exams is given. Criterion A is not met because criterion E fails and only a single custom-measured outcome domain (L2 speaking) was assessed.
    • G

      Graduation Tracking

      • Measurement ended two days after the intervention with no follow-up, so participants were not tracked to graduation; criterion Y is also unmet, which rules G out.
      • "Second, because the repeat task was administered only two days after the initial task, readers should be cautious in generalizing the findings to scenarios where tasks are repeated more than once." (p. 118)
      • Relevant Quotes: 1) "The three groups repeated the same oral task after two days." (Abstract, p. 102) 2) "Second, because the repeat task was administered only two days after the initial task, readers should be cautious in generalizing the findings to scenarios where tasks are repeated more than once." (p. 118) Detailed Analysis: Criterion G requires tracking participants until graduation from their educational stage. Here, all outcome measurement concluded two days after the intervention, with a third session only for covariate testing. There is no mention of any longer-term follow-up or of tracking the undergraduates through to degree completion. A search of the citing literature for this article (via Semantic Scholar, checked 2026-07-27) and of the authors' other publications (Duong et al., 2021a, 2021b; Duong & Le, 2022) found no subsequent paper tracking this same 90-participant cohort toward graduation; the related Duong and Le (2022) study uses a different, smaller sample (40 Vietnamese university students) and is not a follow-up of this cohort. In addition, since the prerequisite criterion Y is not met, G cannot be met. Criterion G is not met because no tracking beyond two days after the intervention is reported, and no follow-up publication tracking this cohort to graduation was found.
    • P

      Pre-Registered

      • The paper contains no mention of any pre-registration, registry, or published protocol for the study.
      • Relevant Quotes: No relevant quotes: the paper contains no statement about trial registration, a registry platform (e.g., OSF, ClinicalTrials.gov, AsPredicted), or a pre-registered protocol or analysis plan anywhere in the text. Detailed Analysis: Criterion P requires that the full study protocol, including hypotheses, methods, and planned analyses, be publicly registered before data collection began. The methodology and statistical analysis sections (pp. 106-111) describe the design and analyses but never reference a registration ID, registry link, or protocol publication, and no registration statement appears elsewhere in the article. An internet search (checked 2026-07-27) for a matching pre-registration by these authors for this study did not surface any entry on OSF, AsPredicted, or a similar registry. Criterion P is not met because no pre-registration of the study is mentioned in the paper or found through an internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.