Video-Based Interaction, Negotiation for Comprehensibility, and Second Language Speech Learning: A Longitudinal Study

Kazuya Saito and Yuka Akiyama

Published:
ERCT Check Date:
DOI: 10.1111/lang.12184
  • L2 languages
  • higher education
  • Asia
  • US
0
  • C

    Participants self-selected into groups based on their stated interest in flyers rather than being randomly assigned, so this is not a randomised trial at any level.

    "We assigned the students interested in the former to the experimental group ... and those interested in the latter to the comparison group ..." (p. 50)

  • E

    Outcomes were measured with a custom picture-description production task and bespoke rating scales rather than a standardised, widely recognised exam.

    "In this study, a timed picture description was adopted from Saito (2015) to measure the participants' pronunciation, fluency, vocabulary, and grammar performance during spontaneous speech." (p. 55)

  • T

    The study tracked participants from pretest through a full academic semester to posttest, exceeding the one-term minimum.

    "This study examined the impact of video-based conversational interaction on the longitudinal development (one academic semester) of second language production ..." (p. 43)

  • D

    The comparison group's size, demographics, activities, and baseline equivalence to the experimental group are explicitly documented.

    "The remaining 15 Japanese NNS learners formed the comparison group and participated in weekly individual vocabulary/grammar activities instead of task-based interaction activities with NSs." (p. 54)

  • S

    There was no school-level randomisation; a single university's students self-selected into condition based on stated interest.

    "We assigned the students interested in the former to the experimental group ... and those interested in the latter to the comparison group ..." (p. 50)

  • I

    The same authors who designed the intervention also recruited participants, trained interlocutors, and administered the study, with no independent evaluator overseeing the trial as a whole.

    "During the orientation in Week 2, NSs received training from the researcher on how to negotiate for comprehensibility ..." (p. 53)

  • Y

    The study tracked participants for only about one academic semester (roughly 12 weeks), well short of the year-long duration required.

    "This study examined the impact of video-based conversational interaction on the longitudinal development (one academic semester) ..." (p. 43)

  • B

    The comparison group received a time-matched (30 minutes/week) alternative educational activity, while the tested resource (NS interaction) was the explicit treatment variable, so time/effort was balanced.

    "The weekly assignments, which typically took 30 minutes to complete at home, were graded and recorded by the researcher." (p. 55)

  • R

    No independent replication of this specific study by a different research team could be identified in the paper or via external search.

  • A

    Criterion A cannot be met because criterion E (standardised exam) is not met, and the study only assessed L2 English speech, not other subjects.

  • G

    Criterion G cannot be met because criterion Y is not met, and no follow-up study tracking this cohort toward graduation was found, either in the paper or via external search.

  • P

    The paper contains no reference to a pre-registered protocol, hypotheses, or analysis plan published before data collection began, and none was found via external search.

Abstract

This study examined the impact of video-based conversational interaction on the longitudinal development (one academic semester) of second language production by college-level Japanese English-as-a-foreign-language learners. Students in the experimental group engaged in weekly dyadic conversation exchanges with native speakers in the United States via telecommunication tools. The native speaker interlocutors were trained to provide interactional feedback (recasts) when the nonnative speakers' utterances hindered successful understanding (i.e., negotiation for comprehensibility). The students in the comparison group received regular foreign language instruction without any interaction with native speakers. The coded video data showed that the experimental students worked on improving all linguistic domains of language, likely in response to their native speaker interlocutors' interactional feedback (recasts, negotiation) during the treatment. The pretest-posttest data of the students' spontaneous production showed that they made significant gains in the dimensions of comprehensibility, fluency, and lexicogrammar but not in those of accentedness and pronunciation.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Participants self-selected into groups based on their stated interest in flyers rather than being randomly assigned, so this is not a randomised trial at any level.
      • "We assigned the students interested in the former to the experimental group ... and those interested in the latter to the comparison group ..." (p. 50)
      • Relevant Quotes: 1) "We recruited participants by distributing two different flyers, one offering conversational activities and the other offering vocabulary/grammar activities. We assigned the students interested in the former to the experimental group (10 males, 5 females; Mage = 19.2 years) and those interested in the latter to the comparison group (6 males, 9 females; Mage = 18.9 years)." (p. 50) 2) "Thirty first- and second-year Japanese undergraduate students majoring in business at a university in Japan participated as volunteers." (p. 50) Detailed Analysis: Criterion C requires a genuine Randomised Controlled Trial with randomisation clearly described at the class level (or a stronger school level), or, under the tutoring exception, at minimum true random assignment at the student level. Here, the paper explicitly describes self-selection: students who expressed interest in "conversational activities" (via flyer) were placed in the experimental group, and those interested in "vocabulary/grammar activities" were placed in the comparison group. This is volunteer self-selection based on stated preference, not random assignment by any investigator- controlled randomisation procedure. No quote anywhere in the Method section describes a random number generator, coin flip, or any randomisation mechanism for allocating participants to condition. Because there is no randomisation at all (let alone at the class or school level), the tutoring exception (which still requires actual randomisation, just permits it at the student level) does not rescue this design either. Criterion C is not met because group assignment was based on participants' self-selected interest (via flyer) rather than any random allocation procedure.
    • E

      Exam-based Assessment

      • Outcomes were measured with a custom picture-description production task and bespoke rating scales rather than a standardised, widely recognised exam.
      • "In this study, a timed picture description was adopted from Saito (2015) to measure the participants' pronunciation, fluency, vocabulary, and grammar performance during spontaneous speech." (p. 55)
      • Relevant Quotes: 1) "In this study, a timed picture description was adopted from Saito (2015) to measure the participants' pronunciation, fluency, vocabulary, and grammar performance during spontaneous speech." (p. 55) 2) "The three picture descriptions were combined and stored as a single wav file for each speaker at the pretest and posttest sessions, resulting in 60 speech samples (30 NNSs × 2 tests)." (p. 56) 3) "[Raters] used a free-moving slider on a computer screen based on a 1,000-point scale to evaluate comprehensibility ... and accentedness ..." (p. 57) 4) "the purpose of the project was to improve their L2 vocabulary and grammar ability with the goal of attaining higher scores on the Test of English for International Communication (TOEIC)." (pp. 54-55) Detailed Analysis: The primary outcome measures in this study are a custom timed picture-description production task (developed by the authors, adapted from one of the authors' own earlier work) scored by novice and expert raters on bespoke 1,000-point sliders for constructs such as comprehensibility, accentedness, segmentals, word stress, intonation, and speech rate, plus researcher-coded lexicogrammar measures. None of these are widely recognised, standardised exams; they are study-specific instruments designed by the research team to capture fine-grained phonetic and linguistic detail. The only standardised, widely recognised exam mentioned (TOEIC) was referenced only as motivational framing for the comparison group's homework, not as an outcome measure actually administered or analysed in the study's results. Criterion E is not met because the outcome measures are custom-designed production and rating instruments, not standardised, widely-recognised exams.
    • T

      Term Duration

      • The study tracked participants from pretest through a full academic semester to posttest, exceeding the one-term minimum.
      • "This study examined the impact of video-based conversational interaction on the longitudinal development (one academic semester) of second language production ..." (p. 43)
      • Relevant Quotes: 1) "This study examined the impact of video-based conversational interaction on the longitudinal development (one academic semester) of second language production ..." (p. 43, Abstract) 2) "After the participants took the pretests in Week 1, they joined weekly 30-minute extracurricular L2 activities outside of their regular EFL syllabus ... between Weeks 2 and 11 ... One week after the last session (in Week 12), the participants took the posttests." (p. 50) Detailed Analysis: The intervention (and comparison activity) began after the Week 1 pretest, with weekly sessions running through Week 11, and the final outcome measurement (posttest) occurring in Week 12. This spans roughly 11-12 weeks (one full academic semester), which comfortably exceeds the minimum one-term (approximately 3-4 months) requirement for criterion T. This criterion concerns duration of tracking, and is independent of whether the design was actually randomised. Criterion T is met because the interval from intervention start (Week 2) to final measurement (Week 12) spans one full academic semester, which is at least one academic term.
    • D

      Documented Control Group

      • The comparison group's size, demographics, activities, and baseline equivalence to the experimental group are explicitly documented.
      • "The remaining 15 Japanese NNS learners formed the comparison group and participated in weekly individual vocabulary/grammar activities instead of task-based interaction activities with NSs." (p. 54)
      • Relevant Quotes: 1) "The remaining 15 Japanese NNS learners formed the comparison group and participated in weekly individual vocabulary/grammar activities instead of task-based interaction activities with NSs." (p. 54) 2) "During the orientation (Week 2), the 15 Japanese students in the comparison group were explicitly told that the purpose of the project was to improve their L2 vocabulary and grammar ability with the goal of attaining higher scores on the Test of English for International Communication (TOEIC) ... the NNSs were asked to practice using a variety of vocabulary and grammar activities, which consisted of vocabulary recall tests based on JACET 8000 ... and fill-in-the-blank grammar questions in Part 5 of the TOEIC ... The weekly assignments, which typically took 30 minutes to complete at home, were graded and recorded by the researcher." (pp. 54-55) 3) "We assigned the students interested in the former to the experimental group (10 males, 5 females; Mage = 19.2 years) and those interested in the latter to the comparison group (6 males, 9 females; Mage = 18.9 years)." (p. 50) 4) "Prior to the project, the two groups were comparable in their global production ... pronunciation ... fluency ... vocabulary ... and grammar domains ..." (p. 62) Detailed Analysis: The paper documents the comparison group's size (n=15), demographics (gender split, mean age), the specific weekly activities they completed (JACET 8000 vocabulary recall and TOEIC-style grammar fill-in-the-blank exercises), the time commitment (30 minutes/week), and confirms via baseline Mann-Whitney tests that the two groups were statistically comparable at pretest across all outcome domains. This constitutes clear, detailed documentation of the control/ comparison group's composition, activities, and baseline equivalence. Criterion D is met because the comparison group's size, demographics, baseline scores, and weekly activities are explicitly and clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • There was no school-level randomisation; a single university's students self-selected into condition based on stated interest.
      • "We assigned the students interested in the former to the experimental group ... and those interested in the latter to the comparison group ..." (p. 50)
      • Relevant Quotes: 1) "We recruited participants by distributing two different flyers, one offering conversational activities and the other offering vocabulary/grammar activities. We assigned the students interested in the former to the experimental group ... and those interested in the latter to the comparison group ..." (p. 50) 2) "Thirty first- and second-year Japanese undergraduate students majoring in business at a university in Japan participated as volunteers." (p. 50) Detailed Analysis: Criterion S requires randomisation at the school (or equivalent institutional) level. This study drew all 30 participants from a single university and assigned them to condition based on their self-reported interest in one of two flyers, not by randomising schools, classes, or even individual students through any formal allocation procedure. There is no mention of multiple schools/institutions being involved or randomised. Criterion S is not met because there was no school-level (or any level) randomisation; participants self-selected their condition, and all participants came from a single institution.
    • I

      Independent Conduct

      • The same authors who designed the intervention also recruited participants, trained interlocutors, and administered the study, with no independent evaluator overseeing the trial as a whole.
      • "During the orientation in Week 2, NSs received training from the researcher on how to negotiate for comprehensibility ..." (p. 53)
      • Relevant Quotes: 1) "During the orientation in Week 2, NSs received training from the researcher on how to negotiate for comprehensibility ..." (p. 53) 2) "The NNSs were required to report to the researcher the date and time of each session." (p. 52) 3) "The weekly assignments ... were graded and recorded by the researcher." (p. 55) 4) "To explore the nature of communicative focus on form during the semester-long L2 interaction activities, a linguistically trained coder watched the video-recorded interactions ..." (p. 53) Detailed Analysis: Throughout the Method section, the study authors (Saito and Akiyama, referred to as "the researcher(s)") are described as directly designing the intervention, training the NS interlocutors, recruiting and orienting participants, administering and grading the comparison group's weekly assignments, and supervising data collection. While independent novice and expert raters were recruited to score the anonymised, randomised speech samples (a safeguard against rating bias), and a "linguistically trained coder" analysed the interaction data, there is no statement that the trial's design, implementation, or overall data collection was conducted by a body independent of the intervention's designers. The same research team that created the task-based interaction protocol and recast training also ran the entire study end to end. Criterion I is not met because the study was designed, implemented, and administered by the same research team throughout, with no independent third-party conducting the trial or overseeing data collection as a whole.
    • Y

      Year Duration

      • The study tracked participants for only about one academic semester (roughly 12 weeks), well short of the year-long duration required.
      • "This study examined the impact of video-based conversational interaction on the longitudinal development (one academic semester) ..." (p. 43)
      • Relevant Quotes: 1) "After the participants took the pretests in Week 1, they joined weekly 30-minute extracurricular L2 activities ... between Weeks 2 and 11 ... One week after the last session (in Week 12), the participants took the posttests." (p. 50) 2) "This study examined the impact of video-based conversational interaction on the longitudinal development (one academic semester) ..." (p. 43) Detailed Analysis: The entire study, from pretest to posttest, spanned only about 12 weeks (one academic semester), not the roughly 9-10 month academic year (or at least 75% of it, approximately 7 months) required by criterion Y. There is no indication in the paper of any tracking beyond Week 12. Criterion Y is not met because outcomes were measured after only about 12 weeks (one semester), far short of 75% of an academic year.
    • B

      Balanced Control Group

      • The comparison group received a time-matched (30 minutes/week) alternative educational activity, while the tested resource (NS interaction) was the explicit treatment variable, so time/effort was balanced.
      • "The weekly assignments, which typically took 30 minutes to complete at home, were graded and recorded by the researcher." (p. 55)
      • Relevant Quotes: 1) "Fifteen of the Japanese L2 learners constituted the experimental group and participated in dyadic interaction with NS interlocutors ... over one academic semester (nine sessions in total). In each 60-minute session, the participants interacted with each other in English for the first half of the session ... The methodology and results for the 30 minutes of interaction in English ... are reported here." (p. 51) 2) "While the Japanese students in the comparison group did vocabulary/grammar exercise activities, those in the experimental group engaged in task-based conversation activities with their NS partners ..." (p. 50) 3) "Between Weeks 3 and 11, the NNSs were asked to practice using a variety of vocabulary and grammar activities ... The weekly assignments, which typically took 30 minutes to complete at home, were graded and recorded by the researcher." (p. 55) Detailed Analysis: Applying the decision tree: additional resources are present (the experimental group received access to trained NS conversation partners via video-conferencing technology, which the comparison group did not). However, this extra resource (NS interaction) is explicitly the treatment variable under investigation -- the study's entire purpose is to test whether video-based negotiated interaction with NSs (versus an alternative-content activity) improves L2 oral ability. Critically, the comparison group was not left idle ("business as usual" only); it received a matched weekly time commitment (30 minutes/week, over the same 9-week window) of an alternative structured educational activity (vocabulary/grammar exercises). This means the two conditions were balanced in overall educational engagement time even though the specific resource being tested (NS interaction) was, by design, exclusive to the experimental group. Criterion B is met because the additional resource under investigation (NS interaction) is the explicit treatment variable, and the comparison group received a time-matched alternative educational activity (30 minutes/week) rather than no activity at all.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study by a different research team could be identified in the paper or via external search.
      • Relevant Quotes: 1) No quotes in the paper reference an independent replication of this specific study by a different research team. 2) "In the other part of the project, the participants reversed roles ... The details of this project are reported in another manuscript." (Note 3, p. 69) -- this refers to a companion analysis by the same authors, not an independent replication. Detailed Analysis: An internet search for independent replications of this specific study (Saito & Akiyama, 2017, "Video-Based Interaction, Negotiation for Comprehensibility, and Second Language Speech Learning") did not identify any peer-reviewed study by a different research team that independently reproduced this exact design (video-based dyadic NS-NNS interaction versus vocabulary/grammar comparison group, with the same production-test battery) in a different context. Searches of Google Scholar's citing-literature list surfaced papers that cite this study (e.g., Young, 2018, explicitly noting that "replication of this study may... lead to different results" without itself being a replication; Canals et al., 2025, on video-based CMC corrective feedback, addressing instructor feedback rather than dyadic negotiation; Gokgoz-Kurt, 2023, on fluency and visual/acoustic cues; and Isbell, Dixon, & Plonsky, 2022, discussing the study within a research-methods handbook), none of which constitute an independent replication. Related follow-on work overlapping in authorship with Akiyama -- e.g., a longitudinal study titled "Effects of Video-Based Interaction on the Development of Second Language Listening Comprehension Ability," and "Learner beliefs and corrective feedback in telecollaboration" (cited in the paper itself as Akiyama, in press) -- does not constitute independent replication because it shares an author with the original study. Criterion R is not met because no independent replication of this specific study by a different research team was found.
    • A

      All-subject Exams

      • Criterion A cannot be met because criterion E (standardised exam) is not met, and the study only assessed L2 English speech, not other subjects.
      • Relevant Quotes: 1) "In this study, a timed picture description was adopted from Saito (2015) to measure the participants' pronunciation, fluency, vocabulary, and grammar performance during spontaneous speech." (p. 55) Detailed Analysis: Per the specification, criterion A requires criterion E (standardised exam-based assessment) to be met as a prerequisite; since E was not met (the outcomes were assessed via a custom picture-description task and bespoke rating scales, not a standardised exam), criterion A automatically fails regardless of subject coverage. Separately, the study also only measures L2 English speech outcomes and does not assess any other academic subject. Criterion A is not met because criterion E is not met, and additionally only L2 English speech outcomes (not other main subjects) were assessed.
    • G

      Graduation Tracking

      • Criterion G cannot be met because criterion Y is not met, and no follow-up study tracking this cohort toward graduation was found, either in the paper or via external search.
      • Relevant Quotes: 1) "One week after the last session (in Week 12), the participants took the posttests." (p. 50) 2) No mention anywhere in the paper of any follow-up beyond the Week 12 posttest, or of tracking participants toward graduation. Detailed Analysis: Per the specification, criterion G requires criterion Y (year duration) to be met as a prerequisite; since Y was not met (tracking stopped after about 12 weeks), criterion G automatically fails. An internet search for follow-up publications by Saito and/or Akiyama tracking this same cohort of 30 Japanese undergraduates toward graduation did not surface any such study. Related follow-on publications sharing authorship with this paper (e.g., a longitudinal study on video-based interaction and L2 listening comprehension, and "Learner beliefs and corrective feedback in telecollaboration," Akiyama, in press) pursue different research questions (listening comprehension development, learner beliefs) rather than graduation tracking of this cohort, so they do not satisfy this criterion. Criterion G is not met because criterion Y is not met, and no follow-up study tracking this cohort toward graduation was found in the paper or via external search.
    • P

      Pre-Registered

      • The paper contains no reference to a pre-registered protocol, hypotheses, or analysis plan published before data collection began, and none was found via external search.
      • Relevant Quotes: 1) No quotes anywhere in the paper reference a pre-registration platform, registry ID, or registration date for the study protocol. Detailed Analysis: A thorough review of the Method, Results, and supplementary materials descriptions in the paper (Appendices S1-S3, covering error correction scripts and rating training materials) reveals no mention of a registered protocol, hypotheses, or analysis plan published before data collection began (Week 1, pretests). The paper does carry an "Open Materials" badge for the IRIS repository, but this concerns post hoc sharing of study materials, not pre-registration of the protocol before the study began. An internet search of the Open Science Framework (OSF) and general web sources for a pre-registration by Kazuya Saito or Yuka Akiyama relating to this study did not surface any registry entry. Criterion P is not met because no pre-registration of the study protocol is referenced anywhere in the paper or found via external search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.