Optimizing Fluency Training for Speaking Skills Transfer: Comparing the Effects of Blocked and Interleaved Task Repetition

Yuichi Suzuki

Published:
ERCT Check Date:
DOI: 10.1111/lang.12433
  • L2 languages
  • higher education
  • Asia
0
  • C

    Individual students were randomly assigned to the blocked or interleaved condition at a single university, not entire classes or schools, and the tutoring/personal-teaching exception does not cleanly apply to this self-administered practice program.

    "They were randomly assigned to either a blocked-practice condition (n = 24) or an interleaved-practice condition (n = 26)."

  • E

    Fluency was measured via researcher-elicited oral narrations of six-frame cartoons, analyzed acoustically with PRAAT, not via a standardized exam-based assessment.

    "A different set of two prompts-street (Oba, 2018) and airport (Suzuki, 2011)-was used for the pretest and posttest."

  • T

    Training lasted only 3 consecutive days and the posttest occurred just 1 day after the final training session, far short of a full academic term.

    "All the participants returned to the computer lab for the posttest 1 day after the last training session."

  • D

    The control group's size, recruitment source, and baseline proficiency scores are explicitly reported and statistically compared with the two training groups.

    "The learners in the control group (n = 18), who took the pretest and posttest only, were recruited from the same classes as those assigned to the two experimental groups."

  • S

    Randomization occurred among individual students at a single university, not among schools.

    "They were randomly assigned to either a blocked-practice condition (n = 24) or an interleaved-practice condition (n = 26)."

  • I

    The single author designed, conducted, and analyzed the entire study, assisted only by research assistants and coders under his direction, with no independent third-party evaluator.

    "I am grateful to Atsushi Miura, Misaki Kuratsubo, Miyu Koyama, Taeko Hosaka, and Kazuma Arai for their dedicated assistance in data collection and coding."

  • Y

    Because criterion T (Term Duration) is not met, this stronger criterion cannot be met either; in any case the study lasted only about a week.

    "All the participants returned to the computer lab for the posttest 1 day after the last training session."

  • B

    The extra practice time given to the two training groups relative to the control group is itself the explicit treatment variable under investigation (comparing schedules of the same total practice amount), and the two training conditions received identical amounts of time and material.

    "the participants assigned to the blocked-practice condition performed the same narrative three times a day... whereas those in the interleaved-practice condition performed three different narratives on each of the three days."

  • R

    No independent replication of this specific 3-day, cartoon-narration blocked-versus-interleaved design was found; a related but distinct 2023 study by a different team used different tasks, a longer duration, and reported partly contrary results.

  • A

    Because criterion E (Exam-based Assessment) is not met, this stronger criterion cannot be met either; moreover, only speaking fluency was measured, not other subjects.

    "the study reported here focused on utterance fluency"

  • G

    Because criterion Y (Year Duration) is not met, this stronger criterion cannot be met; the study tracked participants for about a week with no long-term or graduation follow-up.

    "All the participants returned to the computer lab for the posttest 1 day after the last training session."

  • P

    The paper contains no statement of a pre-registered protocol, registry ID, or registration date.

Abstract

In an exploration of the effects of task-repetition practice on fluency development, English-as-a-foreign language learners performed three oral narrative tasks involving six-frame cartoons for 3 consecutive days. They engaged in task-repetition practice under either a blocked (Day 1: A-A-A; Day 2: B-B-B; Day 3: C-C-C) or an interleaved (Day 1: A-B-C; Day 2: A-B-C; Day 3: A-B-C) task repetition schedule. The results yielded by a posttest involving new six-frame cartoons indicated that blocked practice resulted in greater fluency development (faster articulation rate and shorter mid-clause pause duration) than did interleaved practice. Moreover, the learners in the blocked-practice group tended to pause more frequently at clause boundaries. Blocked practice also led to significantly longer mean length of run and higher phonation/time ratio during training, although this advantage failed to transfer to meaningful pretest- posttest changes. These dynamic fluency developmental patterns are discussed to elucidate the underlying proceduralization in L2 speech processes.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Individual students were randomly assigned to the blocked or interleaved condition at a single university, not entire classes or schools, and the tutoring/personal-teaching exception does not cleanly apply to this self-administered practice program.
      • "They were randomly assigned to either a blocked-practice condition (n = 24) or an interleaved-practice condition (n = 26)."
      • Relevant Quotes: 1) "The participants (aged 18-22 years) were recruited for the current study through announcements in regular EFL classes at a university in Japan." (p. 294) 2) "They were randomly assigned to either a blocked-practice condition (n = 24) or an interleaved-practice condition (n = 26)." (p. 294) 3) "the participants were instructed to follow the prescribed 3-day fluency training program outside of the lab (e.g., in a quiet place such as at home)." (p. 296) 4) "However, the participants assigned to the blocked-practice condition performed the same narrative three times a day ... whereas those in the interleaved-practice condition performed three different narratives on each of the three days." (p. 296) Detailed Analysis: The unit of randomization here is the individual volunteer student, recruited from several regular EFL classes at one university, not an intact class or school. Under a literal reading this fails the class-level (or stronger) requirement, since classes and schools were never the randomized unit. A possible mitigating factor is that the actual training was carried out individually and privately by each participant, at home, using a personal recorder and printed booklet, with no shared classroom activity during which contamination between conditions could occur (participants never practiced together or observed each other's materials). This substantially reduces the specific contamination risk the criterion exists to guard against. However, the ERCT exception is worded specifically for interventions "designed for personal teaching like tutoring," implying an active one-to-one teaching relationship; this study is instead a self-administered, booklet-guided practice regimen with no tutor or instructor personalizing delivery. Because the paper gives no indication that the intervention was framed as tutoring or one-to-one teaching, and randomization was performed at the student level across mixed classes at a single site with no class- or school-level allocation described, I judge this to fall outside the exception on a literal reading. Criterion C is not met because randomization occurred at the individual-student level and the paper does not frame the intervention as personal tutoring, so the exception is not clearly applicable.
    • E

      Exam-based Assessment

      • Fluency was measured via researcher-elicited oral narrations of six-frame cartoons, analyzed acoustically with PRAAT, not via a standardized exam-based assessment.
      • "A different set of two prompts-street (Oba, 2018) and airport (Suzuki, 2011)-was used for the pretest and posttest."
      • Relevant Quotes: 1) "A different set of two prompts-street (Oba, 2018) and airport (Suzuki, 2011)-was used for the pretest and posttest. As in the training prompts, the pretest and posttest prompts also consisted of six-panel stories." (p. 295) 2) "Using the free sound-analysis software PRAAT (Boersma & Weenink, 2016), three trained coders first identified the filled and unfilled (silent) pauses of at least 200 milliseconds duration..." (p. 297) 3) "the following nine measures were selected... 1. mean length of run..., 2. articulation rate..., 3. phonation/time ratio..." (p. 297-298) Detailed Analysis: The outcome measures in this study are acoustic fluency indices (articulation rate, pause duration and frequency, mean length of run, phonation/time ratio, repetitions, and repairs) derived from timed, recorded oral narrations of picture-story prompts that were adopted from prior fluency research (Heaton, 1996; de Jong & Vercellotti, 2016) and coded with a phonetic-analysis tool. These are not standardized, widely recognized exams; they are researcher-selected performance tasks and lab-coded speech metrics designed specifically for fluency research, not a state-wide, national, or otherwise standardized exam instrument as required by the ERCT standard. Criterion E is not met because the study relied on custom oral-narrative elicitation tasks and acoustic coding rather than a standardized, exam-based assessment.
    • T

      Term Duration

      • Training lasted only 3 consecutive days and the posttest occurred just 1 day after the final training session, far short of a full academic term.
      • "All the participants returned to the computer lab for the posttest 1 day after the last training session."
      • Relevant Quotes: 1) "the study reported here focused on utterance fluency... In the current study, English-as-a- foreign-language (EFL) learners engaged in a 3-day fluency training program outside their regular English classes." (p. 286) 2) "In the pretest, all participants were tested individually in a computer lab 1 week prior to the training session." (p. 295-296) 3) "All the participants returned to the computer lab for the posttest 1 day after the last training session." (p. 296) Detailed Analysis: The entire study spans roughly 1 week pretest lead- in, 3 consecutive days of training, and a posttest administered only 1 day after training ends. This interval is measured in days, not the 3-4 months that constitute an academic term. There is no indication of any longer follow-up. Criterion T is not met because the interval from intervention start to outcome measurement spans only about a week, far short of one academic term.
    • D

      Documented Control Group

      • The control group's size, recruitment source, and baseline proficiency scores are explicitly reported and statistically compared with the two training groups.
      • "The learners in the control group (n = 18), who took the pretest and posttest only, were recruited from the same classes as those assigned to the two experimental groups."
      • Relevant Quotes: 1) "The learners in the control group (n = 18), who took the pretest and posttest only, were recruited from the same classes as those assigned to the two experimental groups." (p. 294) 2) "The means (standard deviations) were 47.96 (8.00), 44.62 (9.58), and 46.93 (6.70) for the blocked-practice, interleaved-practice, and control groups, respectively. A one-way ANOVA revealed that there were no significant differences among the three groups, F(2, 61) = 1.02, p = .37, eta^2 = .03." (p. 294) 3) "The participants in the control group were not part of the random group assignment because only the blocked-practice and interleaved-practice group members volunteered to engage in the training outside their regular classes." (Note 4, p. 316) Detailed Analysis: The paper documents the control group's size (n = 18), recruitment source (same classes as the treatment groups), what they received (pretest and posttest only, no training), and their baseline proficiency (a dedicated proficiency test with means/SDs reported and statistically compared to the two training groups, showing no significant baseline differences). This level of detail is sufficient to assess comparability of the control group at baseline and to confirm that they received no intervention. Criterion D is met because the control group's size, source, treatment, and baseline characteristics are clearly and quantitatively documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomization occurred among individual students at a single university, not among schools.
      • "They were randomly assigned to either a blocked-practice condition (n = 24) or an interleaved-practice condition (n = 26)."
      • Relevant Quotes: 1) "The participants (aged 18-22 years) were recruited for the current study through announcements in regular EFL classes at a university in Japan." (p. 294) 2) "They were randomly assigned to either a blocked-practice condition (n = 24) or an interleaved-practice condition (n = 26)." (p. 294) Detailed Analysis: All participants were drawn from a single university, and randomization was performed at the level of the individual student rather than across multiple schools or institutions. There is no cluster of schools involved anywhere in the design. Criterion S is not met because there is no school-level randomization; the study involves individual students at one institution.
    • I

      Independent Conduct

      • The single author designed, conducted, and analyzed the entire study, assisted only by research assistants and coders under his direction, with no independent third-party evaluator.
      • "I am grateful to Atsushi Miura, Misaki Kuratsubo, Miyu Koyama, Taeko Hosaka, and Kazuma Arai for their dedicated assistance in data collection and coding."
      • Relevant Quotes: 1) "This study was supported by a Grant-in-Aid for Scientific Research (KAKENHI) from the Japan Society for the Promotion of Science (JP18K12)." (p. 285, front matter) 2) "I am grateful to Atsushi Miura, Misaki Kuratsubo, Miyu Koyama, Taeko Hosaka, and Kazuma Arai for their dedicated assistance in data collection and coding." (p. 285, front matter) 3) "To ensure that the participants performed the fluency training as indicated, a research assistant sent a daily reminder to them using a chat application on their smart phones." (p. 296) Detailed Analysis: The paper is single-authored (Yuichi Suzuki), who designed the study, the training materials protocol, and the statistical analysis plan. The individuals credited (coders, a research assistant sending reminders) worked under the author's direction as assistants rather than as an independent evaluation team separate from the designer of the intervention. There is no statement of an external, independent organization conducting or overseeing data collection or analysis separately from the author who designed the training conditions being compared. Criterion I is not met because the same author who designed the blocked/interleaved training conditions also conducted and analyzed the study, with only dependent research assistants, not an independent evaluation team.
    • Y

      Year Duration

      • Because criterion T (Term Duration) is not met, this stronger criterion cannot be met either; in any case the study lasted only about a week.
      • "All the participants returned to the computer lab for the posttest 1 day after the last training session."
      • Relevant Quotes: 1) "In the current study, English-as-a-foreign- language (EFL) learners engaged in a 3-day fluency training program outside their regular English classes." (p. 286) 2) "All the participants returned to the computer lab for the posttest 1 day after the last training session." (p. 296) Detailed Analysis: Per the ERCT rules, Y automatically fails when T is not met. Substantively, the total tracked period (1-week pretest lead-in, 3 days of training, and a 1-day-later posttest) is on the order of 10 days, nowhere near 75% of an academic year (~9-10 months). Criterion Y is not met both because T is not met and because the actual duration is far short of a year.
    • B

      Balanced Control Group

      • The extra practice time given to the two training groups relative to the control group is itself the explicit treatment variable under investigation (comparing schedules of the same total practice amount), and the two training conditions received identical amounts of time and material.
      • "the participants assigned to the blocked-practice condition performed the same narrative three times a day... whereas those in the interleaved-practice condition performed three different narratives on each of the three days."
      • Relevant Quotes: 1) "The study sample included 68 English learners at a Japanese university who engaged in oral narrative tasks using six-frame cartoons three times a day for three consecutive days. Specifically, the sequence of three types of cartoons was manipulated while keeping the task variation equal, allowing the effects of the blocked-practice condition... and interleaved- practice condition... to be compared." (p. 292) 2) "These booklets were created for both intervention groups. However, the participants assigned to the blocked-practice condition performed the same narrative three times a day... whereas those in the interleaved-practice condition performed three different narratives on each of the three days." (p. 296) 3) "The learners in the control group (n = 18), who took the pretest and posttest only, were recruited from the same classes as those assigned to the two experimental groups." (p. 294) Detailed Analysis: Applying the criterion B decision tree: extra practice time/resources are present relative to the passive control group (EXTRA_RESOURCES_PRESENT = true), and this additional narrative-repetition practice is explicitly the treatment variable being tested (RESOURCES_ARE_TREATMENT = true) rather than a supplementary add-on to some other core intervention, so the branch resolves to "met" regardless of whether the control matches the resource. The central comparison of this study is between the blocked and interleaved conditions, which receive identical amounts of practice time, an identical number of sessions (nine narrations total across three days), and identical materials/instructions, differing only in the sequencing (order) of the same three story prompts. Since the amount of training resource is explicitly held constant across these two active conditions, and the research question is specifically about the effect of task sequencing (not amount of resource), the two active arms are balanced by design. Regarding the passive control group, the additional practice time given to the two active groups is not a supplementary add-on to some other core intervention; the practice itself, and specifically its scheduling, is the entire treatment variable being tested. The control group's lack of any training is therefore the natural "business as usual" baseline needed to assess whether task repetition training works at all, analogous to cases where the additional resource is integral to and inseparable from the treatment being evaluated. Criterion B is met because the two active conditions received identical time/resources (isolating the sequencing variable), and the control group's lack of training reflects the fact that training itself is the explicit treatment being tested.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific 3-day, cartoon-narration blocked-versus-interleaved design was found; a related but distinct 2023 study by a different team used different tasks, a longer duration, and reported partly contrary results.
      • Relevant Quotes: 1) "Given the small number of participants in this study, further research should be conducted to elucidate the effectiveness of blocked and interleaved practice in L2 speaking training, instead of drawing definitive conclusions from the findings yielded by this single experiment." (p. 310, self-acknowledged as a single experiment) Detailed Analysis: An internet search (2026) for follow-up or replication work identified a related paper: Zhang, M., Yi, N., & Zhou, D. (2023), "The Effects of Task Repetition Schedules on L2 Fluency Enhancement," Languages (MDPI), 8(4), 252. This is confirmed to be an independent research team unaffiliated with Yuichi Suzuki. The study involved 90 university freshmen assigned to blocked, interleaved, and control groups using PRAAT-based fluency coding. However, that study used different practice tasks (problem-solving tasks rather than six-frame cartoon narrations), a substantially longer training period (about three weeks rather than three days), and reported that the interleaved-repetition group was advantageous compared with the blocked-repetition group for most fluency measures (except silent pause numbers) - the opposite pattern from the current paper. This is a conceptually related but methodologically distinct study addressing the same general research question rather than an independent replication of this specific study's design, procedure, and claims. No study was identified that replicates the specific 3-day/6-frame-cartoon blocked-versus- interleaved design and reproduces its findings. Criterion R is not met because no independent replication of this specific study was found; the closest related work (Zhang et al., 2023) used a different task, duration, and found a different result pattern.
    • A

      All-subject Exams

      • Because criterion E (Exam-based Assessment) is not met, this stronger criterion cannot be met either; moreover, only speaking fluency was measured, not other subjects.
      • "the study reported here focused on utterance fluency"
      • Relevant Quotes: 1) "the study reported here focused on utterance fluency (e.g., Skehan, 2003), defined as objective features of utterance such as speed (e.g., articulation rate), breakdown (e.g., pauses), and repair fluency (e.g., repetitions)." (p. 286) Detailed Analysis: Per the ERCT rules, criterion A automatically fails when criterion E is not met, which is the case here. Substantively, the study measured only L2 speaking fluency via a narrow set of acoustic indices from an oral narrative task; no other subjects or skill domains (e.g., reading, writing, listening, grammar, other academic subjects) were assessed. Criterion A is not met both because E is not met and because only a single narrow outcome domain (speech fluency) was measured.
    • G

      Graduation Tracking

      • Because criterion Y (Year Duration) is not met, this stronger criterion cannot be met; the study tracked participants for about a week with no long-term or graduation follow-up.
      • "All the participants returned to the computer lab for the posttest 1 day after the last training session."
      • Relevant Quotes: 1) "All the participants returned to the computer lab for the posttest 1 day after the last training session." (p. 296) 2) "because the posttest was administered only 1 day after the treatment, it is crucial to investigate the durability of training transfer effects in future research using delayed posttests." (p. 313, Limitations) Detailed Analysis: Per the ERCT rules, G automatically fails when Y is not met, which is the case here. The authors themselves flag the single, immediate posttest (1 day after training) as a limitation and call for future delayed-posttest research; there is no mention of any tracking toward graduation, nor any indication of a follow-up publication tracking this cohort longer-term. An internet search (2026) for subsequent papers by Yuichi Suzuki tracking this same cohort of 68 Japanese university learners (e.g., toward course or degree completion) did not surface any such follow-up publication. This is unsurprising given the study population (adult university volunteers in a short, out-of-class fluency-training add-on) does not lend itself to graduation tracking in the sense intended by the ERCT standard (K-12 or similar staged educational progression). Criterion G is not met both because Y is not met and because the study explicitly used only a 1-day-later posttest with no long-term follow-up, and no follow-up graduation-tracking publication was found.
    • P

      Pre-Registered

      • The paper contains no statement of a pre-registered protocol, registry ID, or registration date.
      • Relevant Quotes: (No quotes are available: the manuscript, including its Method, Statistical Analysis, and Open Research Badges sections, contains no reference to a trial registry, a pre-registered protocol, or a registration date.) Detailed Analysis: The article carries an "Open Materials" badge for sharing training materials via the IRIS repository, but this is distinct from pre-registration of hypotheses, methods, and analysis plans before data collection. No registry name, ID, or timing information is mentioned anywhere in the text, notes, or supplementary information list. An internet search (2026) for a pre-registration record for this study (e.g., on OSF or similar registries) likewise did not surface any such record. Criterion P is not met because there is no evidence anywhere in the paper, or found via internet search, of a pre-registered protocol.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.