Abstract
To examine the effects of task repetition with different schedules, English-as-a-foreign-language classroom learners performed the same oral narrative task six times under three different schedules. They narrated the same six-frame cartoon story (a) six times consecutively in one class (massed practice), (b) three times at the beginning and at the end of a class (short-spaced practice), and (c) three times as a part of two classes 1 week apart (long-spaced practice). The results yielded by an immediate posttest using a novel cartoon showed that massed practice reduced breakdown fluency (mid-clause and clause-final pauses) the most. However, the participants in the massed-practice group showed degraded speed (slower articulation rate) and repair fluency (more verbatim repetition). The effects of repetition schedule seem limited on a 1-week delayed posttest involving a novel cartoon. Yet, when participants narrated the same practiced cartoon 1 week later, massed practice also resulted in more verbatim repetition.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- The study is a quasiexperiment in which four intact classes were assigned to conditions without randomisation, so it is not a class-level RCT.
- "This was a classroom-based research employing a quasiexperimental research design with four intact English classes at a Japanese university."
Relevant Quotes:
1) "This was a classroom-based research employing a quasiexperimental research design with four intact English classes at a Japanese university." (p. 540)
2) "They were recruited from four intact English classes, which were assigned to the massed (n = 20), the short-spaced (n = 23), the long-spaced (n = 21), or the control group (n = 15)." (p. 541)
3) "Second, because random assignment of participants to each condition was not feasible in this classroom research, individual differences such as cognitive aptitudes should have been controlled at least statistically." (p. 556)
Detailed Analysis:
Criterion C requires a randomised controlled trial with clearly described randomisation at the class level (or stronger). Although the treatment conditions were allocated at the level of whole intact classes (which would prevent within-class contamination), the paper explicitly describes the design as "quasiexperimental" and states that random assignment "was not feasible in this classroom research." No randomisation procedure of any kind (of classes or of students) is described; four pre-existing classes were simply assigned to the four conditions. The intervention is whole-class oral narration practice, not personal one-to-one tutoring, so the tutoring exception does not apply. Without random assignment, the study is not an RCT at all, so the class-level RCT requirement cannot be satisfied.
Criterion C is not met because the four intact classes were assigned to conditions without any described randomisation, and the authors themselves label the design quasiexperimental.
-
E
Exam-based Assessment
- Outcomes were custom researcher-coded oral fluency measures from picture-narration tasks, not standardised exams (TOEIC was used only to describe baseline proficiency).
- "the following seven fluency measures were coded to cover speed, breakdown, and repair fluency (Skehan, 2003)"
Relevant Quotes:
1) "the following seven fluency measures were coded to cover speed, breakdown, and repair fluency (Skehan, 2003)" (p. 544)
2) "Two picture prompts (Chase and Surprise) were used for the fluency training in the present study. The prompts were adopted from Heaton (1996) and have been used in many studies on L2 oral production" (p. 541)
3) "Three six-frame picture prompts (Bicycle, Bus, and Race) that were unfamiliar to the participants were used in the pretest and posttests" (p. 542)
4) "The speech samples were annotated using a free sound analysis software PRAAT (Boersma & Weenink, 2016)." (p. 543)
5) "Based on their mean score (M = 421.84, SD = 98.88) on a standardized English test (Test of English for International Communication [TOEIC]), their English proficiency was estimated to fall between A2 (elementary) and B1 (intermediate) level" (p. 541)
Detailed Analysis:
Criterion E requires that outcomes be measured with widely recognised standardised exams. The outcomes in this study were seven researcher-coded utterance fluency measures (articulation rate, pause durations and frequencies, repetitions, repairs) extracted from oral narrations of picture prompts, annotated with PRAAT by the research team. These are custom research measures aligned with the intervention, not standardised exams. The only standardised test mentioned, TOEIC, was used solely to describe baseline proficiency, not as an outcome measure.
Criterion E is not met because outcomes were custom researcher-coded fluency measures from picture-narration tasks rather than a standardised exam.
-
T
Term Duration
- Outcomes were measured only up to one week after a one-to-two-week training phase, far short of a full academic term.
- "One week after the last training session (Week 3 or Week 4), a delayed posttest with a new narrative task was administered for examining the long-lasting transfer of fluency training."
Relevant Quotes:
1) "In Week 1, all participants took a pretest" (p. 542)
2) "In Week 2, the training phase started for the three experimental groups." (p. 542)
3) "One week after the last training session (Week 3 or Week 4), a delayed posttest with a new narrative task was administered for examining the long-lasting transfer of fluency training." (p. 543)
4) "The goal of this short-term classroom intervention study was elucidating the effects of different task repetition schedules on the development of L2 utterance fluency." (p. 537)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Here the intervention began in Week 2, the immediate posttest followed the sixth narration (Week 2 or Week 3 depending on group), and the final delayed posttest occurred one week after the last training session (Week 3 or Week 4). The entire study, from intervention start to final measurement, spanned at most about two to three weeks, and the authors themselves describe it as a "short-term classroom intervention study." This is far shorter than one academic term.
Criterion T is not met because the interval from intervention start to the final measurement was only about two to three weeks, far less than one academic term.
-
D
Documented Control Group
- The control group's size, origin, baseline proficiency and pretest scores, and conditions (tests only) are clearly documented.
- "Participants in the control group only took pretest, immediate posttest, and delayed posttest."
Relevant Quotes:
1) "The study sample consisted of 79 first-year students at a private university in Japan who had been studying English as a foreign language for at least 6 years before entering university. They were recruited from four intact English classes, which were assigned to the massed (n = 20), the short-spaced (n = 23), the long-spaced (n = 21), or the control group (n = 15)." (p. 541)
2) "According to one-way ANOVA results, there were no statistically significant differences in the TOEIC scores among the four groups, F(3, 75) = 1.57, p = .20, η2 = .059." (p. 541)
3) "Participants in the control group only took pretest, immediate posttest, and delayed posttest." (p. 543)
4) "According to the one-way ANOVAs, there was no significant main effect of condition for any fluency pretest measures (p > .10), with the exception of clause-final pause duration (p = .01)." (p. 557, Note 5)
Detailed Analysis:
Criterion D requires that the control group be documented with respect to who they are, their baseline performance, and what they received. The paper specifies the control group's size (n = 15), its origin (one of four intact first-year English classes at a Japanese private university), its baseline proficiency (TOEIC comparison across the four groups showing no significant differences), baseline fluency pretest scores (compared across groups, with descriptive statistics reported in Appendix F), and exactly what the control group did (only the pretest, immediate posttest, and delayed posttest, with no fluency training). This provides adequate documentation of the control group's composition, baseline performance, and conditions.
Criterion D is met because the control group's size, source, baseline TOEIC and pretest fluency scores, and treatment (tests only, no training) are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Allocation involved four intact classes within one university, so there was no school-level randomisation.
- "The study took place in four intact English classes where the second author of this article was the instructor."
Relevant Quotes:
1) "They were recruited from four intact English classes, which were assigned to the massed (n = 20), the short-spaced (n = 23), the long-spaced (n = 21), or the control group (n = 15)." (p. 541)
2) "The study took place in four intact English classes where the second author of this article was the instructor." (p. 542)
Detailed Analysis:
Criterion S requires randomisation among schools (institutions or implementing units). This study took place within a single Japanese private university, and allocation was among four intact classes taught by the same instructor, with no randomisation described at any level. There is no school-level assignment of any kind.
Criterion S is not met because the study involved four intact classes within a single university and no school-level randomisation was performed.
-
I
Independent Conduct
- The authors designed, delivered (the second author was the classroom instructor), and analysed the intervention with no independent evaluators.
- "The study took place in four intact English classes where the second author of this article was the instructor."
Relevant Quotes:
1) "The study took place in four intact English classes where the second author of this article was the instructor." (p. 542)
2) "This study was supported by Yokohama Academic Foundation for the first author and Grant-in-Aid for Scientific Research (KAKENHI) from Japan Society for the Promotion of Science (19K13274) to the second author." (p. 536)
3) "We are also grateful to Atsushi Miura, Chisaki Inaba, Miyu Koyama, Misaki Kuratsubo, Naoya Nishiura, and Kohta Shimada for their dedicated assistance in data coding." (p. 536)
Detailed Analysis:
Criterion I requires that the study be conducted independently from those who designed the intervention. In this study, the two authors designed the fluency-training intervention, and the second author personally delivered it as the classroom instructor of all four classes. Data coding assistants are acknowledged, but they worked under the authors' direction; there is no statement of any external evaluation team or third-party oversight of data collection, analysis, or conclusions.
Criterion I is not met because the intervention was designed, delivered (by the second author as instructor), and analysed by the same research team with no independent oversight.
-
Y
Year Duration
- The study lasted only a few weeks in total, far short of the required 75% of an academic year (and T is not met).
- "One week after the last training session (Week 3 or Week 4), a delayed posttest with a new narrative task was administered"
Relevant Quotes:
1) "In Week 2, the training phase started for the three experimental groups." (p. 542)
2) "One week after the last training session (Week 3 or Week 4), a delayed posttest with a new narrative task was administered" (p. 543)
3) "Yet, as the distributed practice effects were evident even with six task repetitions on the immediate posttest in the current study, a longer intervention study (e.g., over one semester) may be needed to demonstrate durability of its effects." (p. 555)
Detailed Analysis:
Criterion Y requires that outcomes be tracked for at least 75% of an academic year after the intervention begins. Because criterion T (term duration) is already not met, this stronger criterion automatically fails. The entire study lasted at most about four weeks from pretest to delayed posttest, and the authors themselves note that a longer study "over one semester" would be needed, confirming the short duration.
Criterion Y is not met because the study spanned only a few weeks, nowhere near 75% of an academic year, and criterion T is also not met.
-
B
Balanced Control Group
- The experimental groups received identical amounts of practice within regular class time, and the task-repetition practice itself was the treatment variable tested against a business-as-usual control.
- "Practically, establishing distributed practice effects for certain aspects of L2 learning can help maximize the outcome of repeated practice without changing the total practice time."
Relevant Quotes:
1) "It took about a total of 45 minutes for performing the narrative six times. For the remaining 45 minutes, participants engaged in regular class activities (i.e., reading a passage with comprehension questions and a dictation task of the passage), which was not relevant to the training task." (p. 542)
2) "Practically, establishing distributed practice effects for certain aspects of L2 learning can help maximize the outcome of repeated practice without changing the total practice time." (p. 537)
3) "Participants in the control group only took pretest, immediate posttest, and delayed posttest." (p. 543)
4) "Constant time limit was imposed throughout" (p. 542)
Detailed Analysis:
Criterion B requires that intervention and control conditions be balanced in time and resources unless the extra input is itself the treatment variable. Applying the decision tree: extra resources are present, since the three experimental groups received about 45 minutes of extra speaking-practice time per training class that the control group did not receive. However, this additional practice time is the primary treatment variable under investigation: the whole study is designed to test how different temporal distributions of the same amount of task-repetition practice (massed vs. short-spaced vs. long-spaced) affect fluency, framed explicitly as testing distribution "without changing the total practice time" among the three active conditions. Among the three experimental groups themselves (the primary contrast of interest), practice time was identical: every group performed the same narrative task six times with the same 90-second planning and 120-second performance windows ("constant time limit was imposed throughout"), and only the temporal distribution of those six repetitions differed. The extra practice given to the experimental groups relative to control replaced ordinary class activities during regular class time (no extra budget, materials, or out-of-class time was added), and repeated speaking practice is itself the intervention being tested against a business-as-usual, no-training control. This matches the "resources are the treatment variable" branch of the decision tree, under which the control group may remain business-as-usual by design.
Criterion B is met because the three experimental schedules were exactly matched on practice time (isolating the variable of interest, spacing), and the additional practice itself, delivered within regular class time, is the treatment variable explicitly tested against a business-as-usual control, consistent with the updated criterion B definition.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study by a different research team was found in the paper or via external search (related studies by Kakitani & Kormos, 2024, and Tabari et al., 2025, differ in design/modality or are not independent).
- "To the best of our knowledge, Bui et al. (2019) conducted the first and only study in which the task repetition interval was systematically manipulated in research design to investigate the distributed practice effects in speaking tasks."
Relevant Quotes:
1) "To the best of our knowledge, Bui et al. (2019) conducted the first and only study in which the task repetition interval was systematically manipulated in research design to investigate the distributed practice effects in speaking tasks." (p. 540)
2) "The current study extends the line of investigation into repeated engagement of the same speaking task under different schedules (e.g., Bui et al., 2019; Suzuki, 2021b)." (p. 537)
Detailed Analysis:
Criterion R requires that this specific study be independently replicated by a different team in a peer-reviewed journal. The paper itself positions the work as novel, extending the single prior study (Bui et al., 2019), which predates it and used a different design (two repetitions under five intervals: 0-, 1-, 3-, 7-, and 14-day), so it cannot count as a replication of this study. An internet search of the citation index (Semantic Scholar lists 31 citing records for this article, DOI 10.1017/S0272263121000358) for subsequent independent replications of this specific six-repetition, three-schedule EFL classroom design found related but non-replicating work: Kakitani and Kormos (2024), examining 1-day versus 7-day spacing effects on L2 fluency development with a different sample (116 Japanese learners) and a different design; and Tabari, Lee, and Hanzawa (2025), "Effects of massed and spaced task repetitions on L2 writing task performance and task emotions," which targets written rather than oral production and is co-authored by Hanzawa, one of the original authors, so it is not independent. No independent peer-reviewed replication of this particular oral-fluency, three-schedule classroom study by a different research team was found.
Criterion R is not met because no independent peer-reviewed replication of this specific study by a different research team was found in the paper or via external search.
-
A
All-subject Exams
- Criterion E is not met and only L2 oral fluency was measured, so all-subject standardised assessment is absent.
- "Third, as the current study focused on fluency changes, other speech aspects such as complexity and accuracy (appropriateness) were not analyzed."
Relevant Quotes:
1) "the following seven fluency measures were coded to cover speed, breakdown, and repair fluency (Skehan, 2003)" (p. 544)
2) "Third, as the current study focused on fluency changes, other speech aspects such as complexity and accuracy (appropriateness) were not analyzed." (p. 556)
Detailed Analysis:
Criterion A requires standardised exam-based assessment across all main subjects, and criterion E is an explicit prerequisite. Criterion E is not met (outcomes were custom fluency codings, not standardised exams), so criterion A automatically fails. Moreover, the study measured only oral fluency in English speaking; no other subjects, or even other aspects of English proficiency, were assessed.
Criterion A is not met because criterion E fails and only a single narrow outcome domain (L2 oral fluency) was measured.
-
G
Graduation Tracking
- Measurement stopped one week after training with no tracking of the first-year students to graduation (Y is not met, and no follow-up publication tracking this cohort was found).
- "One week after the last training session (Week 3 or Week 4), a delayed posttest with a new narrative task was administered"
Relevant Quotes:
1) "One week after the last training session (Week 3 or Week 4), a delayed posttest with a new narrative task was administered" (p. 543)
2) "Second, if it was postponed further, participants could have improved their speaking ability outside of this intervention." (p. 557, Note 2)
Detailed Analysis:
Criterion G requires tracking participants until graduation from their educational stage, and criterion Y is a prerequisite. Criterion Y is not met, so G automatically fails. Measurement ended one week after training; the first-year university participants were not followed to graduation. An internet search for subsequent publications by Suzuki and/or Hanzawa tracking this same cohort (the 79 first-year students recruited from four intact classes) was conducted; no follow-up paper tracking this specific cohort toward graduation was found. Other work by the same authors identified in the citation record (e.g., Suzuki, 2021a, 2021b; Hanzawa, 2021; Tabari, Lee, & Hanzawa, 2025) uses different samples, tasks, or modalities rather than extending follow-up of this cohort.
Criterion G is not met because tracking ended one week after the training phase, prerequisite criterion Y is not met, and no follow-up publication tracking this cohort to graduation was found.
-
P
Pre-Registered
- The paper reports an Open Materials badge but no pre-registration of hypotheses, methods, or analyses before data collection; no registry entry was found via search.
Relevant Quotes:
1) "The experiment in this article earned an Open Materials badge for transparent practices. The materials are available at https://www.iris-database.org/iris/app/home/detail?id= york%3a939256" (p. 536)
2) "All test prompts are available in the IRIS digital repository of data collection instruments (Marsden et al., 2016)." (p. 542)
Detailed Analysis:
Criterion P requires pre-registration of the full study protocol (hypotheses, methods, planned analyses) before data collection began. The paper documents open materials (an Open Materials badge and deposit of prompts/materials in the IRIS repository), which concerns transparency of research instruments, not pre-registration of hypotheses or analysis plans. No registry (e.g., OSF, ClinicalTrials.gov, AsPredicted), registration ID, or registration date is mentioned anywhere in the paper. An internet search of the publisher's article page confirmed the badges attached to this article are an Open Access badge (CC-BY) and an Open Materials badge only; no Pre-registered Design/Analysis Plan badge or external trial-registry entry was found for this study.
Criterion P is not met because no pre-registered protocol or registry entry is reported or found; only open materials are documented.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.