The effects of interleaved and blocked corpus-based practice on L2 pragmatic development

Ying Zhang

Published:
ERCT Check Date:
DOI: 10.1017/S0272263123000062
  • L2 languages
  • higher education
  • China
  • digital assessment
0
  • C

    Randomisation was at the individual student level rather than at the class or school level, and the computerized practice intervention does not qualify for the tutoring exception.

    "Then, the 63 learners were randomly assigned to two groups to practice the two target pragmatic features."

  • E

    Outcomes were assessed with multimedia discourse completion tasks and a rubric custom-designed for this study, not with a widely recognised standardised exam.

    "To address the drawbacks of traditional DCTs, MMDCTs ... were designed for the current study."

  • T

    The interval from intervention start to the final delayed posttest was only about five to seven weeks, which is shorter than one full academic term.

    "both groups completed a delayed posttest five weeks later."

  • D

    The comparison group's size, demographic characteristics, baseline pretest performance, and exact treatment conditions are documented in detail.

    "no noticeable difference in pragmatic accuracy of the target speech acts on the pretest between the IP group (M = 39.98, SD = 3.71) and the BP group (M = 39.01, SD = 4.03)"

  • S

    The study randomised individual students within one university rather than randomising schools or institutional units.

    "Then, the 63 learners were randomly assigned to two groups to practice the two target pragmatic features."

  • I

    The single author designed, delivered, and evaluated the intervention herself, with no independent evaluation team or third-party oversight.

    "During Week 2, all students received two hours of pragmatics instruction on making requests and suggestions delivered by the researcher."

  • Y

    The study tracked outcomes for only about five weeks after the intervention, far short of 75% of an academic year, and criterion T is not met.

    "both groups completed a delayed posttest five weeks later."

  • B

    Both groups received identical instruction, tasks, feedback, and practice time, with only the sequencing of the same 20 tasks differing as the treatment variable.

    "In total, the practice session included 20 tasks (i.e., 10 practice tasks x 2 target speech acts)."

  • R

    The paper describes itself as the first study of its kind, and a citation search found no independent peer-reviewed replication of this specific experiment by a different research team.

    "The present study is the first to explore the effects of blocked practice and interleaved practice on L2 pragmatic development."

  • A

    Criterion E is unmet and the study assessed only English pragmatics outcomes with custom instruments, not standardised exams across all main subjects.

    "this study examined both accuracy and fluency of L2 learners' pragmatic behaviors."

  • G

    Measurement stopped at a five-week delayed posttest, with no tracking of the sophomore participants through to graduation found in the author's later publications, and criterion Y is not met.

    "both groups completed a delayed posttest five weeks later."

  • P

    No pre-registration statement, registry identifier, or registration date is mentioned in the paper, and none was found via an external registry search.

Abstract

A handful of second/foreign language (L2) studies have examined the effects of practice schedules and reported the advantage of interleaved practice (i.e., practice multiple skills simultaneously) over blocked practice (i.e., practice one skill first and then proceed to the next one). However, no studies in the realm of L2 pragmatics have explored this theme. This study investigated the influence of interleaved corpus-based practice and blocked corpus-based practice on L2 pragmatic development. Sixty-three L2 learners of English from a university in China received instruction on two pragmatic features: suggestions and requests. After the instruction, they were randomly assigned to an interleaved-practice group (n = 31) or a blocked-practice group (n = 32). Results from multimedia discourse completion tasks on the immediate and delayed posttests showed facilitative and long-term effects of interleaved practice on pragmatic accuracy. Moreover, the results revealed positive and durable influence of blocked practice on fluency. Implications are discussed.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was at the individual student level rather than at the class or school level, and the computerized practice intervention does not qualify for the tutoring exception.
      • "Then, the 63 learners were randomly assigned to two groups to practice the two target pragmatic features."
      • Relevant Quotes: 1) "After the instruction, they were randomly assigned to an interleaved-practice group (n = 31) or a blocked-practice group (n = 32)." (p. 812, Abstract) 2) "Then, the 63 learners were randomly assigned to two groups to practice the two target pragmatic features. The interleaved-practice (IP) group consisted of 31 learners, whereas the blocked-practice (BP) group consisted of 32 learners." (p. 818) 3) "Sixty-three sophomores (35 females, 28 males) from EFL classes at a public university in China participated in this study." (p. 818) 4) "After the pretest, the learners were randomly assigned to an IP group or a BP group." (p. 819) Detailed Analysis: Criterion C requires randomisation at the class level (or stronger school level), with an exception for one-to-one tutoring or personal-teaching interventions. The paper clearly states that the 63 individual learners, drawn from EFL classes at a single public university, were randomly assigned to the two practice conditions. The unit of randomisation was the individual student, not intact classes or schools. The intervention (computerized corpus-based discourse completion practice completed individually at computers) is not a personal tutoring or one-to-one teaching intervention, so the tutoring exception does not apply. Individual randomisation of students drawn from the same classes risks the contamination that class-level randomisation is designed to prevent. Criterion C is not met because randomisation was conducted at the individual student level within a single university, and no valid tutoring exception applies.
    • E

      Exam-based Assessment

      • Outcomes were assessed with multimedia discourse completion tasks and a rubric custom-designed for this study, not with a widely recognised standardised exam.
      • "To address the drawbacks of traditional DCTs, MMDCTs ... were designed for the current study."
      • Relevant Quotes: 1) "To address the drawbacks of traditional DCTs, MMDCTs (i.e., encompassing text-formatted prompts as well as dialogues presented aurally and visually via pictures) were designed for the current study." (p. 822) 2) "Moreover, three equivalent versions of the tests were designed for the pretest, immediate posttest, and delayed posttest to avoid practice effects." (p. 823) 3) "With respect to the design of the target test items, first, each of them was designed based on the authentic data related to making requests or suggestions from MICASE (Simpson et al., 2002)." (p. 823) 4) "a four-level rubric (see Table 3) was adopted to measure the two groups' speech acts of making requests and suggestions" (p. 824) 5) "They were intermediate English learners based on their scores (M = 84.39, SD = 3.95) on the TOEFL (Test of English as a Foreign Language), which they took three months prior to this study." (p. 818) Detailed Analysis: Criterion E requires that outcomes be measured with widely recognised standardised exams, not instruments specially designed for the study. The outcome measures here were multimedia discourse completion tasks (MMDCTs) that were "designed for the current study" by the researcher, based on the MICASE corpus, and scored with a custom four-level pragmatic accuracy rubric plus a manually calculated speech rate measure. Although the standardised TOEFL is mentioned, it was used only to describe participants' baseline proficiency, not as an outcome measure. The researcher-built MMDCTs, however carefully piloted and reliability-checked, are custom study-specific instruments closely aligned with the intervention content, which is exactly the situation the criterion is designed to guard against. Criterion E is not met because outcomes were measured with custom-designed MMDCTs and a researcher-developed rubric rather than a recognised standardised exam.
    • T

      Term Duration

      • The interval from intervention start to the final delayed posttest was only about five to seven weeks, which is shorter than one full academic term.
      • "both groups completed a delayed posttest five weeks later."
      • Relevant Quotes: 1) "During Week 1, all participants completed a background questionnaire. During Week 2, all students received two hours of pragmatics instruction on making requests and suggestions delivered by the researcher." (p. 819) 2) "The entire practice session (i.e., 20 practice items) lasted one hour" (p. 820) 3) "After the practice session, both groups completed an immediate posttest. Moreover, to explore the retention of the benefits of the interleaved or blocked practice, both groups completed a delayed posttest five weeks later." (p. 820) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. In this study, the intervention consisted of two hours of instruction in Week 2 followed by a one-hour practice session, with an immediate posttest right after practice and a delayed posttest five weeks later. Even counting from the background questionnaire in Week 1 to the final delayed posttest, the total span of the study is roughly seven weeks, well short of a 3-4 month academic term. The standard allows short interventions but insists on term-long follow-up tracking from intervention start, which this five-week delayed posttest does not provide. Criterion T is not met because the final outcome measurement occurred only about five weeks after the practice intervention, far less than one full academic term.
    • D

      Documented Control Group

      • The comparison group's size, demographic characteristics, baseline pretest performance, and exact treatment conditions are documented in detail.
      • "no noticeable difference in pragmatic accuracy of the target speech acts on the pretest between the IP group (M = 39.98, SD = 3.71) and the BP group (M = 39.01, SD = 4.03)"
      • Relevant Quotes: 1) "Sixty-three sophomores (35 females, 28 males) from EFL classes at a public university in China participated in this study. Their ages ranged from 19.5 to 22.6 years (Mage = 20.7). They were intermediate English learners based on their scores (M = 84.39, SD = 3.95) on the TOEFL" (p. 818) 2) "The interleaved-practice (IP) group consisted of 31 learners, whereas the blocked-practice (BP) group consisted of 32 learners." (p. 818) 3) "In contrast, the BP group first engaged in 10 practice tasks pertinent to making requests and then participated in 10 practice tasks involving making suggestions." (pp. 819-820) 4) "The results from an independent samples t-test demonstrated no noticeable difference in pragmatic accuracy of the target speech acts on the pretest between the IP group (M = 39.98, SD = 3.71) and the BP group (M = 39.01, SD = 4.03), t(61) = .991, p = .326, d = .250." (pp. 826-827) 5) "Moreover, no significant difference in fluency of the target pragmatic features was found on the pretest between the IP group (M = 2.69, SD = .32) and the BP group (M = 2.63, SD = .35), t(61) = .786, p = .435, d = .198." (p. 827) Detailed Analysis: Criterion D requires detailed documentation of the comparison group: who they are, their baseline performance, and the treatment they received. This study uses a comparison-group design in which the blocked-practice (BP) group serves as the comparison condition for the interleaved-practice (IP) group. The paper documents the BP group's size (n = 32), the demographic profile of the sample (age, gender, TOEFL-based proficiency including speaking and listening subscores, no residence abroad), the exact treatment the BP group received (the same instruction plus 20 practice tasks in blocked order), and its baseline performance on both outcome measures (pretest means, SDs, and confidence intervals for pragmatic accuracy and fluency), together with statistical checks confirming baseline equivalence, normality, and homogeneity of variances. This constitutes clear, quotable documentation of the comparison group's characteristics, size, and conditions, sufficient for proper comparison. Criterion D is met because the comparison (blocked-practice) group's size, demographics, baseline performance, and exact treatment are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • The study randomised individual students within one university rather than randomising schools or institutional units.
      • "Then, the 63 learners were randomly assigned to two groups to practice the two target pragmatic features."
      • Relevant Quotes: 1) "Sixty-three sophomores (35 females, 28 males) from EFL classes at a public university in China participated in this study." (p. 818) 2) "Then, the 63 learners were randomly assigned to two groups to practice the two target pragmatic features." (p. 818) Detailed Analysis: Criterion S requires randomisation among schools or equivalent implementing institutions. This study took place at a single public university in China, and randomisation was performed at the individual student level within that one institution. No schools, sites, or institutional units were randomised, and only one institution participated. Criterion S is not met because randomisation occurred at the individual student level within a single university, not among schools or institutional units.
    • I

      Independent Conduct

      • The single author designed, delivered, and evaluated the intervention herself, with no independent evaluation team or third-party oversight.
      • "During Week 2, all students received two hours of pragmatics instruction on making requests and suggestions delivered by the researcher."
      • Relevant Quotes: 1) "During Week 2, all students received two hours of pragmatics instruction on making requests and suggestions delivered by the researcher." (p. 819) 2) "The present study is the first to explore the effects of blocked practice and interleaved practice on L2 pragmatic development." (p. 813) 3) "Three trained raters with PhD degrees in second language studies scored the data independently." (pp. 824-825) 4) "Competing interests. The author declares none." (p. 833) Detailed Analysis: Criterion I requires that the study be conducted independently from those who designed the intervention. This is a single-author study in which the same researcher designed the practice materials and MMDCTs, delivered the pragmatics instruction personally, administered the treatment, and conducted the statistical analyses. Although three trained raters scored the outcome data independently and interrater reliability was high, rating by hired raters is not the same as an independent third-party evaluation team: the design, implementation, data collection, and analysis all remained under the control of the intervention designer, and no external oversight or evaluation agency is mentioned anywhere in the paper. Criterion I is not met because the same single researcher designed the intervention, delivered the instruction, and analysed the outcomes without any independent third-party evaluation.
    • Y

      Year Duration

      • The study tracked outcomes for only about five weeks after the intervention, far short of 75% of an academic year, and criterion T is not met.
      • "both groups completed a delayed posttest five weeks later."
      • Relevant Quotes: 1) "During Week 2, all students received two hours of pragmatics instruction on making requests and suggestions delivered by the researcher." (p. 819) 2) "both groups completed a delayed posttest five weeks later." (p. 820) Detailed Analysis: Criterion Y requires that outcomes be measured at least 75% of a full academic year (roughly 9-10 months) after the intervention begins. As established under criterion T, the entire study - background questionnaire, instruction, practice, immediate posttest, and delayed posttest - spanned only about seven weeks. Because criterion T (one term) is not met, the stricter year-duration criterion Y automatically fails as well, per the prompt's dependency rule. Criterion Y is not met because the roughly seven-week study period is far shorter than 75% of an academic year, and criterion T is already unmet.
    • B

      Balanced Control Group

      • Both groups received identical instruction, tasks, feedback, and practice time, with only the sequencing of the same 20 tasks differing as the treatment variable.
      • "In total, the practice session included 20 tasks (i.e., 10 practice tasks x 2 target speech acts)."
      • Relevant Quotes: 1) "All learners received explicit instruction germane to how to make pragmalinguistically and sociopragmatically appropriate requests and suggestions in English." (p. 818) 2) "Inspired by the design of the practice schedules in Suzuki's (2021) study, at the practice stage of the current study, the IP group partook in intermixing practice (i.e., a-b-a-b-a-b ...)." (p. 819) 3) "In contrast, the BP group first engaged in 10 practice tasks pertinent to making requests and then participated in 10 practice tasks involving making suggestions." (pp. 819-820) 4) "It took approximately three minutes for the participants to complete each practice item. The entire practice session (i.e., 20 practice items) lasted one hour" (p. 820) 5) "In total, the practice session included 20 tasks (i.e., 10 practice tasks x 2 target speech acts)." (p. 822) Detailed Analysis: Applying the criterion B decision tree: did the intervention add extra time, materials, or budget relative to the comparison condition? No. Both groups received the identical two-hour pragmatics instruction, the identical set of 20 computerized corpus-based practice tasks (10 requests and 10 suggestions), the identical exemplar-response feedback after each task, and the identical one-hour practice session length. The only difference between conditions was the ordering of the same practice items (interleaved a-b-a-b versus blocked a...a-b...b), which is precisely the treatment variable under investigation. Since no extra time, materials, or budget were given to either group (EXTRA_RESOURCES_PRESENT is false), the balance requirement is trivially satisfied at branch 1 of the decision tree, regardless of the fact that ordering itself is the manipulated variable. Criterion B is met because both groups received identical instruction, identical practice tasks, and identical time on task, with only the ordering of practice items differing as the treatment variable, so no additional resources requiring balancing were introduced.
  • Level 3 Criteria

    • R

      Reproduced

      • The paper describes itself as the first study of its kind, and a citation search found no independent peer-reviewed replication of this specific experiment by a different research team.
      • "The present study is the first to explore the effects of blocked practice and interleaved practice on L2 pragmatic development."
      • Relevant Quotes: 1) "The present study is the first to explore the effects of blocked practice and interleaved practice on L2 pragmatic development." (p. 813) 2) "The present study is the first to investigate the effects of practice schedules (i.e., interleaved practice versus blocked practice) on accuracy and fluency of L2 learners' pragmatic performance." (p. 817) 3) "Crucially, no studies in the domain of L2 pragmatics have investigated the effects of practice schedules on pragmatic competence" (p. 813) Internet Search Findings: A citation search for this paper (DOI 10.1017/ S0272263123000062) via Semantic Scholar returned four citing works as of this check: (a) Heidari (2026), "The L2 proficiency puzzle: spaced and massed practice in deliberate learning of English collocations," which studies vocabulary collocation learning, not pragmatics or this design; (b) Zhang (2025), "Incidental L2 pragmatics learning through playing a massively multiplayer online role-playing game," a different pragmatics study by the same author using an MMORPG-based design with 169 learners, not a replication of the interleaved-versus-blocked corpus-based DCT design; (c) Zhang (2024), "The Bidirectionality of Pragmatic Transfer in Chinese English Language Learners' Compliment Responses," which examines compliment-response transfer and proficiency, unrelated to practice schedules; and (d) Zhang (2023), "The Influence of Game-Enhanced Communication on EFL Learners' Pragmatic Competence in Compliment Responses," also on compliment responses via gaming. None of these citing works, nor any other paper located through this search, attempts to independently reproduce the interleaved-versus-blocked corpus-based practice design on requests and suggestions reported in this paper. Detailed Analysis: Criterion R requires independent replication of the study by a different research team in a different context, published in a peer-reviewed journal. The paper explicitly and repeatedly positions itself as the first study of practice schedules in L2 pragmatics, so no prior replication of this design exists. Related studies cited in the paper (e.g., Nakata & Suzuki, 2019; Suzuki et al., 2022; Suzuki, 2021) examine interleaving in L2 grammar, speaking, or vocabulary, not the pragmatics-focused corpus-based design tested here, and thus are precursors rather than replications. The citation search conducted for this verification confirms that, more than three years after publication, no independent research team has published a replication of this specific interleaved-versus-blocked corpus-based pragmatics practice experiment; the only closely related subsequent pragmatics work is by the same author, using different methods (gaming-based input) and different outcome behaviours (compliment responses), which does not constitute a reproduction of this study. Criterion R is not met because the study self-identifies as the first of its kind and, following an internet-backed citation search, no independent replication of this specific experiment by a different research team has been published.
    • A

      All-subject Exams

      • Criterion E is unmet and the study assessed only English pragmatics outcomes with custom instruments, not standardised exams across all main subjects.
      • "this study examined both accuracy and fluency of L2 learners' pragmatic behaviors."
      • Relevant Quotes: 1) "this study examined both accuracy and fluency of L2 learners' pragmatic behaviors." (p. 817) 2) "To assess pragmatic accuracy (i.e., declarative pragmatic knowledge), ... a four-level rubric (see Table 3) was adopted to measure the two groups' speech acts of making requests and suggestions" (p. 824) 3) "To measure fluency (i.e., procedural pragmatic knowledge) of the learners' target speech acts, the current study focused on speed fluency assessed by speech rate." (p. 825) Detailed Analysis: Criterion A requires standardised exam-based assessment across all main subjects, and by the prompt's dependency rule it cannot be met if criterion E is not met. Criterion E fails here because outcomes were measured with custom-designed MMDCTs, so criterion A automatically fails. Moreover, the study measured only a narrow slice of one subject area - pragmatic accuracy and fluency for two English speech acts - and assessed no other academic subjects. Although this is higher education where a specialised focus can sometimes be justified, the prerequisite of standardised exam-based assessment is still unmet. Criterion A is not met because criterion E is not met and only a single narrow language-pragmatics outcome was assessed, with no standardised exams across subjects.
    • G

      Graduation Tracking

      • Measurement stopped at a five-week delayed posttest, with no tracking of the sophomore participants through to graduation found in the author's later publications, and criterion Y is not met.
      • "both groups completed a delayed posttest five weeks later."
      • Relevant Quotes: 1) "both groups completed a delayed posttest five weeks later." (p. 820) 2) "all of them intended to obtain their master's or doctoral degrees in the United States after graduating from college." (p. 818) 3) "Future studies may recruit L2 learners from various L1 backgrounds." (p. 833) Internet Search Findings: The same citation search used for criterion R (Semantic Scholar, DOI 10.1017/S0272263123000062, four citing works identified) was reviewed for any follow-up publication by Ying Zhang tracking this cohort of 63 sophomores toward graduation. The subsequent Zhang papers located - "Incidental L2 pragmatics learning through playing a massively multiplayer online role-playing game" (2025, n = 169), "The Bidirectionality of Pragmatic Transfer in Chinese English Language Learners' Compliment Responses" (2024, n = 68), and "The Influence of Game-Enhanced Communication on EFL Learners' Pragmatic Competence in Compliment Responses" (2023, n = 105) - all involve different, larger participant samples and different outcome measures (gaming-based compliment-response pragmatics), and none references following up the original 63-participant interleaved/blocked practice cohort. No paper tracking this specific cohort to graduation was found. Detailed Analysis: Criterion G requires tracking participants until graduation from their educational stage, and by the prompt's dependency rule it cannot be met if criterion Y is not met. Criterion Y fails, so criterion G automatically fails. Substantively, measurement ended with the delayed posttest five weeks after practice; the sophomore participants were not followed to the end of their degree programmes, and the internet search conducted for this verification found no follow-up publication tracking this cohort to graduation. Criterion G is not met because tracking ended five weeks after the intervention with no follow-up to graduation found in a search of the author's subsequent publications, and criterion Y is unmet.
    • P

      Pre-Registered

      • No pre-registration statement, registry identifier, or registration date is mentioned in the paper, and none was found via an external registry search.
      • Relevant Quotes: 1) "(Received 11 July 2022; Revised 30 December 2022; Accepted 04 February 2023)" (p. 812) 2) "SPSS 26 was used to conduct the statistical analyses in the current study." (p. 825) 3) "Competing interests. The author declares none." (p. 833) Internet Search Findings: The paper contains no registry name, identifier, or registration date anywhere in its text. A search for a corresponding pre-registration record (e.g., on OSF or a similar registry) using the study title, author name, and submission window (received 11 July 2022) did not surface any matching registration. No positive evidence of a pre-registered protocol for this study was found. Detailed Analysis: Criterion P requires that the full study protocol, including hypotheses, methods, and planned analyses, be registered on a public registry before data collection began. The paper contains no mention of pre-registration, no registry name (e.g., ClinicalTrials.gov, OSF, AsPredicted), no registration ID, and no registration date anywhere in the methods, acknowledgments, or declarations, and the internet search conducted for this verification did not locate any external registration record for this study either. Without any quoted or externally located evidence of a pre-registered protocol, the criterion cannot be satisfied. Criterion P is not met because the paper contains no reference to any pre-registration of the study protocol, and no external pre-registration record was found.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.