L1 and L2 use in pre-task planning and task repetition: The role of proficiency

Liping Chen, Yuichi Suzuki, and Craig Lambert

Published:
ERCT Check Date:
DOI: 10.1093/applin/amaf059
  • L2 languages
  • higher education
  • China
0
  • C

    No quote confirms class- or school-level randomisation; assignment to planning conditions is only described as "matched," and the unit involved was individually recruited university volunteers, not classes or schools.

    "The participants in the four groups were matched in terms of their CET and EIT to ensure initial comparability of the groups on these features (Appendix S1)." (p. 7)

  • E

    Outcomes were measured with custom-coded temporal speech fluency metrics (speech rate, pausing, repair) derived from an oral narrative task, not with any standardised exam.

    "Four dependent variables (speech rate, within- clause pausing, between-clause pausing, and repair) were calculated following the operationalization of fluency measures in Table 1." (p. 8)

  • T

    The entire study, from planning to the fourth task performance, took place within a single roughly 35-minute session, far short of one academic term.

    "The opportunity to participate in a 35-min research project on speaking English was offered to all students at the university." (p. 7)

  • D

    The control (NP) condition's treatment is briefly stated, but no explicit sample size, demographics, or baseline data specific to that group are quoted in the main text.

    "The control group did not receive time to plan." (p. 7)

  • S

    Randomisation, to the extent it occurred, was among individually recruited university volunteers at a single institution, not among schools.

    "A total of 128 Chinese EFL learners (61 female, 67 male) from a university in mainland China voluntarily participated in the study." (p. 6)

  • I

    The same research team designed the study, collected the data (with two RAs), transcribed and verified the recordings, and conducted the analysis; no independent evaluators are described.

    "Two RAs manually transcribed all task performances, with the first author verifying the transcripts." (p. 8)

  • Y

    Since criterion T (term duration) is not met, and the whole procedure occurred within a single roughly 35-minute session, criterion Y is necessarily not met.

    "The opportunity to participate in a 35-min research project on speaking English was offered to all students at the university." (p. 7)

  • B

    The extra planning time given to the L1P, L2P, and L1P/L2P groups is itself the treatment variable under investigation, so the NP control group's lack of planning time represents the intended business-as-usual baseline rather than an unaddressed imbalance.

    "This study compares four conditions for the language of planning variable: (1) L1P, (2) L2P, (3) combined L1P and L2P (L1P/L2P), and (4) NP as a control." (p. 6)

  • R

    No independent replication of this specific study is mentioned or was found; internet searches confirm this is a very recent (2025) study extending prior work rather than reporting or being the subject of any published reproduction.

    "The current study extends this line of inquiry by comparing speech production under four conditions: L1P, L2P, combined L1P/L2P (50% L1, 50% L2), and no planning (NP), a control condition." (p. 5)

  • A

    Since criterion E (exam-based assessment) is not met, criterion A is automatically not met.

    "Four dependent variables (speech rate, within- clause pausing, between-clause pausing, and repair) were calculated following the operationalization of fluency measures in Table 1." (p. 8)

  • G

    Since criterion Y (year duration) is not met, criterion G is automatically not met, and no follow-up publications tracking this cohort toward graduation were found through internet searches.

    "Finally, participants returned to separate rooms to perform the oral narrative task four times with RAs." (p. 8)

  • P

    No pre-registration of the study protocol, registry platform, or registration date is mentioned anywhere in the paper, and none was found through internet searches of the publisher's page.

Abstract

This is an investigation of the interplay between collaborative pre-task planning language, task repetition, and L2 proficiency in oral narrative task performance. A total of 128 EFL learners engaged in paired collaborative planning under one of four conditions: L1, L2, combined L1-L2, and no planning. Subsequently, each participant individually completed an oral narrative task four times. Task performance was analyzed for utterance fluency, and L2 proficiency was measured using an elicited imitation task. Results indicated that while the language that learners used to plan did not significantly affect fluency, its interaction with learner proficiency was crucial. More proficient L2 learners benefited most from L1 planning, particularly in formulation and monitoring processes, whereas learners with lower proficiency struggled in this condition. Task repetition also enhanced fluency across all conditions, partially mitigating initial proficiency-related differences. The study highlights the importance of learners' individual differences, in this case proficiency level, in task-based language teaching. Findings challenge a one-size-fits-all approach to task implementation and support the judicious incorporation of L1 during pre-task planning, particularly for more proficient learners.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • No quote confirms class- or school-level randomisation; assignment to planning conditions is only described as "matched," and the unit involved was individually recruited university volunteers, not classes or schools.
      • "The participants in the four groups were matched in terms of their CET and EIT to ensure initial comparability of the groups on these features (Appendix S1)." (p. 7)
      • Relevant Quotes: 1) "A total of 128 Chinese EFL learners (61 female, 67 male) from a university in mainland China voluntarily participated in the study." (p. 6) 2) "Before the experiment, all participants booked their sessions individually using an online form. They saw only the time slots available and chose one suitable to their schedule. The two participants who selected the same time slot in the schedule were paired together for the project." (p. 7) 3) "The participants in the four groups were matched in terms of their CET and EIT to ensure initial comparability of the groups on these features (Appendix S1)." (p. 7) Detailed Analysis: Criterion C requires a clearly described randomisation process at the class (or stronger school) level, unless the intervention is a one-to-one tutoring exception. Here, participants were individually recruited university volunteers who self-selected a time slot; pairing of partners was incidental (based on matching time-slot choice), not a description of how pairs/individuals were allocated to the four planning conditions (L1P, L2P, L1P/L2P, NP). The only description of group formation uses the word "matched," which indicates comparability was verified statistically (via ANOVA on CET and EIT) rather than confirming a randomisation procedure. No quote anywhere in the paper states that classes, schools, or even individuals were "randomly assigned" to planning conditions. Furthermore, even if assignment had been random, it occurred among individually recruited volunteer students at a single university, not at the class or school level, and the tutoring exception does not apply since this is a group-planning task, not personal tutoring. Criterion C is not met because there is no quoted evidence of class-level (or stronger) randomisation, and the described unit of participation is individual student volunteers.
    • E

      Exam-based Assessment

      • Outcomes were measured with custom-coded temporal speech fluency metrics (speech rate, pausing, repair) derived from an oral narrative task, not with any standardised exam.
      • "Four dependent variables (speech rate, within- clause pausing, between-clause pausing, and repair) were calculated following the operationalization of fluency measures in Table 1." (p. 8)
      • Relevant Quotes: 1) "Participants completed a picture-cued oral narrative task without linguistic input. The task used a 60-frame picture story from the children's book 'Mr. I' (Trondheim 2007)." (p. 7) 2) "Unfilled pauses of 25 ms or more ... were then identified using PRAAT 6.203 (Boersma and Weenink 2023), and the position of each unfilled pause was coded. Filled pauses and repairs were manually coded in the transcripts." (p. 8) 3) "Four dependent variables (speech rate, within- clause pausing, between-clause pausing, and repair) were calculated following the operationalization of fluency measures in Table 1." (p. 8) Detailed Analysis: The outcome measures are researcher-coded temporal fluency indices (speech rate, within-/between- clause pausing, repair), derived from transcribed and phonetically analysed (PRAAT) recordings of a picture-narration task. This is not a standardised, widely recognised exam of the kind required by criterion E (e.g. a national or state exam); it is a bespoke laboratory measurement protocol built specifically for fluency research. L2 proficiency itself was measured with an elicited imitation test, which, while a recognised research instrument in SLA, is used here as a covariate/ moderator rather than the primary outcome, and is likewise not a standardised curricular exam. Criterion E is not met because the primary outcome measures are custom fluency codings from an oral narrative task, not standardised exam-based assessments.
    • T

      Term Duration

      • The entire study, from planning to the fourth task performance, took place within a single roughly 35-minute session, far short of one academic term.
      • "The opportunity to participate in a 35-min research project on speaking English was offered to all students at the university." (p. 7)
      • Relevant Quotes: 1) "The opportunity to participate in a 35-min research project on speaking English was offered to all students at the university." (p. 7) 2) "Participants first completed EIT tests in separate rooms. They then completed the 10-min collaborative planning session together in the same room. ... Finally, participants returned to separate rooms to perform the oral narrative task four times with RAs." (p. 8) 3) "Before the first task performance, they were informed that there would be a three-minute time limit on each task performance." (p. 8) Detailed Analysis: Criterion T requires that outcomes be measured at least one academic term after the intervention begins. Here, the "intervention" (pre-task planning) and all four outcome measurements (task performances 1 through 4) occur within a single approximately 35-minute laboratory session on one day. There is no follow-up beyond this single sitting. Criterion T is not met because the interval between the intervention and outcome measurement is a single same-day session, far shorter than one academic term.
    • D

      Documented Control Group

      • The control (NP) condition's treatment is briefly stated, but no explicit sample size, demographics, or baseline data specific to that group are quoted in the main text.
      • "The control group did not receive time to plan." (p. 7)
      • Relevant Quotes: 1) "The L1P group used Chinese (L1) for both the question and the table sections. The L2P group used English (L2) for both the question and the table sections. The combined L1P/L2P group used Chinese for the questions section during the first 5 min, then switched to English for the table section in the second 5 min. The control group did not receive time to plan." (p. 7) 2) "The participants in the four groups were matched in terms of their CET and EIT to ensure initial comparability of the groups on these features (Appendix S1). One-way analysis of variance (ANOVA) revealed no significant main effect of group for EIT score, F(1, 50)=0.737, P=.875, and for CET score, F(1, 88)=1.258, P=.521." (p. 7) Detailed Analysis: The paper does state what the control (NP) group experienced (no planning time) and reports that groups were statistically comparable on proficiency measures (CET, EIT) with no significant differences. However, the standard requires a detailed description of the control group's characteristics and size. The main text does not quote the specific sample size of the NP group, nor its demographic breakdown (age, gender) separate from the overall sample of 128; this detail is deferred to "Appendix S1," which is not part of the reviewed text. The documentation available in the main article is therefore only partial. Criterion D is not met because the control group's size and demographic/baseline characteristics are not directly quoted or detailed in the main text, only referenced to an inaccessible appendix.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation, to the extent it occurred, was among individually recruited university volunteers at a single institution, not among schools.
      • "A total of 128 Chinese EFL learners (61 female, 67 male) from a university in mainland China voluntarily participated in the study." (p. 6)
      • Relevant Quotes: 1) "A total of 128 Chinese EFL learners (61 female, 67 male) from a university in mainland China voluntarily participated in the study." (p. 6) 2) "Announcements were made at the university at which the study was conducted in mainland China. The opportunity to participate in a 35-min research project on speaking English was offered to all students at the university." (p. 7) Detailed Analysis: Criterion S requires randomisation among schools (or equivalent institutions), not individuals. This study recruited volunteer undergraduates from a single university via an open announcement, and the described pairing/grouping process operates at the level of individual student volunteers, not separate institutions. Criterion S is not met because the study involves a single university and individually recruited participants, not multiple randomised schools.
    • I

      Independent Conduct

      • The same research team designed the study, collected the data (with two RAs), transcribed and verified the recordings, and conducted the analysis; no independent evaluators are described.
      • "Two RAs manually transcribed all task performances, with the first author verifying the transcripts." (p. 8)
      • Relevant Quotes: 1) "Data were collected from one pair of participants at a time in quiet rooms on campus by the first researcher and two RAs." (p. 7) 2) "Two RAs manually transcribed all task performances, with the first author verifying the transcripts." (p. 8) 3) "Of the 256 task performances, 52 (20%) were double coded by the first author and two RAs." (p. 8) Detailed Analysis: Criterion I requires that data collection and analysis be conducted independently of the team that designed the intervention/study. Here, the first author (who designed the study) directly collected data alongside RAs she supervised, verified the transcripts herself, and was one of the coders. There is no statement of an external, independent agency or blinded third-party evaluator being responsible for data collection or analysis. Criterion I is not met because the designing research team itself collected, transcribed, verified, and analysed the data, with no independent third-party conduct described.
    • Y

      Year Duration

      • Since criterion T (term duration) is not met, and the whole procedure occurred within a single roughly 35-minute session, criterion Y is necessarily not met.
      • "The opportunity to participate in a 35-min research project on speaking English was offered to all students at the university." (p. 7)
      • Relevant Quotes: 1) "The opportunity to participate in a 35-min research project on speaking English was offered to all students at the university." (p. 7) 2) "Finally, participants returned to separate rooms to perform the oral narrative task four times with RAs." (p. 8) Detailed Analysis: Per the ERCT standard, if criterion T is not met, criterion Y is automatically not met. Independently of that rule, the data also plainly show the entire study, planning through the fourth performance, took place in a single same-day session of about 35 minutes, nowhere near the required ~75% of an academic year. Criterion Y is not met, both because criterion T failed and because the study duration was a single short session.
    • B

      Balanced Control Group

      • The extra planning time given to the L1P, L2P, and L1P/L2P groups is itself the treatment variable under investigation, so the NP control group's lack of planning time represents the intended business-as-usual baseline rather than an unaddressed imbalance.
      • "This study compares four conditions for the language of planning variable: (1) L1P, (2) L2P, (3) combined L1P and L2P (L1P/L2P), and (4) NP as a control." (p. 6)
      • Relevant Quotes: 1) "This study compares four conditions for the language of planning variable: (1) L1P, (2) L2P, (3) combined L1P and L2P (L1P/L2P), and (4) NP as a control." (p. 6) 2) "The control group did not receive time to plan." (p. 7) 3) "This worksheet was designed to structure participants' discourse during the 10-min collaborative planning session." (p. 7) Detailed Analysis: Applying the updated criterion B decision procedure: the additional resource in this study is the 10 minutes of collaborative pre-task planning time (and accompanying worksheet) given to the L1P, L2P, and L1P/L2P groups (EXTRA_RESOURCES_PRESENT = true). This extra planning time is not a negligible amount (10 minutes out of a 35-minute session), so the negligible-difference branch does not apply. However, the presence or absence of planning time is precisely the independent variable ("language of planning") that the study is explicitly designed to test (RESOURCES_ARE_TREATMENT = true), with NP explicitly framed as "a control" condition representing the no-planning baseline. This matches the exception in the updated standard: when a study explicitly tests the impact of an additional resource (here, planning time/language) as the primary treatment variable, the control group may receive the standard, resource-free baseline. The NP condition is exactly this kind of business-as-usual comparison, not an incidental or unaddressed confound, and the design is not a within-subjects design requiring a separate baseline check. Criterion B is met because the additional planning time is the explicit treatment variable under study, and the NP control group's lack of planning time is the intended, clearly stated business-as- usual baseline.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study is mentioned or was found; internet searches confirm this is a very recent (2025) study extending prior work rather than reporting or being the subject of any published reproduction.
      • "The current study extends this line of inquiry by comparing speech production under four conditions: L1P, L2P, combined L1P/L2P (50% L1, 50% L2), and no planning (NP), a control condition." (p. 5)
      • Relevant Quotes: 1) "The current study extends this line of inquiry by comparing speech production under four conditions: L1P, L2P, combined L1P/L2P (50% L1, 50% L2), and no planning (NP), a control condition." (p. 5) Detailed Analysis: Criterion R requires independent replication of this specific study by a different research team, published in a peer-reviewed outlet. The paper itself frames its four-condition design (L1P, L2P, L1P/L2P, NP) as a novel extension of prior planning- language research (e.g. Lambert et al. 2021), not as a replication of an existing study. The article was published as an advance-access piece on 22 September 2025. Internet searches (web search for the title, authors, and DOI, and for any citing or replicating work) found only the original publication itself, a publication announcement, and unrelated pre-task-planning literature; no independent replication of this exact design and cohort was found, which is unsurprising given the paper's very recent publication date relative to this check (July 2026). Criterion R is not met because this is an original, recently published study with no evidence of independent replication found through internet searches.
    • A

      All-subject Exams

      • Since criterion E (exam-based assessment) is not met, criterion A is automatically not met.
      • "Four dependent variables (speech rate, within- clause pausing, between-clause pausing, and repair) were calculated following the operationalization of fluency measures in Table 1." (p. 8)
      • Relevant Quotes: 1) "Four dependent variables (speech rate, within- clause pausing, between-clause pausing, and repair) were calculated following the operationalization of fluency measures in Table 1." (p. 8) Detailed Analysis: Per the ERCT standard, criterion A requires criterion E to be met as a prerequisite. The study measures only oral fluency on a single narrative task via custom temporal coding, not standardised exams in any subject, let alone across all main subjects. There is also no assessment of other curricular subjects (e.g. mathematics, general language arts beyond speaking fluency) alongside the target measure. Criterion A is not met because criterion E is not met, and no all-subject standardised assessment is reported.
    • G

      Graduation Tracking

      • Since criterion Y (year duration) is not met, criterion G is automatically not met, and no follow-up publications tracking this cohort toward graduation were found through internet searches.
      • "Finally, participants returned to separate rooms to perform the oral narrative task four times with RAs." (p. 8)
      • Relevant Quotes: 1) "Finally, participants returned to separate rooms to perform the oral narrative task four times with RAs." (p. 8) Detailed Analysis: Per the ERCT standard, if criterion Y is not met, criterion G is automatically not met. There is also no mention anywhere in the paper of any follow-up data collection after the single-session fourth task performance, let alone tracking participants to graduation. Internet searches for follow-up papers by the same author team tracking this cohort found no such publications; only the original article and an unrelated body of prior work by these authors on planning and task repetition were located. Criterion G is not met because criterion Y failed and no graduation or long-term follow-up tracking is reported or discoverable.
    • P

      Pre-Registered

      • No pre-registration of the study protocol, registry platform, or registration date is mentioned anywhere in the paper, and none was found through internet searches of the publisher's page.
      • Relevant Quotes: 1) "Supplementary data is available at Applied Linguistics online." (p. 17) 2) "Funding: The preparation of this article was supported by a research grant (YQWH2022004) that the first author received from Research Center for langauge and Culture Studies in International Oil and Gas Regions, the School of Foreign Languages, Southwest Pertoleum University." (p. 17) Detailed Analysis: No statement referencing a pre-registration registry (e.g. OSF, AsPredicted, a trial registry), registration ID, or registration date appears anywhere in the paper, including in the funding, supplementary data, or methods sections that were reviewed. A check of the publisher's landing page for this article likewise surfaced no reference to pre-registration, a study registry, or trial registration. Criterion P is not met because there is no quoted or discoverable evidence of a pre-registered protocol.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.