The Poster Carousel in the ESL Classroom: What Happens to Learners' L2 Fluency During Same and Parallel-Task Repetition?

Ann-Marie Hunter

Published:
ERCT Check Date:
DOI: 10.1002/tesq.3257
  • L2 languages
  • adult education
  • UK
0
  • C

    Intact classes (not individual students within a class) were randomly assigned to the two task-repetition conditions, so the unit of randomisation was the class.

    "Intact classes of EFL learners at a private English language school in Central London were randomly assigned to take part in either a same-task repetition (STR) poster carousel training session (n = 24) or a parallel-task repetition (PTR) carousel training session (n = 22) during their normal scheduled English class." (pp. 6-8)

  • E

    Outcomes were custom research measures of utterance fluency derived from manual PRAAT analysis of recorded speech, not any standardised, widely recognised exam.

    "their performances were audio-recorded, transcribed, and analyzed using PRAAT to detail their L2 utterance fluency." (Abstract, p. 1)

  • T

    The entire intervention and all outcome measurement took place within a single scheduled class session, far short of one academic term of follow-up.

    "In both conditions, participants gave three oral performances (a narrative task) ... during their normal scheduled English class." (p. 8)

  • D

    There is no business-as-usual control group, and the comparison (PTR) group's demographics and baseline characteristics are not documented separately from the pooled sample.

    "46 ESL students took part in the study and the age range of participants was 18-42 (M = 20.1 years)." (p. 8)

  • S

    Randomisation occurred among intact classes within a single private language school, not among schools or sites.

    "Intact classes of EFL learners at a private English language school in Central London were randomly assigned..." (pp. 6-8)

  • I

    The sole author designed the two carousel versions, implemented the study, and manually performed the fluency analysis herself, with no independent evaluation team.

    "each extract was annotated manually by the author using the standard textgrid feature." (p. 14)

  • Y

    Since the entire study occurred within one class session, tracking did not remotely approach 75% of an academic year (and criterion T is already not met).

    "In both conditions, participants gave three oral performances (a narrative task) ... during their normal scheduled English class." (p. 8)

  • B

    Both randomised conditions received the same amount of class time and an identical activity structure (three storyboard retellings during a normal lesson); the only difference - same versus different stories - is the treatment contrast itself.

    "In both conditions, participants gave three oral performances (a narrative task) - in the STR group, the three performances related to the same task and in the PTR group they were three different tasks of the same type." (p. 8)

  • R

    No independent, peer-reviewed replication of this specific poster carousel fluency study exists; the paper itself calls for future replications and a targeted citation search found only review/overview articles citing the study, not replications.

    "Manual fluency analysis of classroom data is time-consuming, but replications could be carried out which make use of existing PRAAT functions and specially created scripts to measure fluency automatically." (p. 28)

  • A

    Only L2 speaking fluency was measured with custom instruments; no other subjects were assessed and criterion E is not met, which automatically fails this criterion.

    "their performances were audio-recorded, transcribed, and analyzed using PRAAT to detail their L2 utterance fluency." (Abstract, p. 1)

  • G

    Measurement ended within the single training session with no follow-up of participants, let alone tracking to graduation; a search for follow-up publications by the same author found none, and prerequisite criterion Y is not met.

    "The case study presented here has shown that increased fluency during task repetition may not be expected to translate into longer term increased fluency..." (pp. 27-28)

  • P

    The paper contains no mention of any pre-registration, registry platform, or registered protocol, and a targeted search of registries and citation databases found no pre-registration record for this study.

Abstract

This paper reports on the impact of an English as a Second Language (ESL) speaking activity - the poster carousel - on English learners' second language (L2) fluency. Two versions of the poster carousel were developed to observe the effect of talking about (a) the same poster three times (same-task repetition) or (b) three different posters (parallel-task repetition). 46 ESL learners took part, and their performances were audio-recorded, transcribed, and analyzed using PRAAT to detail their L2 utterance fluency. The findings suggest that learners were more fluent during repeated performances in the same-task repetition poster carousel group. No changes in fluency were observed for the parallel-task repetition poster carousel group. A detailed case study explores these observed fluency increases in the same-task repetition group and generates further hypotheses for empirical exploration. Some observations are made which relate to the selection of fluency measures in L2 fluency research.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Intact classes (not individual students within a class) were randomly assigned to the two task-repetition conditions, so the unit of randomisation was the class.
      • "Intact classes of EFL learners at a private English language school in Central London were randomly assigned to take part in either a same-task repetition (STR) poster carousel training session (n = 24) or a parallel-task repetition (PTR) carousel training session (n = 22) during their normal scheduled English class." (pp. 6-8)
      • Relevant Quotes: 1) "Intact classes of EFL learners at a private English language school in Central London were randomly assigned to take part in either a same-task repetition (STR) poster carousel training session (n = 24) or a parallel-task repetition (PTR) carousel training session (n = 22) during their normal scheduled English class." (pp. 6-8) 2) "A mixed between- and within-subjects design was used to look for fluency changes within-subjects across three task performances and between-subjects to compare changes between performances for the two different groups." (p. 8) 3) "A preliminary power analysis with G-Power recommended a total sample size of 86. This was not a possibility in the current study for reasons of practicality (intact classes) and availability of time for manual analysis." (p. 8, footnote 2) Detailed Analysis: Criterion C requires randomisation at the class level (or stronger) rather than randomising individual students within a single classroom. The paper explicitly states that "intact classes" were randomly assigned to the STR or PTR condition, which means the unit of randomisation was the whole class, preventing within-class contamination between the two conditions. The description of the randomisation process is brief (the exact number of classes and the randomisation method are not reported), but the quote clearly establishes that entire classes, not individual students, were the assigned unit, which is what the criterion requires. The exception for tutoring does not apply here because this is a whole-class activity. Criterion C is met because intact classes were randomly assigned to conditions, satisfying the class-level randomisation requirement.
    • E

      Exam-based Assessment

      • Outcomes were custom research measures of utterance fluency derived from manual PRAAT analysis of recorded speech, not any standardised, widely recognised exam.
      • "their performances were audio-recorded, transcribed, and analyzed using PRAAT to detail their L2 utterance fluency." (Abstract, p. 1)
      • Relevant Quotes: 1) "their performances were audio-recorded, transcribed, and analyzed using PRAAT to detail their L2 utterance fluency." (Abstract) 2) "This study adopts three dependent variables which are traditionally associated with speed, breakdown and repair aspects of utterance fluency, respectively." (p. 7) 3) "articulation rate was calculated as the total number of raw syllables produced in the 1-minute sample from each performance that was analyzed, divided by total speaking time (excluding all pauses >250 ms), and multiplied by 60." (p. 7) 4) "each extract was annotated manually by the author using the standard textgrid feature." (p. 14) 5) "A unique PRAAT script was developed which would generate frequencies and durations for the intervals that had been manually created. This output was then used to calculate the specific fluency measures as outlined above." (p. 14) Detailed Analysis: Criterion E requires that outcomes be measured with a standard, widely recognised, standardised exam rather than instruments created for the study. Here the outcomes were three researcher-defined utterance-fluency metrics (articulation rate, frequency of mid-clause pauses, frequency of repetitions) computed from one-minute speech samples that were manually annotated by the author in PRAAT using a bespoke script. These are research measures specific to this study's analytical framework, not a named standardised examination such as IELTS, TOEFL, or a national curriculum exam. The only standardised reference is the CEFR level used for course placement via an in-house school test, which is itself not a standardised exam and was not an outcome measure. Criterion E is not met because outcomes were measured with custom PRAAT-based fluency metrics rather than a recognised standardised exam.
    • T

      Term Duration

      • The entire intervention and all outcome measurement took place within a single scheduled class session, far short of one academic term of follow-up.
      • "In both conditions, participants gave three oral performances (a narrative task) ... during their normal scheduled English class." (p. 8)
      • Relevant Quotes: 1) "Intact classes of EFL learners ... were randomly assigned to take part in either a same-task repetition (STR) poster carousel training session ... or a parallel-task repetition (PTR) carousel training session ... during their normal scheduled English class." (pp. 6-8) 2) "In both conditions, participants gave three oral performances (a narrative task) - in the STR group, the three performances related to the same task and in the PTR group they were three different tasks of the same type." (p. 8) 3) "Participants were given up to 2 minutes to tell their stories, but they did not always use all the allotted time so a 1-minute extract was taken from the beginning of each performance." (p. 14) 4) "Steps 5 and 6 are repeated until each A has told the same story 3 times to 3 different visitors." (p. 9) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. In this study the intervention was a single poster carousel training session conducted "during their normal scheduled English class," and the outcome data were the three task performances recorded within that same session. There was no delayed post-test and no follow-up measurement of any kind; the interval from intervention start to final measurement is on the order of one lesson, not one term. The paper itself notes that longer-term or transfer benefits were not captured: fluency changes were observed only across immediate repetitions. Criterion T is not met because all outcomes were measured within the same single class session in which the intervention took place, far short of a term.
    • D

      Documented Control Group

      • There is no business-as-usual control group, and the comparison (PTR) group's demographics and baseline characteristics are not documented separately from the pooled sample.
      • "46 ESL students took part in the study and the age range of participants was 18-42 (M = 20.1 years)." (p. 8)
      • Relevant Quotes: 1) "46 ESL students took part in the study and the age range of participants was 18-42 (M = 20.1 years). First languages spoken were: Arabic (3), Kasakh (1), Portuguese (2), Slovenian (1), Serbian (1), Thai (1), Ukrainian (1), French (8), German (7), Spanish (4), Chinese (2), Japanese (4), Korean (4), Swedish (1), Dutch (1), and Italian (5)." (p. 8) 2) "The proficiency level of participants was Intermediate (B1, CEFR), and the school allocated new students to a proficiency level based on their performance on an in-house grammar test, a short oral interview, and a writing sample." (p. 8) 3) "Intact classes of EFL learners ... were randomly assigned to take part in either a same-task repetition (STR) poster carousel training session (n = 24) or a parallel-task repetition (PTR) carousel training session (n = 22)." (pp. 6-8) Detailed Analysis: Criterion D requires detailed documentation of the control group, including demographic information, baseline performance, and the conditions it received. This study has no untreated or business-as-usual control group at all: both arms received a version of the poster carousel intervention (same-task vs parallel-task repetition), so the PTR group is at best an active comparison condition. Even treating the PTR group as the control, the paper reports demographics (age, first languages, proficiency) only for the pooled sample of 46 students, with no breakdown by group, no per-group demographic table, and no baseline equivalence analysis. Group sizes (n = 24 and n = 22) and the PTR procedure are described, and Table 1 reports per-group performance-1 fluency descriptives, but there is no documentation showing that the two intact classes were comparable in composition at baseline, which is the core purpose of this criterion. Criterion D is not met because the study lacks a documented control group: demographics and baseline characteristics are reported only for the whole sample, not for the comparison group specifically, and no business-as-usual control exists.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred among intact classes within a single private language school, not among schools or sites.
      • "Intact classes of EFL learners at a private English language school in Central London were randomly assigned..." (pp. 6-8)
      • Relevant Quotes: 1) "Intact classes of EFL learners at a private English language school in Central London were randomly assigned to take part in either a same-task repetition (STR) poster carousel training session (n = 24) or a parallel-task repetition (PTR) carousel training session (n = 22)." (pp. 6-8) Detailed Analysis: Criterion S requires randomisation at the level of whole schools, sites, or equivalent implementing institutions. This study took place at a single private English language school in Central London, and the randomised units were intact classes within that one school. No multiple schools or sites were involved and no school-level assignment occurred, so the stronger school-level requirement cannot be satisfied. Criterion S is not met because the study randomised classes within one single school rather than randomising schools or sites.
    • I

      Independent Conduct

      • The sole author designed the two carousel versions, implemented the study, and manually performed the fluency analysis herself, with no independent evaluation team.
      • "each extract was annotated manually by the author using the standard textgrid feature." (p. 14)
      • Relevant Quotes: 1) "Through piloting and discussion with class teachers, two new versions of the poster carousel were developed that required learners to create storyboards and talk about extreme life experiences." (p. 8) 2) "each extract was annotated manually by the author using the standard textgrid feature." (p. 14) 3) "A unique PRAAT script was developed which would generate frequencies and durations for the intervals that had been manually created." (p. 14) 4) "A 10% sample of the data was re-examined for syllable count (used in the calculation of articulation rate) and pause boundaries by a trained researcher. The second rater coded 22 samples of speech (roughly 10% of the total data) and Spearman's correlations revealed an exceptionally high reliability score (0.993) between scorers." (pp. 14-15) 5) "Ann-Marie Hunter is a Lecturer in TESOL and Linguistics at York St John University. Her primary research focus is L2 speaking and language teaching." (p. 28, The Author) Detailed Analysis: Criterion I requires that the study be conducted independently of the people who designed the intervention, for example by an external evaluation team. Here the paper is single-authored; the author (with input from class teachers) developed the two poster carousel versions, ran the classroom sessions, and personally carried out the manual PRAAT annotation and analysis of every recording. The only external involvement was a second rater re-checking about 10% of syllable counts and pause boundaries for reliability, which is a quality check, not independent conduct of data collection or analysis. There is no statement of third-party oversight or an external evaluator. Criterion I is not met because the same person designed the intervention, implemented it, and analysed the outcomes with no independent evaluation team.
    • Y

      Year Duration

      • Since the entire study occurred within one class session, tracking did not remotely approach 75% of an academic year (and criterion T is already not met).
      • "In both conditions, participants gave three oral performances (a narrative task) ... during their normal scheduled English class." (p. 8)
      • Relevant Quotes: 1) "Intact classes of EFL learners ... were randomly assigned to take part in either a same-task repetition (STR) poster carousel training session ... during their normal scheduled English class." (pp. 6-8) 2) "Participants were given up to 2 minutes to tell their stories, but they did not always use all the allotted time so a 1-minute extract was taken from the beginning of each performance." (p. 14) 3) "The case study presented here has shown that increased fluency during task repetition may not be expected to translate into longer term increased fluency but may have been invested in other ways, giving rise to increased complexity or accuracy in the longer term, for example." (pp. 27-28) Detailed Analysis: Criterion Y requires outcome tracking covering at least 75% of a full academic year from intervention start. Per the prompt rules, criterion Y automatically fails when criterion T fails, and T is not met here. The intervention and all measurements were confined to a single lesson: three immediate task performances recorded in one session, with no delayed follow-up whatsoever. The tracking interval is therefore minutes to an hour rather than months. Criterion Y is not met because the study spanned only a single class session, nowhere near an academic year, and the prerequisite criterion T also fails.
    • B

      Balanced Control Group

      • Both randomised conditions received the same amount of class time and an identical activity structure (three storyboard retellings during a normal lesson); the only difference - same versus different stories - is the treatment contrast itself.
      • "In both conditions, participants gave three oral performances (a narrative task) - in the STR group, the three performances related to the same task and in the PTR group they were three different tasks of the same type." (p. 8)
      • Relevant Quotes: 1) "In both conditions, participants gave three oral performances (a narrative task) - in the STR group, the three performances related to the same task and in the PTR group they were three different tasks of the same type." (p. 8) 2) "Intact classes of EFL learners ... were randomly assigned to take part in either a same-task repetition (STR) poster carousel training session ... or a parallel-task repetition (PTR) carousel training session ... during their normal scheduled English class." (pp. 6-8) 3) "Each pair creates a six-frame storyboard on poster paper which illustrates the text they were assigned." (STR step 2, p. 9; PTR step 2, p. 11) 4) "In the PTR condition, speakers spoke about three different topics, but the procedure followed ensured that the stimulus stories were naturally counterbalanced among the participants to limit the possibility of a task effect." (p. 11) 5) "Careful piloting of the materials revealed that they generated oral accounts that were similar in length and structure." (p. 14) Detailed Analysis: Applying the Criterion B decision tree from the updated specification: the first check is whether the intervention group received any extra time or budget relative to the comparison group. Here both groups took part during their normal scheduled English class, used the same stimulus texts and storyboard materials, and gave the same number (three) of timed oral performances. Neither condition (STR or PTR) received extra instructional time, money, or materials relative to the other - this is a comparison of two active variants of the same activity, analogous to Example 4 in the specification (active control with comparable time and structure). Since EXTRA_RESOURCES_PRESENT is false, the decision tree returns "met" without needing to evaluate whether resources are integral to the treatment. The single difference between arms - repeating the same story versus telling three different counterbalanced stories - is precisely the treatment variable being tested (same-task vs parallel-task repetition), i.e., integral to the contrast rather than a separable confounding add-on. Criterion B is met because both conditions received identical time, materials, and activity structure, with the only difference being the task-repetition contrast under test, so no extra-resource imbalance exists to evaluate.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent, peer-reviewed replication of this specific poster carousel fluency study exists; the paper itself calls for future replications and a targeted citation search found only review/overview articles citing the study, not replications.
      • "Manual fluency analysis of classroom data is time-consuming, but replications could be carried out which make use of existing PRAAT functions and specially created scripts to measure fluency automatically." (p. 28)
      • Relevant Quotes: 1) "Manual fluency analysis of classroom data is time-consuming, but replications could be carried out which make use of existing PRAAT functions and specially created scripts to measure fluency automatically (De Jong & Wempe, 2009) on a larger scale." (p. 28) 2) "To date, there has not been a detailed empirical exploration of the poster carousel and its impact on utterance fluency." (p. 6) 3) "Adopting a case study methodology, Lynch and Maclean (2000, 2001) showed how the quality of the talk of five L2 English learners improved over the six performances of the poster carousel..." (p. 5) Internet Search Findings (Criterion R): A citation search (Crossref and Semantic Scholar, via DOI 10.1002/tesq.3257) shows the paper has been cited by three later works as of the check date: Kim & Namkung (2026), "Task Variation, Language Production, and Learning Outcomes in Task-Based Language Teaching Research: Insights from Five Decades of TESOL Quarterly" (TESOL Quarterly); Kahng (2026), "Fluency" (TESOL Quarterly, doi 10.1002/tesq.70117), whose abstract describes it as an overview article synthesizing "dominant theoretical frameworks ... varied methodological approaches ... and pedagogical applications" of L2 fluency research rather than a new empirical replication; and Kakitani (2025), "Oral task repetition research via videoconferencing" (Research Methods in Applied Linguistics), a methodological discussion of task-repetition research designs. None of these three citing works report an independent empirical reproduction of Hunter's specific poster carousel same-task/parallel-task comparison in a new sample; they cite the study as prior literature within broader reviews or methods discussions. No dedicated replication study of this specific classroom experiment was located in any available source. Detailed Analysis: Criterion R requires that the study be independently replicated by a different team in a different context and published in a peer-reviewed journal. The paper explicitly positions itself as the first detailed empirical exploration of the poster carousel's impact on utterance fluency, and it frames replication as future work. Prior related work (Lynch & Maclean 2000, 2001) predates this study and used a different, non-experimental case study methodology, so it is background rather than replication. The broader task repetition literature (e.g., Lambert et al., Suzuki et al.) studies related phenomena but does not reproduce this specific randomised classroom comparison of same-task versus parallel-task poster carousels. The internet-backed citation search confirms that, as of the check date, no independent peer-reviewed replication of this 2023 study has appeared; the works that cite it are literature-review or methods-discussion articles, not reproductions of its design. Criterion R is not met because no independent peer-reviewed replication of this specific study exists; the author herself calls for future replications and none have been identified.
    • A

      All-subject Exams

      • Only L2 speaking fluency was measured with custom instruments; no other subjects were assessed and criterion E is not met, which automatically fails this criterion.
      • "their performances were audio-recorded, transcribed, and analyzed using PRAAT to detail their L2 utterance fluency." (Abstract, p. 1)
      • Relevant Quotes: 1) "their performances were audio-recorded, transcribed, and analyzed using PRAAT to detail their L2 utterance fluency." (Abstract) 2) "This study adopts three dependent variables which are traditionally associated with speed, breakdown and repair aspects of utterance fluency, respectively." (p. 7) Detailed Analysis: Criterion A requires standardised exam-based assessment across all main subjects taught, with criterion E as a prerequisite. Criterion E is not met (custom PRAAT fluency measures rather than standardised exams), so criterion A automatically fails. Moreover, the study assessed only one narrow domain - L2 oral fluency in English - and no other subjects or even other aspects of English proficiency (reading, writing, listening) were measured. Although this is a specialised adult language-school context where an exception might be argued for focusing on English, the absence of any standardised exam means the criterion cannot be satisfied. Criterion A is not met because criterion E fails and only a single custom-measured outcome domain (L2 speaking fluency) was assessed.
    • G

      Graduation Tracking

      • Measurement ended within the single training session with no follow-up of participants, let alone tracking to graduation; a search for follow-up publications by the same author found none, and prerequisite criterion Y is not met.
      • "The case study presented here has shown that increased fluency during task repetition may not be expected to translate into longer term increased fluency..." (pp. 27-28)
      • Relevant Quotes: 1) "The case study presented here has shown that increased fluency during task repetition may not be expected to translate into longer term increased fluency but may have been invested in other ways, giving rise to increased complexity or accuracy in the longer term, for example." (pp. 27-28) 2) "Finally, it would be valuable to explore learners' perspectives on their fluency increases during the poster carousel activity. This could be achieved through stimulated recall and post-intervention interviews, for example." (p. 28) Internet Search Findings (Criterion G): A search of Ann-Marie Hunter's publication record (via Semantic Scholar author search) lists her works from 2016 through 2026, including "Aspects of Fluency Across Assessed Levels of Speaking Proficiency" (2020), "Spacing effects on repeated L2 task performance" (2019), and two 2023 and later papers on language-teacher leadership and adult migrant language policy in England. None of these titles indicate a longitudinal follow-up tracking the same 46 adult ESL learners from the poster carousel study, and no separate follow-up paper reporting graduation or longer-term outcomes for this specific cohort was found in any available source. Detailed Analysis: Criterion G requires tracking participants until graduation from their educational stage. Per the prompt rules, G automatically fails when criterion Y fails, and Y is not met here. All measurement occurred within the single carousel session; there was no delayed post-test, no longitudinal follow-up, and no graduation data. The participants were adult students at a private language school studying general English rather than working toward a specific graduation milestone, and the internet search confirms no follow-up publications tracking this cohort exist. Criterion G is not met because no follow-up of any kind occurred after the single session, no subsequent tracking publication was found, and the prerequisite criterion Y also fails.
    • P

      Pre-Registered

      • The paper contains no mention of any pre-registration, registry platform, or registered protocol, and a targeted search of registries and citation databases found no pre-registration record for this study.
      • Relevant Quotes: No quotes mentioning pre-registration exist in the paper. The methodology section describes the design, dependent variables, procedure, materials, and statistical analysis (pp. 6-16) without any reference to a registry. 1) "A preliminary power analysis with G-Power recommended a total sample size of 86. This was not a possibility in the current study for reasons of practicality (intact classes) and availability of time for manual analysis." (p. 8, footnote 2) Internet Search Findings (Criterion P): A search for a pre-registration record associated with this study (via bibliographic/registry search covering common registries such as OSF and AsPredicted) returned no matching record linking Ann-Marie Hunter or the poster carousel study to any registered protocol. Detailed Analysis: Criterion P requires that the full study protocol (hypotheses, methods, planned analyses) be registered on a public registry before data collection began. The paper never mentions ClinicalTrials.gov, OSF, AsPredicted, a trial registry ID, or any pre-registered protocol. The only planning artefact mentioned is an a priori power analysis, which is not a pre-registration. The internet search corroborates the absence of any registry entry for this study. With no quoted evidence of registration or its timing, the criterion cannot be satisfied. Criterion P is not met because no pre-registration of the study protocol is mentioned anywhere in the paper and none was located through internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.