Abstract
The present study investigates the effect of different types of task implementation on teaching L2 interactional sequences. 81 EFL learners were randomly assigned to one of three experimental groups. In the first experimental group (T1-EG, n = 27), implicit instruction appeared during the pre-task phase, while the post-task phase included an explicit focus on forms. The second group (T2-EG, n = 32) received implicit instruction in the target structures and a reactive focus on form during task performance. The third group (PPP-EG, n = 27) followed a presentation - practice - production (PPP) lesson framework. Groups' pragmatic production was measured using written discourse completion tasks. Results showed that in the current study, all three groups reported gains, yet the implicit-explicit condition (T1-EG) appeared to be more beneficial for teaching the interactional sequences than the implicit-only (T2-EG) or the PPP (PPP-EG) framework.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was at the individual student level, not at the class or school level, and the intervention was whole-class teaching, so no tutoring exception applies.
- "In the study, 81 homogenous upper-intermediate EFL learners were randomly assigned to three experimental groups (T1-EG, n = 27, T2-EG, n = 27, and PPP-EG, n = 27)." (p. 509)
Relevant Quotes:
1) "81 EFL learners were randomly assigned to one of three experimental groups." (p. 502, abstract)
2) "In the study, 81 homogenous upper-intermediate EFL learners were randomly assigned to three experimental groups (T1-EG, n = 27, T2-EG, n = 27, and PPP-EG, n = 27). Care was taken to ensure an even number of participants." (p. 509)
3) "They were chosen for the study following convenience sampling, i.e., based on their common level of proficiency and because they were all taught by the present author." (p. 507)
4) "Firstly, since random sampling was not feasible in the present study, readers should judge whether the findings could apply in the case of learners of lower levels of proficiency or in second language contexts." (p. 516)
Detailed Analysis:
Criterion C requires randomisation at the class level (or stronger, school level), unless the intervention is one-to-one tutoring. The paper states that the 81 individual learners "were randomly assigned to one of three experimental groups," which is randomisation at the individual student level, not the class or school level. All participants came from the same secondary school, were taught by the same teacher-researcher, and the three conditions were run in parallel within this single setting, which creates the exact contamination risk that class-level randomisation is designed to prevent. The intervention is whole-class lessons (45-minute lessons with pair and group work), not personal tutoring, so the tutoring exception does not apply.
Criterion C is not met because randomisation was performed at the individual student level within a single school context rather than at the class or school level, and no tutoring exception applies.
-
E
Exam-based Assessment
- Outcomes were measured with a custom, author-made written discourse completion test rather than any recognised standardised exam.
- "The WDCTs were created for the purpose of this study and validated using pilot testing on a comparable group of learners." (p. 507)
Relevant Quotes:
1) "Written discourse completion tasks (WDCTs) were chosen as research instruments in the present study." (p. 507)
2) "The WDCTs were created for the purpose of this study and validated using pilot testing on a comparable group of learners." (p. 507)
3) "As the researcher was also the participants' teacher, the 15 items were adjusted to learners' developmental levels and their past learning experiences (following Mackey, Gass, 2022)." (pp. 507-508)
4) "A maximum of three points was awarded for a correct response to each scenario." (p. 508)
Detailed Analysis:
Criterion E requires the use of standardised, widely recognised exam-based assessments rather than instruments specially designed for the study. The paper explicitly states that the WDCTs "were created for the purpose of this study," were piloted by the author, and were adjusted by the teacher-researcher to the learners' developmental levels. The scoring rubric (0-3 points per scenario) was also devised by the author, with only intra-rater reliability reported (93.3%). This is precisely the custom, researcher-made assessment that the standard identifies as problematic because it may be overly aligned with the intervention. No national, state-wide, or otherwise recognised standardised exam was used at any point.
Criterion E is not met because outcomes were measured with custom written discourse completion tasks created by the author for this study, not with a standardised exam.
-
T
Term Duration
- Outcomes were last measured about three weeks after a short four-lesson intervention, far less than one full academic term of tracking.
- "The necessary data were collected three times: before the intervention (pre-test), within two days after the fourth lesson (post-test), and three weeks later (a delayed post-test)." (p. 511)
Relevant Quotes:
1) "The participants received a series of 4 lessons focused on three interactional sequences." (p. 509)
2) "The necessary data were collected three times: before the intervention (pre-test), within two days after the fourth lesson (post-test), and three weeks later (a delayed post-test)." (p. 511)
3) "Also, the study followed a short intervention of 4 lessons. A longitudinal study of the effects of the three types of instruction might shed more light on their effectiveness." (p. 516)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Here the intervention consisted of only four 45-minute lessons. The post-test was administered within two days of the fourth lesson and the delayed post-test only three weeks later. Even taking the delayed post-test as the final measurement point, the total interval from intervention start to final measurement is on the order of one to two months at most (four lessons plus three weeks), far short of a full academic term. The author himself acknowledges the intervention was short and calls for longitudinal follow-up.
Criterion T is not met because the final outcome measurement occurred only about three weeks after a four-lesson intervention, well short of one full academic term.
-
D
Documented Control Group
- The business-as-usual comparison group (PPP-EG) is documented with its size, shared demographic profile, condition received, and baseline pre-test scores.
- "PPP-EG followed their regular coursebook lessons." (p. 510)
Relevant Quotes:
1) "The participants in the study were 81 Polish secondary/high school learners of English as a foreign language in a town in the north of Poland." (p. 507)
2) "They were all aged 17 at the time of the study. All participants had already had at least six years of compulsory English instruction in their primary school. Their secondary school offered five hours of English per week, and the study took place while the learners were in the third grade." (p. 507)
3) "The participants' level of proficiency at the time of the study can be described as B2+/upper-intermediate (using the CEFR scale) or Advanced Mid (using the ACTFL rating)." (p. 507)
4) "PPP-EG followed their regular coursebook lessons. In the initial phases of the lesson, an interactional sequence was presented to the learners through explicit metapragmatic instruction." (p. 510)
5) "pre-test PPP-EG 28.25 3.25" (Table 1, p. 511)
Detailed Analysis:
Criterion D requires that the comparison/control group be well documented in terms of demographics, baseline performance, and treatment received. This study has no untreated control; instead the PPP-EG, which "followed their regular coursebook lessons", serves as the business-as-usual comparison against the two task-based conditions. The paper documents this group's size (n = 27), the demographic profile shared by the whole homogeneous sample (Polish third-grade secondary students, aged 17, B2+/upper-intermediate, five hours of English per week, taught by the same teacher for over two years), exactly what instruction the group received (explicit metapragmatic presentation, about 20 minutes of language-focused exercises, then one task performance with focus on form), and its baseline pre-test mean and standard deviation in Table 1 (M = 28.25, SD = 3.25). This allows a proper comparison and interpretation of results. All narrative quotes were verified verbatim against the PDF; item 5 is a transcription of the corresponding Table 1 row.
Criterion D is met because the comparison group's size, characteristics, baseline scores, and received condition are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- The trial involved a single school with student-level random assignment, so no school-level randomisation occurred.
- "The participants in the study were 81 Polish secondary/high school learners of English as a foreign language in a town in the north of Poland." (p. 507)
Relevant Quotes:
1) "In the study, 81 homogenous upper-intermediate EFL learners were randomly assigned to three experimental groups (T1-EG, n = 27, T2-EG, n = 27, and PPP-EG, n = 27)." (p. 509)
2) "The participants in the study were 81 Polish secondary/high school learners of English as a foreign language in a town in the north of Poland." (p. 507)
Detailed Analysis:
Criterion S requires randomisation among schools (or equivalent implementing units). This study took place in a single secondary school in northern Poland, with all participants taught by the same teacher-researcher, and randomisation was performed at the individual student level into three groups. No multiple schools or sites were involved, and no school-level assignment of any kind is described.
Criterion S is not met because the study was conducted in a single school with student-level randomisation, not a school-level RCT.
-
I
Independent Conduct
- A single teacher-researcher designed the intervention, taught all groups, created and scored the tests, and analysed the data without any independent oversight.
- "This means that their teacher (the teacher-researcher) had taught them for over two years, having conducted around 250 classes of 45 minutes each." (p. 507)
Relevant Quotes:
1) "They were chosen for the study following convenience sampling, i.e., based on their common level of proficiency and because they were all taught by the present author." (p. 507)
2) "This means that their teacher (the teacher-researcher) had taught them for over two years, having conducted around 250 classes of 45 minutes each." (p. 507)
3) "As the researcher was also the participants' teacher, the 15 items were adjusted to learners' developmental levels and their past learning experiences." (pp. 507-508)
4) "Intra-rater reliability was ensured by scoring each WDCT twice at an interval of two weeks, to ensure that the results were consistent. The reliability yielded was 93.3%." (p. 508)
Detailed Analysis:
Criterion I requires that the study be conducted independently from the designers of the intervention, for example by an external evaluation team. Here the single author is simultaneously the intervention designer, the classroom teacher delivering all three conditions, the creator of the assessment instrument, the sole rater of the WDCTs (only intra-rater reliability is reported, implying one rater), and the analyst. There is no mention of any external evaluators, independent test administrators, blinded raters, or third-party oversight anywhere in the paper.
Criterion I is not met because the same teacher-researcher designed, delivered, scored, and analysed the study with no independent oversight.
-
Y
Year Duration
- The whole study spanned only a few weeks, far short of the required 75% of an academic year of tracking.
- "The necessary data were collected three times: before the intervention (pre-test), within two days after the fourth lesson (post-test), and three weeks later (a delayed post-test)." (p. 511)
Relevant Quotes:
1) "The participants received a series of 4 lessons focused on three interactional sequences." (p. 509)
2) "The necessary data were collected three times: before the intervention (pre-test), within two days after the fourth lesson (post-test), and three weeks later (a delayed post-test)." (p. 511)
Detailed Analysis:
Criterion Y requires outcome tracking covering at least 75% of a full academic year (roughly 9-10 months) from the intervention start. The entire study, from the first of four lessons to the delayed post-test, spans only a few weeks. Since the weaker Term Duration criterion (T) is already not met, this stronger Year Duration criterion cannot be met either, per the prompt's rule that if T is not met then Y is not met.
Criterion Y is not met because the study tracked outcomes for only a few weeks, nowhere near 75% of an academic year.
-
B
Balanced Control Group
- All three arms received the same four regular 45-minute lessons from the same teacher with documented comparable time allocations, and remaining structural differences were the treatment contrast itself.
- "In all three groups (T1-EG, T2-EG, and PPP-EG), each of the first three lessons targeted one interactional sequence, while the fourth lesson aimed to consolidate the three sequences through productive practice activities." (p. 509)
Relevant Quotes:
1) "In all three groups (T1-EG, T2-EG, and PPP-EG), each of the first three lessons targeted one interactional sequence, while the fourth lesson aimed to consolidate the three sequences through productive practice activities." (p. 509)
2) "This took about 15 minutes of the 45-minute lesson." (p. 509)
3) "Such procedure ensured that: T1-EG and T2-EG had the same amount of exposure to the models of the three interactional sequences as measured by time on task (about 10 minutes) ... T1-EG performed each task twice (about 20 minutes spent on task performance, but no focus on form in the while-task stage) ... T2-EG performed each task three times (about 30 minutes of task performance, but no explicit instruction) ... PPP-EG performed the task once (about 10 minutes spent on the task) but received explicit instruction and completed more language-focused activities before task performance" (pp. 510-511)
4) "In the last stage, the learners performed the same tasks as T1-EG and T2-EG in the while-task phase of the lesson." (p. 510)
Detailed Analysis:
Criterion B requires that groups be balanced in time and budget so that outcome differences can be attributed to the intervention itself rather than to extra resources. This study has no true no-treatment control group (see criterion D), so the comparison is among three treatment arms; the closest thing to a "business as usual" comparator is PPP-EG, which followed the regular coursebook lessons. All three arms received the same dosage: a series of four regular 45-minute lessons delivered by the same teacher within the normal five-hours-per-week English course, using the same target sequences and the same core tasks. No arm received extra lesson time, extra materials, technology, or additional budget beyond the standard curriculum; the arms differed only in how the same lesson time was allocated between implicit modelling, task performance, explicit exercises, and feedback - and this allocation difference is precisely the instructional contrast being tested (task implementation type vs the PPP framework), i.e., it is integral to the treatment comparison. The paper explicitly documents the minute-by-minute time budgets to show comparable exposure. Following the decision tree, no extra time or budget was present in any arm relative to the others (all activity occurred within regular scheduled lessons of equal length and count), so the balance requirement is satisfied.
Criterion B is met because all three groups received equal instructional time within regular 45-minute lessons from the same teacher, with differences in internal lesson structure being the treatment contrast itself.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific 2024 study by a different research team in a peer-reviewed journal is reported or known; the only citing works found are by the same author.
Relevant Quotes:
1) "The present study explores the efficacy of different types of task implementation and the more traditional 'presentation – practice – production' (PPP) framework in developing EFL learners' ability to produce interactional sequences..." (p. 503)
2) "Thus, the study discussed in the present paper examines how teaching L2 pragmatics can be expanded through the use of tasks." (p. 506)
Detailed Analysis:
Criterion R requires that this specific study be independently replicated by a different research team in a different context and published in a peer-reviewed journal. The paper itself is presented as addressing a gap ("the two domains have rarely been brought together"), i.e., it is a novel study, and it cites no replication of its own design. The paper was published in June 2024 in Neofilolog.
An internet search for citing works (OpenAlex, cited-by count = 2) identified only two subsequent publications, both by the same author, Tomasz Róg: a 2025 book chapter "The Notion and Scope of Task-Based Language Teaching" (Springer, DOI 10.1007/978-3-031-86566-4_2) and a 2025 article "The impact of task-based and task-supported instruction on the acquisition of L2 hedge phrases" (DOI 10.58221/mosp.v119i3.55699). Neither is an independent replication by a different research team: both are by the original author, and the second examines a different linguistic target (hedge phrases) rather than reproducing this specific three-condition study of interactional sequences. Related prior studies cited in the original paper (e.g., Nguyen, 2008; Li, Ellis, Zhu, 2016) predate this study and examine different designs and targets, so they are precedents, not replications. No independent published replication of this specific three-condition study of task implementation for teaching interactional sequences to Polish EFL learners was found in any available source.
Criterion R is not met because no independent replication of this specific study by a different research team has been identified.
-
A
All-subject Exams
- Criterion E is not met and only L2 pragmatic production was assessed, with no standardised exams covering other main subjects.
- "Groups' pragmatic production was measured using written discourse completion tasks." (p. 502)
Relevant Quotes:
1) "Groups' pragmatic production was measured using written discourse completion tasks." (p. 502, abstract)
2) "The WDCTs used in the present study consist of 15 items. Each interactional sequence targeted in the instruction (i.e., making a recommendation, reaching an agreement through negotiation, and defending a decision) was elicited through five different situations." (p. 507)
Detailed Analysis:
Criterion A requires standardised exam-based assessment across all main school subjects, and criterion E is a prerequisite. Criterion E is not met (the study used a custom WDCT instrument), so criterion A automatically fails. Moreover, the study measured only one narrow outcome - L2 English pragmatic production of three interactional sequences - and assessed no other school subjects (mathematics, Polish, science, etc.) in any form.
Criterion A is not met because criterion E fails and only a single narrow L2 pragmatics outcome was measured, with no other subjects assessed.
-
G
Graduation Tracking
- Measurement ended three weeks after the intervention with no tracking of students through to graduation, and no follow-up publication tracking this cohort was found.
- "The necessary data were collected three times: before the intervention (pre-test), within two days after the fourth lesson (post-test), and three weeks later (a delayed post-test)." (p. 511)
Relevant Quotes:
1) "The necessary data were collected three times: before the intervention (pre-test), within two days after the fourth lesson (post-test), and three weeks later (a delayed post-test)." (p. 511)
2) "Also, the study followed a short intervention of 4 lessons. A longitudinal study of the effects of the three types of instruction might shed more light on their effectiveness." (p. 516)
Detailed Analysis:
Criterion G requires tracking participants until graduation from their educational stage, and criterion Y is a prerequisite. Criterion Y is not met, so G automatically fails. Measurement ended with the delayed post-test three weeks after the intervention, while participants were third-grade secondary school students; no follow-up through secondary school graduation is reported or planned, and the author explicitly calls for longitudinal research as future work.
An internet search for subsequent publications by the same author (OpenAlex citing-works list, cited-by count = 2) found only a 2025 book chapter on the scope of task-based language teaching and a 2025 article on L2 hedge phrase acquisition; neither reports any follow-up tracking of this same 81-student cohort through to graduation. No follow-up publication with graduation tracking for these participants was found in any available source.
Criterion G is not met because tracking stopped three weeks after the intervention, criterion Y is not met, and no follow-up publication tracking these participants to graduation was found.
-
P
Pre-Registered
- The paper contains no mention of pre-registration, a registry ID, or a published protocol of any kind, and no external registry record was found.
Relevant Quotes:
No quotes mentioning pre-registration, a trial registry, a registration ID, or a published protocol appear anywhere in the paper. The methods section describes the participants, instrument, procedure, and statistical analysis (pp. 507-511) without any reference to a registered protocol.
Detailed Analysis:
Criterion P requires that the full study protocol (including hypotheses, methods, and planned analyses) be pre-registered on a public registry before data collection began, with verifiable timing. This paper contains no mention of any registry platform (e.g., ClinicalTrials.gov, OSF, AsPredicted, ISRCTN), no registration number, no protocol paper, and no statement about pre-registration timing. An internet search for a pre-registration record associated with the author or this study found no relevant registry entry. Absent any such evidence, the criterion cannot be satisfied.
Criterion P is not met because the paper contains no reference to any pre-registered protocol or registry entry, and no external registry record was found.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.