Abstract
Task-Based Language Teaching has been developed in response to the teacher-dominated, focus-on-forms methods such as Present, Practice, Produce (PPP). The body of literature is replete with studies examining the learning efficacy of the PPP approach versus TBLT; however, these studies did not use assessment tasks in comparing these two methods. To this end, the present study used an Assessment Task, a Grammaticality Judgment Test (GJT), and an Elicited Imitation Test (EIT) to compare the efficacy of PPP versus TBLT. Thirty-four lower-intermediate English language learners in Iran were randomly assigned to TBLT, PPP, and Control groups. The study results revealed that the performance of TBLT and PPP on the GJT and EIT significantly improved from pre-assessment to post-assessment, while the Control group did not show any significant improvements on any of the tests. Results indicated that only the TBLT group made substantial improvements in TBLA in the post-assessment, while the PPP and Control groups' performance did not significantly improve.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was performed at the level of three whole, intact classes (each assigned entirely to Control, PPP, or TBLT), not among individual students within a single class.
- "The three classes were randomly assigned to three groups of Control, PPP, and TBLT by pulling their names out of a hat." (p. 8)
Relevant Quotes:
1) "The three classes were randomly assigned to three groups of Control, PPP, and TBLT by pulling their names out of a hat." (p. 8)
2) "The participants of the study were chosen from an English Language Institute entitled Parsian Language School in Mazandaran province, located in northern Iran... The study included 18 female and 16 male Iranian English as a Foreign Language (EFL) learners." (p. 7)
3) "One instructor undertook the TBLT treatment, and the other did the PPP treatment." (p. 7)
Detailed Analysis:
The unit of randomisation described in the paper is the intact class, not the individual student within a shared classroom. Three existing classes at the Parsian Language School were each randomly allocated, as whole units, to one of the three conditions (Control, PPP, TBLT) via a lottery method ("pulling their names out of a hat"). Because every student inside a given class received the same condition, and separate instructors delivered the TBLT and PPP treatments to their own distinct classes, there is no risk of within-class contamination between conditions. This satisfies the ERCT requirement that randomisation occur at least at the class level rather than among students sharing one classroom, even though only a single class served each condition.
Final sentence: Criterion C is met because the paper explicitly describes random assignment of entire classes, not individual students within one class, to the study conditions.
-
E
Exam-based Assessment
- All three outcome measures (GJT, EIT, Assessment Task) were researcher-designed or borrowed-from-a-prior-study instruments, not widely recognised standardised exams.
- "The GJT and EIT were borrowed from Li et al.'s (2016) study, as the present study is a quasi-replication of it, but the researcher designed the Assessment Task to specifically distinguish it from other studies in the body of literature." (p. 8)
Relevant Quotes:
1) "The GJT and EIT were borrowed from Li et al.'s (2016) study, as the present study is a quasi-replication of it, but the researcher designed the Assessment Task to specifically distinguish it from other studies in the body of literature." (p. 8)
2) "The GJT was an untimed test including 40 items whose grammaticality was to be judged by students. This test was supposed to assess the students' explicit knowledge of the target feature, i.e., the knowledge of the rules and regulations of the past passive voice in English (See Appendix A)." (p. 9)
3) "The Assessment Task was a narrative task that required the students to read a hypothetical short story about a robbery in Miami, USA (See Appendix C)." (p. 9)
Detailed Analysis:
None of the three instruments used to measure outcomes qualify as widely recognised, standardised exams. The GJT and EIT are experimental instruments originally constructed for a different research study (Li et al., 2016) targeting one narrow grammatical structure (past passive voice), and the Assessment Task is an original narrative rewriting exercise designed by the researcher specifically for this study. There is no indication that any of these are validated, standardised, or widely used assessments (e.g., a national curriculum exam or standardised language proficiency test); they are all purpose-built research instruments.
Final sentence: Criterion E is not met because the study relies exclusively on custom-built or borrowed-from-another-experiment research instruments rather than a standardised, widely-recognised exam.
-
T
Term Duration
- The treatment lasted about one hour (or 30 minutes for PPP) with post-assessment administered in the same or immediately following session, far short of one academic term.
- "Additionally, the one-hour instruction of the study as the treatment was too short to come to a robust finding regarding implicit knowledge and task performance." (p. 14)
Relevant Quotes:
1) "The study used a quasi-experimental design with a format of pre-assessment and immediate post-assessments." (p. 7)
2) "The students in the PPP group went through 30 minutes of instruction on the past passive voice, and then they had to practice the structure through the following activities..." (p. 7)
3) "Additionally, the one-hour instruction of the study as the treatment was too short to come to a robust finding regarding implicit knowledge and task performance." (p. 14)
4) "Afterwards, the narrative task was administered to each group. The instructor ensured that students understood the task's instructions well..." (p. 9)
Detailed Analysis:
The paper explicitly labels its own design as using "immediate post-assessments," with outcomes captured directly after a single, brief treatment session (roughly 30 minutes to one hour of instruction). The authors themselves acknowledge in the Limitations section that the "one-hour instruction... was too short" for robust conclusions. There is no interval of weeks or months, let alone a full academic term (3-4 months), between the start of the intervention and the measurement of outcomes.
Final sentence: Criterion T is not met because outcomes were measured immediately after a single short session rather than after a term-long interval.
-
D
Documented Control Group
- The control group's activity, size, and baseline scores on all three outcome measures are clearly documented.
- "The treatment for the control group consisted of five reading passages taken from a reading book entitled Intermediate Steps to Understanding, authored by Hill (1980)." (p. 8)
Relevant Quotes:
1) "The treatment for the control group consisted of five reading passages taken from a reading book entitled Intermediate Steps to Understanding, authored by Hill (1980). The passages were at the intermediate level of proficiency according to the book. At the end of each of the five passages, there were comprehension questions that students needed to answer. The students were asked to read the passage and then answer comprehension questions. The instructor followed up by reading the passage to students, explaining difficult words, and making sure the students comprehended the passage." (p. 8)
2) Table 2 reports, for the Control group: "GJT ... Control 9 22.33 3.93"; "EIT ... Control 9 17.33 2.87"; "Task ... Control 9 13.61 1.83" (p. 10)
3) "The three classes were randomly assigned to three groups of Control, PPP, and TBLT by pulling their names out of a hat." (p. 8)
Detailed Analysis:
The paper clearly describes what the control group did during the study (reading five passages with comprehension questions, no passive-voice instruction), its size (n = 9), and its baseline (pre-test) performance with means and standard deviations on all three outcome measures (Table 2), plus a pre-assessment Kruskal-Wallis test confirming no significant baseline differences among the three groups. This level of detail allows assessment of comparability between the control and treatment groups at baseline.
Final sentence: Criterion D is met because the control group's activities, size, and baseline scores are explicitly documented in the text and in Table 2.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred among three classes within a single language institute, not across multiple schools.
- "The participants of the study were chosen from an English Language Institute entitled Parsian Language School in Mazandaran province, located in northern Iran." (p. 7)
Relevant Quotes:
1) "The participants of the study were chosen from an English Language Institute entitled Parsian Language School in Mazandaran province, located in northern Iran. Parsian Language School is a privately-owned language institute that teaches the English and French languages." (p. 7)
2) "The three classes were randomly assigned to three groups of Control, PPP, and TBLT by pulling their names out of a hat." (p. 8)
Detailed Analysis:
All participants and all three classes came from a single institute (Parsian Language School). The randomisation unit was the class, not the school/institute, and there was no comparison across multiple schools or institutes. This is a class-level (not school-level) design.
Final sentence: Criterion S is not met because randomisation and implementation were confined to three classes within one single language institute rather than spanning multiple schools.
-
I
Independent Conduct
- The classroom teachers delivered treatments under researcher instruction, but there is no evidence of an independent, third-party team handling data collection or analysis.
- "The teachers were given instructions on how to administer the test and how to do the treatments, especially the task implementation, according to Willis's and Willis's (2007) model." (p. 7)
Relevant Quotes:
1) "Additionally, two instructors at the Parsian institute cooperated in the study. Both instructors were in the midst of their doctoral studies in the field of applied linguistics... The teachers were given instructions on how to administer the test and how to do the treatments, especially the task implementation, according to Willis's and Willis's (2007) model. One instructor undertook the TBLT treatment, and the other did the PPP treatment." (p. 7)
2) "...but the researcher designed the Assessment Task to specifically distinguish it from other studies in the body of literature." (p. 8)
3) "Another critical aspect of the present study was using the classroom's actual teacher. Most of the previous studies on this subject used their researcher as the instructor in the study... which would make their results biased." (p. 14)
Detailed Analysis:
The paper highlights, as a strength, that classroom teachers (rather than the researchers themselves) delivered the treatments, which does reduce one source of bias relative to some prior studies. However, this is not the same as independent conduct of the evaluation: the researcher(s) designed the assessment instruments (including the Assessment Task), instructed the teachers on exactly how to implement both treatments, and there is no mention of an external or third-party team collecting data, scoring tests, or performing the statistical analysis independently of the research team. No disclosure of an outside evaluation body is present.
Final sentence: Criterion I is not met because, although classroom teachers rather than the researchers delivered instruction, the overall study design, instrument development, data collection oversight, and analysis remained in the hands of the research team with no documented independent third-party evaluator.
-
Y
Year Duration
- Because Criterion T (Term Duration) is not met, Criterion Y cannot be met; the actual treatment-to-measurement interval was also far shorter than a year.
- "Additionally, the one-hour instruction of the study as the treatment was too short to come to a robust finding regarding implicit knowledge and task performance." (p. 14)
Relevant Quotes:
1) "The study used a quasi-experimental design with a format of pre-assessment and immediate post-assessments." (p. 7)
2) "Additionally, the one-hour instruction of the study as the treatment was too short to come to a robust finding regarding implicit knowledge and task performance." (p. 14)
Detailed Analysis:
Per the ERCT specific instructions, if Criterion T is not met, Criterion Y is automatically not met, since Y is the stronger version of the same duration requirement. Independently of that rule, the actual interval here (a single treatment session followed by an immediate post-test) is a matter of hours, not the required 75% of an academic year.
Final sentence: Criterion Y is not met both because the weaker Term Duration criterion (T) is not met and because the study's actual duration was only about one hour.
-
B
Balanced Control Group
- All three groups received a comparable single class session with their instructor; the difference between groups is the instructional content itself (the explicit variable under study), not extra time or budget.
- "The treatment for the control group consisted of five reading passages... At the end of each of the five passages, there were comprehension questions that students needed to answer." (p. 8)
Relevant Quotes:
1) "The students in the PPP group went through 30 minutes of instruction on the past passive voice, and then they had to practice the structure through the following activities..." (p. 7)
2) "The treatment for the control group consisted of five reading passages taken from a reading book entitled Intermediate Steps to Understanding, authored by Hill (1980)... At the end of each of the five passages, there were comprehension questions that students needed to answer. The students were asked to read the passage and then answer comprehension questions. The instructor followed up by reading the passage to students, explaining difficult words, and making sure the students comprehended the passage." (p. 8)
3) "The instructor spent the Monday class administering the assessments to the students of the TBLT group... The PPP instructor followed the same procedure for the PPP and Control groups." (p. 9)
Detailed Analysis:
This study's central research question is a direct comparison of three instructional approaches (TBLT, PPP, and a business-as-usual reading/comprehension control) as the explicit treatment variable being tested. All three groups appear to receive one class session with their regular instructor of broadly comparable length (approximately 30 minutes to one hour), with the control group engaged in an active, instructor-led reading-and-comprehension activity rather than being left idle. There is no indication that the TBLT or PPP groups received any additional class time, extra instructional hours, or supplementary budget /materials beyond what the control group received; the groups differ only in the content and pedagogical method of that single session, which is precisely the variable under investigation. Applying the criterion B decision procedure: no extra time or budget resources (EXTRA_RESOURCES_PRESENT) were allocated to the treatment arms relative to the control beyond a difference in instructional content, which is the intervention itself, so the criterion is satisfied at the first branch of the decision tree without needing to invoke the integral-resource exception.
Final sentence: Criterion B is met because the three groups received comparable class time with their own instructor, differing only in instructional content, which is the explicit treatment variable being compared.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study by another research team was found; the paper itself borrows instruments from an earlier, unrelated study rather than being replicated by one.
- "The GJT and EIT were borrowed from Li et al.'s (2016) study, as the present study is a quasi-replication of it..." (p. 8)
Relevant Quotes:
1) "The GJT and EIT were borrowed from Li et al.'s (2016) study, as the present study is a quasi-replication of it, but the researcher designed the Assessment Task to specifically distinguish it from other studies in the body of literature." (p. 8)
Detailed Analysis:
The paper describes itself as a "quasi- replication" of Li et al. (2016), meaning it reuses and extends instruments from an earlier study; this is the reverse of what Criterion R requires. Criterion R asks whether this specific study (Noroozi & Taheri, 2022) has itself been independently replicated by a different research team afterward. An internet citation search (Semantic Scholar/OpenAlex, checked 2026) found ten papers citing this article, including Farahhein, Ab Rahman, and Nawi (2025) on task-supported teaching in blended learning, Baillo, Pradia, Encarnacion, and Fernandez (2025) on TBLT and oral communication skills, and Alshakhi and Albalawi (2024) on task-based language assessment. None of these studies reproduce this paper's specific design (TBLT vs. PPP vs. Control, assessed via the GJT, EIT, and Assessment Task on the English past passive voice with Iranian EFL learners); they use different populations, outcome measures, or research questions. No independent replication of this exact study was located.
Final sentence: Criterion R is not met because there is no evidence that this specific study has been independently replicated by a different research team.
-
A
All-subject Exams
- Because Criterion E is not met, Criterion A is automatically not met; additionally, only a single narrow grammatical feature was assessed, not multiple subjects.
- "This test was supposed to assess the students' explicit knowledge of the target feature, i.e., the knowledge of the rules and regulations of the past passive voice in English (See Appendix A)." (p. 9)
Relevant Quotes:
1) "This test was supposed to assess the students' explicit knowledge of the target feature, i.e., the knowledge of the rules and regulations of the past passive voice in English (See Appendix A)." (p. 9)
2) "The EIT was designed to assess the learners' automated knowledge of the target structure (See Appendix B)." (p. 9)
Detailed Analysis:
Per the ERCT specific instructions, if Criterion E (Exam-based Assessment) is not met, Criterion A is not met either, since A is the stronger version of E. Independently, the study exclusively measures a single linguistic feature (the English past passive voice) rather than assessing multiple core subjects or even multiple language skills; there is no coverage of other subjects at all.
Final sentence: Criterion A is not met because Criterion E is not met and because the study measured only one narrow linguistic feature rather than multiple subjects.
-
G
Graduation Tracking
- Because Criterion Y (Year Duration) is not met, Criterion G is automatically not met; outcomes were also measured immediately after the single treatment session, with no follow-up tracking at all, let alone through graduation.
- "Afterwards, the narrative task was administered to each group. The instructor ensured that students understood the task's instructions well..." (p. 9)
Relevant Quotes:
1) "The post-assessment tests were the same as the pre-assessment tests; the only difference was that the order of the items for the tests of the GJT and EIT was changed for the post-assessment to prevent the practice effect... Afterwards, the narrative task was administered to each group." (p. 9)
2) "There are a couple of limitations to this study. First, the final number of students used in this study was limited..." (p. 14) [no mention of any follow-up or planned tracking]
Detailed Analysis:
Per the ERCT specific instructions, since Criterion Y (Year Duration) is not met, Criterion G is automatically not met, as G is a stronger extension of the same duration requirement. Independently, the paper describes only a pre-test/treatment/immediate post-test design with no subsequent data collection at any later point, and no reference to planned or completed follow-up studies tracking the same cohort. An internet search (Semantic Scholar/OpenAlex, checked 2026) for later publications by Majeed Noroozi and/or Seyyedmohammad Taheri that might track this same cohort of Iranian EFL learners toward graduation did not locate any such follow-up study.
Final sentence: Criterion G is not met both because Criterion Y is not met and because no follow-up or graduation tracking was found in this paper or in any located subsequent publication.
-
P
Pre-Registered
- No statement anywhere in the paper mentions a pre-registered protocol, registry platform, or registration date, and no such registration was located via internet search.
Relevant Quotes:
No quotes referencing pre-registration, a trial registry, or a published a-priori protocol were found anywhere in the manuscript, including the Method, Data collection, Limitations, or Conclusion sections.
Detailed Analysis:
A full review of the paper, including its methods and appendices, reveals no mention of a pre-registration on any public registry (e.g., OSF, ClinicalTrials.gov, ISRCTN, AsPredicted) prior to data collection, nor any statement of hypotheses or analysis plans being published in advance. An internet search for a pre-registration record associated with this study or its authors did not locate any matching entry on common registries.
Final sentence: Criterion P is not met because there is no evidence of a pre-registered protocol anywhere in the paper or in available registries.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.