Abstract
This paper investigated the effect of Content-based Instruction (CBI) on students' English language learning. In so doing, two methods of teaching English, that is, the CBI and the Grammar Translation Method (GTM) were compared with regard to the students' achievement in their final examination and language learning orientation. The subjects consisted of 82 freshmen who were randomly assigned into two groups at Gonabad University of Medical Sciences. To collect data, three instruments were employed: the Nelson test of achievement form 050 C, the Language Learning Orientation Scale (LLOS) questionnaire, and a final achievement test. The data were analyzed using t-test and some correlational analyses. The results indicated that there was no significant difference between the groups regarding the Nelson test and LLOS at the onset of the study, but there was a significant difference between the groups' performance regarding the method of teaching English. In other words, the group taught through the CBI outperformed the one taught through the GTM. Moreover, there was a significant difference in the subjects' language learning orientation after treatment (p=0.038). Some suggestions and implications were put forward for the EFL/ESL teachers to consider.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Allocation followed pre-existing course cohorts chosen by convenience sampling, and no genuine class-level randomisation procedure is described.
- "were randomly assigned into two groups, so that the students in Environmental Health and Anesthesia courses were put into the grammar translation group and those in the Public Health and the Laboratory Sciences courses were placed in the content-based instruction group"
Relevant Quotes:
1) "The subjects consisted of 82 freshmen who were randomly assigned into two groups at Gonabad University of Medical Sciences." (Abstract, p. 2157)
2) "They were, also, of both genders (but mostly females) aged between 18-20 years old and were randomly assigned into two groups, so that the students in Environmental Health and Anesthesia courses were put into the grammar translation group and those in the Public Health and the Laboratory Sciences courses were placed in the content-based instruction group." (p. 2159)
3) "At first, the subjects (N=82) were chosen based on the convenience sampling method and were randomly assigned into two groups to be taught through either grammar translation method (method 1) or content-based instruction method (method 2)." (p. 2160)
Detailed Analysis:
Criterion C requires a properly implemented and clearly described randomisation at the class level or stronger. The paper repeatedly claims the 82 subjects "were randomly assigned into two groups," but the same sentence reveals that the two groups coincide exactly with pre-existing degree cohorts: Environmental Health and Anesthesia students "were put into" the GTM group and Public Health and Laboratory Sciences students "were placed in" the CBI group. Individual students therefore could not have been randomly allocated; allocation followed their pre-existing programme enrolment (convenience sampling of intact course groups). At best this is an assignment of four intact course cohorts to two conditions, but the paper never describes any random procedure (e.g., coin toss, random number generation) for assigning those cohorts to methods, and with only two clusters per arm the claimed randomisation is neither clearly described nor verifiably properly implemented. The intervention is whole-class instruction, not one-to-one tutoring, so the tutoring exception does not apply.
Criterion C is not met because the claimed "random assignment" contradicts the allocation of intact pre-existing course cohorts and no class-level randomisation procedure is described.
-
E
Exam-based Assessment
- Outcomes were measured with a researcher-developed final achievement test, not a widely recognised standardised exam.
- "The final achievement test taken from a test bank developed by the researcher included four sections of vocabulary, grammar, translation, and reading comprehension."
Relevant Quotes:
1) "The final achievement test taken from a test bank developed by the researcher included four sections of vocabulary, grammar, translation, and reading comprehension." (p. 2159)
2) "The Nelson test of achievement form 050 C consists of 50 multiple-choice items measuring the learners' general language proficiency in English... It was administered to examine whether the subjects were homogeneous in terms of their English language proficiency." (p. 2159)
3) "Fourthly, at the end of the term, the FAT was administered along with the LLOS questionnaire (as the post-test)." (p. 2160)
Detailed Analysis:
Criterion E requires that the outcome be measured with a widely recognised standardised exam, not an instrument created for the study. The primary academic outcome here is the final achievement test (FAT), which was explicitly "taken from a test bank developed by the researcher" and tailored to the adopted textbook. The Nelson test (a recognised standardised proficiency test) was used only as a pre-test to check group homogeneity, not as the outcome measure. The second outcome, the LLOS, is a motivation questionnaire, not an exam. Because the achievement outcome was measured with a researcher-developed custom test rather than a standardised exam, the criterion fails.
Criterion E is not met because the outcome was a custom final achievement test built by the researcher, while the standardised Nelson test served only as a baseline check.
-
T
Term Duration
- The intervention ran for a full semester with outcomes measured at the end of the term, satisfying the one-term minimum.
- "the treatment was performed for one semester consisting of 25 one-hour-and-a-half sessions for a three-credit general English course"
Relevant Quotes:
1) "Thirdly, the treatment was performed for one semester consisting of 25 one-hour-and-a-half sessions for a three-credit general English course using the general English textbook whose description was given above." (p. 2160)
2) "Fourthly, at the end of the term, the FAT was administered along with the LLOS questionnaire (as the post-test)." (p. 2160)
3) "They were in their second term of a B.Sc. degree at university taking a three-credit general English course, too." (p. 2159)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. The intervention ran "for one semester" (25 sessions of 90 minutes), with the pre-tests administered in the first session and the post-tests (FAT and LLOS) administered "at the end of the term." A university semester constitutes a full academic term, so the interval from intervention start to outcome measurement spans one complete term.
Criterion T is met because the treatment lasted a full semester and outcomes were measured at the end of that term.
-
D
Documented Control Group
- The GTM comparison group's size, demographics, baseline Nelson and LLOS scores, and instructional condition are documented in detail.
- "Nelson GTM 42 24.8810 6.29054 0.981 / CBI 40 24.8500 5.27476" (Table 5)
Relevant Quotes:
1) "The subjects consisted of 82 freshmen majoring in the fields of Environmental Health (N=21), Laboratory Sciences (N=21), Anesthesia (N=19), and Public Health (N=21) at Gonabad university of medical sciences, Gonabad, Iran. They had previously a three-year experience of learning English at junior high schools and a four-year experience at high school and pre-requisite English Centers in the formal system of education in Iran. They were, also, of both genders (but mostly females) aged between 18-20 years old." (p. 2159)
2) "Nelson GTM 42 24.8810 6.29054 0.981 / CBI 40 24.8500 5.27476" (Table 5, p. 2161)
3) "LLOS GTM 42 107.3810 15.34633 0.915 / CBI 40 107.0250 14.66111" (Table 6, p. 2161)
4) "As for the CBI [sic, describing the GTM], the teaching procedure had the following characteristics: most often the language of instruction was Persian... In case of grammar, it was taught deductively." (p. 2160)
Detailed Analysis:
Criterion D requires detailed documentation of the comparison group: composition, baseline performance, and the treatment it received. The GTM group serves as the control (business- as-usual traditional method). The paper documents group sizes (Table 1), the demographic profile of the sample (age 18-20, mostly female, freshmen in named medical-science fields, seven years of prior English study), and baseline equivalence on both the standardised Nelson proficiency test (p=0.981) and the LLOS motivation scale (p=0.915), reported per group with means and standard deviations. The instructional condition experienced by the GTM group is described in detail (Persian-medium instruction, translation of vocabulary and passages, deductive grammar, textbook exercises). A minor inconsistency exists between Table 1 (GTM N=40, CBI N=42) and the t-test tables (GTM N=42, CBI N=40), but the overall documentation of who the comparison group was, their baseline scores, and what they received is adequate.
Criterion D is met because the comparison group's size, background, baseline scores, and instructional condition are documented.
-
Level 2 Criteria
-
S
School-level RCT
- The entire study took place at one university with allocation among internal course cohorts, so no school-level randomisation occurred.
- "The subjects consisted of 82 freshmen... at Gonabad university of medical sciences, Gonabad, Iran."
Relevant Quotes:
1) "The subjects consisted of 82 freshmen majoring in the fields of Environmental Health (N=21), Laboratory Sciences (N=21), Anesthesia (N=19), and Public Health (N=21) at Gonabad university of medical sciences, Gonabad, Iran." (p. 2159)
2) "were randomly assigned into two groups, so that the students in Environmental Health and Anesthesia courses were put into the grammar translation group and those in the Public Health and the Laboratory Sciences courses were placed in the content-based instruction group." (p. 2159)
Detailed Analysis:
Criterion S requires randomisation among schools or equivalent institutional units. This study took place within a single institution (Gonabad University of Medical Sciences), and allocation occurred among course cohorts inside that one university. No multiple schools or sites were involved, and no school-level randomisation is described.
Criterion S is not met because the study was conducted within a single university with allocation among internal course groups, not among schools or sites.
-
I
Independent Conduct
- The same authors designed the intervention, taught the courses, built the outcome test, and analysed the data with no independent evaluators.
- "the Nelson test was in English; however, the researcher himself was available for any question."
Relevant Quotes:
1) "the Nelson test was in English; however, the researcher himself was available for any question." (p. 2160)
2) "The final achievement test taken from a test bank developed by the researcher included four sections..." (p. 2159)
3) "He is, also, a FACULTY MEMBER at the basic science department of Gonabad University of Medical Sciences, Gonabad, Iran, where he has taught pre-requisite, general English and ESP courses to students at different fields of medical sciences for more than 15 years." (p. 2167)
4) "However, it should be mentioned that the study had one limitation even though the researchers did their best to compensate for it... the researchers tried to adopt some other materials to be used in the CBI group." (p. 2166)
Detailed Analysis:
Criterion I requires that the study be conducted independently from those who designed the intervention. Here the same researchers designed the study, selected and adapted the CBI materials, developed the final achievement test, administered the instruments ("the researcher himself was available for any question"), and, given that the first author is the long-serving English instructor at the study university, evidently delivered the instruction as well. There is no mention of any external evaluation team, third-party data collection, blinded test administration, or independent oversight anywhere in the paper.
Criterion I is not met because the intervention designers themselves implemented the teaching, built the outcome test, and collected and analysed the data with no independent oversight.
-
Y
Year Duration
- Tracking lasted only one semester (about 3-4 months), well below 75% of an academic year.
- "the treatment was performed for one semester consisting of 25 one-hour-and-a-half sessions"
Relevant Quotes:
1) "Thirdly, the treatment was performed for one semester consisting of 25 one-hour-and-a-half sessions for a three-credit general English course." (p. 2160)
2) "Fourthly, at the end of the term, the FAT was administered along with the LLOS questionnaire (as the post-test)." (p. 2160)
Detailed Analysis:
Criterion Y requires that outcomes be measured at least 75% of a full academic year (roughly 9-10 months) after the intervention begins. The tracking period here covers exactly one university semester (about 3-4 months, 25 sessions), with post-tests administered at the end of that same term and no later follow-up of any kind. One semester falls well short of 75% of an academic year.
Criterion Y is not met because the interval from intervention start to final measurement was a single semester, far less than 75% of an academic year.
-
B
Balanced Control Group
- Both groups had identical class time, textbook, and course structure, and the CBI group's extra authentic content materials are integral to operationalising the CBI treatment itself, despite the authors framing the asymmetry as a study limitation.
- "In fact, the main textbook employed for the CBI and GTM groups was the same."
Relevant Quotes:
1) "Thirdly, the treatment was performed for one semester consisting of 25 one-hour-and-a-half sessions for a three-credit general English course using the general English textbook whose description was given above." (p. 2160)
2) "In fact, the main textbook employed for the CBI and GTM groups was the same." (p. 2166)
3) "The research reports, magazine articles, and even some news related to the theme of each lesson were introduced to the students only in the CBI group and they were required to provide a summary or a discussion on them for the following sessions." (p. 2166)
4) "However, it should be mentioned that the study had one limitation even though the researchers did their best to compensate for it." (p. 2166)
5) "Part III. Homework (Section One: vocabulary exercises... Section Four: A. Translation practice consisting of a passage, and B. Some specialized terms to be translated into Persian." (p. 2160, describing the shared textbook homework used by both groups)
Detailed Analysis:
Applying the criterion B decision tree: extra resources are present (the CBI group alone received supplementary authentic materials - research reports, magazine articles, news items - plus an associated summary/discussion task), so the "no-extra-resources" and "negligible-difference" branches do not apply. The question is whether these materials are integral to the treatment being tested or a separable, unbalanced add-on.
Both groups received identical scheduled class time (25 sessions of 90 minutes over one semester), the same three-credit course, and the same core textbook, so the core instructional dosage is matched. CBI is defined earlier in the paper as instruction built around integrating authentic content with language teaching aims; since the shared textbook alone could not supply such content, some additional content-carrying materials were a practical requirement for the CBI condition to instantiate CBI at all, which supports treating them as integral to operationalising the independent variable rather than an arbitrary bonus resource.
However, the authors themselves frame this asymmetry as "one limitation" of the study that they "did their best to compensate for," rather than stating in advance that testing additional authentic materials was the deliberate treatment variable, and the added out-of-class summary/discussion homework was not explicitly matched for the GTM group. This self-described "limitation" framing means the resource difference was not perfectly pre-planned or balanced; it reads as an ad hoc adjustment for an identical-textbook design that the authors themselves regard as imperfect. On balance, because the extra input is content material (not extra class time or budget) and is conceptually inseparable from what CBI as a method requires to be genuinely implemented, this is judged as integral to the CBI treatment package despite being imperfectly executed, rather than a separate, unmatched resource advantage such as extra tutoring time or budget.
Criterion B is met because instructional time, textbook, and course structure were identical across groups, and the extra authentic content materials and associated tasks in the CBI arm are integral to operationalising the content-based intervention being tested, even though the authors themselves flag this as an imperfectly executed compensation for using an identical base textbook.
-
Level 3 Criteria
-
R
Reproduced
- No confirmed independent peer-reviewed replication of this specific CBI-versus-GTM trial was found in the paper or via external search of the citation graph.
- "However, few studies, if any, have been carried out to compare the effect of teaching English through the CBI and the GTM in the fields related to the medical sciences at university level."
Relevant Quotes:
1) "However, few studies, if any, have been carried out to compare the effect of teaching English through the CBI and the GTM in the fields related to the medical sciences at university level." (p. 2158)
2) "Many studies, conducted regarding the CBI, have concentrated on its impact in disciplines like accounting (Chau Ngan, 2011; Malmir, Najafi Sarem, & Ghasemi, 2011), technology (Gaynor, 2013), Spanish (Pessoa, Henry, Donta, Tucker, & Lee, 2007)..." (p. 2158)
Detailed Analysis:
Criterion R requires that this specific study be independently replicated by a different team, in a different context, published in a peer-reviewed journal. The paper itself positions the study as novel ("few studies, if any, have been carried out"), citing only related CBI research in other disciplines that predates it and is not a replication.
An internet search of the citation graph (Semantic Scholar, Crossref) was performed for independent replications of this Gonabad CBI-versus-GTM trial in ESP for medical sciences. Nineteen papers citing Amiri & Hosseini Fatemi (2014) between 2017 and 2024 were identified, covering maritime English, aviation English, business English writing, and various ESP/CBI implementation and perception studies in Indonesia, Bangladesh, Ecuador, Taiwan and elsewhere; none of these reproduce this study's specific design. One candidate paper, Ahmadi-Azad & Kuhi (2018), "ESP Vocabulary Instruction: A Comparison of CBI vs. GTM for Iranian Management Students," compares the same two methods (CBI vs GTM) but in a different population (management students rather than medical-science freshmen) with a vocabulary-specific outcome rather than the Nelson/FAT/LLOS battery used here; full-text access to confirm whether it explicitly frames itself as reproducing this study's design and results was not available through search. No paper was found that explicitly references this study as the original design being replicated, nor one that reproduces the specific Gonabad medical-sciences CBI/GTM/Nelson/FAT/LLOS design with comparable reported results in a peer-reviewed outlet.
Criterion R is not met because no confirmed independent peer-reviewed replication of this specific study's design and results could be identified either in the paper or through external search.
-
A
All-subject Exams
- Criterion E fails and only English was assessed with a custom test, so the all-subject exam requirement cannot be met.
- "The final achievement test taken from a test bank developed by the researcher included four sections of vocabulary, grammar, translation, and reading comprehension."
Relevant Quotes:
1) "To collect data, three instruments were employed: the Nelson test of achievement form 050 C, the Language Learning Orientation Scale (LLOS) questionnaire, and a final achievement test." (Abstract, p. 2157)
2) "The final achievement test taken from a test bank developed by the researcher included four sections of vocabulary, grammar, translation, and reading comprehension." (p. 2159)
Detailed Analysis:
Criterion A requires standardised exam-based assessment of all main subjects, and criterion E is an explicit prerequisite. Criterion E is not met (the outcome was a custom researcher-developed test), so criterion A automatically fails. Furthermore, the study measured outcomes only in English language achievement and motivation; no other subjects in the students' medical-science curricula (anatomy, chemistry, etc.) were assessed, and no rationale invoking the specialised-intervention exception is offered with respect to standardised testing. This is a higher-education ESP setting, so a narrow focus could potentially be argued, but the prerequisite failure of E is decisive.
Criterion A is not met because criterion E fails and only a single subject (English) was assessed with a custom test.
-
G
Graduation Tracking
- Measurement ended with the end-of-semester post-test, and no follow-up toward graduation is reported in this paper or in any subsequent publication by the same authors found via search.
- "Fourthly, at the end of the term, the FAT was administered along with the LLOS questionnaire (as the post-test)."
Relevant Quotes:
1) "Fourthly, at the end of the term, the FAT was administered along with the LLOS questionnaire (as the post-test). And finally, the obtained data were analyzed." (p. 2160)
Detailed Analysis:
Criterion G requires tracking participants until graduation from their educational stage, and criterion Y is a prerequisite: if Y is not met, G is automatically not met. Criterion Y is not met here (tracking lasted one semester), so criterion G fails on that basis alone. Consistently, measurement ended with the post-test at the end of the single semester; the freshmen subjects were years away from completing their B.Sc. degrees, and the paper mentions no follow-up, no planned longitudinal tracking, and no subsequent cohort publications.
An internet search of the citation graph (Semantic Scholar) was performed for later follow-up papers by Amiri and/or Hosseini Fatemi tracking this same Gonabad cohort toward graduation. No such follow-up publication by these authors on this cohort was found among the citing or related literature reviewed.
Criterion G is not met because criterion Y fails and no evidence of graduation tracking was found in this paper or in any subsequent publication by the same authors.
-
P
Pre-Registered
- The paper contains no reference to any pre-registered protocol, registry, or registration date, and none was found through external search.
Relevant Quotes:
No quotes referencing pre-registration, a trial registry, a registration ID, or a published protocol exist anywhere in the paper.
Detailed Analysis:
Criterion P requires that the full study protocol be registered on a public registry before data collection began, with verifiable timing. The paper contains no mention of any registry (e.g., ClinicalTrials.gov, IRCT, OSF), no registration number, no protocol publication, and no statement about hypotheses or analyses being specified in advance. An internet search for a pre-registration record associated with this study or its authors (Amiri, Hosseini Fatemi, Gonabad CBI/GTM trial) found no matching registry entry. The 2014 publication simply reports the study with no transparency infrastructure.
Criterion P is not met because no pre-registration of any kind is referenced in the paper or found through external search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.