Abstract
The grammar translation method and direct method were compared in this study to see how they impacted students' English learning outcomes at secondary school. These results were obtained through the use of an experimental pre-test and post-test control group design. In this study, the participants were all students at Govt. Islamia High School Sambrial District Sialkot. Of the 148 students in 10th grade, 60 students were randomly chosen. This study's second phase used pre-test scores to assign students to experimental and control groups. For the purpose of gathering information about students' academic progress, the MCQ test was created. Six weeks of treatment (Grammar translation) were given to the experimental group. A follow-up test was administered to all participants, experimental and control alike. The collected data were analysed using descriptive and inferential statistics. The experimental and control groups were compared using a T-test. According to the results, students who were taught English through the grammar-translation method performed better on standardised tests than students who were taught the subject directly.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was done at the individual student level within a single school (60 students paired and randomly assigned to two groups), not at the class or school level, and the intervention was whole-group classroom teaching, not tutoring.
- "148 students were enrolled in the course, and 60 were chosen at random. After that, two groups (control and experimental) were formed randomly using random assignment." (p. 351)
Relevant Quotes:
1) "Of the 148 students in 10th grade, 60 students were randomly chosen. This study's second phase used pre-test scores to assign students to experimental and control groups." (Abstract, p. 347)
2) "148 students were enrolled in the course, and 60 were chosen at random. After that, two groups (control and experimental) were formed randomly using random assignment." (p. 351)
3) "Only one school named Govt. Islamia High School Sambrial District Sialkot was selected for this treatment research." (p. 351)
4) "Afterwards, the researcher applied a pre-test (MCQs test) on all selected 60 students and made pairs based on pre-test scores. The researcher divided all the students into 2 groups randomly. The researcher assigned them random names as group A and group B." (p. 351)
Detailed Analysis:
Criterion C requires that randomisation be conducted at the class level or higher (school level), unless the intervention is one-to-one tutoring/personal teaching. The quotes show that a single school was used, 60 individual 10th-grade students were sampled from 148, and individual students were paired on pre-test scores and randomly assigned to the experimental (Grammar Translation Method) or control (Direct Method) group. This is student-level randomisation of pupils drawn from the same school and grade, creating exactly the contamination risk the criterion is designed to prevent: students from the two groups attend the same school and can interact. The intervention is whole-class instruction (sessions of 35-40 minutes taught to a group), not personal tutoring, so the tutoring exception does not apply.
Criterion C is not met because randomisation was carried out at the individual student level within one school, not at the class or school level, and no tutoring exception applies.
-
E
Exam-based Assessment
- Outcomes were measured with a 35-item MCQ test developed by the researcher specifically for this study, not a widely recognised standardised exam.
- "The researcher developed one instrument. The researcher used a test on Students' learning achievements as a pre-test and post-test based on three sub-components of Students' learning Achievement in English." (p. 351)
Relevant Quotes:
1) "For the purpose of gathering information about students' academic progress, the MCQ test was created." (Abstract, p. 347)
2) "The researcher developed one instrument. The researcher used a test on Students' learning achievements as a pre-test and post-test based on three sub-components of Students' learning Achievement in English." (p. 351)
3) "On the development of the MCQs test, it was pilot tested on a small sample to ensure reliability. Only 50 high school students were chosen for the pilot testing." (p. 351)
4) "Each of the 35 MCQs in the test had one correct answer. More specifically, this test was categorised as follows: i) putting correct verb, ii) putting correct spelling, iii) putting correct synonyms, and iv) putting correct grammar." (p. 351)
5) "The test was developed and conducted by the researcher himself." (p. 351)
Detailed Analysis:
Criterion E requires the use of a standard, widely recognised standardised exam rather than an instrument created for the study. The quotes explicitly state that the outcome measure was a 35-item multiple-choice test created by the researcher, derived from the Punjab Textbook Board syllabus content outlines, and pilot tested by the researcher himself. Although the abstract loosely claims students "performed better on standardised tests", the methods make clear the instrument was a custom, researcher-made test, not a recognised external standardised examination (e.g., a board exam). A researcher-built test aligned to the taught content is precisely the biased custom assessment the criterion warns against.
Criterion E is not met because outcomes were assessed with a custom MCQ test developed by the researcher for this study, not a recognised standardised exam.
-
T
Term Duration
- The intervention lasted only about 4-6 weeks (22 days total) with the post-test administered immediately afterwards, far short of a full academic term.
- "Research treatment was about 4 weeks; the time duration for treatment was 35-40 minutes in each session." (p. 352)
Relevant Quotes:
1) "Six weeks of treatment (Grammar translation) were given to the experimental group. A follow-up test was administered to all participants, experimental and control alike." (Abstract, p. 347)
2) "Total Days 22 days ... Pretest-Post Test Days 2 days ... Total SLOS 20 (one day per SLO) ... Days of Research Intervention 22 Days" (Table 2, p. 352)
3) "Each SLO was taught for 1 day. Research treatment was about 4 weeks; the time duration for treatment was 35-40 minutes in each session." (p. 352)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. The abstract states six weeks of treatment, while the methods section and Table 2 describe a 22-day intervention (about 4 weeks), with the pre-test and post-test administered within the same window (2 of the days). The post-test was given immediately at the end of the treatment, so the interval from intervention start to outcome measurement was at most about six weeks. There is no mention of any delayed follow-up measurement one term or more after the start. Whether 4 or 6 weeks, this is well below the 3-4 month minimum required for one academic term.
Criterion T is not met because the interval from intervention start to outcome measurement was only about 4-6 weeks, well short of one full academic term.
-
D
Documented Control Group
- Beyond stating that the control group was taught by the Direct Method, the paper reports no baseline (pre-test) scores or demographic breakdown for the control group and the group sizes are inconsistently reported (30 vs 20) without explanation.
- "Afterwards, the researcher applied a pre-test (MCQs test) on all selected 60 students and made pairs based on pre-test scores. The researcher divided all the students into 2 groups randomly." (p. 351)
Relevant Quotes:
1) "Afterwards, the researcher applied a pre-test (MCQs test) on all selected 60 students and made pairs based on pre-test scores. The researcher divided all the students into 2 groups randomly." (p. 351)
2) "One teacher taught the treatment group by the Grammar Translation method, and the second taught the control group with the traditional Direct Method on specific topics of 10th grade English published by PTB." (p. 351)
3) "Table 1. Direct Method 30, Grammar Translation Method 30, Total 60" (p. 351)
4) "Moreover, Students in 10 grade in Govt. Islamia High School Sambrial were only boy students. As a result, no information about the gender of the participants could be gathered." (p. 349)
5) "Table 3 ... Group N Mean SD; Experimental 20 4.800 .410; Control 20 4.100 .968" (p. 352)
Detailed Analysis:
Criterion D requires detailed documentation of the control group: demographics, baseline performance, and the conditions it received. The paper does describe the control condition (taught the same 10th-grade English topics via the Direct Method) and states the intended group size (30 per group). It also notes participants were all boys in grade 10 at one school. However, although a pre-test was administered and used for pairing, no baseline (pre-test) scores are ever reported for either group, so baseline comparability cannot be assessed. No demographic table is provided. Furthermore, the sampling section reports 30 students per group while every results table reports N=20 per group, an unexplained discrepancy with no account of attrition, leaving the actual composition of the control group unclear. This falls short of the detailed, verifiable control-group documentation the criterion requires.
Criterion D is not met because the control group's baseline performance and demographics are never reported and the reported group sizes are inconsistent (30 vs 20) without explanation.
-
Level 2 Criteria
-
S
School-level RCT
- The study took place in a single school with individual students randomised to groups, so there was no school-level randomisation.
- "Only one school named Govt. Islamia High School Sambrial District Sialkot was selected for this treatment research." (p. 351)
Relevant Quotes:
1) "Only one school named Govt. Islamia High School Sambrial District Sialkot was selected for this treatment research." (p. 351)
2) "148 students were enrolled in the course, and 60 were chosen at random. After that, two groups (control and experimental) were formed randomly using random assignment." (p. 351)
Detailed Analysis:
Criterion S requires randomisation among schools (or equivalent implementing units). This study was conducted entirely within one purposively selected school, and randomisation occurred at the individual student level within that school. No multiple schools or sites were involved and nothing was randomised at the institutional level.
Criterion S is not met because a single school was used and randomisation was at the student level, not the school level.
-
I
Independent Conduct
- The researcher himself developed the module and the test, delivered the treatment, and conducted the analysis, with no independent third-party conduct or oversight.
- "The test was developed and conducted by the researcher himself. ... Researcher himself gave treatment." (p. 351)
Relevant Quotes:
1) "Researcher himself gave treatment." (p. 351)
2) "The test was developed and conducted by the researcher himself." (p. 351)
3) "For standardisation, the researcher hired two mathematics teachers of the same experience, age, and ability to conduct the treatment study." (p. 351)
4) "For standardisation, the researcher himself taught the treatment group by grammar translation method strategies and the control group with traditional lecture method on specific topics of 10th grade English published by PTB." (p. 352)
5) "In the current study, the grammar-translation method Module was developed by the researcher and validated by 6 subject specialists working in the school education department on 16 plus BPS scales." (p. 352)
Detailed Analysis:
Criterion I requires the study to be conducted independently from those who designed the intervention, or at least with documented third-party oversight of data collection and analysis. Here the researcher developed the grammar translation module, developed and administered the outcome test, and (per two of the quoted statements) delivered the treatment himself and analysed the data. The paper contains contradictory statements about whether hired teachers or the researcher taught the groups, but under either version the design, instrumentation, data collection, and analysis all remained with the same author team. Validation of the module by six subject specialists concerns content validity, not independent conduct of the trial. No external evaluator or independent data-collection team is mentioned.
Criterion I is not met because the same researcher designed the intervention, built and administered the test, delivered the treatment, and analysed the results with no independent oversight.
-
Y
Year Duration
- The study spanned only about 4-6 weeks from intervention start to final measurement, nowhere near 75% of an academic year, and criterion T is already not met.
- "Total Days 22 days ... Days of Research Intervention 22 Days" (Table 2, p. 352)
Relevant Quotes:
1) "Six weeks of treatment (Grammar translation) were given to the experimental group." (Abstract, p. 347)
2) "Total Days 22 days ... Pretest-Post Test Days 2 days ... Days of Research Intervention 22 Days" (Table 2, p. 352)
3) "Research treatment was about 4 weeks; the time duration for treatment was 35-40 minutes in each session." (p. 352)
Detailed Analysis:
Criterion Y requires outcomes to be measured at least 75% of a full academic year (roughly 9-10 months) after the intervention begins. The entire study, including pre-test and post-test days, lasted 22 days according to Table 2 (about 4 weeks, or 6 weeks per the abstract). The post-test was administered immediately after the treatment. This is a small fraction of an academic year. Additionally, per the ranking rules, since criterion T (Term Duration) is not met, criterion Y cannot be met.
Criterion Y is not met because the tracking period was only about 4-6 weeks, far below 75% of an academic year.
-
B
Balanced Control Group
- Both groups received the same amount of instructional time (35-40 minute sessions over the same 22-day period) on the same 10th-grade PTB English topics, differing only in teaching method, so no extra time or budget was given to either group and time/resources were balanced.
- "One teacher taught the treatment group by the Grammar Translation method, and the second taught the control group with the traditional Direct Method on specific topics of 10th grade English published by PTB." (p. 351)
Relevant Quotes:
1) "One teacher taught the treatment group by the Grammar Translation method, and the second taught the control group with the traditional Direct Method on specific topics of 10th grade English published by PTB." (p. 351)
2) "For standardisation, the researcher hired two mathematics teachers of the same experience, age, and ability to conduct the treatment study." (p. 351)
3) "Each SLO was taught for 1 day. Research treatment was about 4 weeks; the time duration for treatment was 35-40 minutes in each session." (p. 352)
4) "The researcher controlled extra coaching within the treatment period with the help of the institution's administrator and the student's parents." (p. 351)
Detailed Analysis:
Criterion B requires comparing the nature, quantity, and quality of resources (time, budget, materials, adult support) given to each condition, and asks whether the control condition received a comparable substitute for the intervention's inputs, unless additional resources are explicitly the treatment variable. Applying the decision tree: did the intervention group receive extra time or budget beyond what the control group received? No. Both groups were actively taught the same specific topics of the 10th-grade PTB English curriculum over the identical 22-day treatment window, in sessions of the same 35-40 minute duration, by teachers described as matched on experience, age, and ability (in the version of the design where two teachers were hired; in the alternative statement the same researcher taught both groups, which equally balances the instructor). The only difference between conditions was the pedagogical method (Grammar Translation vs Direct Method), which is the treatment contrast itself, not an additional resource. No extra materials, time, or budget were given to the experimental group, so the "no extra resources present" branch of the decision tree applies and the criterion is trivially satisfied; the researcher also attempted to control for outside coaching during the treatment period.
Criterion B is met because both conditions received equal instructional time on the same curriculum content and differed only in the teaching method being compared, with no additional resources given to either group.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study is reported or found; prior GTM-vs-DM studies cited in the literature review predate this trial and are not replications of it, and an internet search for later citing or replicating studies by independent teams found none.
Relevant Quotes:
1) "(Shejbalová, 2006) Looking at students' acquisition of a second language at an early stage using GTM and DM approaches. According to his experimental group results, GTM outperforms other methods in language acquisition." (p. 349)
2) "Chang (2011) concluded that teaching foreign languages in Taiwan using grammar-translation was his most efficient and effective method." (p. 350)
3) "The Grammar Translation Method was the best method for Bangladeshi students (Mondal, 2012), using a survey study method to gather data from teachers." (p. 350)
Detailed Analysis:
Criterion R requires that this specific study be independently replicated by a different team, in a different context, in a peer-reviewed journal. The paper cites earlier work comparing GTM with other methods (Shejbalova 2006; Chang 2011; Mondal 2012; Walia 2015), but these are prior studies, mostly surveys or different designs, published before this 2022 trial; they cannot be replications of it. A web search for later citations or replication attempts of this specific study (GTM vs Direct Method with 10th-grade boys at Govt. Islamia High School Sambrial, Sialkot, Pakistan, using this MCQ instrument) by an independent team did not surface any such publication. The existence of a broader literature on GTM effectiveness does not constitute an independent reproduction of this study's specific design and findings.
Criterion R is not met because no independent peer-reviewed replication of this specific study exists; cited related works predate the study and are not replications, and no subsequent independent replication was found via internet search.
-
A
All-subject Exams
- Only English outcomes were measured with a custom test, no other school subjects were assessed, and criterion E is not met, which automatically fails criterion A.
- "The researcher used a test on Students' learning achievements as a pre-test and post-test based on three sub-components of Students' learning Achievement in English." (p. 351)
Relevant Quotes:
1) "The researcher used a test on Students' learning achievements as a pre-test and post-test based on three sub-components of Students' learning Achievement in English." (p. 351)
2) "More specifically, this test was categorised as follows: i) putting correct verb, ii) putting correct spelling, iii) putting correct synonyms, and iv) putting correct grammar." (p. 351)
Detailed Analysis:
Criterion A requires measuring impact on all main subjects taught at the educational level, using standardised exams, and explicitly presupposes criterion E. Here, only English language sub-skills (verbs, spelling, synonyms, grammar) were assessed; no other 10th-grade subjects (e.g., mathematics, science, Urdu) were measured. Moreover, the assessment was a custom researcher-made MCQ test, so criterion E is not met, which by the ranking rules makes criterion A automatically not met. No specialised-intervention exception is claimed or justified in the paper.
Criterion A is not met because only English was assessed, using a custom test, and the prerequisite criterion E is not met.
-
G
Graduation Tracking
- Measurement stopped at the post-test immediately after the 4-6 week treatment, with no follow-up tracking to graduation and no follow-up publications found via internet search, and prerequisite criterion Y is not met.
- "A follow-up test was administered to all participants, experimental and control alike." (Abstract, p. 347)
Relevant Quotes:
1) "A follow-up test was administered to all participants, experimental and control alike." (Abstract, p. 347)
2) "Total Days 22 days ... Pretest-Post Test Days 2 days" (Table 2, p. 352)
Detailed Analysis:
Criterion G requires tracking participants until graduation from their educational stage. The only outcome measurement was the post-test administered at the end of the roughly 4-6 week treatment window; the "follow-up test" in the abstract refers to this immediate post-test, not a long-term follow-up. There is no mention of tracking the grade-10 cohort to matriculation (end of secondary school) or of any planned or published follow-up study of this cohort. A web search for subsequent publications by Fahad Izhar or Muhammad Aamir Hashmi tracking this same cohort to graduation did not find any such paper. Additionally, per the ranking rules, criterion G cannot be met because criterion Y is not met.
Criterion G is not met because measurement ended immediately after the short treatment with no tracking of participants to graduation, and no follow-up publication was found.
-
P
Pre-Registered
- The paper contains no mention of any pre-registration, registry, or published protocol for the study, and no registry entry was found via internet search.
Relevant Quotes:
1) "The investigators employed a pre-test and post-test control group design." (p. 350)
2) "This section explains the research materials and procedures used. This chapter covers research design, sampling, instrumentation, data collection, and analysis." (p. 350)
Detailed Analysis:
Criterion P requires that the full study protocol be registered on a public registry before data collection began, with a verifiable link and date. The methods section describes the design, sampling, instrumentation, and analysis, but nowhere in the paper is there any reference to a pre-registration platform (e.g., ClinicalTrials.gov, OSF, AEA registry), a registration ID, or a pre-published protocol. An internet search for a registry entry associated with this study, its authors, or the study site did not find any pre-registration record. Without any quoted evidence of registration prior to data collection, the criterion fails.
Criterion P is not met because no pre-registration or protocol registration is mentioned anywhere in the paper, and none was found through internet search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.