Investigating the Effectiveness of Direct Method and Grammar Translation Method in Teaching Reading Skills in Medical Sciences

Ali Al-Sultan, Faisal Al-Dawli

Published:
ERCT Check Date:
DOI: 10.59628/jhs.v5i3.2261
  • reading
  • L2 languages
  • higher education
  • Asia
0
  • C

    Assignment was made to three intact academic departments (whole cohorts of ~40 students each) rather than to individuals within one shared classroom, satisfying class-level randomisation, although the abstract confusingly labels the design "quasi-experimental."

    "These were then randomly assigned to the three study groups: - First Experimental Group (DM): Anesthetics (f) - Second Experimental Group (GTM): Operations Technology (h) - Control Group (LM): Radiology (g)." (p. 683-684)

  • E

    The reading test was a custom instrument built from a single course textbook unit, not a widely recognised standardised exam.

    "The test format is built in the light of student's book and practice file, (Nursing -1- Book, Oxford English for Careers)." (p. 684)

  • T

    The intervention and measurement together spanned only about two months (42 hours of teaching), well short of a full academic term.

    "The study lasted for two months, three days a week, starting from November 2024 to the end of December 2024, in the second academic term at 21 September UMAS." (p. 686)

  • D

    The control (Lecture Method) group's department, sample size, and pre/post-test descriptive statistics are clearly reported.

    "Control Group (LM): Radiology (g)" (p. 684)

  • S

    Randomisation occurred among departments within a single university, not among separate schools.

    "The first level students and English teachers at 21 September UMAS in the academic year 2023 are the population of this study." (p. 683)

  • I

    The same researcher(s) designed the intervention, delivered/coordinated the teaching, conducted interviews, and analysed the data, with no independent third-party evaluator involved.

    "So, the researcher was only allowed to take notes during the interview." (p. 685)

  • Y

    Because the Term Duration criterion (T) is not met, the stronger Year Duration criterion is also not met; the study covered only about two months.

    "The study lasted for two months, three days a week, starting from November 2024 to the end of December 2024, in the second academic term at 21 September UMAS." (p. 686)

  • B

    All three groups received the same course content, teaching duration, and identical pre/post-tests; only the pedagogical method (not the amount of time, materials, or budget) differed between groups.

    "The materials used in this study contained an English course book as EMP with exactly similar content taught by two different teachers in three different teaching methods, namely DM, GTM and LM." (p. 686)

  • R

    The authors explicitly state this is the first study of its kind, and a dedicated internet search found no independent replication.

    "Geographically, no studies have investigated the GTM and DM with ESP students in Yemen... So, this study is the first to fill this gap through investigating the effect of GTM and DM compared to the traditional method (LM) of teaching reading skills at 21 September university." (p. 682-683)

  • A

    Only reading skills in English were assessed, no other core subjects were measured, and criterion E (a prerequisite) is not met.

    "This study investigates the impact of Grammar Translation Method (GTM) and the Direct Method (DM) on the development of reading skills of the first-year medical students at 21 September University for Medical and Applied Sciences in Sana'a." (p. 677)

  • G

    Since criterion Y is not met and no follow-up beyond the immediate post-test is reported or found in later publications, there is no graduation tracking.

    "Second, the intervention period was limited to 42 hours." (p. 693)

  • P

    No pre-registration of the study protocol is mentioned anywhere in the paper, and an internet search of registry platforms found no matching registration.

Abstract

This study investigates the impact of Grammar Translation Method (GTM) and the Direct Method (DM) on the development of reading skills of the first-year medical students at 21 September University for Medical and Applied Sciences in Sana'a. A total of 120 students enrolled in an English for Medical Purposes (EMP) course were selected through a random sampling technique and divided into three groups, each consisting of 40 participants. The first and second groups are assigned to Direct Method and Grammar Translation Method respectively, and the third (i.e., the control group) to Lecture Method (LM). The study employed a quasi-experimental design with pre- and post-tests over a two-month intervention period. Data were collected through tests and semi-structured interviews. Test results were analyzed quantitively using SPSS and interview data were analyzed qualitatively to capture teachers' perspectives. Findings from both quantitative and qualitative analyses indicated that Grammar Translation Method and Direct Method were more effective than Lecture Method in enhancing reading skills. These gains were evident in key reading domains, including comprehension, vocabulary, and grammatical accuracy. The study recommends further longitudinal research into varied instructional methods to assess their long-term effects on medical students' reading proficiency and to inform EMP curriculum development.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Assignment was made to three intact academic departments (whole cohorts of ~40 students each) rather than to individuals within one shared classroom, satisfying class-level randomisation, although the abstract confusingly labels the design "quasi-experimental."
      • "These were then randomly assigned to the three study groups: - First Experimental Group (DM): Anesthetics (f) - Second Experimental Group (GTM): Operations Technology (h) - Control Group (LM): Radiology (g)." (p. 683-684)
      • Relevant Quotes: 1) "This study used a sample of 120 students divided into three groups. The researcher employed a simple random sampling technique to ensure the sample's validity and representativeness." (p. 683) 2) "To illustrate the process, the researcher provided a simplified example: from a population of 9 departments (N=9), they randomly selected a sample of 3 (n=3) using a lottery method (Kothari, 2004). From the 28 possible combinations, the departments of Anesthetics (f), Radiology (g), and Operations Technology (h) were selected." (p. 683-684) 3) "These were then randomly assigned to the three study groups: - First Experimental Group (DM): Anesthetics (f) - Second Experimental Group (GTM): Operations Technology (h) - Control Group (LM): Radiology (g)." (p. 683-684) 4) "The study employed a quasi-experimental design with pre- and post-tests over a two-month intervention period." (Abstract, p. 677) Detailed Analysis: The unit of assignment described in the methods is the academic department (an intact cohort of ~40 students), not individual students inside one shared classroom. Anesthetics, Radiology, and Operations Technology were each assigned wholesale to one of the three conditions, which is consistent with the intent of the class-level criterion (avoiding within-classroom contamination between conditions). The department sizes reported for these three departments in Table 1 (40, 40, 40) exactly match the n=40 per group used in the study, supporting that this passage genuinely describes the actual assignment rather than only illustrating a generic textbook procedure. However, the methodology's odd framing as "a simplified example" borrowed from Kothari (2004), combined with the abstract's explicit self-description as a "quasi-experimental design," creates real tension about how rigorously random this assignment actually was; this is flagged as a significant documentation weakness that should be treated with caution during verification. The quote confirming the treatment-arm assignment was re-verified against the PDF and reformatted to match the original bulleted list (it spans the page 683/684 boundary) rather than the single semicolon-joined sentence used previously. Final: Criterion C is met because whole departments/cohorts, not individual students within one class, were assigned to conditions, though the paper's own "quasi-experimental" label is a notable inconsistency worth independent scrutiny.
    • E

      Exam-based Assessment

      • The reading test was a custom instrument built from a single course textbook unit, not a widely recognised standardised exam.
      • "The test format is built in the light of student's book and practice file, (Nursing -1- Book, Oxford English for Careers)." (p. 684)
      • Relevant Quotes: 1) "It consists of three types of questions drawn from unit (3), Nursing (1) Book, Oxford English for Careers (Grice, 2011)." (p. 684) 2) "The test format is built in the light of student's book and practice file, (Nursing -1- Book, Oxford English for Careers)." (p. 684) 3) "The test of this study is validated by 10 university professors who have long experience and interest in teaching English... They agreed on the validity and suitability of the tests. After making some revisions, a pilot test was conducted on 20 students before conducting the main study." (p. 684) Detailed Analysis: The ERCT Standard requires a widely recognised, standardised exam rather than a test purpose-built for the study. Here, the pre/post-test was constructed directly from a single unit ("Hospital Admission") of one specific course textbook, with True/False, matching, and multiple-choice items designed by the researcher. Although the instrument underwent expert review by 10 professors and a pilot with 20 students, and Cronbach's Alpha (.868) was reported, this establishes content validity and reliability for this specific researcher-made test, not that it is a standardised, externally validated exam used more broadly. There is no mention of a national, standardised, or widely recognised assessment instrument. Criterion E is not met because the assessment was a custom test built from course materials rather than a standardised exam.
    • T

      Term Duration

      • The intervention and measurement together spanned only about two months (42 hours of teaching), well short of a full academic term.
      • "The study lasted for two months, three days a week, starting from November 2024 to the end of December 2024, in the second academic term at 21 September UMAS." (p. 686)
      • Relevant Quotes: 1) "The study employed a quasi-experimental design with pre- and post-tests over a two-month intervention period." (p. 677) 2) "The study lasted for two months, three days a week, starting from November 2024 to the end of December 2024, in the second academic term at 21 September UMAS." (p. 686) 3) "Second, the intervention period was limited to 42 hours." (p. 693, Study limitations) 4) "The post-test is given to find out the results of the teaching process in the three groups and to know the achievement of learning after the treatment." (p. 685) Detailed Analysis: The ERCT Standard requires the interval from intervention start to the outcome measurement to span at least one academic term (roughly 3-4 months). Here, the entire intervention plus post-testing occurred within a two-month window (November-December 2024), and the authors themselves flag the short 42-hour intervention period as a limitation of the study. The post-test was administered immediately at the conclusion of the two-month teaching period, with no delayed follow-up measurement that would extend the start-to-measurement interval closer to a full term. Criterion T is not met because the documented interval from intervention start to outcome measurement was only about two months, short of the required academic term.
    • D

      Documented Control Group

      • The control (Lecture Method) group's department, sample size, and pre/post-test descriptive statistics are clearly reported.
      • "Control Group (LM): Radiology (g)" (p. 684)
      • Relevant Quotes: 1) "The first and second groups are assigned to Direct Method and Grammar Translation Method respectively, and the third (i.e., the control group) to Lecture Method (LM)." (p. 677) 2) "Control Group (LM): Radiology (g)" (p. 684) 3) "Table (7): Descriptive Statistics of Pre and Post-tests Results of Control Group (LM)- Reading Skill ... Pretest 40 56.5000 7.77900 1.22997 ... Post-test 40 65.8750 8.76235 1.38545" (p. 688) 4) "Table (8) shows a significant difference between the mean scores on the pretest and post-test, with a p-value of .000." (p. 688) Detailed Analysis: The control group (Lecture Method) is clearly identified as the Radiology department's 40 students, and its baseline (pretest) and outcome (post-test) means, standard deviations, and paired t-test results are explicitly tabulated (Table 7 and 8), alongside a description of what the control condition entailed (traditional lecture-based instruction rather than DM or GTM). This level of detail allows a reader to assess the size and baseline comparability of the control group relative to the treatment groups. Criterion D is met because the control group's identity, size, and baseline/outcome statistics are documented in the paper.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred among departments within a single university, not among separate schools.
      • "The first level students and English teachers at 21 September UMAS in the academic year 2023 are the population of this study." (p. 683)
      • Relevant Quotes: 1) "The first level students and English teachers at 21 September UMAS in the academic year 2023 are the population of this study." (p. 683) 2) "...from a population of 9 departments (N=9), they randomly selected a sample of 3 (n=3) using a lottery method (Kothari, 2004)." (p. 683) 3) "First Experimental Group (DM): Anesthetics (f)" (p. 684) 4) "Second Experimental Group (GTM): Operations Technology (h)" (p. 684) 5) "Control Group (LM): Radiology (g)" (p. 684) Detailed Analysis: The ERCT S criterion requires randomisation among whole schools (or equivalent independent institutional units), not among sub-units of a single institution. Here, the entire study was conducted at one university (21 September UMAS), and the unit of assignment was academic department (Anesthetics, Radiology, Operations Technology) within that single institution, not separate schools or universities. This does not meet the stronger school-level requirement. Criterion S is not met because randomisation occurred among departments of a single university rather than among independent schools.
    • I

      Independent Conduct

      • The same researcher(s) designed the intervention, delivered/coordinated the teaching, conducted interviews, and analysed the data, with no independent third-party evaluator involved.
      • "So, the researcher was only allowed to take notes during the interview." (p. 685)
      • Relevant Quotes: 1) "The researcher conducted a semi-structured interview with teachers at 21 September UMAS for the purpose of evaluating the satisfactory level of EMP teachers about GTM and DM of teaching reading skills." (p. 685) 2) "All of them were interviewed in English by the researcher during 4 days." (p. 685) 3) "So, the researcher was only allowed to take notes during the interview." (p. 685) 4) "All of the data in the exams were processed by SPSS software, version 22, and they were analyzed by the researcher statistically." (p. 686) Detailed Analysis: There is no statement anywhere in the paper of an external, independent evaluation team collecting data, administering tests, or analysing results separately from the researcher(s) who designed the study. On the contrary, the text repeatedly describes "the researcher" personally conducting interviews, taking notes, and statistically analysing the exam data. There is no mention of blinded or third-party test administrators, nor of any separation between the team that designed the DM/GTM materials and the team that assessed outcomes. Criterion I is not met because the same researcher team designed, implemented, and evaluated the study with no documented independent oversight.
    • Y

      Year Duration

      • Because the Term Duration criterion (T) is not met, the stronger Year Duration criterion is also not met; the study covered only about two months.
      • "The study lasted for two months, three days a week, starting from November 2024 to the end of December 2024, in the second academic term at 21 September UMAS." (p. 686)
      • Relevant Quotes: 1) "The study lasted for two months, three days a week, starting from November 2024 to the end of December 2024, in the second academic term at 21 September UMAS." (p. 686) 2) "Second, the intervention period was limited to 42 hours." (p. 693) Detailed Analysis: Per the ERCT specification, if criterion T (Term Duration) is not met, criterion Y (Year Duration) is automatically not met. Independently, the documented intervention and measurement window of about two months (42 hours of teaching) falls far short of the requirement that outcomes be measured at least 75% of an academic year (roughly 9-10 months) after the intervention begins. Criterion Y is not met both because T is not met and because the actual tracked duration (about two months) is far shorter than a year.
    • B

      Balanced Control Group

      • All three groups received the same course content, teaching duration, and identical pre/post-tests; only the pedagogical method (not the amount of time, materials, or budget) differed between groups.
      • "The materials used in this study contained an English course book as EMP with exactly similar content taught by two different teachers in three different teaching methods, namely DM, GTM and LM." (p. 686)
      • Relevant Quotes: 1) "The materials used in this study contained an English course book as EMP with exactly similar content taught by two different teachers in three different teaching methods, namely DM, GTM and LM." (p. 686) 2) "So, in this research, the pre-test and post- test are given to the experimental and control groups with the same test and topic." (p. 683) 3) "Student's book, 'Nursing (1): Oxford English for Careers' by Tony Grice, unit three 'Hospital Admission' (2011)." (p. 686, listed as the shared teaching material for all groups) Detailed Analysis: Applying the Criterion B decision procedure: the first check is whether the intervention adds any extra instructional time, materials, or budget to the treatment (DM/GTM) groups relative to the control (LM) group. All three groups used the same textbook unit ("Hospital Admission"), the same pre-/post-test, and were taught over the same two-month, three-days-a-week schedule; the only manipulated variable is the pedagogical technique used to teach that identical content (Direct Method vs. Grammar Translation Method vs. Lecture Method). Since EXTRA_RESOURCES_PRESENT is false (no additional time, budget, or materials are granted to either experimental group beyond what the control group received), the decision tree resolves at its first branch and the balance requirement is trivially satisfied. Criterion B is met because no additional time, budget, or materials were provided to the intervention groups relative to the control group; only the teaching technique varied.
  • Level 3 Criteria

    • R

      Reproduced

      • The authors explicitly state this is the first study of its kind, and a dedicated internet search found no independent replication.
      • "Geographically, no studies have investigated the GTM and DM with ESP students in Yemen... So, this study is the first to fill this gap through investigating the effect of GTM and DM compared to the traditional method (LM) of teaching reading skills at 21 September university." (p. 682-683)
      • Relevant Quotes: 1) "Despite these findings, several gaps remain. Geographically, no studies have investigated the GTM and DM with ESP students in Yemen. Contextually, there is a lack of research on discipline-specific reading skills in medical ESP." (p. 682) 2) "So, this study is the first to fill this gap through investigating the effect of GTM and DM compared to the traditional method (LM) of teaching reading skills at 21 September university." (p. 682-683) Detailed Analysis: The authors themselves position this study as novel and the first of its kind for this specific population (Yemeni EMP/medical students) and design (DM vs. GTM vs. LM together). A targeted internet search (including Google Scholar, using the authors' names and the study's key terms) was conducted to look for independent replications by other research teams; it returned only the original paper itself, with no subsequent replication study identified. Given the paper's very recent (2026) publication and the narrow, highly specific context (21 September UMAS medical students in Yemen), this is expected, but no external reproduction could be confirmed. Criterion R is not met because this is presented as a first, unreplicated study in its specific context, and internet searching did not surface any independent replication.
    • A

      All-subject Exams

      • Only reading skills in English were assessed, no other core subjects were measured, and criterion E (a prerequisite) is not met.
      • "This study investigates the impact of Grammar Translation Method (GTM) and the Direct Method (DM) on the development of reading skills of the first-year medical students at 21 September University for Medical and Applied Sciences in Sana'a." (p. 677)
      • Relevant Quotes: 1) "This study investigates the impact of Grammar Translation Method (GTM) and the Direct Method (DM) on the development of reading skills of the first-year medical students at 21 September University for Medical and Applied Sciences in Sana'a." (p. 677) 2) "It consists of three types of questions drawn from unit (3), Nursing (1) Book, Oxford English for Careers (Grice, 2011)." (p. 684) Detailed Analysis: Per the ERCT specification, criterion A requires criterion E to be met as a prerequisite; since E (standardised exam) is not met here, A cannot be met either. Independently, the study measured only reading skill outcomes (comprehension, vocabulary, grammatical accuracy within reading) in one English course, with no assessment of other main subjects taught to these medical students (e.g. core medical or science subjects), and no explicit justification is given for restricting measurement to this single skill area beyond it being the study's chosen focus. Criterion A is not met because only reading skills were assessed with a non-standardised test, and the prerequisite criterion E is not met.
    • G

      Graduation Tracking

      • Since criterion Y is not met and no follow-up beyond the immediate post-test is reported or found in later publications, there is no graduation tracking.
      • "Second, the intervention period was limited to 42 hours." (p. 693)
      • Relevant Quotes: 1) "The post-test is given to find out the results of the teaching process in the three groups and to know the achievement of learning after the treatment." (p. 685) 2) "The study has some limitations that should be addressed in future research... To address these limitations, future studies should be conducted across a wider range of medical institutions and over a longer duration." (p. 693) 3) "The study recommends further longitudinal research into varied instructional methods to assess their long-term effects on medical students' reading proficiency..." (p. 677, Abstract) Detailed Analysis: Per the ERCT specification, since criterion Y (Year Duration) is not met, criterion G is automatically not met. In addition, the paper contains no description of any follow-up data collection beyond the immediate post-test at the end of the two-month intervention; instead, the authors explicitly call for future "longitudinal research" as a recommendation, confirming that no such long-term or graduation tracking was performed in this study. An internet search for later papers by Ali Al-Sultan or Faisal Al-Dawli that might track this same cohort of students through graduation (given the very recent, 2026 publication date) did not identify any such follow-up publication. Criterion G is not met because there is no follow-up beyond the immediate post-test, the prerequisite criterion Y is not met, and no follow-up publication tracking this cohort could be found online.
    • P

      Pre-Registered

      • No pre-registration of the study protocol is mentioned anywhere in the paper, and an internet search of registry platforms found no matching registration.
      • Relevant Quotes: (No quotes found: the paper contains no reference to a trial registry, registration ID, or registration date anywhere in the text, methodology, or references.) Detailed Analysis: A thorough review of the methodology, procedures, and references sections reveals no mention of any pre-registration platform (e.g. OSF, AsPredicted, ClinicalTrials.gov, ISRCTN) or of hypotheses/ analysis plans having been published before data collection began. The hypotheses are stated in the main paper itself (Section 1) with no indication they were registered in advance. An internet search of OSF and general web sources for a pre-registration by these authors or of this specific study did not find any matching entry, consistent with the paper's own silence on the matter. Criterion P is not met because no evidence of pre-registration is present in the paper or found through internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.