Investigating Metacognitive Think-Aloud Strategy in Improving Saudi EFL Learners' Reading Comprehension and Attitudes

Abdulaziz Ali Al-Qahtani

Published:
ERCT Check Date:
DOI: 10.5539/elt.v13n9p50
  • reading
  • L2 languages
  • K12
  • Asia
0
  • C

    Randomisation was performed at the individual student level within a single secondary school, not at the class or school level, and no tutoring exception applies.

    "The students were assigned to either the experimental or the control group upon a random basis." (p. 55)

  • E

    The study used a custom-designed Reading Comprehension Skills Test and Attitude Scale created and validated by the researcher, not a standardised, widely recognised exam.

    "To achieve that end, the reading texts as well as the Reading Comprehension Skills Test were submitted to the judgment of a group of eleven jury members who agreed on their validity and suitability." (p. 55)

  • T

    Outcomes were measured immediately after a short, twelve-session intervention lasting only a few weeks, far short of a full academic term.

    "The experimental groups received instruction using metacognitive Think-Aloud strategy for twelve sessions with forty minutes each." (p. 55)

  • D

    The control group's size, baseline test scores, and the "business as usual" traditional treatment it received are clearly documented and compared with the experimental group.

    "Students of the experimental group were instructed by using metacognitive Think-Aloud strategy, whereas, the control group received traditional treatment such as skimming and scanning techniques." (Abstract, p. 50)

  • S

    Randomisation occurred at the student level within a single secondary school, not among multiple schools.

    "The sample of the study consisted of forty male students in the first-year from a secondary school in Taif City, Saudi Arabia (n=40)." (p. 55)

  • I

    The study was designed, conducted, and analysed solely by the same researcher, with no evidence of an independent or third-party evaluation team.

    "Based on the researcher's first-hand field experience as an EFL supervisor for 15 years, many secondary school teachers in Taif City, Saudi Arabia are unaware of reading comprehension as a cognitive skill." (p. 54)

  • Y

    Since the term-duration criterion (T) is not met and the intervention lasted only twelve short sessions, the stronger year-duration requirement is also not met.

    "The experimental groups received instruction using metacognitive Think-Aloud strategy for twelve sessions with forty minutes each." (p. 55)

  • B

    Both groups received their normal amount of classroom reading instruction; only the teaching strategy (Think-Aloud versus traditional skimming/scanning) differed, so no additional time or budget was given exclusively to the intervention group.

    "Students of the experimental group were instructed by using metacognitive Think-Aloud strategy, whereas, the control group received traditional treatment such as skimming and scanning techniques." (Abstract, p. 50)

  • R

    No independent replication of this specific study by another research team was found in the paper or via internet search; other cited Think-Aloud studies target different grade levels and contexts.

    "Alaraj, M. (2015). Using Think-Aloud Strategy to improve English reading comprehension for 9th grade students in Saudi Arabia (Doctoral dissertation)."

  • A

    Only reading comprehension (and an attitude scale) were assessed; no other core subjects were measured, and criterion E is also not met.

    "A quantitative study with a quasi-experimental design was implemented through applying two different instruments: Reading Comprehension Skills Test and Attitude Scale towards learning EFL." (Abstract, p. 50)

  • G

    There is no follow-up beyond the immediate post-test, no subsequent tracking papers by the author were found via internet search, and since criterion Y is not met, criterion G cannot be met either.

    "The analysis of data using t-test showed that the experimental group achieved significantly higher scores than the control group on the post-performance of the test of reading comprehension skills..." (p. 56)

  • P

    No pre-registration of the study protocol, hypotheses, or analysis plan is mentioned anywhere in the paper, and no registry record was found via internet search.

Abstract

The current study's objective examines the effectiveness of using a Think-Aloud strategy in improving Saudi EFL learners' reading comprehension and attitudes towards learning. A quantitative study with a quasi-experimental design was implemented through applying two different instruments: Reading Comprehension Skills Test and Attitude Scale towards learning EFL. The study adopts a pre-post control group design where forty students were randomly assigned to either a control or an experimental group. Students of the experimental group were instructed by using metacognitive Think-Aloud strategy, whereas, the control group received traditional treatment such as skimming and scanning techniques. The findings of the study showed that the attitudes and reading comprehension skills of the experimental group improved significantly as opposed to the control group. The study gives more insight into the importance of applying a Think-Aloud strategy in teaching reading comprehension inside EFL educational context. The study also suggests recommendations for EFL teachers to increase the efficiency of applying this strategy through their teaching procedures.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was performed at the individual student level within a single secondary school, not at the class or school level, and no tutoring exception applies.
      • "The students were assigned to either the experimental or the control group upon a random basis." (p. 55)
      • Relevant Quotes: 1) "The study adopts a pre-post control group design where forty students were randomly assigned to either a control or an experimental group." (Abstract, p. 50) 2) "The sample of the study consisted of forty male students in the first-year from a secondary school in Taif City, Saudi Arabia (n=40). The students were assigned to either the experimental or the control group upon a random basis." (p. 55) Detailed Analysis: The paper explicitly states that individual students (n=40) from a single secondary school were randomly assigned to either the experimental or control condition. There is no mention of randomising by class or by school; instead, forty students within the same first-year cohort were split into two groups of twenty on a random basis. This raises the risk of contamination between the experimental and control students, as they attend the same school and could interact. The intervention (teaching a reading strategy) is not a personal one-to-one tutoring intervention, so the "personal teaching" exception described in the ERCT standard does not apply here. Because randomisation occurred at the student level within one school rather than at the class or school level, and no valid exception applies, criterion C is not met.
    • E

      Exam-based Assessment

      • The study used a custom-designed Reading Comprehension Skills Test and Attitude Scale created and validated by the researcher, not a standardised, widely recognised exam.
      • "To achieve that end, the reading texts as well as the Reading Comprehension Skills Test were submitted to the judgment of a group of eleven jury members who agreed on their validity and suitability." (p. 55)
      • Relevant Quotes: 1) "A quantitative study with a quasi-experimental design was implemented through applying two different instruments: Reading Comprehension Skills Test and Attitude Scale towards learning EFL." (Abstract, p. 50) 2) "The study was pretested with a comparison group of 15 male students to determine the reliability of the reading comprehension study. Then, after analyzing the data obtained, three items were omitted from the entire study, and four more items were updated." (p. 55) 3) "To achieve that end, the reading texts as well as the Reading Comprehension Skills Test were submitted to the judgment of a group of eleven jury members who agreed on their validity and suitability. Four jury members suggested some modifications and accordingly modifications have been made." (p. 55) 4) "To assess the empirical validity of the reading comprehension test, Pearson Correlation between the reading comprehension tests (authentic and inauthentic) and TOEFL were measured as (0.68) and (0.63) respectively." (p. 55) Detailed Analysis: The reading comprehension and attitude instruments used in this study were developed specifically by the researcher for this study, piloted with a comparison group, revised based on item analysis, and reviewed by a jury of eleven experts. Although the researcher correlated the custom test with TOEFL scores to argue for its validity, this correlational check does not make the instrument itself a standardised, widely recognised exam; it remains a researcher-designed measure tailored to the intervention's specific reading texts. Because the assessment instruments were custom-built for this study rather than being standard, widely recognised examinations, criterion E is not met.
    • T

      Term Duration

      • Outcomes were measured immediately after a short, twelve-session intervention lasting only a few weeks, far short of a full academic term.
      • "The experimental groups received instruction using metacognitive Think-Aloud strategy for twelve sessions with forty minutes each." (p. 55)
      • Relevant Quotes: 1) "Finally, the experimental groups received instruction using metacognitive Think-Aloud strategy for twelve sessions with forty minutes each. The teaching sessions were 12 sessions including 12 reading texts." (p. 55) 2) "The analysis of data using t-test showed that the experimental group achieved significantly higher scores than the control group on the post-performance of the test of reading comprehension skills..." (p. 56) Detailed Analysis: The intervention consisted of only twelve 40-minute sessions (approximately eight hours of instruction in total), and outcomes were measured via a post-test administered immediately following completion of these sessions. There is no indication that measurement occurred a full academic term (roughly 3-4 months) after the intervention began; instead, the entire intervention-to-measurement window appears to span only a few weeks. Because the interval from intervention start to outcome measurement is far shorter than one academic term, criterion T is not met.
    • D

      Documented Control Group

      • The control group's size, baseline test scores, and the "business as usual" traditional treatment it received are clearly documented and compared with the experimental group.
      • "Students of the experimental group were instructed by using metacognitive Think-Aloud strategy, whereas, the control group received traditional treatment such as skimming and scanning techniques." (Abstract, p. 50)
      • Relevant Quotes: 1) "Students of the experimental group were instructed by using metacognitive Think-Aloud strategy, whereas, the control group received traditional treatment such as skimming and scanning techniques." (Abstract, p. 50) 2) "Table 1. Means, standard deviations, t-value of means and significance of differences of the two groups in the pre-performance on reading comprehension skills test... Cont. 20 31.82 8.83... 38 0.04." (p. 55) 3) "Table 2. Means, standard deviations, t-value of means and significance of differences of the two groups in the pre-performance on the Attitude Scale... Cont. 20 30.01 7.41 38 0.48." (p. 55) 4) "Pre-test data on the reading comprehension skills test showed group equivalence as the t-value (0.04) was insignificant at p ≤ .05 level as in table (1)." (p. 55) Detailed Analysis: The paper documents the control group's size (n=20), provides its pre-test means and standard deviations for both the reading comprehension test and the attitude scale, and confirms baseline equivalence with the experimental group via non-significant t-tests. It also explicitly states what treatment the control group received during the study (traditional skimming and scanning instruction) rather than leaving it undefined. This level of detail is sufficient to assess whether the control group was comparable at baseline and to understand what activities it engaged in. Because the control group's baseline characteristics, size, and treatment are clearly documented, criterion D is met.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred at the student level within a single secondary school, not among multiple schools.
      • "The sample of the study consisted of forty male students in the first-year from a secondary school in Taif City, Saudi Arabia (n=40)." (p. 55)
      • Relevant Quotes: 1) "The sample of the study consisted of forty male students in the first-year from a secondary school in Taif City, Saudi Arabia (n=40). The students were assigned to either the experimental or the control group upon a random basis." (p. 55) Detailed Analysis: The entire study was conducted within a single secondary school, with individual students from that one school randomly split into experimental and control groups. There is no mention of multiple schools being recruited or randomised. Since criterion C (class-level) is already not met, and school-level randomisation is a stronger requirement, this criterion cannot be met either. Because randomisation was neither at the class nor the school level, criterion S is not met.
    • I

      Independent Conduct

      • The study was designed, conducted, and analysed solely by the same researcher, with no evidence of an independent or third-party evaluation team.
      • "Based on the researcher's first-hand field experience as an EFL supervisor for 15 years, many secondary school teachers in Taif City, Saudi Arabia are unaware of reading comprehension as a cognitive skill." (p. 54)
      • Relevant Quotes: 1) "Based on the researcher's first-hand field experience as an EFL supervisor for 15 years, many secondary school teachers in Taif City, Saudi Arabia are unaware of reading comprehension as a cognitive skill." (p. 54) 2) "Based on a pilot study, the researcher found out that most of the students showed poor achievement..." (p. 54) 3) "The acknowledgment is for public school teachers and students for their efforts and contribution in this study and their valuable recommendations." (p. 58) Detailed Analysis: The paper is single-authored, and the same researcher who designed the study (as a long-time EFL supervisor in the target school district) appears to have also run the pilot study, administered the intervention, and analysed the results. There is no statement of an external or independent evaluation team collecting data or analysing outcomes separately from the intervention's designer. Because the same individual designed and conducted the study without any documented independent oversight, criterion I is not met.
    • Y

      Year Duration

      • Since the term-duration criterion (T) is not met and the intervention lasted only twelve short sessions, the stronger year-duration requirement is also not met.
      • "The experimental groups received instruction using metacognitive Think-Aloud strategy for twelve sessions with forty minutes each." (p. 55)
      • Relevant Quotes: 1) "The experimental groups received instruction using metacognitive Think-Aloud strategy for twelve sessions with forty minutes each. The teaching sessions were 12 sessions including 12 reading texts." (p. 55) Detailed Analysis: Per the ERCT standard, if the weaker Term Duration criterion (T) is not met, then the stronger Year Duration criterion (Y) is automatically not met. Independently, the paper's own description of a twelve-session, 40-minutes-per-session intervention with immediate post-testing shows the tracked period is nowhere close to 75% of an academic year. Criterion Y is not met both because criterion T is not met and because the actual intervention-to-measurement period is far shorter than a year.
    • B

      Balanced Control Group

      • Both groups received their normal amount of classroom reading instruction; only the teaching strategy (Think-Aloud versus traditional skimming/scanning) differed, so no additional time or budget was given exclusively to the intervention group.
      • "Students of the experimental group were instructed by using metacognitive Think-Aloud strategy, whereas, the control group received traditional treatment such as skimming and scanning techniques." (Abstract, p. 50)
      • Relevant Quotes: 1) "Students of the experimental group were instructed by using metacognitive Think-Aloud strategy, whereas, the control group received traditional treatment such as skimming and scanning techniques." (Abstract, p. 50) 2) "Finally, the experimental groups received instruction using metacognitive Think-Aloud strategy for twelve sessions with forty minutes each." (p. 55) Detailed Analysis: Applying the Criterion B decision procedure: the first question is whether the intervention added extra time or budget relative to the control group. The paper does not describe any additional class time, budget, or materials given exclusively to the experimental group beyond their regular reading lessons. Instead, both groups appear to receive their normal reading instruction during the same study period; the only described difference is the instructional technique taught (metacognitive Think-Aloud strategy versus traditional skimming and scanning techniques), not extra dosage, time, or resources. Since no extra resources are present, the decision tree resolves to "met" without needing to consider the integral-resource or within-subjects branches. The paper does not explicitly restate the exact session count for the control group, which is a minor documentation gap, but there is no indication of a resource or time imbalance of the kind the criterion is designed to catch (e.g., extra tutoring hours, additional materials, or added budget for the treatment group only). Because no extra time or budget is described as being given to the intervention group relative to the control group, and the difference between conditions is one of instructional method rather than dosage, criterion B is met, though this verdict rests on the absence of any stated imbalance rather than an explicit confirmation of matched session counts.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study by another research team was found in the paper or via internet search; other cited Think-Aloud studies target different grade levels and contexts.
      • "Alaraj, M. (2015). Using Think-Aloud Strategy to improve English reading comprehension for 9th grade students in Saudi Arabia (Doctoral dissertation)."
      • Relevant Quotes: 1) "Alaraj, M. (2015). Using Think-Aloud Strategy to improve English reading comprehension for 9th grade students in Saudi Arabia (Doctoral dissertation)." (References, p. 59) 2) "Alqahtani, M. A. (2015). The effects of Think-Aloud strategy to improve reading comprehension of 6th grade students in Saudi Arabia (Doctoral dissertation)." (References, p. 60) Detailed Analysis: The paper cites related Think-Aloud studies conducted in the Saudi context (Alaraj, 2015, with 9th graders; Alqahtani, 2015, with 6th graders), but these both predate this 2020 paper, target different populations and grade levels, and (being doctoral dissertations rather than applications of the same experimental design to the same first-year secondary cohort in Taif City) do not constitute independent replications of this specific study. A further internet search (Google Scholar, Semantic Scholar, ResearchGate) for citing or related papers that explicitly reproduce this study's design, sample, and measures in a different context by a different, independent research team returned only general reviews and unrelated metacognitive-reading studies on Saudi EFL learners; no independent reproduction of this exact study was identified. Because no independent replication of this specific study was identified either in the paper or via internet search, criterion R is not met.
    • A

      All-subject Exams

      • Only reading comprehension (and an attitude scale) were assessed; no other core subjects were measured, and criterion E is also not met.
      • "A quantitative study with a quasi-experimental design was implemented through applying two different instruments: Reading Comprehension Skills Test and Attitude Scale towards learning EFL." (Abstract, p. 50)
      • Relevant Quotes: 1) "A quantitative study with a quasi-experimental design was implemented through applying two different instruments: Reading Comprehension Skills Test and Attitude Scale towards learning EFL." (Abstract, p. 50) Detailed Analysis: The study's outcomes are limited to reading comprehension skills and attitudes towards learning EFL; no other core school subjects (such as mathematics, science, or other language arts components) were assessed. In addition, since criterion E (Exam-based Assessment using a standardised test) is not met, the ERCT standard specifies that criterion A cannot be met either. Because only a single, non-standardised reading/attitude measure was used and criterion E is not met, criterion A is not met.
    • G

      Graduation Tracking

      • There is no follow-up beyond the immediate post-test, no subsequent tracking papers by the author were found via internet search, and since criterion Y is not met, criterion G cannot be met either.
      • "The analysis of data using t-test showed that the experimental group achieved significantly higher scores than the control group on the post-performance of the test of reading comprehension skills..." (p. 56)
      • Relevant Quotes: 1) "The analysis of data using t-test showed that the experimental group achieved significantly higher scores than the control group on the post-performance of the test of reading comprehension skills, since t-value (9.44) is significant at p ≤ .01 level." (p. 56) 2) "7. Limitations and Study Forward... The following topics are suggested as areas that need further investigations:" (p. 58) Detailed Analysis: Outcomes were measured only once, immediately following the twelve-session intervention; there is no mention of any subsequent follow-up, let alone tracking through to graduation. The "Limitations and Study Forward" section proposes entirely new future studies with different populations (e.g., female students, slow learners), not a planned longitudinal follow-up of this cohort. An internet search for later publications by Abdulaziz Ali Al-Qahtani (Google Scholar, ResearchGate) that might track this same cohort of forty first-year secondary students through graduation found no such follow-up study; the author's subsequent visible work does not report longer-term tracking of this sample. Per the ERCT standard, since criterion Y is not met, criterion G is automatically not met as well. Because there is no tracking beyond the immediate post-test, no follow-up publication was found, and criterion Y is not met, criterion G is not met.
    • P

      Pre-Registered

      • No pre-registration of the study protocol, hypotheses, or analysis plan is mentioned anywhere in the paper, and no registry record was found via internet search.
      • Relevant Quotes: (No statement referencing a study registry, registration ID, or pre-registration date was found anywhere in the paper.) Detailed Analysis: The paper describes its hypotheses (H1 and H2) directly in the Results section, with no reference to a public trial registry (such as ISRCTN, ClinicalTrials.gov, or OSF Registries), no registration identifier, and no statement about when a protocol was registered relative to data collection. An internet search for a pre-registration record tied to this study or its author (Al-Qahtani, 2020, ELT journal) found no matching entry in any trial or study registry. Because there is no evidence of pre-registration anywhere in the paper or in available registries, criterion P is not met.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.