Abstract
In Egypt most foreign-language audiovisual materials are viewed in the original audio version with the Arabic subtitles. With the increasing access to audiovisual material, EFL learners are more exposed to oral input in English than ever before. However, what has unfairly remained unresolved is the use of subtitles. Thus, the present study represents an effort to empirically examine the effect of feature movies with and without subtitle on listening comprehension of Egyptian EFL learners. The participants of this study were randomly selected from a larger pool of 180 fourth year students from the Department of English Language, Minoufia University, Egypt. 104 participants were randomly chosen and then randomly assigned to the two experimental groups and the control group. The first experimental group viewed the movie with English subtitles (ESG), the second group viewed the movie with Arabic subtitles (ASG), and the control group viewed the movie without subtitles (WSG). After screening the 14 excerpts from 8 movies, 14 multiple-choice listening comprehension tests were administered in order to evaluate their listening comprehensions. Then, each group was asked to complete a questionnaire in order to know their opinions about the way the movie was presented. In the last session, all participants sat on the listening section from the TOEFL test. The results of the data analysis revealed that for Multiple Choice tests, the subtitles groups outperformed the WSG, and ASG performed better than the ESG; however, on the TOEFL test, the analysis of groups' performance revealed a better mean score for the ESG compared to other groups.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was carried out at the level of individual students rather than classes or schools, and no tutoring-style exception applies.
- "One hundred and four participants were randomly assigned to the two experimental groups and the control group, each one consists of 40 participants." (p. 280)
Relevant Quotes:
1) "The participants of this study were randomly selected from a larger pool of 180 fourth year students from the Department of English Language, Minoufia University, Egypt. 104 participants were randomly chosen and then randomly assigned to the two experimental groups and the control group." (p. 275)
2) "One hundred and four participants were randomly assigned to the two experimental groups and the control group, each one consists of 40 participants." (p. 280)
3) "One hundred and four participants were chosen and then randomly assigned to the two experimental groups and the control group." (p. 281)
Detailed Analysis:
The unit of randomisation described throughout the paper is the individual student, drawn from a single pool of fourth-year English majors at one university department. There is no mention of classes, sections, or schools being randomised as intact units, and no indication that contamination between conditions was structurally prevented via class- or school-level assignment. The intervention (group movie viewing sessions with different subtitle conditions) is not a personal, one-to-one tutoring intervention, so the tutoring exception to the class-level requirement does not apply.
Criterion C is not met because randomisation occurred at the student level within one institution, with no applicable exception.
-
E
Exam-based Assessment
- Alongside researcher-made comprehension tests, the study also administered the TOEFL listening section (Institutional TOEFL Program), a widely recognised standardised exam, as a pre/post measure.
- "In the last session, all participants sat on the listening section from the TOEFL test." (p. 275)
Relevant Quotes:
1) "In order to examine the participants' understanding of the movies, 14 comprehension tests were constructed." (p. 281)
2) "The listening part from a comprehensive English Language Proficiency Test was used in order to determine the participants' baseline knowledge of their listening skill prior to the exposure to the excerpts of the movies." (p. 280)
3) "The second test administered in this study was the TOEFL listening test as a pre and post-test." (p. 282)
4) "Then, the participants received a training session on taking the TOEFL (Test of English as a Foreign Language) ITP (Institutional TOEFL Program) during the fall semester, 2015." (p. 281)
5) "In the last session, all participants sat on the listening section from the TOEFL test." (p. 275)
Detailed Analysis:
The study used two separate instruments. The 14 multiple-choice comprehension tests tied to the movie excerpts were purpose-built by the researcher for this study and are not standardised. However, the paper describes a pre/post measure initially labelled a "comprehensive English Language Proficiency Test" and later explicitly identified in the Data Analysis section (quote 3) and abstract as the TOEFL ITP listening section, a widely recognised, standardised international exam. Because the criterion asks only whether a standardised exam-based assessment was used somewhere in the study (not whether it drove the headline finding), the use of the TOEFL ITP listening section as an outcome measure satisfies this requirement, notwithstanding the additional use of a custom instrument.
Criterion E is met because the TOEFL ITP listening section, a widely recognised standardised assessment, was administered as a pre- and post-test measure.
-
T
Term Duration
- The whole study, including the intervention and follow-up measurement, spanned only about six weeks, far short of a full academic term.
- "The experiment lasted for six weeks, the participants took the pre-test in the first week and at the end of the experiment, they took the same test they took in the first week as a post-test..." (p. 281)
Relevant Quotes:
1) "The experiment lasted for six weeks, the participants took the pre-test in the first week and at the end of the experiment, they took the same test they took in the first week as a post-test in order to see if viewing movies for four weeks affected the participants' listening comprehension." (p. 281)
2) "Then the three groups viewed the 14 parts of the movies over 4 weeks in three different conditions..." (p. 281)
Detailed Analysis:
The intervention itself (viewing movie excerpts) ran for four weeks, and the full study, including pre-test and post-test administration, spanned approximately six weeks. The ERCT standard requires outcomes to be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Six weeks is well under half of this minimum, and there is no year-long duration that would substitute for this requirement.
Criterion T is not met because the interval from intervention start to final measurement was only about six weeks, far short of one academic term.
-
D
Documented Control Group
- The control (no-subtitles) group's size, exact treatment condition, and statistically confirmed baseline equivalence with the other groups are documented.
- "there is no significant difference among the performance of the three groups on the pre-test [F = 1.100, p = 0.470 (hence >.05)]. This proves that the three groups were homogenous..." (p. 284)
Relevant Quotes:
1) "...and the control group viewed the movie without subtitles (WSG)." (p. 275)
2) "the control group viewed the scenes without subtitles. Each time the participants viewed four to three scenes." (p. 281)
3) "there is no significant difference among the performance of the three groups on the pre-test [F = 1.100, p = 0.470 (hence >.05)]. This proves that the three groups were homogenous and have similar background from the beginning of the study..." (p. 284)
4) "WSG Pair 1 POSTTEST-PRETEST 3.14706 ... 4.109 33 .000" (Table 3, p. 284) [indicating n=34 analysed in the control/WSG condition]
Detailed Analysis:
The paper clearly specifies the control group's exact condition (viewing the same movie excerpts, for the same duration and number of sessions as the treatment groups, but without any subtitles, i.e. "business as usual" viewing), its size (n=34 analysed, drawn from an initial allocation of 40), and provides a statistical comparison (pre-test ANOVA) confirming the control group did not differ significantly from the treatment groups at baseline. While granular demographic breakdowns (age, gender) specific to the control group alone are not separately tabulated, the combination of exact treatment description, group size, and confirmed baseline equivalence constitutes adequate documentation of the control group under this criterion.
Criterion D is met because the control group's size, treatment condition, and baseline comparability are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred among individual students at a single university, not among schools or institutions.
- "One hundred and four participants were randomly assigned to the two experimental groups and the control group..." (p. 280)
Relevant Quotes:
1) "The participants of this study were randomly selected from a larger pool of 180 fourth year students from the Department of English Language, Minoufia University, Egypt." (p. 275)
2) "One hundred and four participants were randomly assigned to the two experimental groups and the control group, each one consists of 40 participants." (p. 280)
Detailed Analysis:
The entire study was conducted within a single department at a single university (Minoufia University), with individual students as the unit of randomisation. There is no mention of multiple schools or institutions being involved or randomised.
Criterion S is not met because the study was a single-site, student-level RCT, not a school-level RCT.
-
I
Independent Conduct
- The same individual researcher designed, selected materials for, delivered, and evaluated the study, with no independent third-party conducting the trial.
- "The researcher selected seven movies that were produced in 2013 and 2014 in order to make sure that they were not shown on TV yet..." (p. 281)
Relevant Quotes:
1) "The researcher selected seven movies that were produced in 2013 and 2014 in order to make sure that they were not shown on TV yet and minimize the possibility that the participants watched them before." (p. 281)
2) "Moreover, the researcher didn't inform the participants of the movies they were supposed to watch so that they couldn't watch them before the sessions." (p. 281)
3) "In order to examine the participants' understanding of the movies, 14 comprehension tests were constructed." (p. 281) [no attribution to an external team]
4) "Noha Ghoneam, a junior staff member at English language and Literature department, Minoufia University, Egypt." (About the Author, p. 288)
Detailed Analysis:
This is a single-authored paper. The researcher herself selected the movie excerpts, constructed the comprehension tests, ran the sessions, and analysed the results. There is no statement anywhere in the paper of an independent or external evaluation team, blinded administrators, or any third-party oversight of data collection or analysis.
Criterion I is not met because the same individual who designed the study also implemented and evaluated it, with no independent conduct.
-
Y
Year Duration
- Since criterion T (Term Duration) is not met, and the total study duration (about six weeks) is far short of 75% of an academic year, criterion Y is not met either.
- "The experiment lasted for six weeks..." (p. 281)
Relevant Quotes:
1) "The experiment lasted for six weeks, the participants took the pre-test in the first week and at the end of the experiment, they took the same test they took in the first week as a post-test..." (p. 281)
Detailed Analysis:
Per the ERCT specification, if the weaker Term Duration criterion (T) is not met, the stronger Year Duration criterion (Y) cannot be met either. Independently, the study's six-week span is a small fraction of the required 75% of an academic year (~9-10 months).
Criterion Y is not met because criterion T is not met and the study duration is far shorter than a year.
-
B
Balanced Control Group
- All three groups received identical exposure time and materials (the same movie excerpts over the same four weeks); only the presence/language of subtitles differed, so no extra resources were introduced that would require balancing.
- "Then the three groups viewed the 14 parts of the movies over 4 weeks in three different conditions: Experimental group 1 viewed the scenes with Arabic subtitles, Experimental group 2 viewed the scenes with English subtitles and the control group viewed the scenes without subtitles." (p. 281)
Relevant Quotes:
1) "Then the three groups viewed the 14 parts of the movies over 4 weeks in three different conditions: Experimental group 1 viewed the scenes with Arabic subtitles, Experimental group 2 viewed the scenes with English subtitles and the control group viewed the scenes without subtitles. Each time the participants viewed four to three scenes." (p. 281)
2) "Immediately after screening the parts of the movies, multiple-choice listening comprehension tests were administered to the students... The same procedure was followed for each group for all the 4 consecutive sessions." (p. 281)
Detailed Analysis:
Applying the Criterion B decision tree: the intervention variable under study is the subtitle condition itself (Arabic, English, or none), applied to otherwise identical movie excerpts, viewing durations, session counts, and comprehension/proficiency tests across all three groups. No group received additional viewing time, materials, teacher support, or budget beyond the shared sessions (EXTRA_RESOURCES_PRESENT is false), so the resource-balance requirement is trivially satisfied without needing to invoke the integral-resource exception.
Criterion B is met because no group received additional time, materials, or budget beyond the shared viewing sessions; the only manipulated variable was the subtitle condition itself.
-
Level 3 Criteria
-
R
Reproduced
- No evidence of independent replication of this specific study by another research team is present in the paper or found via internet search.
Relevant Quotes:
(No quotes found referencing replication of this specific study by an independent team.)
Detailed Analysis:
This is a small, single-author study conducted at one Egyptian university and published in 2015. An internet search was carried out for subsequent papers by other author teams that independently replicate this specific design (three-condition subtitle comparison with Egyptian EFL university students, TOEFL ITP as pre/post measure). The search surfaced only conceptually related but non-replicating subtitling studies by other authors and contexts (e.g. Latifi et al., 2011, cited in the paper itself, and other unrelated subtitling papers indexed on academic repositories), none of which reproduce this specific study's design, population, or results. No independent replication of Ghoneam (2015) was identified.
Criterion R is not met because no independent replication of this specific study was found.
-
A
All-subject Exams
- Only listening comprehension in English was assessed; no other academic subjects were measured, and no specialised-intervention exception is stated.
- "Only the listening part of the test was administered to the learners due to the fact that watching movies basically requires listening abilities." (p. 280-281)
Relevant Quotes:
1) "Only the listening part of the test was administered to the learners due to the fact that watching movies basically requires listening abilities." (p. 280-281)
2) "In order to examine the participants' understanding of the movies, 14 comprehension tests were constructed." (p. 281)
Detailed Analysis:
The study exclusively measures one sub-skill (listening comprehension) within one subject (English as a foreign language). No other core subjects, or even other language sub-skills (reading, writing, speaking), were assessed. While the paper offers a rationale for restricting testing to the listening section (movies "basically require listening abilities"), this justifies narrowing within-instrument scope rather than constituting the kind of specialised vocational/upper- secondary exception described in the standard.
Criterion A is not met because only listening comprehension was assessed, with no coverage of other main subjects or a qualifying specialised-intervention exception.
-
G
Graduation Tracking
- Since criterion Y (Year Duration) is not met, and there is no tracking of participants beyond the six-week study period, criterion G is not met.
Relevant Quotes:
(No quotes describing any follow-up or tracking of participants beyond the six-week study period were found.)
Detailed Analysis:
Per the ERCT specification, since the weaker Year Duration criterion (Y) is not met, the stronger Graduation Tracking criterion (G) cannot be met either. Independently, the paper reports no follow-up of participants after the final post-test in the sixth week, let alone tracking through to graduation. An internet search for subsequent publications by Noha Sobhy Ghoneam tracking the same cohort of participants did not surface any follow-up study; no such papers could be located.
Criterion G is not met because criterion Y is not met and no post-study follow-up is reported or found.
-
P
Pre-Registered
- No statement of a pre-registered protocol, registry platform, or registration date is present anywhere in the paper, and no registry entry was found online.
Relevant Quotes:
(No quotes referencing pre-registration, a registry platform, or a registration date were found anywhere in the paper.)
Detailed Analysis:
The paper contains no methods statement, disclosure, or reference indicating that the study protocol, hypotheses, or analysis plan were registered on any public registry prior to data collection. Given the study's nature (a small classroom-based EFL listening experiment published in 2015, before pre-registration became common practice in applied linguistics/education research), an internet search for a registry entry did not identify any pre-registration record for this study.
Criterion P is not met because no evidence of pre-registration is present in the paper or was found online.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.