Uni-, bi-, and multimodal mobile-assisted listening: Differential effects of app mode on EFL listening comprehension and recognition

Marwa Hafour

Published:
ERCT Check Date:
DOI: 10.64152/10125/73637
  • L2 languages
  • higher education
  • Africa
  • EdTech app
  • mobile learning
0
  • C

    Randomisation was done at the individual student level within one university cohort, not at class or school level, and the tutoring exception does not apply.

    "Subsequently, using systematic random sampling, the eligible participants were randomly assigned into 3 groups practicing mobile-assisted listening in 3 modalities: unimodal (n = 35), bimodal (n = 35), and multimodal (n = 37)." (p. 7)

  • E

    Outcomes were measured with the widely recognised standardised Cambridge B1 PET listening test (two parallel forms), with only six added recognition items as a minor modification.

    "Two parallel forms of modified Cambridge B1 PET listening tests were administered as pre-and posttests (Appendix A)." (p. 7)

  • T

    Outcomes were measured at the end of a 10-week intervention, which is shorter than the required full academic term of roughly 3-4 months.

    "Intervention procedures spanned 10 weeks, during which they practiced modality-specific mobile-assisted listening both in-and out-of-class." (p. 8)

  • D

    The study has no control group at all - all three randomised arms were active treatments - so a documented control group is absent.

    "it is recommended to replicate the study at hand on a larger sample of learners in other language learning contexts at different language levels for longer intervention spans using different research designs (e.g., adding control/ comparison groups)." (p. 21)

  • S

    Individual students within a single university were randomised, so there was no school-level randomisation.

    "Subsequently, using systematic random sampling, the eligible participants were randomly assigned into 3 groups practicing mobile-assisted listening in 3 modalities: unimodal (n = 35), bimodal (n = 35), and multimodal (n = 37)." (p. 7)

  • I

    The single author designed, delivered, monitored, and analysed the intervention herself with no independent evaluation team or third-party oversight.

    "In-class listening practice (which lasted 45 minutes per session) was guided and monitored by the instructor (the researcher)" (p. 8)

  • Y

    The study tracked outcomes for only 10 weeks, far less than 75% of an academic year.

    "Intervention procedures spanned 10 weeks, during which they practiced modality-specific mobile-assisted listening both in-and out-of-class." (p. 8)

  • B

    All three comparison arms received equivalent practice time, apps, and instructor support, with only the app modality (the treatment variable) differing between groups.

    "Following the in-class activities, students were instructed to engage in independent listening practice using the same app(s) for a minimum of 15 minutes on two separate occasions during the same week." (p. 9)

  • R

    The study is presented as a novel first comparison, and a citation search found no independent peer-reviewed replication of this specific design by another team.

    "it is recommended to replicate the study at hand on a larger sample of learners in other language learning contexts at different language levels for longer intervention spans" (p. 21)

  • A

    Only English listening was assessed; no other main subjects or skill areas were measured with standardised exams.

    "Each test comprised 4 parts, 25 items: 19 multiple-choice items assessing listening comprehension and 6 fill-in-the-gap items basically assessing listening recognition." (p. 7)

  • G

    Participants were only tracked to the immediate posttest after 10 weeks, with no follow-up until graduation found in the paper or in a citation search of subsequent publications.

    "Intervention procedures spanned 10 weeks, during which they practiced modality-specific mobile-assisted listening both in-and out-of-class." (p. 8)

  • P

    The paper contains no reference to a pre-registered protocol, registry ID, or registration date.

Abstract

Mobile apps are becoming part and parcel of our daily lives. Hence, this study examined the differential effects of app modes on listening comprehension and recognition. From a pool of Egyptian EFL sophomores, 107 students were randomly assigned into 3 groups practicing mobile-assisted listening in 3 modalities: Unimodal (n = 35), Bimodal (n = 35), and Multimodal (n = 37). Following sequential explanatory mixed-method design, scores on pre-post listening tests and responses to closed/open-ended perceptions survey questions were analyzed using Two-way Mixed ANOVA, Linear Regression, and inductive thematic analysis. While all three modalities demonstrated comparable effectiveness in improving listening comprehension, performance variations were observed in listening recognition. Both the uni- and multimodal groups surpassed the bimodal group in listening recognition, while exhibiting similar levels of listening recognition improvement. As such, apps focusing on comprehension exercises demonstrated greater efficacy in triggering lower-level cognitive processes (listening recognition) than those exclusively targeting word recognition. Further, listening recognition was found to be a weak predictor of listening comprehension. Participants' perceptions consolidated these propositions. They offered some app functionality and user-experience enhancement suggestions. They advocated for incorporating more personalized, gamified, and interactive tools (e.g., shadowing, translation, and speed control tools), coupled with increased content and exercise variety (stepping beyond multiple-choice and fill-in-the-blank formats).

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was done at the individual student level within one university cohort, not at class or school level, and the tutoring exception does not apply.
      • "Subsequently, using systematic random sampling, the eligible participants were randomly assigned into 3 groups practicing mobile-assisted listening in 3 modalities: unimodal (n = 35), bimodal (n = 35), and multimodal (n = 37)." (p. 7)
      • Relevant Quotes: 1) "From a pool of Egyptian EFL sophomores, 107 students were randomly assigned into 3 groups practicing mobile-assisted listening in 3 modalities: Unimodal (n = 35), Bimodal (n = 35), and Multimodal (n = 37)." (p. 1, Abstract) 2) "Subsequently, using systematic random sampling, the eligible participants were randomly assigned into 3 groups practicing mobile-assisted listening in 3 modalities: unimodal (n = 35), bimodal (n = 35), and multimodal (n = 37)." (p. 7) 3) "From a pool of convenient 273 Egyptian sophomore students (studying English as a Foreign Language), 107 students were deemed eligible for the study." (p. 7) 4) "In-class listening practice (which lasted 45 minutes per session) was guided and monitored by the instructor (the researcher), with the exercises being practiced in a whole-class mode wherein the instructor's mobile device was projected onto the class screen and speakers." (p. 8) Detailed Analysis: The unit of randomisation was the individual student: 107 eligible sophomores from a single pool at one university were individually assigned to the three modality groups via systematic random sampling. There is no indication that intact classes or schools were the unit of assignment; on the contrary, the participants were drawn from the same cohort taught by the same instructor. The tutoring/personal-teaching exception does not apply: the intervention is group-based mobile app practice, including whole-class in-class sessions, not one-to-one tutoring. Because peers from the same cohort were placed in different arms while attending the same courses, cross-group contamination (e.g., students trying apps of another modality) cannot be excluded. Criterion C is not met because randomisation was performed at the individual student level within a single cohort, not at the class or school level, and no valid exception applies.
    • E

      Exam-based Assessment

      • Outcomes were measured with the widely recognised standardised Cambridge B1 PET listening test (two parallel forms), with only six added recognition items as a minor modification.
      • "Two parallel forms of modified Cambridge B1 PET listening tests were administered as pre-and posttests (Appendix A)." (p. 7)
      • Relevant Quotes: 1) "Two parallel forms of modified Cambridge B1 PET listening tests were administered as pre-and posttests (Appendix A). Each test comprised 4 parts, 25 items: 19 multiple-choice items assessing listening comprehension and 6 fill-in-the-gap items basically assessing listening recognition." (p. 7) 2) "Six more direct fill-in-the gap items were added to assess listening recognition of the audio content in Part 4 of the test. These added word recognition items, while still being based on the audio content of Part 4 of the original test, were the only modification to the original tests." (p. 7) 3) "Their CEFR (Common European Framework of Reference for Languages) level was B1." (p. 7) Detailed Analysis: The outcome measure was the Cambridge B1 PET (Preliminary English Test) listening test, a widely recognised, internationally standardised examination aligned with the CEFR, and appropriate for the participants' B1 proficiency level. Two parallel official forms were used as pre- and posttests. The only modification was the addition of six fill-in-the-gap word-recognition items based on the audio content of Part 4 of the original test; the listening comprehension outcome (19 multiple-choice items) and the original six recognition items come directly from the standardised instrument. The assessment was therefore not a custom test designed for the study but a recognised standardised exam with a minor, clearly documented extension. Criterion E is met because outcomes were measured with parallel forms of the standardised Cambridge B1 PET listening test, with only a minor documented modification.
    • T

      Term Duration

      • Outcomes were measured at the end of a 10-week intervention, which is shorter than the required full academic term of roughly 3-4 months.
      • "Intervention procedures spanned 10 weeks, during which they practiced modality-specific mobile-assisted listening both in-and out-of-class." (p. 8)
      • Relevant Quotes: 1) "Intervention procedures spanned 10 weeks, during which they practiced modality-specific mobile-assisted listening both in-and out-of-class." (p. 8) 2) "Note. ... a Week 1 was an orientation session." (Table 2, p. 11) 3) "Despite being adequate to yield significant results, the 10-week intervention was not prolonged enough for listening recognition practice to fully automate lower-level cognitive processes of aural recognition." (p. 21) Detailed Analysis: The intervention lasted 10 weeks (including a week-1 orientation), and the posttest was administered at the end of the intervention, per the pre-post design described in the Method and Figure 1. Ten weeks is approximately 2.3-2.5 months, which falls short of the ERCT definition of a term as a semester or equivalent (approximately 3-4 months). No additional delayed follow-up measurement beyond the immediate posttest is reported that would extend the interval from intervention start to outcome measurement to a full term. The author herself flags the 10-week span as a limitation. Criterion T is not met because the interval from intervention start to outcome measurement was only about 10 weeks, shorter than one full academic term (3-4 months).
    • D

      Documented Control Group

      • The study has no control group at all - all three randomised arms were active treatments - so a documented control group is absent.
      • "it is recommended to replicate the study at hand on a larger sample of learners in other language learning contexts at different language levels for longer intervention spans using different research designs (e.g., adding control/ comparison groups)." (p. 21)
      • Relevant Quotes: 1) "the eligible participants were randomly assigned into 3 groups practicing mobile-assisted listening in 3 modalities: unimodal (n = 35), bimodal (n = 35), and multimodal (n = 37)." (p. 7) 2) "The participants were EFL sophomore students aged 18-19. Their CEFR (Common European Framework of Reference for Languages) level was B1." (p. 7) 3) "After assigning the participants to the 3 groups, the researcher verified that the three groups were homogeneous using Levene's test of homogeneity of variances (as described in the Data Analysis section)." (p. 7) 4) "it is recommended to replicate the study at hand on a larger sample of learners in other language learning contexts at different language levels for longer intervention spans using different research designs (e.g., adding control/ comparison groups)." (p. 21) Detailed Analysis: The study has no control group at all: all three arms (unimodal, bimodal, multimodal) received an active mobile-assisted listening treatment, and the author explicitly recommends "adding control/ comparison groups" as future research, confirming that no business-as-usual or untreated control condition existed. Criterion D requires a well-documented control group with demographics, baseline performance, and conditions received; with no control group, such documentation cannot exist. Furthermore, even for the comparison arms, only pooled sample characteristics (age 18-19, B1 level) and a Levene's homogeneity test are reported; no per-group demographic or baseline descriptive table is provided. Criterion D is not met because the study contains no control group (all three arms were active treatments) and therefore no documented control group data for comparison.
  • Level 2 Criteria

    • S

      School-level RCT

      • Individual students within a single university were randomised, so there was no school-level randomisation.
      • "Subsequently, using systematic random sampling, the eligible participants were randomly assigned into 3 groups practicing mobile-assisted listening in 3 modalities: unimodal (n = 35), bimodal (n = 35), and multimodal (n = 37)." (p. 7)
      • Relevant Quotes: 1) "Subsequently, using systematic random sampling, the eligible participants were randomly assigned into 3 groups practicing mobile-assisted listening in 3 modalities: unimodal (n = 35), bimodal (n = 35), and multimodal (n = 37)." (p. 7) 2) "From a pool of convenient 273 Egyptian sophomore students (studying English as a Foreign Language), 107 students were deemed eligible for the study. The author was also the instructor of their undergraduate courses: Technology-enhanced Language Learning and Active Learning Strategies." (p. 7) Detailed Analysis: Criterion S requires randomisation among schools or equivalent institutional units. This study drew all participants from a single institution (one pool of Egyptian sophomores taught by the author) and randomised individual students into the three modality groups. No multiple schools, sites, or centres were involved, and no school-level assignment is described anywhere in the paper. Criterion S is not met because randomisation occurred at the individual student level within a single institution, not among schools or sites.
    • I

      Independent Conduct

      • The single author designed, delivered, monitored, and analysed the intervention herself with no independent evaluation team or third-party oversight.
      • "In-class listening practice (which lasted 45 minutes per session) was guided and monitored by the instructor (the researcher)" (p. 8)
      • Relevant Quotes: 1) "The author was also the instructor of their undergraduate courses: Technology-enhanced Language Learning and Active Learning Strategies." (p. 7) 2) "In-class listening practice (which lasted 45 minutes per session) was guided and monitored by the instructor (the researcher), with the exercises being practiced in a whole-class mode wherein the instructor's mobile device was projected onto the class screen and speakers." (p. 8) 3) "the instructor employed two strategies: (a) requesting screenshots/screen recordings from some students and (b) randomly selecting other students for synchronous instructor-guided screen-sharing sessions to monitor and support their out-of-class practice." (p. 9) 4) "Having made an extensive review of the listening apps, the author found that they fall under 3 categories" (p. 2) Detailed Analysis: This is a single-author study in which the same person designed the intervention framework (classifying the apps and building the modality-specific practice framework), taught the participants as their course instructor, delivered and monitored the intervention, collected the data, and analysed the results. There is no external evaluation team, no third-party data collection, and no statement of independent oversight anywhere in the paper. A peer reviewer assisted only with coding a subset of qualitative data, which does not constitute independent conduct of the trial. Criterion I is not met because the intervention designer, instructor, data collector, and analyst were the same single author with no independent oversight.
    • Y

      Year Duration

      • The study tracked outcomes for only 10 weeks, far less than 75% of an academic year.
      • "Intervention procedures spanned 10 weeks, during which they practiced modality-specific mobile-assisted listening both in-and out-of-class." (p. 8)
      • Relevant Quotes: 1) "Intervention procedures spanned 10 weeks, during which they practiced modality-specific mobile-assisted listening both in-and out-of-class." (p. 8) 2) "Despite being adequate to yield significant results, the 10-week intervention was not prolonged enough for listening recognition practice to fully automate lower-level cognitive processes of aural recognition." (p. 21) Detailed Analysis: Criterion Y requires outcome measurement at least 75% of an academic year (roughly 7-9 months) after intervention start. Here the entire study, from intervention start to posttest, spanned only 10 weeks (about 2.3-2.5 months), with no delayed follow-up. This is far below 75% of any academic year definition. In addition, per the prompt rule, criterion T (Term Duration) is not met, which automatically means criterion Y cannot be met. Criterion Y is not met because tracking lasted only about 10 weeks, far short of 75% of an academic year, and criterion T is also unmet.
    • B

      Balanced Control Group

      • All three comparison arms received equivalent practice time, apps, and instructor support, with only the app modality (the treatment variable) differing between groups.
      • "Following the in-class activities, students were instructed to engage in independent listening practice using the same app(s) for a minimum of 15 minutes on two separate occasions during the same week." (p. 9)
      • Relevant Quotes: 1) "After assigning the participants to their modality-specific group, the intervention sessions commenced with an in-class mobile-based orientation session familiarizing the participants in each group with their designated modality-specific listening practice apps." (p. 8) 2) "In-class listening practice (which lasted 45 minutes per session) was guided and monitored by the instructor (the researcher)" (p. 8) 3) "Following the in-class activities, students were instructed to engage in independent listening practice using the same app(s) for a minimum of 15 minutes on two separate occasions during the same week." (p. 9) 4) "Table 2 provides a detailed account of the intervention framework of mobile-based, modality-specific listening practice (the frequency and type of app-based practice per session/week)." (p. 9) 5) "b Some apps were in more than one mode; participants were guided to practice listening in their assigned mode only." (Table 2 note, p. 11) Detailed Analysis: This is a three-arm active comparison design with no untreated arm. All three groups received the same intervention framework: a week-1 orientation, weekly 45-minute in-class guided sessions, and a minimum of two 15-minute independent out-of-class practice sessions per week over the same 10-week period, often even using the same apps restricted to the assigned mode. The only manipulated difference between groups was the app modality (comprehension-only, recognition-only, or combined exercises), which is precisely the treatment variable under study. Time on task, materials (free mobile apps), instructor guidance, and monitoring were comparable across all arms, so no arm received extra educational time or budget beyond the others. Applying the decision procedure: no arm received additional time/budget relative to the others (EXTRA_RESOURCES_PRESENT is effectively false across the three compared arms), so the criterion is satisfied on that basis alone; even if the modality-specific apps were considered an "extra resource" relative to a hypothetical business-as-usual baseline, that resource is the explicit treatment variable being tested (RESOURCES_ARE_TREATMENT), which also satisfies the criterion. Criterion B is met because all three randomised arms received equivalent time, apps, and instructor support, with the app modality itself being the only difference and the explicit treatment variable.
  • Level 3 Criteria

    • R

      Reproduced

      • The study is presented as a novel first comparison, and a citation search found no independent peer-reviewed replication of this specific design by another team.
      • "it is recommended to replicate the study at hand on a larger sample of learners in other language learning contexts at different language levels for longer intervention spans" (p. 21)
      • Relevant Quotes: 1) "To the researcher's best knowledge, previous studies have not been able to comprehensively address the topic of comparing the different modalities and apps, nor their impact on all listening facets." (p. 2) 2) "it is recommended to replicate the study at hand on a larger sample of learners in other language learning contexts at different language levels for longer intervention spans using different research designs (e.g., adding control/ comparison groups)." (p. 21) Internet Search for Independent Replication: A citation search (Google Scholar) for this article shows 2 citing works. (a) Hafour, Aladini, & Alsabbagh (2026), "EFL Learners' Agency, Perceptions, and Macro/Micro-Listening Processes in Mobile vs. Adaptive AI-Driven Aural Experiences," International Journal of Computer-Assisted Language Learning and Teaching, 16(1), 1-28 - authored by the same lead researcher, so this cannot count as an independent replication even though its abstract snippet indicates it "compares learners' agency, perceptions, and macro/micro-listening skills in different aural training modalities: mobile-assisted language learning and intelligent mobile-..." (i.e. a different comparison: MALL vs. adaptive AI, not a repeat of the uni-/bi-/multimodal design). (b) Indah, Masruddin, & Furwana (2026), "Mobile-Assisted Listening Learning: The Effect of the Learn English Listening Application on Indonesian EFL Students," FOSTER: Journal of English Language Teaching, 7(2), 425-438. Its abstract states: "A quantitative approach with a pre-experimental one-group pre-test and post-test design was employed. The participants consisted of 20 students selected through purposive sampling." This is an unrelated, non-RCT, single-group study of a different app by a different team; it cites Hafour's paper only in passing and does not replicate this study's design, sample, or context. No independent, peer-reviewed replication of this specific study was found in any available source. Detailed Analysis: The paper positions itself as the first study to compare uni-, bi-, and multimodal mobile-assisted listening, explicitly noting that prior research has not addressed this comparison, and it calls for future replication. Published in 2025, no independent replication of this specific three-modality comparison by a different research team in a different context is referenced in the paper, and related earlier MALL listening studies (e.g., Jia & Hew, 2022; Tai & Chen, 2024) test different interventions and designs, so they do not constitute replications of this trial. The citation search confirms that, as of this check, no independent team has replicated this study's specific design. Criterion R is not met because no independent peer-reviewed replication of this specific study exists; the paper itself describes the comparison as novel and calls for replication, and a citation search found no such replication.
    • A

      All-subject Exams

      • Only English listening was assessed; no other main subjects or skill areas were measured with standardised exams.
      • "Each test comprised 4 parts, 25 items: 19 multiple-choice items assessing listening comprehension and 6 fill-in-the-gap items basically assessing listening recognition." (p. 7)
      • Relevant Quotes: 1) "Each test comprised 4 parts, 25 items: 19 multiple-choice items assessing listening comprehension and 6 fill-in-the-gap items basically assessing listening recognition." (p. 7) 2) "While enrolled in academic courses such as Reading, Writing, Conversation, and Listening, and Phonetics, these students faced a significant practical limitation: a lack of real consistent listening practice throughout their studies." (p. 7) Detailed Analysis: Criterion A requires standardised assessment across all main subjects taught at the educational level. This study measured only English listening (comprehension and recognition); no other subjects or even other English skills (reading, writing, speaking) were assessed with standardised exams. The participants were university EFL students enrolled in multiple language courses (Reading, Writing, Conversation, Listening, Phonetics), yet outcomes in those areas were not measured. The paper offers no explicit rationale invoking the specialised-intervention exception, and even within the EFL programme only one skill area was tested. Criterion A is not met because only listening outcomes were assessed, without standardised measurement of the other main subjects or skills in the participants' programme.
    • G

      Graduation Tracking

      • Participants were only tracked to the immediate posttest after 10 weeks, with no follow-up until graduation found in the paper or in a citation search of subsequent publications.
      • "Intervention procedures spanned 10 weeks, during which they practiced modality-specific mobile-assisted listening both in-and out-of-class." (p. 8)
      • Relevant Quotes: 1) "Intervention procedures spanned 10 weeks, during which they practiced modality-specific mobile-assisted listening both in-and out-of-class." (p. 8) 2) "The limitations of the current study serve as implicit suggestions for future research. Accordingly, it is recommended to replicate the study at hand on a larger sample of learners in other language learning contexts at different language levels for longer intervention spans using different research designs" (p. 21) Internet Search for Follow-up/Graduation-Tracking Papers: A citation search (Google Scholar) for this article identified 2 citing works, including a 2026 paper by the same lead author: Hafour, Aladini, & Alsabbagh (2026), "EFL Learners' Agency, Perceptions, and Macro/Micro-Listening Processes in Mobile vs. Adaptive AI-Driven Aural Experiences," International Journal of Computer-Assisted Language Learning and Teaching, 16(1), 1-28. Its available abstract snippet indicates it "compares learners' agency, perceptions, and macro/micro- listening skills in different aural training modalities: mobile-assisted language learning and intelligent mobile-...". This is a distinct comparative study (MALL vs. adaptive AI-driven listening), not a description of continued tracking of the original 107-student cohort toward graduation, and no indication was found that it follows the same participants. No paper reporting graduation-tracking of the original cohort was found in any available source. Detailed Analysis: Measurement ended with the posttest at the close of the 10-week intervention while participants were sophomores. There is no follow-up of the cohort through to graduation from their degree programme, no mention of planned follow-up studies tracking these participants, and the author's own call for "longer intervention spans" confirms tracking was short. Additionally, per the prompt rule, criterion Y is not met, which automatically means criterion G is not met. Criterion G is not met because tracking stopped at the immediate posttest after 10 weeks, criterion Y is unmet, and no follow-up publication tracking the same cohort to graduation was found via citation search.
    • P

      Pre-Registered

      • The paper contains no reference to a pre-registered protocol, registry ID, or registration date.
      • Relevant Quotes: 1) "This study was conducted in line with the mixed-method design, more specifically the sequential explanatory design. As such, the quantitative data collection and analysis were followed by a qualitative phase, as illustrated in Figure 1." (p. 7) 2) "An online survey (Appendix B) was prepared to assess the perceptions of the participants in the multimodal group." (p. 7) Detailed Analysis: The paper contains no mention of any pre-registration of the study protocol, hypotheses, or analysis plan on any registry platform (e.g., ClinicalTrials.gov, OSF, AsPredicted, ISRCTN). The Method section describes the design, instruments, and analyses, but provides no registry identifier, no registration date, and no reference to a published protocol. Since the paper itself gives no registry name or ID to verify, no registry lookup could be performed; with no evidence of registration before data collection, the criterion fails. Criterion P is not met because no pre-registration statement, registry link, or registration date appears anywhere in the paper.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.