The impact of mobile-assisted project-based learning on developing EFL students' speaking skills

Hassane Benlaghrissi, L. Meriem Ouahidi

Published:
ERCT Check Date:
DOI: 10.1186/s40561-024-00303-y
  • L2 languages
  • K12
  • Africa
  • project-based learning
  • EdTech app
  • mobile learning
0
  • C

    Randomisation was at the individual student level within a single school for group-based classroom instruction, not at the class or school level, and the tutoring exception does not apply.

    "Based on the test's results, 91 students constituted the study participants who were randomly assigned to three groups: the experimental group (N = 31; F = 20, M = 11), the first control group (N = 29; F = 16, M = 13), and the second control group (N = 31; F = 18, M = 13)." (pp. 5-6)

  • E

    Outcomes were measured with a speaking test designed by the researchers based on students' textbooks, scored with a locally modified IELTS rubric, which is a custom instrument rather than a widely recognised standardised exam.

    "A speaking test was used as a pre- and post-test to evaluate students' speaking scores. The researchers designed the test based on students' textbooks and graded it out of 20." (p. 10)

  • T

    The intervention began in February and outcomes were measured in the last week of May of the same semester, an interval of roughly 3.5 months, which covers at least one full academic term.

    "The study was carried out for one semester in the academic year 2021-2022. It began immediately in the second semester, in the second week of February, with the pre-test, and continued to the last week of May with the post-test." (p. 6)

  • D

    Both control groups are documented in detail, including sizes, gender and age composition, baseline speaking scores, and precise descriptions of the instruction each control condition received.

    "In contrast, the first control group was taught speaking through project-based learning method (PBL), and the second control group was taught speaking conventionally using the ECRIF model (Encounter, Clarify, Remember and Internalise, and fluently use)." (p. 6)

  • S

    The study randomised individual students within a single school, so no school-level (or institution-level) randomisation took place.

    "All the students belonged to the same school and had almost the same English level." (p. 6)

  • I

    The authors designed the intervention, taught the classes as the teacher-researcher, designed and administered the tests, and analysed the data themselves, with no independent third-party evaluation.

    "The interviews took place in the teacher-researcher room where he usually teaches his classes... The interviews were recorded using the teacher's mobile phone voice recorder." (pp. 10-11)

  • Y

    The study lasted one semester (mid-February to late May, about 3.5 months), which is well short of 75% of an academic year.

    "The study was carried out for one semester in the academic year 2021-2022. It began immediately in the second semester, in the second week of February, with the pre-test, and continued to the last week of May with the post-test." (p. 6)

  • B

    All three groups received active instruction on the same units with comparable class time; the PBL control did identical projects without phones, and the mobile phone use that differentiates the experimental group is the integral treatment variable being tested.

    "Therefore, students in the PBL group received the same project stages, learning objectives, and project tasks. However, the only difference between this group and the experimental group was the use of technology through mobile phones." (pp. 8-9)

  • R

    Neither the paper nor an internet search identified any independent, peer-reviewed replication of this specific mobile-assisted project-based learning trial by a different research team.

  • A

    Only English speaking skills were assessed with a custom test, so no standardised all-subject assessment took place, and the prerequisite criterion E is also unmet.

    "Two instruments were employed to collect data: a speaking pre- and post-test to evaluate the three groups' oral proficiency and a 5-Likert scale survey..." (p. 1, Abstract)

  • G

    Measurement ended with the post-test in May at the close of the semester, with no tracking of participants to graduation and no follow-up publications on this cohort; prerequisite criterion Y is also unmet.

    "The study was carried out for one semester in the academic year 2021-2022... and continued to the last week of May with the post-test." (p. 6)

  • P

    The paper contains no mention of any pre-registration or registry, and an external search found no registration record for this trial.

Abstract

Combining mobile-assisted language learning (MALL) with project-based learning (PBL) might be the potential framework for enhancing EFL learners' speaking skills. However, only a few studies have scrutinised the impact of modern technologies on project work. More importantly, investigating how MALL, as a new field within ICT with unique pedagogical affordances, and PBL can enhance learners' speaking skills is still lacking in the literature. Accordingly, this study examines how integrating MALL through mobile phones and PBL, defined as mobile-assisted project-based learning or mobile-assisted projects, improves Moroccan secondary school students' speaking performance. A true experimental study was conducted with 91 students assigned randomly to one experimental group and two control groups. The experimental group received instruction through mobile-assisted projects over one semester. In contrast, participants in the first control group taught speaking through project-based learning, and participants in the second control group received traditional teaching. Two instruments were employed to collect data: a speaking pre- and post-test to evaluate the three groups' oral proficiency and a 5-Likert scale survey to detect the experimental group participants' experience and attitudes toward the implementation. Based on independent sample t tests and paired sample t tests (SPSS-26), it was found that instruction through mobile-assisted projects was considerably more effective than project-based learning and conventional teaching in enhancing learners' overall speaking performance and sub-skills: fluency and coherence, lexical resource, grammatical range and accuracy, and pronunciation. Further, the results of the attitude post-questionnaire demonstrated a very high positive perception of the participants toward the implementation. As a result, these findings confirm the pedagogical role of combining MALL with PBL as an innovative mode of instruction in enhancing EFL learners' speaking performance.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was at the individual student level within a single school for group-based classroom instruction, not at the class or school level, and the tutoring exception does not apply.
      • "Based on the test's results, 91 students constituted the study participants who were randomly assigned to three groups: the experimental group (N = 31; F = 20, M = 11), the first control group (N = 29; F = 16, M = 13), and the second control group (N = 31; F = 18, M = 13)." (pp. 5-6)
      • Relevant Quotes: 1) "A true experimental study was conducted with 91 students assigned randomly to one experimental group and two control groups." (p. 1, Abstract) 2) "Based on the test's results, 91 students constituted the study participants who were randomly assigned to three groups: the experimental group (N = 31; F = 20, M = 11), the first control group (N = 29; F = 16, M = 13), and the second control group (N = 31; F = 18, M = 13)." (pp. 5-6) 3) "All the students belonged to the same school and had almost the same English level." (p. 6) 4) "The current study adopted a pretest-posttest equivalent group design experiment with one independent and dependent variable." (p. 6) Detailed Analysis: Criterion C requires randomisation of entire classes (or stronger, entire schools) to prevent contamination between treatment and control participants. The paper explicitly states that 91 individual students from the same school were randomly assigned to the three groups. The unit of randomisation is therefore the individual student, not the class. The intervention is group-based classroom instruction (projects carried out in class groups over a semester), not personal one-to-one tutoring, so the tutoring exception that would allow student-level randomisation does not apply. Because all participants attended the same school and the experimental group visibly used mobile phones for projects, contamination between groups is plausible and exactly the risk this criterion is designed to prevent. Criterion C is not met because students, not classes or schools, were the unit of randomisation and no valid exception applies.
    • E

      Exam-based Assessment

      • Outcomes were measured with a speaking test designed by the researchers based on students' textbooks, scored with a locally modified IELTS rubric, which is a custom instrument rather than a widely recognised standardised exam.
      • "A speaking test was used as a pre- and post-test to evaluate students' speaking scores. The researchers designed the test based on students' textbooks and graded it out of 20." (p. 10)
      • Relevant Quotes: 1) "A speaking test was used as a pre- and post-test to evaluate students' speaking scores. The researchers designed the test based on students' textbooks and graded it out of 20." (p. 10) 2) "The scoring rubric was based on the International English Language Testing System (IELTS). However, the band descriptors were modified with the help of two English language supervisors to meet the target context." (p. 12) 3) "For validity reasons, the speaking test was given to eight field experts: four university professors of applied linguistics, two English language supervisors, and two experienced secondary school teachers of English." (p. 13) 4) "Two raters rated students' recordings to prove the speaking test's reliability." (p. 14) Detailed Analysis: Criterion E requires that outcomes be measured with a standard, widely recognised standardised exam rather than an instrument specially created for the study. Here the speaking pre-/post-test was explicitly "designed" by the researchers themselves from the students' textbooks. Only the scoring rubric borrowed from IELTS band descriptors, and even those descriptors were modified for the local context. The interview questions themselves (Appendix 1) are custom. Ad-hoc expert validation (CVI = 0.885) and inter-rater reliability checks do not convert a researcher-made instrument into a recognised standardised exam. No national or international standardised test was administered. Criterion E is not met because the outcome measure was a custom, researcher-designed speaking test rather than a recognised standardised exam.
    • T

      Term Duration

      • The intervention began in February and outcomes were measured in the last week of May of the same semester, an interval of roughly 3.5 months, which covers at least one full academic term.
      • "The study was carried out for one semester in the academic year 2021-2022. It began immediately in the second semester, in the second week of February, with the pre-test, and continued to the last week of May with the post-test." (p. 6)
      • Relevant Quotes: 1) "The study was carried out for one semester in the academic year 2021-2022. It began immediately in the second semester, in the second week of February, with the pre-test, and continued to the last week of May with the post-test." (p. 6) 2) "Beginning in the third week of February, the three groups participated in the study using different instructional methods." (p. 6) 3) "During one semester of experimental treatment, mobile phone-assisted project participants were required to prepare five projects (1 for each two weeks followed by a week off)." (p. 8) 4) "The experimental group received instruction through mobile-assisted projects over one semester." (p. 1, Abstract) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. The intervention started in the third week of February 2022 and the post-test was administered in the last week of May 2022 (week 17 of the treatment schedule in Table 2). The interval from intervention start to outcome measurement is therefore roughly 3 to 3.5 months, and the authors explicitly frame the study as lasting "one semester", which corresponds to the standard's definition of a term. The measurement occurred at the end of this full semester of treatment, satisfying the term-long tracking requirement. Criterion T is met because outcomes were measured one full semester (about 3.5 months) after the intervention began.
    • D

      Documented Control Group

      • Both control groups are documented in detail, including sizes, gender and age composition, baseline speaking scores, and precise descriptions of the instruction each control condition received.
      • "In contrast, the first control group was taught speaking through project-based learning method (PBL), and the second control group was taught speaking conventionally using the ECRIF model (Encounter, Clarify, Remember and Internalise, and fluently use)." (p. 6)
      • Relevant Quotes: 1) "Based on the test's results, 91 students constituted the study participants who were randomly assigned to three groups: the experimental group (N = 31; F = 20, M = 11), the first control group (N = 29; F = 16, M = 13), and the second control group (N = 31; F = 18, M = 13)." (pp. 5-6) 2) "All the students belonged to the same school and had almost the same English level. They were all studying English as a second foreign language... The participants' ages ranged from 14 to 18 years. The study participants are presented in Table 1." (p. 6) 3) "In contrast, the first control group was taught speaking through project-based learning method (PBL), and the second control group was taught speaking conventionally using the ECRIF model (Encounter, Clarify, Remember and Internalise, and fluently use)." (p. 6) 4) "It is worth noting that the experimental group used their mobile phones while the two control groups received instruction without using their mobile phones or any other mobile device." (p. 6) 5) "For the PBL group, the mean score was 8.52, with a standard deviation of 3.007. On the other hand, the ECRIF group's mean score (N = 31) was 8.94, and its standard deviation was 3.214." (p. 16, Table 10) Detailed Analysis: Criterion D requires clear documentation of the control group's composition, baseline performance, and treatment conditions. The paper documents both control groups: Table 1 gives group sizes, gender split, and age distribution; Table 10 and Tables 13-14 give baseline (pre-test) speaking scores overall and by sub-skill, with homogeneity confirmed by placement test, normality tests, and Levene's tests; and the Methodology section describes exactly what each control group received (paper-based projects with identical stages and tasks for the PBL group, Table 4; conventional ECRIF speaking lessons for the second group, Table 5), explicitly noting they used no mobile devices. Criterion D is met because the control groups' demographics, baseline scores, and instructional conditions are documented in detail.
  • Level 2 Criteria

    • S

      School-level RCT

      • The study randomised individual students within a single school, so no school-level (or institution-level) randomisation took place.
      • "All the students belonged to the same school and had almost the same English level." (p. 6)
      • Relevant Quotes: 1) "Based on the test's results, 91 students constituted the study participants who were randomly assigned to three groups..." (pp. 5-6) 2) "All the students belonged to the same school and had almost the same English level." (p. 6) 3) "The study participants were 10th-grade Moroccan EFL public secondary school students." (p. 5) Detailed Analysis: Criterion S requires randomisation among schools or equivalent implementing institutions. This study took place in a single Moroccan public secondary school, and the 91 participating students were individually randomised into the three conditions. No multiple schools, sites, or institutional units were involved or randomised. Criterion S is not met because randomisation occurred at the individual student level within one school, not among schools.
    • I

      Independent Conduct

      • The authors designed the intervention, taught the classes as the teacher-researcher, designed and administered the tests, and analysed the data themselves, with no independent third-party evaluation.
      • "The interviews took place in the teacher-researcher room where he usually teaches his classes... The interviews were recorded using the teacher's mobile phone voice recorder." (pp. 10-11)
      • Relevant Quotes: 1) "The researchers designed the test based on students' textbooks and graded it out of 20. It was in the form of interview questions between the teacher-researcher and students..." (p. 10) 2) "The interviews took place in the teacher-researcher room where he usually teaches his classes. Each interview lasted 3 to 8 min, depending on each student's speaking ability. The interviews were recorded using the teacher's mobile phone voice recorder." (pp. 10-11) 3) "In designing the speaking lesson plans, the researcher used one of the current frameworks that support oral proficiency: the ECRIF framework..." (p. 9) 4) "The questionnaire was designed using Sphinx software and administered by the researchers at the end of the academic year using the face-to-face method." (p. 14) 5) "This study is part of the corresponding author's PhD thesis. Both authors have contributed, read, and approved the final version of the article submitted for publication." (p. 28, Author contributions) Detailed Analysis: Criterion I requires the study to be conducted independently from those who designed the intervention, or at minimum documented third-party oversight of data collection and analysis. Here the first author was simultaneously the designer of the mobile-assisted project model, the classroom teacher delivering all three conditions, the developer and administrator of the speaking pre-/post-test, and the analyst. Outcome interviews were conducted and recorded by the teacher-researcher himself, who was necessarily unblinded to group assignment. Two external raters scored recordings and outside experts validated instruments, but rating and validation assistance is not the independent conduct or external evaluation the criterion requires; there is no statement of any independent evaluation team. Criterion I is not met because the same authors designed, delivered, assessed, and analysed the intervention without independent oversight.
    • Y

      Year Duration

      • The study lasted one semester (mid-February to late May, about 3.5 months), which is well short of 75% of an academic year.
      • "The study was carried out for one semester in the academic year 2021-2022. It began immediately in the second semester, in the second week of February, with the pre-test, and continued to the last week of May with the post-test." (p. 6)
      • Relevant Quotes: 1) "The study was carried out for one semester in the academic year 2021-2022. It began immediately in the second semester, in the second week of February, with the pre-test, and continued to the last week of May with the post-test." (p. 6) 2) "During one semester of experimental treatment, mobile phone-assisted project participants were required to prepare five projects (1 for each two weeks followed by a week off)." (p. 8) 3) "After one semester of implementing mobile-assisted projects, the three groups set for the post-test to assess their performance..." (p. 18) Detailed Analysis: Criterion Y requires outcomes to be measured at least 75% of a full academic year (roughly 9-10 months, so at least about 7 months) after the intervention begins. The tracking interval here runs from mid-February to the last week of May 2022, i.e., approximately 3.5 months - a single semester. This is far below the 75%-of-a-year threshold, and no longer-term follow-up measurement is reported. Criterion Y is not met because the interval from intervention start to final measurement was only about one semester (roughly 3.5 months).
    • B

      Balanced Control Group

      • All three groups received active instruction on the same units with comparable class time; the PBL control did identical projects without phones, and the mobile phone use that differentiates the experimental group is the integral treatment variable being tested.
      • "Therefore, students in the PBL group received the same project stages, learning objectives, and project tasks. However, the only difference between this group and the experimental group was the use of technology through mobile phones." (pp. 8-9)
      • Relevant Quotes: 1) "Similarly, students in the first control group (PBL) were required to carry out five paper-based projects relating to the same units without using their mobile phones or any technological device." (p. 8) 2) "Therefore, students in the PBL group received the same project stages, learning objectives, and project tasks. However, the only difference between this group and the experimental group was the use of technology through mobile phones." (pp. 8-9) 3) "The second control group was conventionally taught the same speaking content. In each unit, two sessions were devoted to speaking." (p. 9) 4) "The independent variable (X1) was the integration of MALL and PBL, defined as mobile-assisted projects." (p. 6) 5) "Before the intervention, the researcher ensured that all the experimental group participants had mobile phones." (p. 6) 6) "Accordingly, this study examines how integrating MALL through mobile phones and PBL, defined as mobile-assisted project-based learning or mobile-assisted projects, improves Moroccan secondary school students' speaking performance." (p. 1, Abstract) Detailed Analysis: Criterion B requires comparing the time, budget, and materials given to intervention and control conditions. Following the decision tree: extra resources are present in the experimental condition - students used their own mobile phones and apps (WhatsApp, YouTube, voice recorders, a smart TV connection), including some out-of-class practice "anytime and anywhere". However, both control conditions were active: the PBL group carried out the same five projects with identical stages, objectives, and tasks on the same biweekly schedule, only paper-based; the ECRIF group received the same speaking content over two sessions per unit. In-class instructional time was therefore comparable across groups. The differentiating resource - mobile phone technology - is explicitly the independent variable being tested ("The independent variable (X1) was the integration of MALL and PBL"), i.e., it is integral to the intervention rather than a separable, confounding add-on, and students used their own existing phones so no meaningful budget difference arose. Under the standard's exception (resources are the treatment variable, matched active controls), the design is acceptably balanced. Note clearly: the additional resource in the experimental group is mobile phone/app access and associated out-of-class use, and this is the defined treatment. Criterion B is met because the controls were active with comparable instructional time and the mobile technology difference is the integral treatment variable being tested.
  • Level 3 Criteria

    • R

      Reproduced

      • Neither the paper nor an internet search identified any independent, peer-reviewed replication of this specific mobile-assisted project-based learning trial by a different research team.
      • Relevant Quotes: 1) "However, only a few studies have scrutinised the impact of modern technologies on project work. More importantly, investigating how MALL, as a new field within ICT with unique pedagogical affordances, and PBL can enhance learners' speaking skills is still lacking in the literature." (p. 1, Abstract) 2) "Thus, future research should include prominent participants in diverse settings with different age groups and speaking proficiency levels." (p. 25, Limitations) Detailed Analysis: Criterion R requires independent replication of the study by a different research team in a different context, published in a peer-reviewed journal. The paper itself positions the study as novel, filling a gap in the literature, and calls for future research to confirm its findings - implying no replication existed at publication. An internet search (July 2026) for replications of this mobile-assisted project-based learning speaking trial found the original article on Springer, ERIC, ResearchGate, and Studocu, and related MALL/PBL studies by other teams (e.g. a 2025 paper by Lyu and Bidin, and other project-based-learning speaking studies), but these merely cite the original study as background literature; no published study by an independent team replicates this specific intervention (mobile-assisted projects vs. PBL vs. ECRIF control, same design, Moroccan secondary EFL context) was found. Related studies on MALL or PBL separately do not constitute replication of this particular mobile-assisted projects trial. Criterion R is not met because no independent, peer-reviewed replication of this specific study was found in the paper or through external search.
    • A

      All-subject Exams

      • Only English speaking skills were assessed with a custom test, so no standardised all-subject assessment took place, and the prerequisite criterion E is also unmet.
      • "Two instruments were employed to collect data: a speaking pre- and post-test to evaluate the three groups' oral proficiency and a 5-Likert scale survey..." (p. 1, Abstract)
      • Relevant Quotes: 1) "Two instruments were employed to collect data: a speaking pre- and post-test to evaluate the three groups' oral proficiency and a 5-Likert scale survey to detect the experimental group participants' experience and attitudes toward the implementation." (p. 1, Abstract) 2) "First, a limited number of participants targeted only one English skill." (p. 25, Limitations) 3) "Researchers should also investigate other English skills, such as writing, to confirm the present study's findings with other skills' findings." (p. 25) Detailed Analysis: Criterion A requires standardised exam-based assessment of all main subjects taught at the educational level, with criterion E as a prerequisite. Criterion E is not met (the speaking test is custom), so criterion A automatically fails. Substantively, the study measured only one sub-domain (speaking) of one subject (English as a foreign language); no other school subjects (mathematics, science, Arabic, etc.) were assessed, and the authors themselves acknowledge targeting "only one English skill". This is not a specialised vocational context that would justify the exception. Criterion A is not met because only EFL speaking was assessed, with a non-standardised instrument, and criterion E is unmet.
    • G

      Graduation Tracking

      • Measurement ended with the post-test in May at the close of the semester, with no tracking of participants to graduation and no follow-up publications on this cohort; prerequisite criterion Y is also unmet.
      • "The study was carried out for one semester in the academic year 2021-2022... and continued to the last week of May with the post-test." (p. 6)
      • Relevant Quotes: 1) "The study was carried out for one semester in the academic year 2021-2022. It began immediately in the second semester, in the second week of February, with the pre-test, and continued to the last week of May with the post-test." (p. 6) 2) "Week 17: Took the speaking post-test. Responded to the attitude post-questionnaire." (Table 2, p. 7) 3) "Thus, future research should include prominent participants in diverse settings with different age groups and speaking proficiency levels." (p. 25, Limitations) Detailed Analysis: Criterion G requires participants to be followed until graduation from their educational stage. Per the prompt rules, criterion G cannot be met when criterion Y is not met, and Y fails here. The participants were 10th-grade students; data collection ended with the post-test and attitude questionnaire in week 17 (last week of May 2022), long before their secondary school graduation. No follow-up tracking is described in this paper. An internet search (July 2026) for subsequent papers by Benlaghrissi and Ouahidi tracking this same cohort of 10th-grade students was conducted; it found two other papers by the same authors (a 2023 vocabulary/Flashcard World study and a 2024 WhatsApp-based writing study), but both examine different interventions and outcome skills rather than following up this speaking-intervention cohort toward graduation. No graduation-tracking follow-up study on this specific cohort was found. Criterion G is not met because tracking stopped at the end-of-semester post-test, no follow-up publication tracking this cohort to graduation was found, and the prerequisite Y criterion is unmet.
    • P

      Pre-Registered

      • The paper contains no mention of any pre-registration or registry, and an external search found no registration record for this trial.
      • Relevant Quotes: 1) "The data collected and analysed for this research is available in SPSS format. It will be released for private use upon request." (p. 28, Availability of data and materials) 2) "No funding was received to assist with the preparation of this manuscript." (p. 28, Funding) Detailed Analysis: Criterion P requires the full study protocol to be registered on a public registry before data collection began. The paper contains no reference to any registry (e.g., ClinicalTrials.gov, OSF, AEA, ISRCTN), no registration ID, and no pre-registration statement in the methods, declarations, or data availability sections. An internet search (July 2026) for a registration record associated with this trial, its authors, or its DOI likewise returned no pre-registration record on any registry. With no evidence of registration, let alone registration prior to the February 2022 data collection, the criterion fails. Criterion P is not met because no pre-registration of the study protocol is mentioned in the paper or discoverable externally.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.