An Experimental Study of the Effectiveness of Role-play in Improving Fluency in Jordanian EFL Students' Speaking Skills

Luqman M. Rababah

Published:
ERCT Check Date:
DOI: 10.5430/wjel.v15n4p30
  • L2 languages
  • higher education
  • Asia
0
  • C

    Randomization was done at the individual student level within a classroom-based (not one-to-one tutoring) intervention, so the class-level requirement is not satisfied.

    "The research randomly allocated 50 Jordanian EFL students to experimental or control groups." (p. 33)

  • E

    Outcomes were measured with a speaking test and questionnaire constructed for this study, not a recognised standardised exam.

    "The research employed a speaking test and a questionnaire. Speaking fluency was tested before and after the intervention." (p. 33)

  • T

    The intervention and follow-up lasted only six weeks, with post-tests immediately after, which is far shorter than one academic term.

    "Fifty intermediate EFL students were studied for six weeks." (Abstract, p. 30)

  • D

    The control group is described only by size and mean test scores, with no demographic or baseline characteristics, and the paper explicitly labels its own descriptive and inferential results tables as fictional and hypothetical.

    "This table is fictional and based on the outcomes data." (p. 33)

  • S

    Randomisation occurred at the individual student level within a single university, with no schools or institutions being randomised.

    "The research randomly allocated 50 Jordanian EFL students to experimental or control groups." (p. 33)

  • I

    A single author designed, delivered, and evaluated the intervention with no external or third-party evaluation team.

    "Dr. Luqman Rababah was responsible for the entire study." (p. 36)

  • Y

    The whole study lasted six weeks, which is far below 75% of an academic year; since criterion T fails, Y fails as well.

    "Fifty intermediate EFL students were studied for six weeks." (Abstract, p. 30)

  • B

    Both groups received classroom instruction and speaking practice over the same six-week window with no documented extra time or budget for the experimental group; the role-play activity is the treatment variable itself rather than a separable added resource.

    "Both groups got classroom instruction and speaking practice; however, the experimental group also participated in role-plays." (Abstract, p. 30)

  • R

    No independent, peer-reviewed replication of this specific six-week Jordanian role-play trial was found in the paper or via external search; the study was published only in February 2025, leaving little time for replication.

    "However, to my understanding, empirical evidence is scarce to examine the effects of role-play on NSr/Flu in the Jordanian EFL environment." (p. 33)

  • A

    Only speaking fluency in English was assessed with a custom test; no other core subjects were measured and criterion E is not met, which also fails A.

    "Pre- and post-tests assessed participants' speaking fluency." (Abstract, p. 30)

  • G

    Measurement stopped at the six-week post-test with no tracking to graduation; criterion Y is not met, which also fails G, and no follow-up publications tracking this cohort were found.

    "To see whether role-play improves speaking fluency long-term, future studies might examine this." (p. 35)

  • P

    The paper contains no mention of any pre-registered protocol, registry, or registration date, and no external registry record for this study was found.

Abstract

This research examined whether role-play exercises improved Jordanian EFL students' speaking fluency. Fifty intermediate EFL students were studied for six weeks. The experimental and control groups were randomly assigned. Both groups got classroom instruction and speaking practice; however, the experimental group also participated in role-plays. Pre- and post-tests assessed participants' speaking fluency. The research compares the fluency levels of a role-play group with a typical classroom teaching group. The experimental and control groups differed significantly in speaking fluency development. Role-playing improved speaking fluency more than the control group. These findings imply that role-playing may improve Jordanian EFL students' speaking fluency. Role-play games help students to utilize the target language spontaneously and naturally, improving speaking fluency and confidence. Thus, EFL teachers should use role-play to develop students' speaking abilities.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomization was done at the individual student level within a classroom-based (not one-to-one tutoring) intervention, so the class-level requirement is not satisfied.
      • "The research randomly allocated 50 Jordanian EFL students to experimental or control groups." (p. 33)
      • Relevant Quotes: 1) "The research randomly allocated 50 Jordanian EFL students to experimental or control groups." (p. 33) 2) "Fifty intermediate EFL students were studied for six weeks. The experimental and control groups were randomly assigned." (Abstract, p. 30) 3) "Role-play sessions simulated real-world events like ordering food or negotiating a deal. Pair-based exercises increased target-language communication." (p. 33) Detailed Analysis: Criterion C requires randomisation at the class level (or the stronger school level), unless the intervention is one-to-one tutoring or personal teaching, in which case student-level randomisation is acceptable. The paper states that 50 individual students were randomly allocated to experimental or control groups; there is no mention of intact classes, sections, or schools being the unit of randomisation. The intervention itself is a classroom-based, group/pair-based role-play technique delivered during instruction, not personal tutoring, so the tutoring exception does not apply. Randomising individual students into treatment and control within the same instructional environment creates exactly the contamination risk that criterion C is designed to prevent. No quote describing class-level or school-level allocation could be found anywhere in the paper. All quotes above were re-verified as verbatim against the source PDF. Criterion C is not met because students, not classes or schools, were the unit of randomisation and the intervention is not one-to-one tutoring.
    • E

      Exam-based Assessment

      • Outcomes were measured with a speaking test and questionnaire constructed for this study, not a recognised standardised exam.
      • "The research employed a speaking test and a questionnaire. Speaking fluency was tested before and after the intervention." (p. 33)
      • Relevant Quotes: 1) "The research employed a speaking test and a questionnaire. Speaking fluency was tested before and after the intervention. Monologues, dialogues, and role-play s comprised the test." (p. 33) 2) "The monologue, dialogue, and role-play sections asked participants to talk for two minutes on a subject, five minutes with the examiner, and five minutes with their partner. Transcribed oral examinations were analyzed." (p. 33) 3) "Pre- and post-tests assessed participants' speaking fluency." (Abstract, p. 30) Detailed Analysis: Criterion E requires that outcomes be measured with a widely recognised standardised exam (e.g., a national exam, IELTS, TOEFL, or similar) rather than an instrument constructed by the researchers for the study. The paper describes an ad hoc speaking test consisting of monologue, dialogue, and role-play sections, scored on a 5-point scale, plus a questionnaire on motivation and self-efficacy. No named standardised test is mentioned anywhere, no validity or reliability evidence is reported, and no rubric provenance is given. The assessment is aligned with the intervention itself (it even contains a role-play section, favouring the experimental group), which is precisely the bias the criterion guards against. Quotes re-verified as verbatim. Criterion E is not met because the study used a custom, study-specific speaking test rather than a recognised standardised exam.
    • T

      Term Duration

      • The intervention and follow-up lasted only six weeks, with post-tests immediately after, which is far shorter than one academic term.
      • "Fifty intermediate EFL students were studied for six weeks." (Abstract, p. 30)
      • Relevant Quotes: 1) "Fifty intermediate EFL students were studied for six weeks." (Abstract, p. 30) 2) "The experimental group had six weeks of role-play , whereas the control group had language training." (p. 33) 3) "The pre- and post-tests measured participants' speaking fluency before and after the intervention." (p. 33) 4) "The short intervention time limits this investigation. To see whether role-play improves speaking fluency long-term, future studies might examine this." (p. 35) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Here the entire study spanned six weeks, and the post-test was administered at the end of that six-week period. Six weeks is roughly half of even a short term and clearly below the 3-4 month threshold. The author explicitly acknowledges "the short intervention time limits this investigation" and calls for future long-term studies, confirming there was no term-long follow-up tracking. Quotes re-verified as verbatim against the source PDF. Criterion T is not met because the interval from intervention start to outcome measurement was only six weeks, well short of one academic term.
    • D

      Documented Control Group

      • The control group is described only by size and mean test scores, with no demographic or baseline characteristics, and the paper explicitly labels its own descriptive and inferential results tables as fictional and hypothetical.
      • "This table is fictional and based on the outcomes data." (p. 33)
      • Relevant Quotes: 1) "The experimental group had six weeks of role-play , whereas the control group had language training." (p. 33) 2) "Both groups got classroom instruction and speaking practice; however, the experimental group also participated in role-plays." (Abstract, p. 30) 3) "Table 1. descriptive data for experimental and control group pre- and post-test speaking scores ... Control 25 2.4 2.7 0.3 0.4" (p. 33) 4) "This table is fictional and based on the outcomes data." (p. 33) 5) "The following table shows hypothetical ANOVA and post-hoc Tukey test results:" (p. 33) Detailed Analysis: Criterion D requires detailed documentation of the control group: demographics, baseline performance, and the treatment it received. The paper reports only the control group's size (n = 25), its mean pre-test (2.4) and post-test (2.7) scores, and a one-line statement that it received "language training" / regular classroom teaching. No demographic information (age, gender, institution details, proficiency verification) is provided for the control group, and no baseline equivalence checks are described. CRITICAL INTEGRITY FLAG (re-verified verbatim against the PDF during this verification pass): immediately after Table 1 (the only source of control-group descriptive statistics), the paper states, word for word, "This table is fictional and based on the outcomes data." (p. 33). Immediately before Table 2 (the ANOVA and post-hoc Tukey results that underpin the paper's claim of a statistically significant effect, t(48) = 5.6 / F(1,48) = 31.4), the paper states, word for word, "The following table shows hypothetical ANOVA and post-hoc Tukey test results:" (p. 33). Both quotes were located exactly where cited and are reproduced verbatim; they are not artifacts of OCR or misreading. This means the author explicitly labels the paper's own quantitative results tables as fictional/ hypothetical rather than as real analysis of real collected data. This undermines not only the control-group documentation required by criterion D but the evidentiary basis of the entire Results section (Table 1 descriptive statistics and Table 2 ANOVA/Tukey results, including the headline significance claims repeated in the Abstract, Results, and Discussion). No retraction, correction, or errata notice for this article was found via external search as of this verification. This paper should be flagged for manual review and likely exclusion from the ERCT corpus rather than treated as a normal "not met" on documentation grounds alone, because the self-declared fictional/hypothetical nature of the results tables calls into question whether the reported RCT outcome data are genuine. Criterion D is not met because the control group lacks demographic and baseline documentation, and, more seriously, the paper explicitly and verbatim describes its own descriptive and inferential results tables as fictional and hypothetical.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred at the individual student level within a single university, with no schools or institutions being randomised.
      • "The research randomly allocated 50 Jordanian EFL students to experimental or control groups." (p. 33)
      • Relevant Quotes: 1) "The research randomly allocated 50 Jordanian EFL students to experimental or control groups." (p. 33) 2) "I would like to express our gratitude to Jadara University and its students for their support and participation in this study." (p. 36) Detailed Analysis: Criterion S requires randomisation among schools or equivalent implementing institutions. This study drew 50 students from a single Jordanian university (Jadara University) and randomised them individually into two groups. Only one institution was involved, so school-level randomisation was impossible by design, and no quote anywhere in the paper describes assignment of schools, sites, or centres. Quotes re-verified as verbatim. Criterion S is not met because randomisation was at the student level within one university, not across schools.
    • I

      Independent Conduct

      • A single author designed, delivered, and evaluated the intervention with no external or third-party evaluation team.
      • "Dr. Luqman Rababah was responsible for the entire study." (p. 36)
      • Relevant Quotes: 1) "Dr. Luqman Rababah was responsible for the entire study." (Authors' contributions, p. 36) 2) "Therefore, this investigation explored the role of role-play activities proposed by the language teacher for fluency in spoken English of a sample of university students." (p. 33) 3) "Funding NA ... Competing interests NA" (p. 36) Detailed Analysis: Criterion I requires that the study be conducted independently of the intervention's designers, or at minimum that data collection and analysis be handled by an external party. Here the paper is single-authored and the author contribution statement explicitly says one person "was responsible for the entire study": design of the role-play intervention, its delivery, the construction of the speaking test, data collection, and analysis. There is no mention of external evaluators, blinded raters, an independent statistician, or third-party oversight of any kind. Quotes re-verified as verbatim. Criterion I is not met because the same single author designed the intervention and conducted the entire evaluation without any independent oversight.
    • Y

      Year Duration

      • The whole study lasted six weeks, which is far below 75% of an academic year; since criterion T fails, Y fails as well.
      • "Fifty intermediate EFL students were studied for six weeks." (Abstract, p. 30)
      • Relevant Quotes: 1) "Fifty intermediate EFL students were studied for six weeks." (Abstract, p. 30) 2) "The experimental group had six weeks of role-play , whereas the control group had language training." (p. 33) 3) "The short intervention time limits this investigation." (p. 35) Detailed Analysis: Criterion Y requires outcome tracking covering at least 75% of a full academic year (roughly 7+ months) from intervention start. The study's total span was six weeks, with outcomes measured immediately at the end. This is a small fraction of an academic year. Additionally, the weaker term-duration criterion T is not met, and per the instructions, if T is not met then Y cannot be met. Criterion Y is not met because the six-week study falls far short of 75% of an academic year, and the prerequisite criterion T also fails.
    • B

      Balanced Control Group

      • Both groups received classroom instruction and speaking practice over the same six-week window with no documented extra time or budget for the experimental group; the role-play activity is the treatment variable itself rather than a separable added resource.
      • "Both groups got classroom instruction and speaking practice; however, the experimental group also participated in role-plays." (Abstract, p. 30)
      • Relevant Quotes: 1) "Both groups got classroom instruction and speaking practice; however, the experimental group also participated in role-plays." (Abstract, p. 30) 2) "The experimental group had six weeks of role-play , whereas the control group had language training." (p. 33) 3) "The research compares the fluency levels of a role-play group with a typical classroom teaching group." (Abstract, p. 30) 4) "Role-play sessions simulated real-world events like ordering food or negotiating a deal. Pair-based exercises increased target-language communication." (p. 33) Detailed Analysis (re-applied under the updated criterion B decision procedure): Step 1 (EXTRA_RESOURCES_PRESENT): the paper gives no evidence that the experimental group received more class time, budget, or materials than the control group. Both conditions ran for the same six weeks and both are described as receiving classroom instruction and speaking practice; the only documented difference is the type of speaking activity (role-play vs. "language training"), not its quantity. Since no additional time or budget is shown, the criterion is met on this basis alone. Step 2 (fallback, if resources were considered "extra"): even under a stricter reading where the role-play sessions count as an add-on, they are explicitly framed as the primary treatment variable under test ("The research compares the fluency levels of a role-play group with a typical classroom teaching group"), with the control group continuing an active, business-as-usual "language training" condition rather than receiving nothing. This satisfies the treatment-variable exception. Caveat: the paper's dosage reporting is vague -- it does not state exact weekly minutes for either condition, so exact time-matching cannot be independently verified from the text. This imprecision is noted but does not change the conclusion under either branch of the decision procedure above. Criterion B is met because no extra time or budget for the experimental group is documented, and to the extent the role-play sessions differ from the control's activities, they are the integral treatment variable being tested against an active business-as-usual control.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent, peer-reviewed replication of this specific six-week Jordanian role-play trial was found in the paper or via external search; the study was published only in February 2025, leaving little time for replication.
      • "However, to my understanding, empirical evidence is scarce to examine the effects of role-play on NSr/Flu in the Jordanian EFL environment." (p. 33)
      • Relevant Quotes: 1) "However, to my understanding, empirical evidence is scarce to examine the effects of role-play on NSr/Flu in the Jordanian EFL environment." (p. 33) 2) "This research fills a gap in the literature by showing how role-play exercises affect Jordanian speaking fluency." (p. 31) 3) "To prove role-play's usefulness, a bigger sample size and long-term impacts on speaking fluency might be studied." (p. 35) Detailed Analysis: Criterion R requires independent replication of this specific study by a different team, published in a peer-reviewed journal. The paper itself positions the study as filling a gap, i.e., as novel in the Jordanian EFL context, and cites only prior, distinct role-play studies (e.g., Lahbibi & Farhane 2023 in Morocco; Dwiyanti & Lolita 2023 in Indonesia), which are earlier investigations of the general technique in other contexts, not replications of this trial. External search conducted for this verification (publisher site, ResearchGate, IDEAS/RePEc, Scribd listings, and general web search via search engines) located only re-postings and indexing of this same original 2025 article; no independent, peer-reviewed replication of this specific study by a different research team was found. Given the article was published February 7, 2025, there has been limited time for an independent replication to appear in the literature. Criterion R is not met because no independent peer-reviewed replication of this specific study exists as of this verification.
    • A

      All-subject Exams

      • Only speaking fluency in English was assessed with a custom test; no other core subjects were measured and criterion E is not met, which also fails A.
      • "Pre- and post-tests assessed participants' speaking fluency." (Abstract, p. 30)
      • Relevant Quotes: 1) "Pre- and post-tests assessed participants' speaking fluency." (Abstract, p. 30) 2) "The research employed a speaking test and a questionnaire. Speaking fluency was tested before and after the intervention." (p. 33) 3) "The questionnaire assessed participants' motivation, self-efficacy, and role-play attitudes." (p. 33) Detailed Analysis: Criterion A requires standardised exam-based assessment across all main subjects, and criterion E is an explicit prerequisite. Criterion E is not met (custom speaking test), so A automatically fails. Beyond that, the study measured only English speaking fluency plus attitudinal questionnaire constructs (motivation, self-efficacy); no other university subjects or broader academic outcomes were assessed, and no specialised-intervention justification referencing all- subject coverage is offered. Quotes re-verified as verbatim. Criterion A is not met because criterion E fails and only a single narrow outcome (speaking fluency) was measured.
    • G

      Graduation Tracking

      • Measurement stopped at the six-week post-test with no tracking to graduation; criterion Y is not met, which also fails G, and no follow-up publications tracking this cohort were found.
      • "To see whether role-play improves speaking fluency long-term, future studies might examine this." (p. 35)
      • Relevant Quotes: 1) "The pre- and post-tests measured participants' speaking fluency before and after the intervention." (p. 33) 2) "The short intervention time limits this investigation. To see whether role-play improves speaking fluency long-term, future studies might examine this." (p. 35) 3) "The study's research questions examine whether the role-play intervention affects some components of speaking fluency more than others and whether the observed changes remain over time." (p. 34) Detailed Analysis: Criterion G requires tracking participants until graduation from their educational stage. The study ended with a post-test immediately after the six-week intervention; the author explicitly defers long-term effects to "future studies." Per the criterion-specific instructions, since the prerequisite criterion Y is not met, G cannot be met. An external search for subsequent papers by Luqman Rababah that might track the same cohort of 50 Jordanian EFL students toward graduation was conducted (publisher site, ResearchGate, general web search). The author's other listed publications (e.g., on ChatGPT for thesis writing, online social learning, mobile-assisted listening, virtual reality pronunciation training) address different topics and samples; no follow-up study tracking this specific cohort to graduation was found. I could not find such a follow-up paper; I have not fabricated any quote to fill this gap. Criterion G is not met because tracking stopped at the six-week post-test with no graduation follow-up found in this paper or in subsequent publications, and the prerequisite criterion Y fails.
    • P

      Pre-Registered

      • The paper contains no mention of any pre-registered protocol, registry, or registration date, and no external registry record for this study was found.
      • Relevant Quotes: 1) "Received: June 16, 2024 Accepted: November 4, 2024 Online Published: February 7, 2025" (p. 30) 2) "The Publication Ethics Committee of the Sciedu Press." (Ethics approval, p. 36) 3) "Funding NA" (p. 36) Detailed Analysis: Criterion P requires that the full study protocol (hypotheses, methods, planned analyses) be registered on a public registry before data collection began. The paper contains no reference to any registry (e.g., ClinicalTrials.gov, OSF, AEA registry, ISRCTN), no registration ID, and no registration date. The only ethics-related statement refers to the journal's own publication ethics committee, which is not a study pre-registration. An external search conducted for this verification (publisher site, ResearchGate/IDEAS listings, and general web search covering OSF and related registries) found no pre-registration record associated with this study, its DOI, or its author for this topic. Criterion P is not met because no pre-registration of the study protocol is mentioned in the paper or found externally.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.