Robot-Assisted Instruction of L2 Pragmatics: Effects on Young EFL Learners' Speech Act Performance

Minoo Alemi, Nafiseh Sadat Haeri

Published:
ERCT Check Date:
DOI: 10.64152/10125/44727
  • L2 languages
  • pre-K
  • Asia
0
  • C

    Individual students from one kindergarten course list were randomized into two newly formed groups, which is student-level (not class-level) randomization for a group-teaching (non-tutoring) intervention.

    "These children were then randomly selected and divided into two groups, including RALL (n = 19) and non-RALL groups (n = 19)." (p. 89)

  • E

    Outcomes were measured with a custom pictorial pre/post-test designed by the researchers for this study, not a widely recognised standardised exam.

    "Before starting the lessons, the investigators administered a pictorial pre-test to both groups in order to homogenize the students based on their speech act knowledge." (p. 92)

  • T

    The intervention lasted only four weeks with the post-test administered immediately after the eighth session, far short of one academic term.

    "There were eight one-hour teaching sessions over a period of four weeks for both groups." (p. 86)

  • D

    The control (non-RALL) group's size, composition, baseline performance, and conditions (identical lessons and games without the robot) are clearly documented.

    "In the non-RALL group the lessons included games similar to those in the RALL group, but without the presence of the robot." (p. 86)

  • S

    Randomisation occurred among 38 individual children within a single kindergarten; no schools or sites were randomised.

    "The participants of this study consisted of 38 children from a private kindergarten in Tehran, Iran." (p. 89)

  • I

    The authors designed the intervention, wrote the materials (including the co-authored textbook), programmed the robot, taught both classes, and analysed the data themselves, with no independent evaluators.

    "...the non-RALL (i.e., control group) was managed by the teacher who was also the researcher." (pp. 89-90)

  • Y

    The study tracked outcomes for only four weeks, far below 75% of an academic year (and criterion T is also not met).

    "Participants completed eight one-hour teaching sessions two days a week for four weeks." (p. 93)

  • B

    Both groups received identical lesson plans, games, materials, and eight one-hour sessions; the only difference was the robot itself, which was the explicit treatment variable being tested.

    "The lesson plans were created to teach requests and thanking speech acts, and, aside from the presence (or absence) of NIMA were similar for both groups." (p. 92)

  • R

    No independent replication of this specific RALL pragmatics study by a different research team was found; the authors themselves call for future replication.

    "Furthermore, it is necessary to replicate the results of this study in varying contexts to obtain more robust results." (p. 100)

  • A

    Only English pragmatic performance (two speech acts) was assessed with a custom test; no other subjects were measured, and criterion E is not met, so A automatically fails.

    "The findings revealed a significant difference between the RALL and non-RALL groups' pragmatic performance for thanking and requesting." (p. 86)

  • G

    Measurement stopped immediately after the four-week course with no follow-up or graduation tracking, and criterion Y is not met, so G automatically fails.

    "After eight teaching sessions, a post-test was administered to measure the participants' learning gains." (p. 93)

  • P

    The paper contains no mention of any pre-registered protocol, registry platform, or registration date, and no registry entry for this study was found online.

Abstract

Technology, as a source of instruction, has fulfilled various purposes in foreign language learning environments. During the last decade, Robot-Assisted Language Learning (RALL) has attracted teachers' and researchers' attention due to the look and feel of humanoid robots. However, in the field of pragmatics, studies highlighting the role of RALL have gone relatively unnoticed. To bridge this gap, this study sought to explore the effect of RALL on pragmatic features, including request and thanking speech acts by young Persian-speaking EFL learners. For this aim, 38 preschool children (3 to 6 year-old boys and girls) with no English learning experience were randomly assigned to the RALL (19 students) and non-RALL (19 students) groups. In the RALL group, a humanoid robot was used as an assistant to the teacher to play games, repeat the sentences, and interact with the students. In the non-RALL group the lessons included games similar to those in the RALL group, but without the presence of the robot. There were eight one-hour teaching sessions over a period of four weeks for both groups. Following completion of the lessons in both groups, the results of post-tests were analyzed using an independent sample t-test. The findings revealed a significant difference between the RALL and non-RALL groups' pragmatic performance for thanking and requesting. Based on these findings it can be concluded that RALL instruction was more effective than non-RALL instruction in improving the young learners' performance.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Individual students from one kindergarten course list were randomized into two newly formed groups, which is student-level (not class-level) randomization for a group-teaching (non-tutoring) intervention.
      • "These children were then randomly selected and divided into two groups, including RALL (n = 19) and non-RALL groups (n = 19)." (p. 89)
      • Relevant Quotes: 1) "38 preschool children (3 to 6 year-old boys and girls) with no English learning experience were randomly assigned to the RALL (19 students) and non-RALL (19 students) groups." (Abstract, p. 86) 2) "Before choosing the participants, the kindergarten principal offered the researcher a list of 38 children who were at the first level of the English course. These children were then randomly selected and divided into two groups, including RALL (n = 19) and non-RALL groups (n = 19)." (p. 89) 3) "To collect the data, 38 students were randomly assigned to RALL and non-RALL groups." (p. 93) 4) "The RALL and non-RALL classes took place on different days and lasted eight sessions (two days a week for four weeks) for both groups." (pp. 89-90) Detailed Analysis: The unit of randomisation was the individual child: 38 children from a single kindergarten English-course list were individually randomly assigned to the two conditions. No pre-existing intact classes or schools were randomised. The ERCT C criterion requires randomisation of entire classes (or schools), unless the intervention is personal one-to-one tutoring, in which case student-level randomisation is acceptable. Here the intervention was whole-group classroom teaching (a robot assisting a teacher in front of a class of 19 children), not individual tutoring, so the tutoring exception does not apply. Although the two groups met on different days, which reduces direct contamination, the randomisation itself was at the student level within a single institution rather than at the class or school level. Criterion C is not met because individual students, not intact classes or schools, were randomised, and the intervention was group instruction rather than one-to-one tutoring.
    • E

      Exam-based Assessment

      • Outcomes were measured with a custom pictorial pre/post-test designed by the researchers for this study, not a widely recognised standardised exam.
      • "Before starting the lessons, the investigators administered a pictorial pre-test to both groups in order to homogenize the students based on their speech act knowledge." (p. 92)
      • Relevant Quotes: 1) "Before starting the lessons, the investigators administered a pictorial pre-test to both groups in order to homogenize the students based on their speech act knowledge." (p. 92) 2) "Next, they were randomly assigned to two groups and, at the end of the course, a pictorial post-test in Power Point format was administered to both groups." (p. 92) 3) "In addition, the validity of the tests was examined by two independent English teachers. Since the pictorial test did not meet the requirements for a statistical reliability check, the test's reliability was ensured by asking ten other children, who were previously instructed and familiar with the sentences, to respond to the prompts." (pp. 92-93) 4) "Because all of the children were unable to read and write, both pre-test and post-test were pictorial." (p. 93) Detailed Analysis: The outcome measure was a researcher-made pictorial test in PowerPoint, built around the exact request and thanking sentences taught in the lessons (Tables 1 and 2 list "the request and thanking sentences taught and tested"). The authors themselves note it could not undergo a standard statistical reliability check and instead used an informal check with ten other children. No nationally or internationally recognised standardised exam was used. This is precisely the custom, intervention-aligned assessment that criterion E warns against, since it can inflate apparent effectiveness. Criterion E is not met because the assessment was a custom-made pictorial test designed by the researchers, not a standardised exam.
    • T

      Term Duration

      • The intervention lasted only four weeks with the post-test administered immediately after the eighth session, far short of one academic term.
      • "There were eight one-hour teaching sessions over a period of four weeks for both groups." (p. 86)
      • Relevant Quotes: 1) "There were eight one-hour teaching sessions over a period of four weeks for both groups." (Abstract, p. 86) 2) "Participants completed eight one-hour teaching sessions two days a week for four weeks. The duration was the same for both groups." (p. 93) 3) "After eight teaching sessions, a post-test was administered to measure the participants' learning gains." (p. 93) Detailed Analysis: Criterion T requires the interval from intervention start to outcome measurement to be at least one full academic term (approximately 3-4 months). Here the intervention ran for four weeks (eight one-hour sessions, two per week), and the post-test was given immediately after the final session. The tracking interval from start to measurement was therefore about one month, well below a term. No delayed follow-up measurement is reported. Criterion T is not met because outcomes were measured about four weeks after the intervention started, which is much shorter than one academic term.
    • D

      Documented Control Group

      • The control (non-RALL) group's size, composition, baseline performance, and conditions (identical lessons and games without the robot) are clearly documented.
      • "In the non-RALL group the lessons included games similar to those in the RALL group, but without the presence of the robot." (p. 86)
      • Relevant Quotes: 1) "In the non-RALL group the lessons included games similar to those in the RALL group, but without the presence of the robot." (Abstract, p. 86) 2) "The RALL (i.e., experimental group) had the robot in their class as a teacher assistance (TA) while the non-RALL (i.e., control group) was managed by the teacher who was also the researcher." (pp. 89-90) 3) "The sample included both boys and girls ranging in age from 3 to 6 years old. The children were all preschoolers and unable to read and write in either Farsi (the children's first language) or English. All had little English background and knew only few words of English, a pre-test demonstrated they were unfamiliar with any complete sentences in English." (p. 89) 4) "The scores obtained from the pre-tests were zero showing all the children were uninstructed in English requests and thanking sequences. There was no difference in their pre-test scores." (p. 96) 5) "Non-RALL 19 13.89 1.59 0.36" (Table 3, p. 96) Detailed Analysis: The paper documents the control group in adequate detail: its size (n = 19), demographics (boys and girls aged 3-6, native Persian speakers, preschoolers, no prior English instruction), baseline performance (pre-test scores of zero for all children, confirming homogeneity between groups), and the conditions it received (the same lesson plans, materials, games, and eight one-hour sessions over four weeks, delivered by the teacher without the robot). Descriptive statistics for the control group's post-tests are reported in Tables 3 and 5. This allows a proper comparison between conditions. Criterion D is met because the control group's size, characteristics, baseline scores, and treatment conditions are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred among 38 individual children within a single kindergarten; no schools or sites were randomised.
      • "The participants of this study consisted of 38 children from a private kindergarten in Tehran, Iran." (p. 89)
      • Relevant Quotes: 1) "The participants of this study consisted of 38 children from a private kindergarten in Tehran, Iran." (p. 89) 2) "These children were then randomly selected and divided into two groups, including RALL (n = 19) and non-RALL groups (n = 19)." (p. 89) Detailed Analysis: Criterion S requires randomisation at the level of schools or equivalent implementing institutions. This study took place in a single private kindergarten in Tehran, and randomisation was of individual children into two groups within that one site. No multiple schools, centres, or sites were involved or randomised. Criterion S is not met because the study was conducted in one kindergarten with student-level randomisation, not a school-level RCT.
    • I

      Independent Conduct

      • The authors designed the intervention, wrote the materials (including the co-authored textbook), programmed the robot, taught both classes, and analysed the data themselves, with no independent evaluators.
      • "...the non-RALL (i.e., control group) was managed by the teacher who was also the researcher." (pp. 89-90)
      • Relevant Quotes: 1) "The RALL (i.e., experimental group) had the robot in their class as a teacher assistance (TA) while the non-RALL (i.e., control group) was managed by the teacher who was also the researcher." (pp. 89-90) 2) "The robot was programmed by a team including the researchers, an engineer to program the lesson plans into the robot using the Choregraphe program (Figure 3), and an operator to control the robot in the class." (p. 90) 3) "The teaching materials for this study for both RALL and non-RALL groups were based on the Functional Communication in English textbook by Tajeddin and Alemi (2014) and also stories from the ESL library website." (p. 91) 4) "The researcher chose novice-level sentences for children and tried to find a story for each sentence from the ESL library website." (p. 91) Detailed Analysis: Criterion I requires the study to be conducted independently of the intervention's designers. Here the same research team designed the RALL intervention, based the teaching materials on a textbook co-authored by the first author (Tajeddin & Alemi, 2014), programmed the robot, delivered instruction (the researcher was herself the teacher of the control class), created and administered the tests, and performed the analysis. There is no external evaluation team, third-party data collection, or independent oversight mentioned anywhere in the paper. Only the test validity was "examined by two independent English teachers," which is far from independent conduct of the trial. Criterion I is not met because the intervention designers themselves implemented the intervention, collected the data, and analysed the results without independent oversight.
    • Y

      Year Duration

      • The study tracked outcomes for only four weeks, far below 75% of an academic year (and criterion T is also not met).
      • "Participants completed eight one-hour teaching sessions two days a week for four weeks." (p. 93)
      • Relevant Quotes: 1) "Participants completed eight one-hour teaching sessions two days a week for four weeks. The duration was the same for both groups." (p. 93) 2) "After eight teaching sessions, a post-test was administered to measure the participants' learning gains." (p. 93) Detailed Analysis: Criterion Y requires outcomes to be measured at least 75% of a full academic year (roughly 9-10 months) after the intervention begins. The entire study, from first session to post-test, spanned about four weeks. Additionally, per the criteria-specific rule, since criterion T (Term Duration) is not met, criterion Y cannot be met. Criterion Y is not met because the interval from intervention start to measurement was only about one month, far short of an academic year.
    • B

      Balanced Control Group

      • Both groups received identical lesson plans, games, materials, and eight one-hour sessions; the only difference was the robot itself, which was the explicit treatment variable being tested.
      • "The lesson plans were created to teach requests and thanking speech acts, and, aside from the presence (or absence) of NIMA were similar for both groups." (p. 92)
      • Relevant Quotes: 1) "In the non-RALL group the lessons included games similar to those in the RALL group, but without the presence of the robot. There were eight one-hour teaching sessions over a period of four weeks for both groups." (Abstract, p. 86) 2) "The lesson plans were created to teach requests and thanking speech acts, and, aside from the presence (or absence) of NIMA were similar for both groups." (p. 92) 3) "The above games were performed by NIMA in the RALL class to increase the children's motivation, while the games were performed by the teacher in the non-RALL class." (p. 92) 4) "Participants completed eight one-hour teaching sessions two days a week for four weeks. The duration was the same for both groups." (p. 93) 5) "The only distinction between the two groups was the presence of the robot as an assistant to the teacher in the RALL group." (p. 98) Detailed Analysis: Applying the updated criterion B decision procedure: instructional time and educational activities were carefully matched across conditions (identical lesson plans based on the same textbook and stories, the same games such as pass the ball, Simon says, and mystery bag, the same flash cards and classroom objects, the same teacher, and identical instructional time of eight one-hour sessions over four weeks). The single additional resource present only in the intervention group was the NAO humanoid robot (NIMA), along with the engineering/programming effort behind it, acting as a teaching assistant. This is an extra resource (EXTRA_RESOURCES_PRESENT = true), but the robot IS the treatment variable being evaluated (RESOURCES_ARE_TREATMENT = true): the study's explicit purpose is to "explore the effect of RALL" versus non-RALL instruction, so the robot is integral to, and the defining component of, the intervention being tested rather than a separable confounding add-on of extra time or budget. Per the decision tree, when the extra resource is itself the treatment variable, the control group may remain "business as usual" (here, the same lesson without the robot) without violating the criterion. Instructional time, materials, games, and teacher were otherwise held constant, and the paper explicitly states the only distinction between groups was the robot's presence. Criterion B is met because instructional time, materials, and activities were matched across groups, and the only additional resource (the humanoid robot and its supporting engineering effort) was the explicit treatment variable integral to the intervention being tested, consistent with the exception for resources that are the primary treatment variable.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific RALL pragmatics study by a different research team was found; the authors themselves call for future replication.
      • "Furthermore, it is necessary to replicate the results of this study in varying contexts to obtain more robust results." (p. 100)
      • Relevant Quotes: 1) "Furthermore, it is necessary to replicate the results of this study in varying contexts to obtain more robust results." (p. 100) 2) "Other studies have been done using the same robot for different subjects and all results demonstrate a promising impact on learning (e.g., Alemi & Haeri, 2017a; Alemi, Meghdari, & Haeri, 2017b; Alemi & Bahramipour, 2019)." (p. 89) 3) "However, in the field of pragmatics, studies highlighting the role of RALL have gone relatively unnoticed." (Abstract, p. 86) Detailed Analysis: Criterion R requires independent replication of the study by a different research team in a different context, published in a peer-reviewed journal. The paper explicitly frames itself as filling a gap ("studies highlighting the role of RALL [in pragmatics] have gone relatively unnoticed") and calls for future replication, indicating no replication existed at publication. The related RALL studies cited (Alemi & Haeri 2017a; Alemi, Meghdari, & Haeri 2017b; Alemi & Bahramipour 2019) are by the same author team and address different outcomes (vocabulary, attitudes), not replications of this pragmatics trial. As part of this verification, the paper's citation record was checked via Semantic Scholar (DOI 10.64152/10125/44727), which lists over 50 citing works published between 2020 and 2026, including systematic reviews and meta-analyses of robot-assisted language learning (e.g., "Robot-Assisted Language Learning: A Meta-Analysis," Derakhshan et al., 2024; "Social robots: a meta-analysis of learning outcomes," Karpouzis et al., 2026) and related preschool robot-L2 studies (e.g., "The Influence of Robot Social Behaviors on Second Language Learning in Preschoolers," Yin, Guo, Zheng, Ren, Wang, & Jiang, 2022). None of these citing works is an independent replication of this specific study's design (RALL vs. non-RALL instruction of request and thanking speech acts to young Persian-speaking EFL preschoolers using the same or an equivalent protocol); they either cite it as background/context or study different outcomes, populations, or robot behaviours. No independent, peer-reviewed replication of this particular trial by a different research team was located. Criterion R is not met because no independent, peer-reviewed replication of this specific study by a different research team was found.
    • A

      All-subject Exams

      • Only English pragmatic performance (two speech acts) was assessed with a custom test; no other subjects were measured, and criterion E is not met, so A automatically fails.
      • "The findings revealed a significant difference between the RALL and non-RALL groups' pragmatic performance for thanking and requesting." (p. 86)
      • Relevant Quotes: 1) "The findings revealed a significant difference between the RALL and non-RALL groups' pragmatic performance for thanking and requesting." (Abstract, p. 86) 2) "The following Tables (1) and (2) are the list of request and thanking sentences taught and tested in both the RALL and non-RALL groups" (p. 92) Detailed Analysis: Criterion A requires standardised exam-based assessment of all main subjects taught at the educational level. This study measured only English L2 pragmatic performance (specifically two speech acts: requesting and thanking) using a custom pictorial test. No other areas of the kindergarten curriculum were assessed. Furthermore, the criteria-specific rule states that if criterion E (Exam-based Assessment) is not met, criterion A cannot be met; E fails here because the assessment was custom-made. Criterion A is not met because only a single narrow outcome (two English speech acts) was measured with a non-standardised custom test.
    • G

      Graduation Tracking

      • Measurement stopped immediately after the four-week course with no follow-up or graduation tracking, and criterion Y is not met, so G automatically fails.
      • "After eight teaching sessions, a post-test was administered to measure the participants' learning gains." (p. 93)
      • Relevant Quotes: 1) "After eight teaching sessions, a post-test was administered to measure the participants' learning gains." (p. 93) 2) "In addition, both longitudinal and cross-sectional studies can be done on both children and teenagers using a robot for their studies." (p. 100) Detailed Analysis: Criterion G requires tracking participants until graduation from their educational stage. Here data collection ended with a post-test given immediately after the eighth and final session; the preschool children (aged 3-6) were not followed through the end of kindergarten or beyond. The authors' suggestion that future longitudinal studies "can be done" confirms no long-term tracking occurred. As part of this verification, the citation record of the paper (via Semantic Scholar, DOI 10.64152/10125/44727) and subsequent publications by the same authors (Minoo Alemi, Nafiseh Sadat Haeri) were checked for any follow-up study tracking this specific kindergarten cohort through to graduation. No such follow-up publication was found; the authors' later and related works (e.g., Alemi & Haeri, 2017a; Alemi, Meghdari, & Haeri, 2017b) predate this 2020 study and concern different cohorts and outcomes (greeting/vocabulary, attitudes), not a graduation follow-up of these 38 preschoolers. Additionally, per the criteria-specific rule, criterion G cannot be met because criterion Y is not met. Criterion G is not met because measurement stopped immediately after the four-week intervention with no tracking to graduation, and no follow-up publication tracking this cohort was found.
    • P

      Pre-Registered

      • The paper contains no mention of any pre-registered protocol, registry platform, or registration date, and no registry entry for this study was found online.
      • Relevant Quotes: 1) "All parents' approval for participation in the study was provided through informed consent forms." (p. 90) 2) "To collect the data, 38 students were randomly assigned to RALL and non-RALL groups." (p. 93) Detailed Analysis: Criterion P requires the full study protocol (hypotheses, methods, planned analyses) to be publicly pre-registered before data collection began. The paper describes informed consent and its procedures but makes no mention of any registry (e.g., ClinicalTrials.gov, OSF, AEA registry), no registration ID, and no registration date. As part of this verification, a search was conducted for a pre-registration record associated with this study or its authors (Alemi and Haeri); no registry entry, protocol, or pre-registration reference for this specific RALL pragmatics study was located. There is no evidence anywhere in the paper or online of a pre-registered protocol. Criterion P is not met because no pre-registration statement or registry reference appears in the paper, and none was found through internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.