Automated versus peer assessment: Effects on learners' English public speaking

Chunping Zheng, Xu Chen, Huayang Zhang, Ching Sing Chai

Published:
ERCT Check Date:
DOI:
  • L2 languages
  • higher education
  • China
  • EdTech platform
  • digital assessment
  • formative assessment
0
  • C

    Randomisation was done at the individual student level within a single university course, not at the class or school level, and no tutoring exception applies.

    "The participants were then randomly assigned into two groups, with 26 students in the control group (G1), undertaking self-, peer, and teacher assessment, and another 26 students in the experimental group (G2), experiencing self-, automated, and teacher assessment." (p. 216)

  • E

    Outcomes were measured with researcher-adapted questionnaires and rubric-based teacher/machine ratings on the authors' own platform, not with a widely recognised standardised exam.

    "This study used two questionnaires for pre- and post-intervention surveys on learners' EPS self-efficacy and learning engagement." (p. 216)

  • T

    The intervention and outcome measurement spanned a full 16-week semester (early September to late December 2022), which covers at least one academic term.

    "The quasi-experimental research was carried out over 16 successive weeks between early September and late December 2022, with two hours of teaching time each week." (p. 216)

  • D

    The control group (G1, n = 26) is clearly documented, including its assessment condition, sample demographics, and baseline scores compared statistically with the experimental group.

    "The participants were then randomly assigned into two groups, with 26 students in the control group (G1), undertaking self-, peer, and teacher assessment" (p. 216)

  • S

    Randomisation occurred among individual students within one course at a single university, so there was no school-level randomisation.

    "They were selected for convenience because all the participants attended the same course and consented to participate in the research." (p. 215)

  • I

    The E-platform intervention was developed by a team led by the first author, and the same authors conducted, funded, and analysed the study without any independent third-party evaluation.

    "The E-platform was developed by a research team led by the first author." (p. 213)

  • Y

    The study lasted only 16 weeks (one semester), well short of 75% of a full academic year.

    "The quasi-experimental research was carried out over 16 successive weeks between early September and late December 2022, with two hours of teaching time each week." (p. 216)

  • B

    Both groups received the same course, the same platform, and three assessment modes each (self and teacher plus either peer or automated), so time and resources were balanced and the assessment mode itself was the treatment variable.

    "The control group (G1) undertook self-, peer, and teacher assessment via the platform, while the experimental group (G2) experienced self-, automated, and teacher assessment." (p. 210)

  • R

    No independent replication of this study by a different research team is mentioned in the paper, and an internet citation search found no such replication published since.

  • A

    Criterion E is not met, and only English public speaking outcomes were measured, with no standardised assessment of other core subjects.

    "As for learners' EPS competence, we referred to three constructs, namely speakers' management public speaking anxiety..., speaking and writing competence" (p. 217)

  • G

    Measurement ended with the Week 16 post-test in December 2022, criterion Y is unmet, and no internet-searchable follow-up publication tracks this cohort toward graduation.

    "the three formal speeches were conducted in Week 6, Week 11, and Week 16 during the instruction weeks (from Week 1 to Week 16) in the classroom" (p. 216)

  • P

    The paper contains no mention of pre-registration of the study protocol on any registry before data collection, and no registration record was found via internet search.

Abstract

This quasi-experimental research investigates the employment of a formative assessment platform aided by artificial intelligence in an English public speaking course. The platform integrates deep learning, automatic speech recognition, and automatic writing evaluation. It provides automated assessment and immediate feedback on speakers' public speaking anxiety and their speaking and writing competence. Fifty-two English public speaking learners were randomly assigned to two groups. The control group (G1) undertook self-, peer, and teacher assessment via the platform, while the experimental group (G2) experienced self-, automated, and teacher assessment. The ANCOVA results revealed that students in G1 perceived significantly higher social engagement than those in G2, which indicates that social interaction between learners during peer assessment cannot be substituted by automated assessment. The chi-square analysis showed students' different concerns regarding online formative assessment. While students in G1 showed concerns for peers' qualifications and willingness to provide feedback, students in G2 suggested generating more detailed automated feedback to improve self-learning. No significant differences were found in learners' English public speaking self-efficacy, engagement, or competence. This indicates that automated assessment can serve as an effective strategy for formative assessment and that AI tools may supplement peers as reliable learning companions in the foreseeable future.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was done at the individual student level within a single university course, not at the class or school level, and no tutoring exception applies.
      • "The participants were then randomly assigned into two groups, with 26 students in the control group (G1), undertaking self-, peer, and teacher assessment, and another 26 students in the experimental group (G2), experiencing self-, automated, and teacher assessment." (p. 216)
      • Relevant Quotes: 1) "Fifty-two English public speaking learners were randomly assigned to two groups." (p. 210, Abstract) 2) "The participants were then randomly assigned into two groups, with 26 students in the control group (G1), undertaking self-, peer, and teacher assessment, and another 26 students in the experimental group (G2), experiencing self-, automated, and teacher assessment." (p. 216) 3) "They were selected for convenience because all the participants attended the same course and consented to participate in the research." (p. 215) 4) "This quasi-experimental research investigates the employment of a formative assessment platform aided by artificial intelligence in an English public speaking course." (p. 210, Abstract) Detailed Analysis: Criterion C requires randomisation at the class level or stronger (school level) to prevent contamination between treatment and control participants. In this study all 52 participants attended the same compulsory English public speaking course at one university, and individual students were randomly assigned to the two assessment conditions. This is student-level randomisation within a single course, exactly the design the criterion warns against: students in the peer assessment group and the automated assessment group shared the same classroom environment, so contamination between conditions cannot be excluded. The intervention is a classroom-based formative assessment mode, not personal one-to-one tutoring, so the tutoring exception that would allow student-level randomisation does not apply. Criterion C is not met because students were individually randomised within a single course rather than by class or school, and no tutoring exception applies.
    • E

      Exam-based Assessment

      • Outcomes were measured with researcher-adapted questionnaires and rubric-based teacher/machine ratings on the authors' own platform, not with a widely recognised standardised exam.
      • "This study used two questionnaires for pre- and post-intervention surveys on learners' EPS self-efficacy and learning engagement." (p. 216)
      • Relevant Quotes: 1) "This study used two questionnaires for pre- and post-intervention surveys on learners' EPS self-efficacy and learning engagement. The first is the EPS self-efficacy scale, which measures learners' EPS self-efficacy in four dimensions with 12 items, each on a five-point Likert scale..." (p. 216) 2) "The second questionnaire evaluated learners' engagement in EPS learning and was adapted from the measurements of Wang et al. (2016) and Luan et al. (2023)." (p. 216) 3) "As for learners' EPS competence, we referred to three constructs, namely speakers' management public speaking anxiety..., speaking and writing competence... based on a five-point Likert scale (1 = very anxious/poor, 2 = anxious/fair, 3 = average, 4 = calm/good, 5 = very calm/excellent)" (p. 217) 4) "Three teachers were invited to rate learners' public speaking competence, and their average scores were taken as learners' final scores." (p. 217) 5) "Teachers' average scores of learners' first formal public speeches were used as the pre-test scores to indicate their initial EPS competence, and their scoring of learners' last formal public speeches were used as the post-test scores" (p. 217) Detailed Analysis: Criterion E requires that outcomes be measured with standard, widely recognised standardised exams rather than instruments created or assembled for the study. Here the outcome measures are (a) self-report Likert questionnaires on self-efficacy and engagement adapted from prior research instruments, and (b) teacher ratings of learners' classroom speeches using the authors' own rubric (Appendix A) on their self-developed E-platform. Although the platform incorporates iFLYTEK ASR and the Pigai AWE engine, no recognised standardised examination (e.g., TOEFL, IELTS, CET, or a national curriculum exam) was administered as an outcome measure. Teacher ratings and research questionnaires, however carefully validated, do not constitute standardised exam-based assessment under the ERCT standard. Criterion E is not met because outcomes were assessed with study-specific questionnaires and rubric-based human/machine ratings rather than a recognised standardised exam.
    • T

      Term Duration

      • The intervention and outcome measurement spanned a full 16-week semester (early September to late December 2022), which covers at least one academic term.
      • "The quasi-experimental research was carried out over 16 successive weeks between early September and late December 2022, with two hours of teaching time each week." (p. 216)
      • Relevant Quotes: 1) "The quasi-experimental research was carried out over 16 successive weeks between early September and late December 2022, with two hours of teaching time each week." (p. 216) 2) "The course lasted for 16 weeks with a two-hour class period each week." (p. 213) 3) "As shown in Appendix B, the three formal speeches were conducted in Week 6, Week 11, and Week 16 during the instruction weeks (from Week 1 to Week 16) in the classroom. The first 3 weeks (Week 1 to Week 3) were pre-intervention practice periods" (p. 216) 4) "Teachers' average scores of learners' first formal public speeches were used as the pre-test scores... and their scoring of learners' last formal public speeches were used as the post-test scores" (p. 217) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (about 3-4 months) after the intervention begins. The study ran across a complete 16-week university semester from early September to late December 2022. The formative assessment conditions operated throughout the instruction weeks, and the final outcome measurements (the post-test questionnaires and teacher scoring of the third formal speech) were taken in Week 16, roughly 3.5 to 4 months after the intervention began. A 16-week interval corresponds to one full academic semester/term, satisfying the minimum duration requirement. Criterion T is met because outcomes were measured about 16 weeks (one full semester) after the intervention started.
    • D

      Documented Control Group

      • The control group (G1, n = 26) is clearly documented, including its assessment condition, sample demographics, and baseline scores compared statistically with the experimental group.
      • "The participants were then randomly assigned into two groups, with 26 students in the control group (G1), undertaking self-, peer, and teacher assessment" (p. 216)
      • Relevant Quotes: 1) "The participants were recruited during the 2022-2023 academic year, including 44 female and eight male students (average age 19.25 years). They were sophomore students in the English department of a university in the northern part of China." (p. 215) 2) "The participants were then randomly assigned into two groups, with 26 students in the control group (G1), undertaking self-, peer, and teacher assessment, and another 26 students in the experimental group (G2), experiencing self-, automated, and teacher assessment." (p. 216) 3) "For G1, the means and standard deviations of the pre-test data were 3.06 (SD = 0.72) in language competence (LC), 3.05 (SD = 0.66) in delivery competence (DC), 3.81 (SD = 0.51) in topic competence (TC), and 3.42 (SD = 0.45) in organization competence (OC)." (p. 217) 4) "No significant differences (t = 0.12, n.s. for PSA; t = 0.33, n.s. for SC; t = 0.34, n.s. for WC) were found. Thus, students in both conditions shared similar prior EPS competence when they gave their first speech." (p. 219) Detailed Analysis: Criterion D requires detailed documentation of the control group: who they are, their size, baseline characteristics, and what treatment they received. The paper specifies the control group's size (n = 26), population (sophomore English majors at a university in northern China, recruited 2022-2023, sample of 44 female and 8 male students, mean age 19.25), and its exact condition (self-, peer, and teacher assessment on the same E-platform). Tables 1-3 report pre-test means and standard deviations for the control group on all outcome dimensions (self-efficacy, engagement, and competence), with statistical baseline-equivalence tests showing groups were comparable at the start (only topic-competence self-efficacy differed, which the ANCOVA design controls for). This constitutes adequate documentation of the control group. Criterion D is met because the control group's size, composition, condition, and baseline performance are clearly documented and compared with the experimental group.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred among individual students within one course at a single university, so there was no school-level randomisation.
      • "They were selected for convenience because all the participants attended the same course and consented to participate in the research." (p. 215)
      • Relevant Quotes: 1) "The study was conducted in a compulsory course for EFL learners designed to improve their EPS skills at a comprehensive university in northern China." (p. 213) 2) "They were selected for convenience because all the participants attended the same course and consented to participate in the research. The participants were then randomly assigned into two groups" (pp. 215-216) Detailed Analysis: Criterion S requires randomisation among schools or equivalent institutional units. This study took place within a single course at one university, and individual students were the unit of randomisation. No multiple institutions or sites were involved, and no school-level assignment occurred. Criterion S is not met because randomisation was at the student level within one course at a single university, not at the school level.
    • I

      Independent Conduct

      • The E-platform intervention was developed by a team led by the first author, and the same authors conducted, funded, and analysed the study without any independent third-party evaluation.
      • "The E-platform was developed by a research team led by the first author." (p. 213)
      • Relevant Quotes: 1) "The E-platform was developed by a research team led by the first author. It can conduct automated assessment, self-, peer- and teacher assessment of learners' EPS competence based on their speech videos, audios, and speech drafts (Zheng et al., 2024)." (p. 213) 2) "The maintenance and operations of the platform were funded by the National Natural Science Foundation in China and the Big Data Research and Service Center at the first author's institution" (p. 215) 3) "This research is funded by the National Natural Science Foundation of China (62107005), and the BUPT Research Project for Postgraduate Education (2024Y015)." (p. 223) 4) "Three teachers were invited to rate learners' public speaking competence, and their average scores were taken as learners' final scores." (p. 217) Detailed Analysis: Criterion I requires the evaluation to be conducted independently of the intervention's designers, or at minimum documented third-party oversight of data collection and analysis. Here the intervention platform was explicitly developed by a research team led by the first author, and the same author team designed the study, ran the course-based experiment, collected the data, and performed the analyses. While three PhD-qualified teachers rated the anonymised, randomised speech samples, the paper does not state that these raters or any external agency were independent of the research team, and no external evaluation body is mentioned. There is no disclosure or structural separation comparable to the accepted exception examples (external evaluators or government-led trials). Criterion I is not met because the developers of the intervention platform themselves conducted and analysed the evaluation with no documented independent oversight.
    • Y

      Year Duration

      • The study lasted only 16 weeks (one semester), well short of 75% of a full academic year.
      • "The quasi-experimental research was carried out over 16 successive weeks between early September and late December 2022, with two hours of teaching time each week." (p. 216)
      • Relevant Quotes: 1) "The quasi-experimental research was carried out over 16 successive weeks between early September and late December 2022, with two hours of teaching time each week." (p. 216) 2) "The course lasted for 16 weeks with a two-hour class period each week." (p. 213) 3) "the three formal speeches were conducted in Week 6, Week 11, and Week 16 during the instruction weeks (from Week 1 to Week 16)" (p. 216) Detailed Analysis: Criterion Y requires that outcomes be measured at least 75% of a full academic year (roughly 9-10 months, so at least about 7 months) after the intervention begins. This study started in early September 2022 and completed its final measurements in late December 2022, an interval of about 16 weeks (roughly 4 months). This covers a single semester only, or approximately 40-45% of a standard academic year, and no longer-term follow-up measurement is reported. Criterion Y is not met because the interval from intervention start to final measurement was only one 16-week semester, far below 75% of an academic year.
    • B

      Balanced Control Group

      • Both groups received the same course, the same platform, and three assessment modes each (self and teacher plus either peer or automated), so time and resources were balanced and the assessment mode itself was the treatment variable.
      • "The control group (G1) undertook self-, peer, and teacher assessment via the platform, while the experimental group (G2) experienced self-, automated, and teacher assessment." (p. 210)
      • Relevant Quotes: 1) "The control group (G1) undertook self-, peer, and teacher assessment via the platform, while the experimental group (G2) experienced self-, automated, and teacher assessment." (p. 210, Abstract) 2) "These two procedures of online formative assessment distinguished G1 and G2." (p. 211) 3) "Both groups conducted self-assessment and received teacher assessment on the E-platform." (p. 215) 4) "All the participants were asked to deliver three formal public speaking tasks according to their assigned conditions as indicated in Appendix B." (p. 216) 5) "The two different procedures of formative assessment constitute the independent variable in this study." (p. 216) 6) "In the first three weeks of classroom teaching, peers were trained to provide constructive feedback through the E-platform" (p. 215) Detailed Analysis: Criterion B requires that time and resources be balanced between conditions unless the extra resource is itself the treatment variable. Applying the decision tree: both groups attended the identical 16-week, two-hour-per-week compulsory course, delivered the same three formal speeches, were video-recorded in the same way, used the same E-platform, and each received exactly three modes of formative assessment (self- and teacher assessment in common, plus peer assessment for G1 versus automated assessment for G2). No extra instructional time, materials, or budget was given to either group beyond the other; the only difference is the source of the third feedback stream (peer vs. automated), which is the explicit treatment variable being tested and is delivered within the same course hours for both groups. Since no additional time or budget is present at all beyond this substitution of feedback source, the paper satisfies the simplest branch of the decision tree (no extra resources present), independent of whether the feedback-source difference is deemed integral to the design. Criterion B is met because both conditions received equivalent course time, tasks, and platform access, with the assessment mode (peer versus automated) being the explicit treatment variable and no extra time or budget given to either group.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this study by a different research team is mentioned in the paper, and an internet citation search found no such replication published since.
      • Relevant Quotes: 1) "This paper reports our quasi-experimental and pilot study on an EPS formative assessment platform empowered by AI technologies." (p. 222) 2) "The research team is collaborating with multiple institutions to improve the AI-supported assessment of learners' EPS competence in China, and we are positive that upcoming research would include more representative samples" (p. 222) Detailed Analysis: Criterion R requires independent replication of the study by a different research team in a different context, published in a peer-reviewed journal. The paper describes itself as a pilot study of the authors' own E-platform and mentions only the same team's plans for future multi-institution work, which would not be independent. Related studies cited in the paper (e.g., Chen, 2022; Ochoa & Dominguez, 2020; Dai & Wu, 2023) examine different automated feedback tools and designs and are not replications of this specific trial. An internet search (via Semantic Scholar citation records for this paper, paperId 5a20ed68e294ec064182d07f84c3f11ed955c249) was conducted to check for independent replications published after this paper. As of the check date the paper had a small number of citing works (e.g., meta-analyses/systematic reviews on AI in language education such as Jantakoon et al. and Abreu & Suarez-Perdomo, and studies on related but distinct AI-feedback topics); none of the citing works are independent replications of this specific quasi-experimental design (same E-platform, same G1 peer vs. G2 automated comparison) by a different research team. No independent replication was found. Criterion R is not met because no independent, peer-reviewed replication of this study by a different research team was found in the paper or via internet search.
    • A

      All-subject Exams

      • Criterion E is not met, and only English public speaking outcomes were measured, with no standardised assessment of other core subjects.
      • "As for learners' EPS competence, we referred to three constructs, namely speakers' management public speaking anxiety..., speaking and writing competence" (p. 217)
      • Relevant Quotes: 1) "As for learners' EPS competence, we referred to three constructs, namely speakers' management public speaking anxiety (the lower the PSA score, the better self-management of PSA), speaking and writing competence (the higher the scores, the better)." (p. 217) 2) "To evaluate the effect of automated assessment on learners' EPS self-efficacy, engagement, and competence, we compared two groups of students with two different interventional conditions." (p. 211) Detailed Analysis: Criterion A requires standardised exam-based assessment of all main subjects, and explicitly presupposes criterion E. Since criterion E is not met (no standardised exams were used at all), criterion A automatically fails. Moreover, the study measured outcomes only within the single domain of English public speaking (anxiety, speaking, writing within the EPS course); no other subjects in the students' curriculum were assessed, and no specialised-intervention justification per the exception is offered that would substitute for standardised measures. Criterion A is not met because criterion E fails and outcomes were measured only in the single domain of English public speaking.
    • G

      Graduation Tracking

      • Measurement ended with the Week 16 post-test in December 2022, criterion Y is unmet, and no internet-searchable follow-up publication tracks this cohort toward graduation.
      • "the three formal speeches were conducted in Week 6, Week 11, and Week 16 during the instruction weeks (from Week 1 to Week 16) in the classroom" (p. 216)
      • Relevant Quotes: 1) "The quasi-experimental research was carried out over 16 successive weeks between early September and late December 2022" (p. 216) 2) "their scoring of learners' last formal public speeches were used as the post-test scores to indicate the group differences after the invention in EPS competence" (p. 217) 3) "Future studies may explore the content of peer and automated feedback to look for categories of comments that significantly enhance students' engagement and development in EPS competence." (p. 222) Detailed Analysis: Criterion G requires tracking participants until graduation from their educational stage, and per the instructions it cannot be met when criterion Y is not met. Criterion Y fails here. The participants were university sophomores recruited in the 2022-2023 academic year, and all data collection ended with the Week 16 post-test in late December 2022; no follow-up beyond the semester, and certainly none through university graduation, is described or planned in the paper beyond a general mention of future multi-institution work. An internet search for follow-up publications by the same author team (Chunping Zheng and colleagues, Beijing University of Posts and Telecommunications) tracking this same 52-student cohort toward graduation found no such paper; the only related publication located is the authors' own prior instrument paper (Zheng et al., 2024, on multimodal deep learning assessment of PSA), which is not a graduation follow-up of this cohort. Criterion G is not met because tracking stopped at the end of the 16-week course with no follow-up to graduation found, and criterion Y is also unmet.
    • P

      Pre-Registered

      • The paper contains no mention of pre-registration of the study protocol on any registry before data collection, and no registration record was found via internet search.
      • Relevant Quotes: 1) "At the beginning of the course, informed of the research procedure and research purposes, all the participants signed consent forms and agreed to voluntarily participate in the research." (p. 216) 2) "Appendix A (available via Open Science Framework) shows the specific evaluation dimensions for the three constructs" (p. 215) Detailed Analysis: Criterion P requires that the full study protocol (hypotheses, methods, planned analyses) be registered on a public registry before data collection began, with a verifiable registration date. The paper mentions informed consent and that appendices (evaluation-dimension materials) are shared via the Open Science Framework, but sharing supplementary materials after the fact is not the same as protocol pre-registration. No registry name, registration ID, or registration date is reported anywhere in the paper, and there is no statement that hypotheses and analysis plans were registered before the study started in September 2022. An internet search for an OSF registration or other registry entry (e.g., ClinicalTrials.gov-style education registries) under the authors' names or this study's title did not locate any pre-registration record; only the paper's own Appendix A/B supplementary-materials postings on OSF were identified, which are not a pre-registered protocol. Criterion P is not met because no pre-registration of the study protocol is reported in the paper or found online.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.