Personalized language learning with an LLM chatbot: effects of immediate vs. delayed corrective feedback

Alireza M. Kamelabad, Beatrice Turano, Mattias Lundin, Gabriel Skantze

Published:
ERCT Check Date:
DOI: 10.3389/feduc.2026.1703664
  • L2 languages
  • higher education
  • EU
  • EdTech website
0
  • C

    Randomization was at the individual student level, but the intervention is one-to-one tutoring by a chatbot positioned as a personal English tutor, so the tutoring exception applies and student-level RCT is acceptable.

    "We designed our chatbot to simulate written conversation practices between an English instructor (chatbot) and an English as a Foreign Language (EFL) learner (user), where the teacher corrects the user's grammatical mistakes."

  • E

    Outcomes were measured with a custom assessment built from 25 sentences taken from each participant's own chat utterances, not a standardised, widely recognised exam.

    "The participants were given 25 sentences extracted from their utterances and were instructed to identify and correct any mistakes they found."

  • T

    The intervention comprised 12 chatbot sessions over roughly four weeks with outcomes measured immediately after the last session, which is far shorter than one academic term.

    "Participants completed 12 sessions in total, conducting three per week for 4 weeks"

  • D

    The study has no control group by design, and the two feedback-timing arms lack documented per-group demographics and baseline performance, so no properly documented comparison group exists.

    "Furthermore, the absence of a control group, dictated by our study's focus on comparing ICF and DCF conditions, constrains the breadth of our findings."

  • S

    Randomisation occurred at the individual student level within four classes at a single institution, not at the school level.

    "Participants were from four classes (A, B, C, D) and pseudo-randomly assigned to either of the two conditions."

  • I

    The same research team designed, built, administered, and analysed the chatbot intervention with no external or third-party evaluation.

    "AM: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing."

  • Y

    The whole study, from first session to final assessment, lasted about one month, far below 75% of an academic year.

    "Participants completed 12 sessions in total, conducting three per week for 4 weeks"

  • B

    Both groups received identical chatbot sessions, identical session dosage, and the same system-generated amount of feedback, differing only in feedback timing, which is the treatment variable itself.

    "Since the same system was used for error detection, the amount of feedback generation was the same between the two conditions, with the only difference being the time when the feedback was given to the participant (immediately in the conversation, delayed in a report)."

  • R

    No independent replication of this specific chatbot feedback-timing RCT by a different team has been published; related studies exist but do not reproduce this study's design, and a citation-graph search found no replication.

  • A

    Only English grammar knowledge was assessed, with a custom instrument, so neither the standardised-exam prerequisite nor the all-subjects coverage is satisfied.

    "The participants were given 25 sentences extracted from their utterances and were instructed to identify and correct any mistakes they found."

  • G

    Measurement stopped immediately after the final chat session with no follow-up, and since criterion Y is not met this criterion cannot be met either.

    "After finishing their final chat session, participants were directed to the language assessment page."

  • P

    The paper contains no mention of a pre-registered protocol or registry entry made before data collection; the only external repository (OSF) was created in January 2023, after data collection had already finished.

Abstract

The emergence of Large Language Models (LLMs) has opened new possibilities for language learning through conversational interaction with chatbots. Yet, little empirical evidence exists on how students experience such interactions and how corrective feedback should be provided. Research suggests that immediate corrective feedback is generally more effective than delayed feedback. Nevertheless, learners' perception of this effectiveness and their preferences for feedback timing, particularly in the domain of Computer-Assisted Language Learning (CALL), remain underexplored. This study investigates the feasibility of providing immediate feedback and examines the impact of feedback timing on user experience and grammar learning gains in English. An in-the-wild experiment was conducted with 66 L2 English learners, who integrated chatbot sessions into their English course as an extracurricular activity over one semester. Participants were randomly assigned to two groups receiving feedback either during or after the conversation. Findings reveal no significant difference in learning gains, but immediate feedback enhanced user experience, leading to overall positive perceptions of the chatbot. Additionally, we explore users' perceptions of the chatbot's social role and personality, offering a roadmap for future enhancements. These results provide valuable insights into the potential of LLMs and chatbots for language learning.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomization was at the individual student level, but the intervention is one-to-one tutoring by a chatbot positioned as a personal English tutor, so the tutoring exception applies and student-level RCT is acceptable.
      • "We designed our chatbot to simulate written conversation practices between an English instructor (chatbot) and an English as a Foreign Language (EFL) learner (user), where the teacher corrects the user's grammatical mistakes."
      • Relevant Quotes: 1) "A between-subject design was employed, comprising two conditions designed to evaluate the influence of correction timing on language learning and interactional dynamics (ICF vs. DCF as defined in Section 3.1.2). Participants were from four classes (A, B, C, D) and pseudo-randomly assigned to either of the two conditions. This strategy ensured a balanced distribution of participants across the four classes and mitigated potential biases." (p. 11) 2) "Participants were randomly assigned to two groups receiving feedback either during or after the conversation." (Abstract, p. 1) 3) "We designed our chatbot to simulate written conversation practices between an English instructor (chatbot) and an English as a Foreign Language (EFL) learner (user), where the teacher corrects the user's grammatical mistakes." (p. 8) 4) "The chatbot was introduced to participants as 'Alex' without specifying gender or age, and positioned as a native English-speaking tutor." (p. 11) 5) "Additionally, all sessions were conducted outside of the school environment, allowing participants to work at their own pace." (p. 12) Detailed Analysis: The unit of randomisation was the individual student, not the class or school: students drawn from four classes were pseudo-randomly assigned to the Immediate or Delayed feedback condition. Ordinarily this would fail criterion C. However, the ERCT standard provides an exception: if the intervention is designed for personal teaching like tutoring, student-level randomisation is acceptable. Here the intervention is inherently one-to-one: each learner individually converses with a chatbot that acts as a personal native English-speaking tutor, and all sessions took place individually outside the school environment. The feedback (immediate or delayed) is delivered privately by the system to each learner within their own sessions, so cross-group contamination in the classroom-instruction sense is minimal. The randomisation process (pseudo-random assignment balanced across the four classes) is described. Criterion C is met because the intervention is one-to-one chatbot tutoring, which qualifies for the personal-teaching exception allowing student-level randomisation.
    • E

      Exam-based Assessment

      • Outcomes were measured with a custom assessment built from 25 sentences taken from each participant's own chat utterances, not a standardised, widely recognised exam.
      • "The participants were given 25 sentences extracted from their utterances and were instructed to identify and correct any mistakes they found."
      • Relevant Quotes: 1) "The duration of the experiment may have been insufficient for participants to apply their learned knowledge to the entire English grammar. Consequently, a comprehensive language test was not administered before and after the experiment. Instead, participants' own mistakes served as a preliminary assessment." (p. 11) 2) "The participants were given 25 sentences extracted from their utterances and were instructed to identify and correct any mistakes they found. Of the 25 sentences provided, roughly 15 included errors made by the participant, while the remaining sentences were correct. Participants earned one point for each mistake they successfully corrected." (p. 11) 3) "During the experiment, we recorded the information about all the grammatical mistakes that the users made including the wrong word, whether the correction was provided, and the type of error. This is then used as a measure of language learning outcomes in the analysis." (p. 12) 4) "Prior to the experiment, participants self-assessed their English proficiency levels as A2 (n = 1), B1 (n = 14), B2 (n = 26), C1 (n = 11), and C2 (n = 2)." (p. 11) Detailed Analysis: Criterion E requires that outcomes be measured with a standard, widely recognised standardised exam rather than an instrument created for the study. Here the learning outcome measures were (a) a bespoke error-correction task assembled from each participant's own chat utterances and (b) counts of grammatical errors logged during the chat sessions, analysed with a GLMM. Both are researcher-constructed, study-specific measures; the authors explicitly state that a comprehensive language test was not administered. English proficiency levels (CEFR) were only self-assessed for description, not measured by a standardised exam. Criterion E is not met because outcomes were assessed with a custom study-specific instrument, not a recognised standardised exam.
    • T

      Term Duration

      • The intervention comprised 12 chatbot sessions over roughly four weeks with outcomes measured immediately after the last session, which is far shorter than one academic term.
      • "Participants completed 12 sessions in total, conducting three per week for 4 weeks"
      • Relevant Quotes: 1) "In this study, we present an in-the-wild experiment where 66 students interacted with an LLM-powered chatbot over 12 sessions distributed over a one-month period." (p. 3) 2) "Participants completed 12 sessions in total, conducting three per week for 4 weeks (the timeline and sessions shown in the dashboard to the users can be seen in Figure 4)." (p. 12) 3) "After finishing their final chat session, participants were directed to the language assessment page." (p. 12) 4) Figure 4 shows the session timeline running from [2022-11-21] to [2022-12-18]. (p. 13) 5) "Additionally, although the experiment lasted for a month, this duration and the type of activity may not be sufficient to generalize the learning outcomes across all aspects of language acquisition." (p. 17) Detailed Analysis: Criterion T requires the interval from intervention start to outcome measurement to be at least one full academic term (approximately 3-4 months). The intervention started on 21 November 2022 and the final sessions ended on 18 December 2022, with the language assessment administered immediately after the final chat session. The full start-to-measurement interval is therefore about four weeks (one month). Although the abstract describes integration "over one semester," the methodology unambiguously documents 12 sessions across 4 weeks with measurement right at the end. One month is well short of a term. Criterion T is not met because outcomes were measured about four weeks after the intervention began, far less than one academic term.
    • D

      Documented Control Group

      • The study has no control group by design, and the two feedback-timing arms lack documented per-group demographics and baseline performance, so no properly documented comparison group exists.
      • "Furthermore, the absence of a control group, dictated by our study's focus on comparing ICF and DCF conditions, constrains the breadth of our findings."
      • Relevant Quotes: 1) "Furthermore, the absence of a control group, dictated by our study's focus on comparing ICF and DCF conditions, constrains the breadth of our findings." (p. 17) 2) "Among the analyzed participants, 24 were assigned to the Immediate Corrective Feedback (ICF) condition and 30 to the Delayed Corrective Feedback (DCF) condition; this imbalance resulted from differential attrition (see Section 5.5)." (p. 10) 3) "The analyzed sample comprised 43 males and 11 females, with ages ranging from 18 to 23 years (x = 20.59, s = 0.98). All participants were native Czech speakers, apart from one Chinese speaker." (pp. 10-11) 4) "Figure 3 shows the balanced distribution of participants by English level across conditions." (p. 11) 5) "Primarily, our approach to measuring language learning gains, specifically the reliance on participants' grammatical mistakes without a pre-test for comparison, presents challenges in conclusively capturing learning gains." (p. 17) Detailed Analysis: Criterion D requires a well-documented control group, including demographics, baseline performance, and treatments received. This study deliberately has no untreated control group; both arms received the chatbot intervention and differ only in feedback timing. Even treating the DCF arm as the comparison group, documentation is incomplete: demographics (gender, age, L1) are reported only for the pooled sample, not per condition; there was no baseline language pre-test, so baseline performance of either group is undocumented (only self-assessed CEFR levels are shown by condition in Figure 3); and the authors themselves flag the absence of a control group and of a pre-test as limitations. Criterion D is not met because no control group exists and per-group demographic and baseline performance documentation is lacking.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred at the individual student level within four classes at a single institution, not at the school level.
      • "Participants were from four classes (A, B, C, D) and pseudo-randomly assigned to either of the two conditions."
      • Relevant Quotes: 1) "Participants were from four classes (A, B, C, D) and pseudo-randomly assigned to either of the two conditions." (p. 11) 2) "The study enrolled 66 bachelor students from five diverse fields: Computer Science, Economics and Management, Applied Informatics, Informatics and Management, and Information Technology." (p. 10) 3) "Participants were randomly assigned to two groups receiving feedback either during or after the conversation." (Abstract, p. 1) Detailed Analysis: Criterion S requires randomisation among schools or equivalent implementing institutions. Here all participants were bachelor students drawn from four classes in what appears to be a single institution, and assignment to the ICF or DCF condition was made at the individual student level. No schools, sites, or centers were randomised. The tutoring exception that satisfies criterion C does not extend to the stronger school-level requirement of criterion S. Criterion S is not met because randomisation was at the student level within one institution, not among schools.
    • I

      Independent Conduct

      • The same research team designed, built, administered, and analysed the chatbot intervention with no external or third-party evaluation.
      • "AM: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing."
      • Relevant Quotes: 1) "We developed a web-application chatbot utilizing the OpenAI GPT-3 (Brown et al., 2020) API, implemented in Python 3.8 and JavaScript, deployed on our private server via Python Flask, and with data stored in a database." (p. 8) 2) "AM: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing. BT: Data curation, Formal analysis, Methodology, Software, Writing – original draft, Writing – review & editing. ML: Software, Writing – original draft. GS: Funding acquisition, Supervision, Writing – review & editing." (p. 18) 3) "The experimenter explained the form's questions and remained available to answer any clarifying questions." (p. 12) 4) "The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest." (p. 18) Detailed Analysis: Criterion I requires the study to be conducted independently from the designers of the intervention, or at minimum with documented third-party oversight of data collection and analysis. Here the authors developed the chatbot system themselves, deployed it on their own server, administered the experiment in the classes, collected the data, and performed the formal analysis, as the author contribution statement makes explicit. No external evaluation agency, independent enumerators, or third-party oversight is mentioned anywhere in the paper. The conflict-of-interest declaration addresses financial ties, not evaluator independence from the intervention design. Criterion I is not met because the intervention designers themselves conducted the data collection and analysis without independent oversight.
    • Y

      Year Duration

      • The whole study, from first session to final assessment, lasted about one month, far below 75% of an academic year.
      • "Participants completed 12 sessions in total, conducting three per week for 4 weeks"
      • Relevant Quotes: 1) "66 students interacted with an LLM-powered chatbot over 12 sessions distributed over a one-month period." (p. 3) 2) "Participants completed 12 sessions in total, conducting three per week for 4 weeks." (p. 12) 3) Figure 4 shows sessions scheduled from [2022-11-21] through [2022-12-18]. (p. 13) 4) "After finishing their final chat session, participants were directed to the language assessment page." (p. 12) Detailed Analysis: Criterion Y requires outcomes to be measured at least 75% of an academic year (roughly 9-10 months, so about 7+ months) after the intervention begins. The intervention and measurement here span approximately four weeks in November-December 2022, with the post-assessment immediately after the final session and no later follow-up. Since the weaker term-duration criterion T is already not met, this stronger year-duration criterion cannot be met, and the one-month tracking interval is an order of magnitude shorter than required. Criterion Y is not met because the tracking interval was about one month, far below 75% of an academic year.
    • B

      Balanced Control Group

      • Both groups received identical chatbot sessions, identical session dosage, and the same system-generated amount of feedback, differing only in feedback timing, which is the treatment variable itself.
      • "Since the same system was used for error detection, the amount of feedback generation was the same between the two conditions, with the only difference being the time when the feedback was given to the participant (immediately in the conversation, delayed in a report)."
      • Relevant Quotes: 1) "The corrections in the DCF are produced with the same algorithm as ICF and saved in the background. At the end of the chat session they are shown to the user as a summary report of their mistakes. Since the same system was used for error detection, the amount of feedback generation was the same between the two conditions, with the only difference being the time when the feedback was given to the participant (immediately in the conversation, delayed in a report)." (p. 8) 2) "Each chat session was considered complete once the user generated 1,000 characters to ensure consistent chat exposure across participants. Participants completed 12 sessions in total, conducting three per week for 4 weeks" (p. 12) 3) "A between-subject design was employed, comprising two conditions designed to evaluate the influence of correction timing on language learning and interactional dynamics (ICF vs. DCF as defined in Section 3.1.2)." (p. 11) Detailed Analysis: Criterion B requires the comparison groups to be balanced in time and resources unless the extra resource is itself the treatment variable. In this two-arm design there is no untreated business-as-usual group to balance against: both arms used the same chatbot, completed the same 12 sessions with the same 1,000-character completion requirement per session (ensuring equal exposure time), and the same algorithm generated the same volume of corrective feedback in both arms. The only difference is the timing of feedback delivery (during the conversation vs. in an end-of-session report), which is precisely the treatment contrast being tested. Following the decision tree, no extra time or budget was given to one arm over the other, so the balance requirement is satisfied. Criterion B is met because both conditions received identical time, materials, and feedback volume, with only the experimentally manipulated feedback timing differing.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific chatbot feedback-timing RCT by a different team has been published; related studies exist but do not reproduce this study's design, and a citation-graph search found no replication.
      • Relevant Quotes: 1) "Crucially, prior research on CF timing has focused exclusively on human tutors, yet the psychological dynamics of human-AI interaction may differ fundamentally." (p. 2) 2) "Regarding feedback timing, Qun (2025) examined immediate vs. delayed feedback in online L2 education and found that both significantly enhanced motivation and learning outcomes compared to no feedback, though no substantial differences emerged between immediate and delayed conditions - a finding that aligns with our own results." (p. 7) 3) "This study aimed to bridge the gap in language learning methodologies by leveraging the potential of LLMs for offering real-time, personalized CF, thus paving the way for pioneering the use of chatbots in language education." (p. 15) Detailed Analysis: Criterion R requires that the study be independently replicated by a different research team in a different context and published in a peer-reviewed journal. This paper was published in February 2026 and presents itself as pioneering work on immediate vs. delayed corrective feedback delivered by an LLM chatbot. The authors position the study as filling a gap, indicating no prior equivalent trials. The related study by Qun (2025) examined feedback timing in online L2 education generally, but it is not a replication of this specific LLM-chatbot in-the-wild RCT with its design, system, and population. To check for reproduction after publication, a citation search (Semantic Scholar, July 2026) for papers citing this DOI (10.3389/feduc.2026.1703664) was conducted. As of the verification date, the only citing work located is Wu (2026), "AI-Empowered English Speaking Teaching and Learning: A Scoping Review of Chinese Core Journal Studies (2022-2026)", a scoping literature review, not an empirical replication of this RCT. Given the paper was published only about five months before this check, insufficient time has elapsed for an independent peer-reviewed replication of this specific design to plausibly exist yet, and none was found. Criterion R is not met because no independent, peer-reviewed replication of this specific study exists.
    • A

      All-subject Exams

      • Only English grammar knowledge was assessed, with a custom instrument, so neither the standardised-exam prerequisite nor the all-subjects coverage is satisfied.
      • "The participants were given 25 sentences extracted from their utterances and were instructed to identify and correct any mistakes they found."
      • Relevant Quotes: 1) "This study investigates the feasibility of providing immediate feedback and examines the impact of feedback timing on user experience and grammar learning gains in English." (Abstract, p. 1) 2) "The participants were given 25 sentences extracted from their utterances and were instructed to identify and correct any mistakes they found." (p. 11) 3) "We categorized all the error types made by the participants based on the classifications specified in Bryant et al. (2017)." (p. 14) Detailed Analysis: Criterion A requires standardised exam-based assessment of all main subjects taught at the relevant educational level, and explicitly treats criterion E as a prerequisite. Since criterion E is not met (the outcome measure was a custom error-correction task), criterion A automatically fails. In addition, only one domain - English grammar - was assessed; no other subjects of the participants' bachelor programs (e.g., their core informatics or economics courses) were measured, and no specialised-intervention justification referencing standardised exams is offered. Criterion A is not met because criterion E fails and only a single subject (English grammar) was assessed with a custom instrument.
    • G

      Graduation Tracking

      • Measurement stopped immediately after the final chat session with no follow-up, and since criterion Y is not met this criterion cannot be met either.
      • "After finishing their final chat session, participants were directed to the language assessment page."
      • Relevant Quotes: 1) "After finishing their final chat session, participants were directed to the language assessment page. ... The experiment concluded once both forms were completed." (p. 12) 2) "Long-term engagement with chatbot systems could illuminate sustained language learning outcomes and user perceptions, inviting further exploration into dialogue strategies, conversation depth, and chatbot personification's impact on learning and engagement." (p. 18) Detailed Analysis: Criterion G requires following participants until graduation from their educational stage, and per the ERCT specification it cannot be met when criterion Y fails, which is the case here. The experiment ended immediately after the fourth week of sessions when participants completed the language assessment and qualitative form. The bachelor students were not tracked to the end of their degree; the authors instead list long-term engagement as future work. A citation search (Semantic Scholar, July 2026) for papers citing this DOI and for other work by the corresponding author (Kamelabad) was conducted to look for a possible follow-up publication tracking this cohort further; no such follow-up paper was found, only an unrelated scoping review by Wu (2026) cites this study. This is consistent with the short, extracurricular nature of the intervention, which was not designed as a longitudinal graduation-tracking study. Criterion G is not met because tracking ceased immediately after the one-month intervention with no graduation follow-up, and criterion Y is also not met.
    • P

      Pre-Registered

      • The paper contains no mention of a pre-registered protocol or registry entry made before data collection; the only external repository (OSF) was created in January 2023, after data collection had already finished.
      • Relevant Quotes: 1) "The datasets presented in this study can be found in online repositories. The data and analysis can be found at: OSF (https://osf.io/m9qgf/) and the code of the web-app of the project can be accessed at: GitHub." (p. 18) 2) "Ethical approval was not required for the studies involving humans because according to the local regulations of the research ethics, the type of research conducted by this study, does not require an ethical approval..." (p. 18) Detailed Analysis: Criterion P requires that the full study protocol, including hypotheses, methods, and planned analyses, be registered on a public registry before data collection began, with quoted evidence of registration and timing. The paper never mentions pre-registration, a trial registry, a registration ID, or a registration date. The OSF link is described only as a repository for data, analysis, and supplementary material shared with the publication, not as a pre-registration made before the November-December 2022 data collection. To verify this independently, the OSF project (osf.io/m9qgf) was queried directly (OSF API, July 2026): its "date_created" field is 2023-01-30, i.e. after the November-December 2022 data-collection window, and it is recorded as a regular project rather than a registration. This confirms the OSF page is a post-hoc data/code repository rather than a prospective pre-registration, corroborating the paper's own text. No pre-registration entry for this study was located on common registries. Criterion P is not met because no pre-registration of the study protocol before data collection is reported or found, and the linked OSF repository was created after data collection concluded.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.