Parent-led vs. AI-guided dialogic reading: Evidence from a randomized controlled trial in children's e-book context

Feiwen Xiao, Ellen Wenting Zou, Jiaju Lin, Zhaohui Li, Dandan Yang

Published:
ERCT Check Date:
DOI: 10.1111/bjet.13615
  • reading
  • L2 languages
  • kindergarten
  • K12
  • China
  • parent involvement
  • EdTech app
  • mobile learning
0
  • C

    Randomisation was at the individual-child level, but the intervention (reading one-on-one with either an AI agent or a parent) is inherently a personal, one-to-one activity that cannot be delivered at classroom scale, so the personal-teaching/tutoring exception to the class-level criterion applies.

    "Children participants (N=67) were randomly assigned to (1) the experimental group (CA group), where they (N=37) read an e-book story with Mia, the CA, on touchscreen tablets, and (2) the control group (N=30) (Parent group), where they read the same e-book story using touchscreen tablets but with their parents." (p. 1793)

  • E

    The study used researcher-assembled vocabulary, comprehension, and retelling tests adapted from a prior unpublished/in-house study, not a widely recognised standardised exam.

    "Adapted from (Yang et al., 2022), this assessment consisted of seven question sets to evaluate children's recall and inference of story content." (p. 1794)

  • T

    The entire intervention was a single reading session of 20-40 minutes, with the primary outcomes measured immediately afterward and again only 5 days later, far short of one academic term.

    "Each child participated in a single reading session, lasting approximately 20-40 min, depending on the child's pace." (p. 1793)

  • D

    The control (parent-led) group is clearly documented, including its size, procedure, materials, and detailed demographic comparison with the experimental group.

    "the control group (N=30) (Parent group), where they read the same e-book story using touchscreen tablets but with their parents." (p. 1793)

  • S

    The study was not conducted at the school level; it took place individually in children's homes via Zoom with no involvement of schools or classes at all.

    "All the reading sessions and assessments took place at children's homes, were administered remotely via Zoom, and were transcribed for further analysis." (p. 1794)

  • I

    The same research team that designed the Storio CA system also conducted the trial, coded engagement, and analysed results, with no mention of an independent or third-party evaluator.

    "The first and second authors reviewed video recordings and systematically annotated instances of each behaviour using a coding scheme." (p. 1796)

  • Y

    Since criterion T (Term Duration) is not met, criterion Y is automatically not met; in any case, the study tracked children for only 5 days after a single session.

    "Delay Test (5 days later)" (Figure 2, p. 1793)

  • B

    Both groups read the identical e-book for the same duration using the same tablets and received the same scripted questions at matching story time-points; the only manipulated variable was whether the reading partner was the CA or the parent, which is precisely the treatment variable being tested.

    "At the same point where Mia asked about Little Oak's appearance, parents were prompted: 'Ask your child: 'What does Little Oak look like in the water?''." (p. 1794)

  • R

    This is a newly published exploratory study; an internet search of citing literature found no evidence of independent replication by another research team.

    "This study is among the first to investigate how generative AI can be seamlessly integrated into children's daily media use and tailored to enrich their language learning experiences." (p. 1801)

  • A

    Since criterion E is not met, criterion A cannot be met; in addition, only English vocabulary, story comprehension, and retelling were assessed, not other core subjects.

    "The following research questions guide the evaluation of this application: 1. How does reading with a CA-embedded e-book impact children's English vocabulary, comprehension, and retelling?" (p. 1791)

  • G

    Since criterion Y is not met, criterion G is automatically not met; the study only tracked outcomes 5 days after a single session, and no follow-up publication tracking this cohort toward graduation was found.

    "Delay Test (5 days later)" (Figure 2, p. 1793)

  • P

    No pre-registration of the study protocol, hypotheses, or analysis plan is mentioned anywhere in the paper, and no matching registration record was found via Crossref or registry search.

Abstract

Large language model (LLM)-based conversational agents (CAs), with their advanced generative capabilities and human-like conversational interfaces, can serve as reading partners for children during dialogic reading and have shown promise in enhancing children's comprehension and conversational skills. However, there is limited research on the efficacy of LLM-based bilingual CAs in children's language acquisition in English as a Foreign Language (EFL) contexts. This randomized controlled trial study investigated the effectiveness of LLM-powered CAs compared with traditional parent-child shared reading in promoting engagement and improving learning outcomes among children with EFL. An interactive e-book featuring a LLM-powered CA was developed to engage children in dialogic reading through questioning and scaffolding. Sixty-seven children, aged 5 to 8, were randomly assigned to either an experimental (AI-led) group or a control (parent-led) group. The study found that children in the experimental group outperformed the control group in reading comprehension, with comparable benefits in vocabulary acquisition and story retelling, both immediately and in delayed tests. In the meantime, this study unpacks children's different engagement patterns when reading with the CA versus reading with their parents. Children reading with the CA demonstrated higher behavioural engagement and visual attention, while those in the parent-led group showed greater affective engagement and narrative-relevant vocalizations. The findings highlighted insights into the potential of LLM-powered CAs in children's language acquisition and suggested key design implications for developing better CAs for children from multilingual backgrounds.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was at the individual-child level, but the intervention (reading one-on-one with either an AI agent or a parent) is inherently a personal, one-to-one activity that cannot be delivered at classroom scale, so the personal-teaching/tutoring exception to the class-level criterion applies.
      • "Children participants (N=67) were randomly assigned to (1) the experimental group (CA group), where they (N=37) read an e-book story with Mia, the CA, on touchscreen tablets, and (2) the control group (N=30) (Parent group), where they read the same e-book story using touchscreen tablets but with their parents." (p. 1793)
      • Relevant Quotes: 1) "Children participants (N=67) were randomly assigned to (1) the experimental group (CA group), where they (N=37) read an e-book story with Mia, the CA, on touchscreen tablets, and (2) the control group (N=30) (Parent group), where they read the same e-book story using touchscreen tablets but with their parents." (p. 1793) 2) "Each child participated in a single reading session, lasting approximately 20-40 min, depending on the child's pace." (p. 1793) 3) "All the reading sessions and assessments took place at children's homes, were administered remotely via Zoom..." (p. 1794) Detailed Analysis: The quotes confirm randomisation occurred at the individual child level: each of the 67 children was independently assigned to either read with the CA or with their own parent. Taken at face value, this is student-level randomisation within a single sample, which the base ERCT criterion would normally consider "not met," since it raises no direct concern of cross-classroom contamination but does not meet the literal class-level requirement. However, this study is not a classroom- or curriculum-delivered intervention at all: it is a one-off, one-to-one reading session conducted separately with each child in their own home (or via Zoom), with either their own parent or an AI conversational agent as the sole partner. There is no group or class in which the intervention could plausibly be delivered collectively; the reading experience is inherently individual, one adult/AI paired with one child at a time, structurally analogous to one-to-one tutoring rather than a classroom-based educational programme. The ERCT standard explicitly carves out an exception for interventions "designed for personal teaching like tutoring," under which "even normal student-level RCT is considered OK." Because the design of this study could not have been implemented as a class-level or school-level intervention without fundamentally changing its nature (parents reading one-on-one with their own children at home), the tutoring exception is judged to apply here, and no contamination risk arises from the individual-level randomisation. Criterion C is met via the personal-teaching/tutoring exception, since the intervention is inherently a one-to-one reading activity rather than a classroom-deliverable programme.
    • E

      Exam-based Assessment

      • The study used researcher-assembled vocabulary, comprehension, and retelling tests adapted from a prior unpublished/in-house study, not a widely recognised standardised exam.
      • "Adapted from (Yang et al., 2022), this assessment consisted of seven question sets to evaluate children's recall and inference of story content." (p. 1794)
      • Relevant Quotes: 1) "Story comprehension test. Adapted from (Yang et al., 2022), this assessment consisted of seven question sets to evaluate children's recall and inference of story content." (p. 1794) 2) "English story vocabulary test. This test was used to assess children's word knowledge of the story-related vocabulary after the intervention. The vocabulary assessment consisted of two parts: Receptive Vocabulary and Expressive Vocabulary, each containing 25 items... (Author, Year)." (p. 1794) 3) "Story retelling test. This assessment aimed to evaluate the children's ability to recall and narrate the story (Yang et al., 2022)." (p. 1795) 4) "The target words... were drawn directly from the story and generally ranged from one to three syllables." (p. 1794) Detailed Analysis: All three outcome measures (story comprehension test, English story vocabulary test, story retelling test) were custom-built instruments tied specifically to the single story ("The Story of an Orange Oakleaf") used in this study, adapted from a related prior study by overlapping authors (Yang et al., 2022) rather than drawn from any nationally or internationally recognised, independently validated standardised exam (e.g., a national curriculum test or a commercially standardised language-proficiency battery). The vocabulary items were deliberately drawn from the story's own text, and the comprehension/retelling items were built around this specific narrative's plot points. While the authors report strong internal consistency and inter-rater reliability (Cronbach's alpha, Cohen's Kappa), reliability statistics do not substitute for standardisation: these are researcher-designed, study-specific instruments, precisely the kind of custom assessment the ERCT standard's E criterion is meant to flag as a risk of inflated, intervention-aligned results. Criterion E is not met because the outcome measures are custom, story-specific instruments rather than standardised, widely recognised exams.
    • T

      Term Duration

      • The entire intervention was a single reading session of 20-40 minutes, with the primary outcomes measured immediately afterward and again only 5 days later, far short of one academic term.
      • "Each child participated in a single reading session, lasting approximately 20-40 min, depending on the child's pace." (p. 1793)
      • Relevant Quotes: 1) "This study features a randomized controlled trial with repeated measures... Each child participated in a single reading session, lasting approximately 20-40 min, depending on the child's pace." (p. 1793) 2) "Children underwent a vocabulary test before the reading session, immediately after, and 5 days later, along with a Story Comprehension Test and Story Retelling Test." (p. 1794) 3) Figure 2 "RCT workflow" shows: Pre-Reading Vocabulary Test -> single reading session -> Post-Reading Test -> Delay Test (5 days later). (p. 1793) Detailed Analysis: The intervention itself is a single reading session lasting only 20 to 40 minutes. The only follow-up beyond the immediate post-test is a delayed test conducted 5 days later. There is no continued exposure or tracking across weeks or months; the entire "intervention start to final measurement" window is under a week. This is dramatically shorter than the one academic term (roughly 3-4 months) required by the ERCT T criterion, and the paper gives no indication of any longer follow-up period. Criterion T is not met because the intervention and its measurement window span only a single session plus a 5-day delay, far short of one academic term.
    • D

      Documented Control Group

      • The control (parent-led) group is clearly documented, including its size, procedure, materials, and detailed demographic comparison with the experimental group.
      • "the control group (N=30) (Parent group), where they read the same e-book story using touchscreen tablets but with their parents." (p. 1793)
      • Relevant Quotes: 1) "the control group (N=30) (Parent group), where they read the same e-book story using touchscreen tablets but with their parents." (p. 1793) 2) "In the Parent group, before reading, parents received a tutorial on using embedded prompt cues in the e-book, which reminded them to ask the same questions Mia posed at the same time point with the CA group." (p. 1794) 3) "T-tests and chi-squared analyses indicated no significant demographic differences between the groups, confirming the comparability of the final samples (Appendix E)." (p. 1796) 4) Appendix E, "Demographic characteristic per group," lists child age, gender, English level, CA use, and parent age, education, English level, and reading fluency separately for the Control group (N=30) and Experimental group (N=37). (pp. 1812-1813) Detailed Analysis: The paper provides a detailed procedural description of the control (Parent-led) condition, including what materials were used (the same e-book story on the same touchscreen tablets), the exact procedure (parents prompted to ask the same questions as Mia, at the same story time-points), and confirms statistically that demographics did not differ significantly between groups. Appendix E further breaks down detailed child and parent characteristics (age, gender, English proficiency, chatbot use, parental education and reading fluency) separately per group, allowing a reader to fully assess baseline comparability. Criterion D is met because the control group's size, procedure, materials, and baseline characteristics are thoroughly documented and shown to be comparable to the experimental group.
  • Level 2 Criteria

    • S

      School-level RCT

      • The study was not conducted at the school level; it took place individually in children's homes via Zoom with no involvement of schools or classes at all.
      • "All the reading sessions and assessments took place at children's homes, were administered remotely via Zoom, and were transcribed for further analysis." (p. 1794)
      • Relevant Quotes: 1) "All the reading sessions and assessments took place at children's homes, were administered remotely via Zoom, and were transcribed for further analysis." (p. 1794) 2) "All the child-parent dyads were recruited using snowball sampling within the researchers' social network." (p. 1795) 3) "Children participants (N=67) were randomly assigned to (1) the experimental group (CA group)... and (2) the control group (N=30) (Parent group)..." (p. 1793) Detailed Analysis: There is no school, classroom, or educational institution involved in this study's design at all. Participants were individually recruited child-parent dyads via snowball sampling, and the reading sessions were conducted remotely at each family's home. Randomisation was performed at the individual child level, not at the level of any school or educational site. Since no school-level randomisation occurred (and indeed no school is part of the study design), the stronger S criterion cannot be satisfied. Criterion S is not met because the study did not involve any school-level randomisation or school-based implementation; it was conducted individually in children's homes.
    • I

      Independent Conduct

      • The same research team that designed the Storio CA system also conducted the trial, coded engagement, and analysed results, with no mention of an independent or third-party evaluator.
      • "The first and second authors reviewed video recordings and systematically annotated instances of each behaviour using a coding scheme." (p. 1796)
      • Relevant Quotes: 1) "Storio is an e-book enhanced by a bilingual conversational agent (CA) named Mia, designed to support dialogic reading for young learners. The design is built upon extensive research..." (p. 1791-1792) 2) "To further ensure the system's usability and credibility for intervention, the two educators reviewed the final implementation and design. We also conducted a pilot test with six children." (p. 1792) 3) "The first and second authors reviewed video recordings and systematically annotated instances of each behaviour using a coding scheme." (p. 1796) 4) "Children's engagement was coded from videos of their reading sessions by two researchers." (p. 1795) Detailed Analysis: The paper describes the same research team as having designed the Storio system and CA (Mia), run the pilot testing, conducted the RCT reading sessions and interviews, and personally coded/annotated the engagement videos ("the first and second authors reviewed video recordings"). There is no statement anywhere in the methods, acknowledgments, or funding sections indicating that an independent or third-party organisation collected data, administered assessments, or performed the analysis independently of the intervention's designers. The "two educators" mentioned reviewed only the CA's design/ usability, not the trial's conduct or analysis. This concentration of design, delivery, coding, and analysis within the same author team is exactly the scenario the I criterion is meant to flag. Criterion I is not met because the same team that designed the CA intervention also conducted the trial, coded the engagement data, and analysed the results, with no independent evaluator involved.
    • Y

      Year Duration

      • Since criterion T (Term Duration) is not met, criterion Y is automatically not met; in any case, the study tracked children for only 5 days after a single session.
      • "Delay Test (5 days later)" (Figure 2, p. 1793)
      • Relevant Quotes: 1) "Each child participated in a single reading session, lasting approximately 20-40 min, depending on the child's pace." (p. 1793) 2) Figure 2, "RCT workflow": Post-Reading Test immediately after the session, then "Delay Test (5 days later)". (p. 1793) Detailed Analysis: Per the ERCT standard's own dependency rule, since criterion T (Term Duration) is not met, criterion Y (Year Duration) cannot be met either. Independently, the total measured interval from intervention start to final measurement is a single reading session plus a 5-day delay, nowhere near 75% of an academic year (approximately 9-10 months). Criterion Y is not met, both because the weaker T criterion is not met and because the study's total tracking window (a single session plus 5 days) is far shorter than the required academic-year threshold.
    • B

      Balanced Control Group

      • Both groups read the identical e-book for the same duration using the same tablets and received the same scripted questions at matching story time-points; the only manipulated variable was whether the reading partner was the CA or the parent, which is precisely the treatment variable being tested.
      • "At the same point where Mia asked about Little Oak's appearance, parents were prompted: 'Ask your child: 'What does Little Oak look like in the water?''." (p. 1794)
      • Relevant Quotes: 1) "the control group (N=30) (Parent group), where they read the same e-book story using touchscreen tablets but with their parents." (p. 1793) 2) "In the Parent group, before reading, parents received a tutorial on using embedded prompt cues in the e-book, which reminded them to ask the same questions Mia posed at the same time point with the CA group. At the same point where Mia asked about Little Oak's appearance, parents were prompted: 'Ask your child: 'What does Little Oak look like in the water?''." (p. 1794) 3) "In the CA group, children completed a tutorial on e-book navigation and practised interacting with Mia for 5 min before reading the e-book." (p. 1793) 4) "Each child participated in a single reading session, lasting approximately 20-40 min, depending on the child's pace." (p. 1793) Detailed Analysis: Applying the ERCT decision procedure for criterion B: the study's explicit research goal is to compare an AI conversational-agent reading partner against a parent-led reading partner using the very same story, the very same 10 dialogic questions at matching story time-points, and the same tablets and duration for both groups. The only substantive difference in resources is that the CA group received a short (5-minute) tutorial on interacting with Mia, while the Parent group's parents received a tutorial on using the embedded prompt cues; these are functionally equivalent, matched onboarding steps rather than an imbalance favouring one condition. No extra instructional time, materials, or budget was given to either group beyond what is intrinsic to having a CA versus a parent as the reading partner. Because the "additional resource" here (an LLM-powered conversational agent) is precisely the treatment variable under investigation, and the control condition (parent-led reading) represents the natural "business-as-usual" comparator for dialogic reading, with matched time, materials, and script, criterion B is satisfied under the standard's treatment-as-resource exception, and no separate resource imbalance is introduced beyond that treatment contrast. Criterion B is met because both conditions received matched time, materials, and story content, and the only substantive difference (CA vs. parent as reading partner) is the explicit treatment variable being tested.
  • Level 3 Criteria

    • R

      Reproduced

      • This is a newly published exploratory study; an internet search of citing literature found no evidence of independent replication by another research team.
      • "This study is among the first to investigate how generative AI can be seamlessly integrated into children's daily media use and tailored to enrich their language learning experiences." (p. 1801)
      • Relevant Quotes: 1) "This study presents a novel generative AI-powered CA integrated into children's e-book reading experience and evaluates its effectiveness in children's language development compared with parent-guided reading in the EFL context. This study is among the first to investigate how generative AI can be seamlessly integrated into children's daily media use and tailored to enrich their language learning experiences. Additionally, it is the first to examine the impact of bilingual CAs on children's bilingual literacy development in a multilingual context." (p. 1801) 2) "This exploratory RCT offers additional evidence that contributes to prior research." (p. 1801) Detailed Analysis: The authors themselves explicitly frame this as an exploratory, first-of-its-kind study ("among the first," "the first to examine"). Published in mid-2025, an internet search (via OpenAlex/Semantic Scholar citation records, 16 citing works as of July 2026) was conducted to check for independent replications. All 16 citing papers found (e.g., studies on AI-dialogic scaffolding in Saudi EFL contexts, AI-mediated paired reading in Swedish classrooms, AI-assisted dialogic reading for bilingual vocabulary, and conversational AI in children's home literacy learning, all 2025-2026) discuss, build on, or apply ideas from this study, but none reproduce this specific Storio/Mia bilingual e-book design with a parent-led comparison condition to test whether its results replicate. The cited prior literature (Xu, Vigil, et al., 2022; Cheng et al., 2024) investigates related but distinct CA-reading designs (e.g., scripted human experimenters, different populations, monolingual CAs), not independent replications of this specific bilingual, LLM-powered Storio/Mia design and its parent-led comparison. Criterion R is not met because no independent replication of this specific study exists; the paper is itself a novel, first exploratory trial, and no reproduction has appeared in the citing literature as of this check.
    • A

      All-subject Exams

      • Since criterion E is not met, criterion A cannot be met; in addition, only English vocabulary, story comprehension, and retelling were assessed, not other core subjects.
      • "The following research questions guide the evaluation of this application: 1. How does reading with a CA-embedded e-book impact children's English vocabulary, comprehension, and retelling?" (p. 1791)
      • Relevant Quotes: 1) "The following research questions guide the evaluation of this application: 1. How does reading with a CA-embedded e-book impact children's English vocabulary, comprehension, and retelling? How does it compare to reading the e-book with a parent?" (p. 1791) 2) "Story comprehension test," "English story vocabulary test," and "Story retelling test" are the only three named outcome measures reported in the Measures section. (pp. 1794-1795) Detailed Analysis: Per the ERCT standard, criterion A requires that criterion E (standardised exam-based assessment) be satisfied as a prerequisite; since E was not met (the outcome measures were custom, story-specific instruments), A cannot be met either. Independently, the study's outcome measures are confined entirely to English-language literacy skills tied to a single story (vocabulary, comprehension, retelling); no other core school subjects (e.g., mathematics, science, social studies) were assessed, nor is there any stated rationale for a specialised, single-subject focus of the kind the ERCT exception allows for upper-secondary/ vocational contexts. Criterion A is not met, both because criterion E is not satisfied and because only English-language outcomes were measured, with no other core subjects assessed.
    • G

      Graduation Tracking

      • Since criterion Y is not met, criterion G is automatically not met; the study only tracked outcomes 5 days after a single session, and no follow-up publication tracking this cohort toward graduation was found.
      • "Delay Test (5 days later)" (Figure 2, p. 1793)
      • Relevant Quotes: 1) Figure 2, "RCT workflow," shows the final data collection point as "Delay Test (5 days later)." (p. 1793) 2) "Future research should investigate performance across multiple sessions and replicate the study with a longer intervention to determine whether the benefits endure once the novelty has worn off." (p. 1804) Detailed Analysis: Per the ERCT standard's dependency rule, since criterion Y (Year Duration) is not met, criterion G (Graduation Tracking) is automatically not met. Substantively, the paper's own limitations section explicitly recommends future work with "a longer intervention" and tracking "across multiple sessions," implicitly acknowledging that this study did not follow children over any extended period, let alone until graduation from their current educational stage. An internet search of the citing literature (16 citing works as of July 2026, via OpenAlex/Semantic Scholar) found no follow-up publication by Xiao, Zou, Lin, Li, or Yang tracking this same cohort of 67 children, let alone until graduation. Criterion G is not met because the weaker Y criterion is not met, and the study's tracking ended just 5 days after a single reading session with no long-term or graduation follow-up found anywhere.
    • P

      Pre-Registered

      • No pre-registration of the study protocol, hypotheses, or analysis plan is mentioned anywhere in the paper, and no matching registration record was found via Crossref or registry search.
      • Relevant Quotes: 1) "FUNDING INFORMATION: This research project did not receive any funding." (p. 1805) 2) "CONFLICT OF INTEREST STATEMENT: The authors declare no conflicts of interest." (p. 1805) 3) "DATA AVAILABILITY STATEMENT: The data used in this study will not be shared to public." (p. 1805) 4) "ETHICS APPROVAL STATEMENT: This study was approved by the Institutional Review Board (IRB) of Penn State University." (p. 1805) Detailed Analysis: The paper's end-matter includes funding, conflict of interest, data availability, and ethics approval statements, but at no point does it mention a pre-registration platform (e.g., OSF, AsPredicted, ClinicalTrials.gov), a pre-registration ID, or a pre-registration date. There is no reference anywhere in the methods or elsewhere to hypotheses, methods, or analysis plans having been published before data collection began. A Crossref metadata check for this DOI returned no clinical-trial-number or registration field, and no matching registration record was found in a registry search. The absence of any such statement, in a paper that is otherwise careful to document ethics approval and conflicts of interest, indicates the study protocol was not pre-registered. Criterion P is not met because no pre-registration statement, registry ID, or date is provided anywhere in the paper, and no independent registration record was located online.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.