Improving Outcomes for English Learners Through Technology: A Randomized Controlled Trial

David Harper, Anita R. Bowles, Lauren Amer, Nick B. Pandža, Jared A. Linck

Published:
ERCT Check Date:
DOI: 10.1177/23328584211025528
  • L2 languages
  • K12
  • US
  • blended learning
  • EdTech platform
  • digital assessment
1
  • C

    Randomisation was conducted at the school level, which is stronger than and automatically satisfies the class-level requirement.

    "Random assignment was done at the school level."

  • E

    The study used the TELL, a standardised, externally validated Pearson assessment aligned to multiple states' language standards.

    "we evaluated the software's effect on student achievement as measured by the Pearson Test of English Language Learning (TELL; Bonk, 2016)."

  • T

    Outcomes were measured about 8 months after the intervention began, far exceeding one academic term.

    "Pretesting was conducted in late August and early September 2017, while posttesting was completed in early May 2018."

  • D

    The control group's curriculum, size, and demographic characteristics are clearly documented in the text and Table 1.

    "Control students continued with the district's standard English curriculum, which consisted of vocabulary development protocols..."

  • S

    Randomisation was performed at the level of whole schools (four vs. four), satisfying the school-level RCT requirement.

    "Random assignment was done at the school level."

  • I

    Rosetta Stone employees (the intervention's designer and funder) are co-authors and remained involved throughout the study, and the "external" consultants who ran the analysis were hired directly by Rosetta Stone rather than by an independent body.

    "Three of the authors work for the funding company, Rosetta Stone. The other two authors were hired by the company from the University of Maryland as external consultants for this research project."

  • Y

    Outcomes were tracked from late August 2017 to early May 2018, covering the substantial majority of the full 2017-2018 academic year.

    "This RCT evaluated the effectiveness of a software intervention, Rosetta Stone Foundations, for ELs in Grades 6 to 8 over the course of one school year."

  • B

    Both groups received the same total daily ESL instructional time; the software replaced existing curriculum content rather than adding extra resources to the treatment group.

    "Taking into account that treatment and control groups had nearly identical attendance rates and received the same amount of daily ESL instruction, we believe that the students received equal amounts of ESL instruction over the course of the year."

  • R

    No independent, peer-reviewed replication of this specific study was found in a citation search; the authors describe replication as a direction for future research.

  • A

    Only English language proficiency (TELL) was assessed; no other main subjects (e.g., math, science) were measured, and no vocational/upper secondary exception applies.

  • G

    Data collection ended with the end-of-year posttest; no tracking toward graduation was reported in the paper, and no follow-up publication by the same authors on this cohort was found via internet search.

  • P

    The paper contains no reference to a pre-registered protocol, registry ID, or registration date, and none was found via internet search.

Abstract

English learners (ELs) in K-12 schools must acquire English while simultaneously mastering content knowledge. Educational technology may support students' learning through the affordance of individualized language practice. The current randomized controlled trial intervention study examined the effects of Rosetta Stone Foundations software on English learning among middle school ELs. The study took place in Grades 6 to 8 of an urban U.S. school district (N = 221). Predictors of interest included time of testing (pretest vs. posttest) and software usage, and covariates included grade level, sex, and attendance. Additionally, socioeconomic status and home language were accounted for due to sample homogeneity. Multilevel models indicated that treatment group students showed larger gains than control group students on oral/aural outcomes. These results indicate that the software intervention enables individualized practice that can produce proficiency-related gains over and above the typical classroom curriculum.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was conducted at the school level, which is stronger than and automatically satisfies the class-level requirement.
      • "Random assignment was done at the school level."
      • Relevant Quotes: 1) "Eight public schools in a large urban school district in Arizona participated in the study during the 2017–2018 school year. Random assignment was done at the school level." (p. 2) 2) "Prior to assigning groups to condition, schools were randomly assigned to two groups... From these two groups, one group of four schools was randomly assigned to the treatment group, and the other group of four schools was assigned to the control group." (p. 2) 3) "One variable considered for random assignment was school type, as there were four middle schools and four K–8 schools. To create two groups, two schools of each type... were assigned to a group..." (p. 2) Detailed Analysis: The paper explicitly states that randomisation was conducted at the school level, with entire schools (stratified by school type) assigned to either the treatment or control condition. Since the ERCT Standard treats school-level RCTs as a stronger form of randomisation that automatically satisfies the weaker class-level requirement, criterion C is met by virtue of the stronger S criterion being met. Criterion C is met because randomisation occurred at the school level, which exceeds the class-level requirement.
    • E

      Exam-based Assessment

      • The study used the TELL, a standardised, externally validated Pearson assessment aligned to multiple states' language standards.
      • "we evaluated the software's effect on student achievement as measured by the Pearson Test of English Language Learning (TELL; Bonk, 2016)."
      • Relevant Quotes: 1) "we evaluated the software's effect on student achievement as measured by the Pearson Test of English Language Learning (TELL; Bonk, 2016)." (p. 2) 2) "The TELL is aligned to state standards on English language development... and has three different test types." (p. 2) 3) "external studies have confirmed that the TELL aligns closely with English Language Development Standards in Arizona (Stevens et al., 2015b), California (Stevens et al., 2015a), Texas (Frantz & Bailey, 2016), and the World-Class Instructional Design and Assessment (WIDA) Consortium (Stevens, 2015)..." (p. 7) 4) "Correlations were .70 or greater for the four language skill domain scores and the overall score, indicating good reliability..." (p. 7) Detailed Analysis: The study used the TELL, a commercially published, norm-referenced diagnostic test from Pearson that is explicitly validated and cross-walked against multiple U.S. states' official English Language Development standards and the WIDA consortium standards. It was not created by the study authors for this study; it is an independently developed, widely used, standardised assessment with documented reliability and validity evidence. Criterion E is met because the outcome measure (TELL) is a widely recognised, externally validated standardised assessment, not a custom-built test.
    • T

      Term Duration

      • Outcomes were measured about 8 months after the intervention began, far exceeding one academic term.
      • "Pretesting was conducted in late August and early September 2017, while posttesting was completed in early May 2018."
      • Relevant Quotes: 1) "Pretesting was conducted in late August and early September 2017, while posttesting was completed in early May 2018." (p. 7) 2) "Eight public schools... participated in the study during the 2017–2018 school year." (p. 2) Detailed Analysis: The intervention and measurement period ran from late August/early September 2017 (pretest) to early May 2018 (posttest), an interval of roughly 8 months. This far exceeds the minimum one academic term (roughly 3-4 months) required by the ERCT Standard. Because the stronger Y (Year Duration) criterion is also met (see below), T is automatically satisfied as well. Criterion T is met because the interval between the start of the intervention and the outcome measurement spans approximately 8 months, well beyond a single academic term.
    • D

      Documented Control Group

      • The control group's curriculum, size, and demographic characteristics are clearly documented in the text and Table 1.
      • "Control students continued with the district's standard English curriculum, which consisted of vocabulary development protocols..."
      • Relevant Quotes: 1) "Control students continued with the district's standard English curriculum, which consisted of vocabulary development protocols such as those by Marzano and Pickering (2005) and Frayer et al. (1969), as well as exercises inspired by Kagan and Kagan (2009)." (p. 7) 2) "Achieve 3000... a software program focusing on literacy, was also commonly used in both treatment and control classrooms. In an end-of-year survey of teachers and paraprofessionals at all school sites, 6/8 (75%) respondents from the control group mentioned Achieve 3000..." (p. 7) 3) "Table 1. Number of Students Included in Analysis, by Grade, Gender, and Condition" showing Control totals of 33, 16, 23 students across Grades 6-8 (Table 1, p. 4) 4) "students were demographically similar across groups (see Table 1)." (p. 4) Detailed Analysis: The paper provides a clear description of what the control group received (standard district curriculum with named vocabulary and cooperative- learning protocols, plus the commonly used Achieve 3000 literacy software), along with baseline demographic breakdowns by grade and gender and a statement of demographic similarity between conditions. This level of detail allows readers to assess the comparability of intervention and control groups. Criterion D is met because the control group's composition, size, and instructional conditions are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was performed at the level of whole schools (four vs. four), satisfying the school-level RCT requirement.
      • "Random assignment was done at the school level."
      • Relevant Quotes: 1) "Random assignment was done at the school level." (p. 2) 2) "Eight public schools in a large urban school district in Arizona participated in the study during the 2017–2018 school year." (p. 2) 3) "From these two groups, one group of four schools was randomly assigned to the treatment group, and the other group of four schools was assigned to the control group." (p. 2) Detailed Analysis: The unit of randomisation was explicitly the school (a public K-8 or middle school), not individual classes or students. Eight schools were stratified by type and split into two groups of four, one group randomly assigned to treatment and the other to control. This directly satisfies the ERCT Standard's definition of a school-level RCT. Criterion S is met because whole schools, not classes or students, were the unit of randomisation.
    • I

      Independent Conduct

      • Rosetta Stone employees (the intervention's designer and funder) are co-authors and remained involved throughout the study, and the "external" consultants who ran the analysis were hired directly by Rosetta Stone rather than by an independent body.
      • "Three of the authors work for the funding company, Rosetta Stone. The other two authors were hired by the company from the University of Maryland as external consultants for this research project."
      • Relevant Quotes: 1) "Three of the authors work for the funding company, Rosetta Stone. The other two authors were hired by the company from the University of Maryland as external consultants for this research project." (Acknowledgments, p. 18) 2) "The external consultants developed the analysis plan, performed the analysis, and were responsible for writing the results. The consultants verified the claims made in the discussion as well as the overall article." (p. 18) 3) "District personnel recruited retired teachers to administer both the pretest and posttest. The retired teachers were blind to group assignment." (p. 7) 4) "Researchers monitored usage throughout the school year and responded to requests from teachers, paraprofessionals, or instructional support specialists... These interactions were similar to those carried out by a client manager for a normal client." (p. 8) Detailed Analysis: Criterion I requires that the evaluation be conducted independently from the designers/ providers of the intervention. Here, the intervention (Rosetta Stone Foundations software) was designed, funded, and marketed by Rosetta Stone, and three of the five authors are Rosetta Stone employees who remain full co-authors of the study. The two "external consultants" who performed the analysis and wrote up the results were not an independent third party (such as a government body or an unaffiliated evaluation firm) but were themselves hired directly by Rosetta Stone specifically for this project, which still leaves the intervention's own funder/designer in control of who conducts the evaluation. Test administration itself was independent (retired teachers recruited by the school district, blind to condition), which is a positive sign for measurement integrity, but this does not extend to the overall research team, since Rosetta Stone staff actively monitored software usage and fielded requests throughout the year in a role the paper itself likens to a "client manager," and remained co-authors responsible for the surrounding narrative. This falls short of the kind of independence documented in stronger cases (e.g., government-led trials or studies where the provider is shown to have no role at all in personnel involved in the study), because the developer/funder's own staff were embedded in the research team rather than excluded from it. Criterion I is not met because the intervention's designer and funder (Rosetta Stone) had its own employees as co-authors and actively involved in the study, and the "external consultants" who performed the analysis were themselves engaged directly by that same company rather than by an independent third party.
    • Y

      Year Duration

      • Outcomes were tracked from late August 2017 to early May 2018, covering the substantial majority of the full 2017-2018 academic year.
      • "This RCT evaluated the effectiveness of a software intervention, Rosetta Stone Foundations, for ELs in Grades 6 to 8 over the course of one school year."
      • Relevant Quotes: 1) "Eight public schools... participated in the study during the 2017–2018 school year." (p. 2) 2) "Pretesting was conducted in late August and early September 2017, while posttesting was completed in early May 2018." (p. 7) 3) "This RCT evaluated the effectiveness of a software intervention, Rosetta Stone Foundations, for ELs in Grades 6 to 8 over the course of one school year." (p. 14, Discussion) 4) "the results provide strong evidence that educational technology can drive positive gains in second language oral and aural proficiency outcomes over even a single academic year." (p. 15) Detailed Analysis: The study tracked the full 2017-2018 school year, with pretesting in late August/early September 2017 and posttesting in early May 2018, an interval of roughly 8 to 8.5 months. Given a typical U.S. academic year of about 9-10 months, this covers well over 75% of a full academic year, satisfying the ERCT Standard's Year Duration threshold. The authors themselves frame the study as spanning "one school year." Criterion Y is met because outcomes were tracked across essentially the full 2017-2018 academic year, comfortably exceeding the 75% threshold.
    • B

      Balanced Control Group

      • Both groups received the same total daily ESL instructional time; the software replaced existing curriculum content rather than adding extra resources to the treatment group.
      • "Taking into account that treatment and control groups had nearly identical attendance rates and received the same amount of daily ESL instruction, we believe that the students received equal amounts of ESL instruction over the course of the year."
      • Relevant Quotes: 1) "In Arizona, in the 2017–2018 school year, ELs received 4 hours of daily instruction that were divided into four 1-hour blocks... For the treatment group, the software was incorporated into the oral English conversation portion of the state's English language proficiency requirements. Control students continued with the district's standard English curriculum..." (p. 7) 2) "Taking into account that treatment and control groups had nearly identical attendance rates and received the same amount of daily ESL instruction, we believe that the students received equal amounts of ESL instruction over the course of the year." (p. 7) 3) "Achieve 3000... was also commonly used in both treatment and control classrooms." (p. 7) 4) "The primary difference between groups is the usage of Rosetta Stone Foundations in the treatment group's oral English conversation block." (p. 7) Detailed Analysis: Re-applying the current criterion B decision procedure (checking for extra time/budget first, then integral-vs-non-integral resources): the software was substituted into an existing instructional block (the oral English conversation hour) that both groups already had, rather than adding extra instructional time or budget to the treatment group. Both groups received the same total 4 hours of daily ESL instruction, had nearly identical attendance, and Achieve 3000 (a supplementary literacy tool) was used similarly in both conditions. Since EXTRA_RESOURCES_PRESENT is false (no additional time or budget was documented for the treatment group beyond replacing existing curriculum content in an equivalent time slot), the decision tree resolves to "met" at the first branch without needing to reach the integral-resource sub-steps. Criterion B is met because the treatment group did not receive additional instructional time or budget relative to the control group; the software simply replaced business-as-usual content within an existing, equally sized instructional block.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent, peer-reviewed replication of this specific study was found in a citation search; the authors describe replication as a direction for future research.
      • Relevant Quotes: 1) "Future research should seek to replicate the current study in other states and school districts that have different curricula and less homogenous EL populations." (p. 15) 2) No quote in the paper references any prior or contemporaneous independent replication of this specific study's design or findings. Detailed Analysis: The authors explicitly frame independent replication as a direction for future research rather than something already accomplished. An internet search (OpenAlex and Semantic Scholar citation records for DOI 10.1177/23328584211025528, OpenAlex work W3175329995) identified approximately 14 works citing this study, including "Trends and challenges of implementing randomized controlled trials in English language education: a systematic review" (Sijali, Poudel, & Dahal, 2026), "Exploring the impact of a CALL tool for emergent bilinguals" (Feroce, Liu, & Chattergoon, 2025), "Mobile language app learners' self-efficacy increases after using generative AI" (Kittredge et al., 2025), and several unrelated studies on augmented reality for learning disabilities and academic-performance modeling. None of these constitute an independent, peer- reviewed replication of this specific Rosetta Stone Foundations RCT (school-level randomisation, TELL outcome measure, middle-school EL population in an Arizona district): the citing works are either systematic reviews mentioning this study, or studies of different tools, different populations, or unrelated topics that merely cite it. Criterion R is not met because no independent, peer-reviewed replication of this specific study was identified via citation search, and the authors themselves describe replication as future work.
    • A

      All-subject Exams

      • Only English language proficiency (TELL) was assessed; no other main subjects (e.g., math, science) were measured, and no vocational/upper secondary exception applies.
      • Relevant Quotes: 1) "we evaluated the software's effect on student achievement as measured by the Pearson Test of English Language Learning (TELL; Bonk, 2016)." (p. 2) 2) "For the sixth- to eighth-grade band, the TELL calculates an overall score, four domain scores [Listening, Speaking, Reading, Writing], and six subskill scores..." (p. 6) 3) No quotes anywhere in the paper describe assessment of mathematics, science, social studies, or any subject outside of English language proficiency. Detailed Analysis: The study exclusively measures English language proficiency (listening, speaking, reading, writing, and related subskills) via the TELL. It does not assess outcomes in other main subjects such as mathematics, science, or social studies for this general middle-school EL population. While the intervention is specifically an English-language instructional tool, the ERCT Standard's exception for narrow, single-subject assessment is intended for "highly specialised interventions in upper secondary or vocational education," which does not describe this middle-school (Grades 6-8), general EL population context. No rationale is given in the paper for why other core subjects were not assessed to check for possible negative spillover effects (e.g., on math or science instructional time). Criterion A is not met because only English language proficiency was assessed, with no measurement of other main subjects and no qualifying exception rationale provided.
    • G

      Graduation Tracking

      • Data collection ended with the end-of-year posttest; no tracking toward graduation was reported in the paper, and no follow-up publication by the same authors on this cohort was found via internet search.
      • Relevant Quotes: 1) "Pretesting was conducted in late August and early September 2017, while posttesting was completed in early May 2018." (p. 7) 2) "Future research should seek to replicate the current study in other states and school districts that have different curricula and less homogenous EL populations. Additionally, further study of software interventions that focus on oral/aural English skills is needed." (p. 15) 3) No quotes describe any follow-up of participants beyond the end of the 2017-2018 school year, nor any tracking through to graduation from middle school or high school. Detailed Analysis: The study's data collection ends with the May 2018 posttest at the conclusion of the single school year under investigation. The authors call for future replication and further study rather than describing any planned or completed longer-term follow-up of this cohort. An internet search for subsequent papers by the same authors was conducted (OpenAlex author record for Anita R. Bowles, OpenAlex ID A5039531564, 28 listed works; also checked citation records for David Harper and the paper's Semantic Scholar/OpenAlex entries). No publication by any of the five authors (Harper, Bowles, Amer, Pandža, or Linck) tracking this same Arizona EL cohort beyond the 2017-2018 posttest, or reporting graduation outcomes for these students, was found. Bowles's most recent listed works after 2021 are unrelated to this cohort (e.g., mobile app/AI research from 2025). Criterion G is not met because tracking stopped at the end-of-year posttest, with no follow-up toward graduation reported in the paper or identified via internet search for later publications by the same authors.
    • P

      Pre-Registered

      • The paper contains no reference to a pre-registered protocol, registry ID, or registration date, and none was found via internet search.
      • Relevant Quotes: 1) No quotes anywhere in the paper (including the Method, Procedure, Analysis, or Acknowledgments sections) mention a study registry, a registration ID, or a date of pre-registration. 2) The paper's only forward-looking statements about the analytic approach are internal descriptions of the PCA and multilevel modeling strategy in the Analysis section (pp. 8-11), not a reference to an externally logged, dated protocol. Detailed Analysis: Criterion P requires documented evidence that the full study protocol, hypotheses, and analysis plan were registered on a public platform before data collection began. This paper contains no reference to any pre-registration platform (e.g., OSF, AsPredicted, ClinicalTrials.gov equivalents for education) or registration date. An internet search of the OSF registries and general web search for a pre-registration record tied to this study's title, authors, or the Rosetta Stone Foundations intervention did not locate any registered protocol for this specific RCT. Without such a reference, there is no way to confirm the analysis plan was fixed in advance of data collection. Criterion P is not met because no pre-registration reference, ID, or date is provided anywhere in the paper, and none was located via internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.