Abstract
Stakeholders and researchers in higher education have long debated the consequences of English-medium instruction (EMI); a key assumption of EMI is that student's academic learning through English should be at least as good as learning through their first language (usually the national language). This study addressed the following question: "What is the impact from English-medium instruction on students' academic performance in an online learning environment?" "Academic performance" was measured in two ways: number of correctly answered test questions and through-put/drop-out rate. The study adopted an experimental design involving a large group (n = 2,263) randomized control study in a programming course. Student participants were randomly allocated to an English-medium version of the course (the intervention group) or a Swedish-medium version of the course (the control group). The findings were that students enrolled on the English-medium version of the course answered statistically significantly fewer test questions correctly; the EMI students also dropped out from the course to a statistically significantly higher degree compared to students enrolled on the Swedish version of the course. The conclusion of this study is thus that EMI may, under certain circumstances, have negative consequences for students' academic performance.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomization was performed at the individual student level within a single online course, not at class or school level, and the intervention is not one-to-one tutoring, so the class-level RCT requirement is not satisfied.
- "At this stage of the enrollment process the applicants were randomly assigned to the EMI or the SMI version of the course (the enrollment cohort was split into two groups using a 50/50 allocation ratio)." (p. 9)
Relevant Quotes:
1) "A central distinguishing feature of a randomized control study is the random assignment of 'units', in this case students, to the intervention and control groups following certain inclusion and exclusion criteria." (p. 8)
2) "At this stage of the enrollment process the applicants were randomly assigned to the EMI or the SMI version of the course (the enrollment cohort was split into two groups using a 50/50 allocation ratio). Use was made of a randomization feature built into the course management software; no allocation concealment mechanism was used." (p. 9)
3) "The course was designed for self-study, with no planned-for in-person teacher-student interaction, but students had the option to contact teachers via e-mail if they experienced problems." (p. 7)
Detailed Analysis:
The unit of randomisation was explicitly the individual student: applicants were randomly split 50/50 between the English-medium (EMI) and Swedish-medium (SMI) versions of the same online course. There was no class-level or school-level allocation; all participants were drawn from one national applicant pool into a single MOOC. The ERCT exception permitting student-level randomisation applies only to personal teaching such as one-to-one tutoring; this intervention is a self-paced online course with no teacher interaction, which is self-study rather than personal tutoring, so the exception does not apply. While the self-paced individual format arguably limits contamination, the standard requires class-level or stronger randomisation (or the tutoring exception), and neither condition holds here.
Criterion C is not met because students, not classes or schools, were the unit of randomisation and the tutoring exception does not apply.
-
E
Exam-based Assessment
- Outcomes were measured with the course's own 42 summative module-test questions, a course-specific instrument rather than a widely recognised standardised exam.
- "We used the number of correctly answered questions, an integer between 0 and 42 for each student, as our first indication of academic performance." (p. 11)
Relevant Quotes:
1) "Each of the eight course modules concluded with a module test with summative questions (similar in appearance to the example in Figure 1, but without the option to request hints)." (p. 7)
2) "One measure of academic performance was students' academic knowledge ... as indicated by correct versus incorrect answers on the summative test questions in the modules. The number of questions in each module test was between two and eight. The total number of summative questions was 42." (p. 11)
3) "We used the number of correctly answered questions, an integer between 0 and 42 for each student, as our first indication of academic performance. All scoring of the test questions was automated, the assessment of outcomes was thus blinded to group allocation." (p. 11)
4) "A second kind of academic performance was retention/drop-out, i.e., the proportion of students who were active versus those dropping out (for whatever reason)." (p. 11)
Detailed Analysis:
The primary knowledge outcome was the number of correctly answered summative questions embedded in the eight module tests of the course itself (adapted from the Stanford Open Learning Initiative "Principles of Computing" course). This is a course-internal assessment created for/with the teaching materials, not a national, state-wide, or otherwise widely recognised standardised examination. No validity or reliability evidence for the test as a standardised instrument is reported. The second outcome, retention/drop-out, is an administrative measure, not an exam at all. Although scoring was automated and blinded, that does not make the instrument a standardised exam.
Criterion E is not met because outcomes were measured with custom course module tests rather than a recognised standardised exam.
-
T
Term Duration
- The study was a short self-paced summer MOOC with outcomes collected within the course modules, and no interval of at least one academic term between intervention start and outcome measurement is documented.
- "In 2020, the course was run as a self-paced eight-module Massive Open Online Course (MOOC), offered in an English (EMI) as well as a Swedish version (SMI)." (p. 7)
Relevant Quotes:
1) "In 2020, the course was run as a self-paced eight-module Massive Open Online Course (MOOC), offered in an English (EMI) as well as a Swedish version (SMI)." (p. 7)
2) "Prospective students could sign up from June 11, 2020, until July 30, 2020." (footnote 3, p. 8)
3) "The learning context targeted by this study was an online introductory course in programming. Since it was launched, the objective of the course has been to level the (knowledge) playing field in preparation for university programming courses..." (p. 6)
4) "First, it must be acknowledged that this is a short-term study, providing a snapshot of EMI as realized in a single course; future work is required to establish whether the results reported here are stable over time and extend to academic programs (though, admittedly, this would involve a number of practical challenges relating to the experimental design)." (p. 20)
5) The course website is "https://www.sommarprogrammering. se/" (footnote 1, p. 6), i.e., "summer programming".
Detailed Analysis:
The intervention was a self-paced eight-module summer preparatory MOOC taken before university entry, with sign-up between June 11 and July 30, 2020. Outcomes (module test answers and drop-out) were collected within the course itself as students progressed through the modules; there was no follow-up measurement at a later point. The authors themselves describe the work as "a short-term study" and a "snapshot". No quotes document an interval of at least one full academic term (roughly 3-4 months) between intervention start and outcome measurement, and the summer course format makes such a duration implausible. The standard allows short interventions only if follow-up tracking extends at least a term after the start, which did not occur here.
Criterion T is not met because outcomes were measured during a short self-paced summer course with no term-long follow-up documented.
-
D
Documented Control Group
- The control group (Swedish-medium version) is clearly documented with its size, baseline demographics in Table 1, and its condition, an identical course differing only in language of instruction.
- "For the avoidance of doubt, the EMI and the SMI versions of the course were identical, the Swedish version being a direct translation of the English original..." (p. 8)
Relevant Quotes:
1) "Since the course was offered in an English (EMI) as well as a Swedish (SMI) version, and since what we wanted to investigate was the impact from using EMI, the two versions of the course were assigned 'intervention group' (EMI) and 'control group' (SMI) status respectively." (p. 8)
2) "For the avoidance of doubt, the EMI and the SMI versions of the course were identical, the Swedish version being a direct translation of the English original, making the study of the independent variable in this design - the medium of instruction - free from 'noise'." (p. 8)
3) "Table 1: Baseline demographic data for study participants in the Swedish and English version of the course." (p. 10; reports gender, age, and mean years of schooling for each group)
4) "ALL STUDENTS (n = 2,263, of which 1,134 were in the SMI version, and 1,129 were in the EMI version..." (p. 10)
5) "...students who made an attempt to answer at least one of the module test questions) amounted to 815 students. The remainder (n = 1,448) were considered drop-out students (649 in the SMI course and 799 in the EMI course)." (p. 10)
Detailed Analysis:
The paper explicitly designates the Swedish-medium (SMI) version as the control group and documents it in detail: group sizes at each stage (1,134 in the analysis sample, 485 active), baseline demographics (gender, age bands, mean years of schooling) reported per group in Table 1, the inclusion criteria applied to all participants (B1 reading in both Swedish and English), and the exact condition the control group experienced (an identical course, directly translated, differing only in the medium of instruction). A participant flow diagram (Figure 2) further documents allocation. This is sufficient to assess comparability of the control group.
Criterion D is met because the control group's size, baseline characteristics, and condition are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred at the individual student level within one online course; no schools or institutional units were randomised.
- "A central distinguishing feature of a randomized control study is the random assignment of 'units', in this case students, to the intervention and control groups..." (p. 8)
Relevant Quotes:
1) "A central distinguishing feature of a randomized control study is the random assignment of 'units', in this case students, to the intervention and control groups following certain inclusion and exclusion criteria." (p. 8)
2) "At this stage of the enrollment process the applicants were randomly assigned to the EMI or the SMI version of the course (the enrollment cohort was split into two groups using a 50/50 allocation ratio)." (p. 9)
Detailed Analysis:
The study randomised individual applicants to two versions of a single MOOC run by the researchers' institutions. There was no randomisation of schools, campuses, sites, or any other institutional unit implementing the intervention. Criterion S requires randomisation among schools or equivalent implementing units, which plainly did not occur.
Criterion S is not met because randomisation was at the student level, not the school or site level.
-
I
Independent Conduct
- The authors themselves translated/adapted the course, ran the trial, and analysed the data, with no external or third-party evaluation team documented.
- "Colleagues of ours (professors of computer science) working with the Open Learning Initiative at Stanford allowed us to use their newly developed Principles of Computing course, which we translated to Swedish..." (p. 6)
Relevant Quotes:
1) "Colleagues of ours (professors of computer science) working with the Open Learning Initiative at Stanford allowed us to use their newly developed Principles of Computing course, which we translated to Swedish (both the Swedish and English version of the course was, however, culturally adapted...)" (p. 6)
2) "In our case, we adopted an experimental design involving a large parallel group randomized control study." (p. 8)
3) "For all statistical calculations we used STATA 17..." (p. 11)
4) "All scoring of the test questions was automated, the assessment of outcomes was thus blinded to group allocation." (p. 11)
Detailed Analysis:
The intervention in this trial is the medium of instruction operationalised through two course versions that the author team itself prepared (they translated and culturally adapted the course). The same author team designed the experiment, recruited and randomised participants, and performed all analyses; no external evaluation agency, independent data collectors, or third-party oversight is mentioned anywhere in the paper. Automated, allocation- blinded scoring reduces measurement bias but does not constitute independent conduct in the sense of the standard, which requires evaluation independent of those who designed the intervention. The original course content came from Stanford OLI, but the tested manipulation (the Swedish translation and the comparison) was created and evaluated by the same authors.
Criterion I is not met because the same team that created the intervention conditions conducted and analysed the study without documented independent oversight.
-
Y
Year Duration
- Since criterion T is not met and the study was a short summer MOOC snapshot, tracking clearly did not span 75% of an academic year.
- "First, it must be acknowledged that this is a short-term study, providing a snapshot of EMI as realized in a single course..." (p. 20)
Relevant Quotes:
1) "In 2020, the course was run as a self-paced eight-module Massive Open Online Course (MOOC)..." (p. 7)
2) "Prospective students could sign up from June 11, 2020, until July 30, 2020." (footnote 3, p. 8)
3) "First, it must be acknowledged that this is a short-term study, providing a snapshot of EMI as realized in a single course; future work is required to establish whether the results reported here are stable over time and extend to academic programs." (p. 20)
Detailed Analysis:
The instruction and outcome measurement were confined to a single self-paced summer preparatory course in 2020, and the authors explicitly characterise the study as short-term. There is no follow-up extending to 75% of an academic year (roughly 9-10 months) after intervention start. Additionally, per the ranking instructions, since criterion T (Term Duration) is not met, criterion Y cannot be met.
Criterion Y is not met because the study tracked outcomes only over a short summer course, far below 75% of an academic year.
-
B
Balanced Control Group
- Both groups received identical courses with the same time, content, and resources, differing only in the medium of instruction, so inputs were fully balanced.
- "For the avoidance of doubt, the EMI and the SMI versions of the course were identical, the Swedish version being a direct translation of the English original, making the study of the independent variable in this design - the medium of instruction - free from 'noise'." (p. 8)
Relevant Quotes:
1) "Save for the language of instruction, the EMI and SMI versions of the course were identical." (p. 7)
2) "For the avoidance of doubt, the EMI and the SMI versions of the course were identical, the Swedish version being a direct translation of the English original, making the study of the independent variable in this design - the medium of instruction - free from 'noise'." (p. 8)
3) "However, there is no obvious reason to believe that the EMI course should disincentivize students in this regard more so than the SMI course (the course design was identical - the only differentiating factor was the medium of instruction)." (p. 19)
Detailed Analysis:
Applying the criterion B decision tree: the intervention group received no extra time, budget, materials, or support relative to the control group. Both groups took the same eight-module self-paced MOOC with the same formative questions, module tests, and optional e-mail teacher contact; the only difference was the language in which the identical content was presented (English vs. Swedish). Since no additional resources were present in either arm (EXTRA_RESOURCES_PRESENT = false), the balance requirement is trivially satisfied; this is close to an ideal balanced design for isolating the medium-of-instruction effect.
Criterion B is met because both conditions received identical educational inputs, differing only in the language of instruction.
-
Level 3 Criteria
-
R
Reproduced
- No independent published replication of this specific randomized EMI-versus-SMI course experiment was found in the paper or via external search; the authors themselves call for future replication.
- "The design of the present study lends itself to a replication in the same kind of online learning context, but in another discipline; such a replication would add to the external validity of the research findings." (p. 20)
Relevant Quotes:
1) "To the best of our knowledge, only three earlier studies of EMI have adopted randomized allocation of students and used control groups as part of their research design...: Roussel et al. (2017), Tatzl and Messnarz (2013), and Vinke (1995)..." (p. 19)
2) "The design of the present study lends itself to a replication in the same kind of online learning context, but in another discipline; such a replication would add to the external validity of the research findings." (p. 20)
3) "Future research replicating the current study could introduce elements of language support as a variable in the research design to investigate this further." (p. 18)
Detailed Analysis:
The earlier randomized EMI studies the authors cite (Roussel et al. 2017; Tatzl and Messnarz 2013; Vinke 1995) predate this trial and used different designs, populations, and contexts; they are antecedents, not replications of this study. The authors explicitly frame replication as future work. A re-verification internet search conducted on 2026-07-27 for independent replications of this specific randomized Swedish/English MOOC experiment again found only related but non-replicating observational EMI studies (e.g., Garcia-Alvarez de Perea and Ramirez-Garcia 2024 on EMI accounting students, and other attitude/QCA studies on EMI), but no independent team has published a peer-reviewed replication of this randomized course-language experiment.
Criterion R is not met because no independent peer-reviewed replication of this specific study was found.
-
A
All-subject Exams
- Only programming/computing knowledge was assessed with a custom course test, so with criterion E unmet and no other subjects measured, the all-subject exam requirement fails.
- "One measure of academic performance was students' academic knowledge (this included, to varying degrees, content knowledge, procedural knowledge, and conditional knowledge of programming)..." (p. 11)
Relevant Quotes:
1) "One measure of academic performance was students' academic knowledge (this included, to varying degrees, content knowledge, procedural knowledge, and conditional knowledge of programming) as indicated by correct versus incorrect answers on the summative test questions in the modules." (p. 11)
2) "A second limitation is that this study covered a single subject/discipline--programming/computer science (broadly speaking)--and earlier research has indicated that there are considerable disciplinary differences..." (p. 20)
Detailed Analysis:
Criterion A requires standardised exam-based assessment across all main subjects. First, criterion E is not met (custom course tests), which by the instructions means criterion A automatically fails. Second, the study explicitly assessed only one subject, programming/computer science, in a single preparatory course; no other subjects were measured. Although the course is a specialised preparatory offering, the assessment used was not a standardised exam, so the specialisation exception cannot rescue the criterion.
Criterion A is not met because criterion E fails and only a single subject was assessed with a custom test.
-
G
Graduation Tracking
- Measurement ended with the course's own module tests, with no tracking of participants to graduation, and criterion Y is unmet which also fails this criterion.
- "First, it must be acknowledged that this is a short-term study, providing a snapshot of EMI as realized in a single course..." (p. 20)
Relevant Quotes:
1) "First, it must be acknowledged that this is a short-term study, providing a snapshot of EMI as realized in a single course; future work is required to establish whether the results reported here are stable over time and extend to academic programs (though, admittedly, this would involve a number of practical challenges relating to the experimental design)." (p. 20)
2) "Since it was launched, the objective of the course has been to level the (knowledge) playing field in preparation for university programming courses..." (p. 6)
Detailed Analysis:
Participants were prospective university students taking a preparatory summer MOOC; measurement stopped with the course module tests and drop-out counts. The authors state that whether effects extend over time or into academic programs is left to future work, i.e., no cohort tracking to any graduation point occurred. A re-verification internet search conducted on 2026-07-27 for follow-up publications by the same author team tracking this cohort toward graduation found none. In addition, per the ranking instructions, criterion G cannot be met when criterion Y is not met, and Y is not met here.
Criterion G is not met because no tracking beyond the course itself, let alone to graduation, was performed.
-
P
Pre-Registered
- The paper contains no mention of a pre-registered protocol or registry entry, and an external search found no pre-registration for this trial.
Relevant Quotes:
1) "Use was made of a randomization feature built into the course management software; no allocation concealment mechanism was used." (p. 9)
2) "In our case, we adopted an experimental design involving a large parallel group randomized control study." (p. 8)
Detailed Analysis:
The methods section describes the design, enrollment (June 11 to July 30, 2020), randomisation, and analysis, but nowhere mentions a registry (e.g., ClinicalTrials.gov, ISRCTN, OSF, AEA registry), a registration ID, or a pre-specified analysis plan. A re-verification internet search conducted on 2026-07-27 for a pre-registration of this trial by Balter, Kann, Mutimukwe, and Malmstrom again found no registry entry. Without quoted evidence of registration before data collection began, the criterion cannot be satisfied.
Criterion P is not met because no pre-registration of the study protocol is reported in the paper or discoverable externally.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.