Abstract
Given the accumulating evidence about audiovisual input as a valuable resource from which knowledge of multiword expressions (MWEs) can be built up incidentally, the next inquiry arises as to what can be done to promote MWE uptake from this resource. Despite the increasing popularity of bilingual subtitles as a form of on-screen text, their effectiveness for the incidental acquisition of MWEs relative to other subtitling forms has not yet been examined. A total of 89 L2 learners were randomly assigned to three experimental groups that differed in terms of the subtitling condition (captions, bilingual subtitles, L1 subtitles) under which they viewed an input video containing target MWEs twice. They were administered a pretest and a posttest to gauge their improvement in MWE knowledge at the level of form recognition and meaning recall. An operation-span task was employed to measure their working memory capacity. Results from generalized linear mixed-effects models revealed that bilingual subtitles had an advantage over captions and L1 subtitles in facilitating MWE meaning uptake, and they were as effective as captions in promoting MWE form uptake. Working memory played a predictive role in the uptake of novel MWEs, with a greater weight observed in bilingual subtitled viewing.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was carried out at the individual student level within a single participant pool, not at the class or school level, and the intervention was not one-to-one tutoring.
- "A total of 89 L2 learners were randomly assigned to three experimental groups that differed in terms of the subtitling condition (captions, bilingual subtitles, L1 subtitles)..." (p. 1)
Relevant Quotes:
1) "A total of 89 L2 learners were randomly assigned to three experimental groups that differed in terms of the subtitling condition (captions, bilingual subtitles, L1 subtitles) under which they viewed an input video containing target MWEs twice." (p. 1, Abstract)
2) "Prior to the treatment, they were randomly assigned with their unique participant ID numbers into three experimental groups." (p. 4)
3) "One-hundred and six Chinese learners of English from various academic backgrounds at a university in the United Kingdom initially participated in this study." (p. 4)
Detailed Analysis:
Criterion C requires randomisation at the class level or stronger (school level), unless the intervention is personal tutoring or one-to-one teaching. In this study, individual volunteer participants recruited from various academic backgrounds at one university were randomly assigned by their participant ID numbers to the three subtitling conditions. The unit of randomisation is therefore the individual student, not intact classes or schools. The intervention (watching a subtitled sitcom video) is not personal tutoring or one-to-one teaching, so the tutoring exception does not apply. Although individual randomisation carries a low contamination risk in this lab-style viewing design, the criterion as defined requires class-level or school-level assignment. All quotes were re-checked against the PDF and are verbatim.
Criterion C is not met because participants were randomised individually, not by class or school, and no valid exception applies.
-
E
Exam-based Assessment
- Outcomes were measured with researcher-made form-recognition and meaning-recall MWE tests created for this study, not with any widely recognised standardised exam.
- "In the form-recognition test (see Figure 2 for an example and Appendix B for all test items), the participants were asked to select a conventional idiomatic word combination in the English language from four options..." (p. 6)
Relevant Quotes:
1) "A form-recognition test and a meaning-recall test were administered in the pretest and posttest, with the order of items randomized to reduce potential testing effects." (p. 6)
2) "In the form-recognition test (see Figure 2 for an example and Appendix B for all test items), the participants were asked to select a conventional idiomatic word combination in the English language from four options: one correct answer, two lures, and an 'I don't know' option." (p. 6-7)
3) "The meaning-recall test (see Figure 3 for an example) required the participants to write down the meanings of the given MWEs in English or their L1." (p. 7)
4) "The Cronbach's alpha values of the form-recognition and meaning-recall tests were .79 and .82 in the pilot study, and .76 and .84 in the main study, respectively, indicating acceptable reliability." (p. 7)
Detailed Analysis:
Criterion E requires that outcomes be measured with standard, widely recognised standardised exams rather than instruments custom-built for the study. Here, the primary outcome measures are a form-recognition test and a meaning-recall test that the researcher constructed specifically around the 22 target MWEs appearing in the treatment video, with lures created by substituting components of the target MWEs. These are bespoke research instruments aligned to the intervention content, the exact situation the criterion is designed to guard against. Standardised instruments mentioned in the paper (IELTS scores, the Updated Vocabulary Levels Test) were used only to describe participant proficiency and as covariates, not as outcome measures. All quotes were re-checked against the PDF and are verbatim.
Criterion E is not met because the outcome measures were custom-made tests specifically designed for this study rather than standardised exams.
-
T
Term Duration
- The whole study spanned only four weeks, with the viewing treatment and the unannounced posttest occurring in the same session in Week 4, far short of one academic term.
- "This study was conducted over a four-week period in two sessions..." (p. 7)
Relevant Quotes:
1) "This study was conducted over a four-week period in two sessions (see Figure 4 for a visual diagram of the research procedures)." (p. 7)
2) "In the first session (Week 1), participants gave their informed consent and completed a battery of tasks: WM test, UVLT, MWE pretest, and background questionnaire. The treatment session took place in Week 4..." (p. 9)
3) "After the treatment, the participants completed a comprehension test and an unannounced MWE posttest." (p. 9)
4) "First, the retention of MWEs acquired under the different subtitling conditions was unknown due to the lack of a delayed posttest." (p. 13)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (about 3-4 months) after the intervention begins. In this study, the intervention consisted of watching a 30-minute video twice in a single treatment session in Week 4, and the posttest was administered immediately after viewing in the same session. The interval from intervention start to outcome measurement is therefore essentially zero, and even the full study window (pretest to posttest) covers only four weeks. The author explicitly acknowledges the absence of any delayed posttest, so there is no term-long follow-up tracking. All quotes were re-checked against the PDF and are verbatim.
Criterion T is not met because outcomes were measured immediately after a single-session intervention within a four-week study period, far short of one academic term.
-
D
Documented Control Group
- The comparison (active control) conditions are documented with group sizes, demographics, and baseline measures of proficiency, vocabulary knowledge, and working memory, with statistical checks confirming baseline comparability.
- "Results of the independent-samples Kruskal-Wallis test (for IELTS scores due to the violation of homogeneity of variances) and the one-way ANOVA analyses showed no significant group differences in IELTS scores... prior vocabulary knowledge... and WM capacity... (see Appendix D for descriptive statistics)." (p. 9)
Relevant Quotes:
1) "An attrition of 17 participants occurred due to either technical issues or lack of attendance in the viewing session, with 89 participants remaining in the three groups: Group 1 (n = 32), Group 2 (n = 29), Group 3 (n = 28)." (p. 4)
2) "The remaining participants (22 males and 67 females) were aged 20-31 years (M = 24.41 years, SD = 2.14) and had been learning English for a minimum of 11 years." (pp. 4-5)
3) "Their mean overall IELTS score was 6.70 (SD = 0.59), which approximately corresponded to B2 to C1 levels in the Common European Framework of Reference for Languages." (p. 5)
4) "Results of the independent-samples Kruskal-Wallis test (for IELTS scores due to the violation of homogeneity of variances) and the one-way ANOVA analyses showed no significant group differences in IELTS scores (χ2(2) = 0.367, p = 0.832), prior vocabulary knowledge (F(2, 86) = 0.348, p = 0.707), and WM capacity (F(2, 86) = 0.551, p = 0.578; see Appendix D for descriptive statistics)." (p. 9)
5) "Fifth, the study did not satisfactorily include a group exposed to the video without on-screen text, nor a control group that did not receive any viewing treatment." (p. 13)
Detailed Analysis:
Criterion D requires that the control/comparison group be well-documented in terms of size, demographics, baseline performance, and the conditions it received. This study has no no-treatment control; instead the three subtitling conditions act as active comparison groups for one another (e.g., the L1 subtitles and captions groups serve as the comparison for the bilingual subtitles group). Each group's size is reported, the sample's demographics (age, gender, years of English study) are described, and Appendix D provides per-group descriptive statistics on IELTS proficiency, prior vocabulary knowledge (UVLT), and working memory, with formal tests confirming no significant baseline differences between groups. What each comparison group received is precisely specified (the identical video viewed twice with captions or L1 subtitles). This level of documentation permits proper comparison and interpretation, which is the purpose of the criterion, even though the authors note the absence of a no-viewing control group as a limitation.
Criterion D is met because the comparison conditions are clearly documented with group sizes, baseline characteristics, and exact descriptions of what each group received.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred at the individual participant level within a single university, so no school-level randomisation took place.
- "Prior to the treatment, they were randomly assigned with their unique participant ID numbers into three experimental groups." (p. 4)
Relevant Quotes:
1) "One-hundred and six Chinese learners of English from various academic backgrounds at a university in the United Kingdom initially participated in this study." (p. 4)
2) "Prior to the treatment, they were randomly assigned with their unique participant ID numbers into three experimental groups." (p. 4)
Detailed Analysis:
Criterion S requires randomisation among schools or equivalent implementing institutions. This study recruited individual volunteers from a single UK university and randomised them individually into three conditions. No schools, sites, or institutional units were randomised; the entire experiment took place within one institution (online via Zoom or in a campus study lounge). This is clearly a student-level, single-site RCT. All quotes were re-checked against the PDF and are verbatim.
Criterion S is not met because randomisation was at the individual student level within one university, not at the school level.
-
I
Independent Conduct
- The single author designed the materials, conducted the experiment, and analysed the data as part of his master's dissertation, with no independent third-party evaluation.
- "This article was part of the author's master's dissertation at University of Oxford." (p. 13)
Relevant Quotes:
1) "This article was part of the author's master's dissertation at University of Oxford. The author would like to thank Robert Woore for his supervision and constructive feedback, and is grateful to the participants whose valuable contributions aided in the completion of this research." (p. 13)
2) "Based on multiple screenings, six engaging five-minute excerpts where a good number of MWEs occurred were selected and assembled into a single 30-minute video using Video Studio Pro 2018..." (p. 5)
3) "The MWE meaning recall test was independently scored by two Chinese-speaking raters who had been trained on the meanings of target items." (p. 9)
Detailed Analysis:
Criterion I requires that the study be conducted independently from those who designed the intervention. Here, a single author designed the intervention materials (selecting and assembling the video, creating the three subtitle versions, developing the tests), recruited and ran the participants, and performed the statistical analyses, all as part of his own master's dissertation. The only element of external involvement is the use of two trained raters to score the meaning-recall test and supervision by a dissertation supervisor, which does not constitute an independent evaluation team or third-party oversight of data collection, analysis, or conclusions. All quotes were re-checked against the PDF and are verbatim.
Criterion I is not met because the same person designed, conducted, and analysed the study with no independent evaluators.
-
Y
Year Duration
- The study lasted only four weeks with an immediate posttest, so it falls far short of 75 percent of an academic year, and the weaker term criterion T is also unmet.
- "This study was conducted over a four-week period in two sessions..." (p. 7)
Relevant Quotes:
1) "This study was conducted over a four-week period in two sessions (see Figure 4 for a visual diagram of the research procedures)." (p. 7)
2) "The treatment session took place in Week 4, at which point the participants had been randomly assigned to three groups. The three groups watched the video twice using one of three forms of subtitles: captions, bilingual subtitles, or L1 subtitles. After the treatment, the participants completed a comprehension test and an unannounced MWE posttest." (p. 9)
Detailed Analysis:
Criterion Y requires outcome measurement at least 75 percent of an academic year (roughly 9-10 months) after the intervention begins. The entire study, from pretest to posttest, spanned four weeks, and the outcome was measured immediately after the single viewing session. Additionally, per the ranking instructions, criterion Y cannot be met when criterion T (term duration) is not met, and T failed here. All quotes were re-checked against the PDF and are verbatim.
Criterion Y is not met because the tracking interval was effectively immediate within a four-week study, nowhere near an academic year.
-
B
Balanced Control Group
- All three groups received identical time and materials - the same 30-minute video watched twice in the same session - with only the subtitle format differing, so inputs were fully balanced across conditions.
- "The three groups watched the video twice using one of three forms of subtitles: captions, bilingual subtitles, or L1 subtitles." (p. 9)
Relevant Quotes:
1) "The three groups watched the video twice using one of three forms of subtitles: captions, bilingual subtitles, or L1 subtitles." (p. 9)
2) "Three different forms of subtitles were added to the video using SrtEdit (PortableSoft, 2012) and Video Studio Pro 2018 (Corel Corporation, 2018), which resulted in three versions of the video (see Figure 1)." (p. 5)
3) "Each participant who wore earphones/headphones accessed the cloud-based research platform Gorilla Experiment Builder (Anwyl-Irvine et al., 2020) on their laptops/desktops through a personalized link tailored to the experimental design." (p. 7)
4) "Based on multiple screenings, six engaging five-minute excerpts where a good number of MWEs occurred were selected and assembled into a single 30-minute video..." (p. 5)
Detailed Analysis:
Criterion B requires that time and resources be balanced across conditions unless extra resources are themselves the treatment variable. In this three-arm design, all groups received exactly the same educational input in terms of quantity and quality: the same 30-minute video, watched twice, in the same session, delivered through the same platform under the same procedural constraints. The only difference between conditions was the form of on-screen text (L2 captions, L1 subtitles, or bilingual subtitles), which is precisely the treatment contrast under investigation. No group received additional instructional time, materials, or budget relative to the others. Applying the current criterion B decision tree: EXTRA_RESOURCES_PRESENT is false, since no arm received extra time or budget beyond the other arms, so the balance requirement is trivially satisfied and the criterion is met without needing to reach the integral-resource or within-subjects branches. It should be noted that the design lacks a no-treatment arm entirely (this is assessed separately under criterion D), but among the randomised comparison groups the inputs are fully matched. All quotes were re-checked against the PDF and are verbatim.
Criterion B is met because all three randomised conditions received identical time and materials, differing only in the subtitle format that constitutes the treatment variable.
-
Level 3 Criteria
-
R
Reproduced
- This 2025 study describes itself as the first of its kind, and a fresh internet search in July 2026 still found no independent published replication of this specific experiment.
- "This study was the first attempt to examine the efficacy of bilingual subtitles on incidental MWE acquisition relative to monolingual subtitles." (p. 11)
Relevant Quotes:
1) "This study was the first attempt to examine the efficacy of bilingual subtitles on incidental MWE acquisition relative to monolingual subtitles." (p. 11)
2) "Further, it was the first study to explore the impact of WM on MWE learning under different subtitling conditions..." (p. 11)
3) "Additionally, it was the first endeavor to compare L1 subtitles and captions for MWE acquisition." (p. 11)
Detailed Analysis:
Criterion R requires that the study be independently replicated by a different research team in a different context, published in a peer-reviewed journal. The paper repeatedly emphasises its novelty as the first study of bilingual subtitles for MWE acquisition, so no replication is referenced within it. A renewed internet search conducted in July 2026 (searching for citations of this paper, replications of its design, and recent bilingual-subtitle/MWE studies) again found only related but distinct studies: work on the sequential use of L1 and bilingual subtitles for single-word vocabulary (e.g., Yuan et al., 2025, British Journal of Educational Psychology), a study on textual enhancement in bilingual-subtitled viewing by a different "Li" (Jinyang Li, Kunming University of Science and Technology, Journal of Language and Education, 2026), and an unrelated approximate replication of a different paper (Majuddin, Siyanova-Chanturia, & Boers, 2021) published in Language Teaching (2026). None of these is an independent replication of this specific experiment comparing captions, L1 subtitles, and bilingual subtitles for MWE form and meaning uptake with repeated viewing by the same author (Wangyin Kenneth Li). Given the paper's recency (2025) and the absence of any matching replication as of this July 2026 check, no independent reproduction of this particular study exists yet. All quotes were re-checked against the PDF and are verbatim.
Criterion R is not met because no independent peer-reviewed replication of this specific study was found, including after a fresh internet search.
-
A
All-subject Exams
- Only knowledge of English multiword expressions was measured with custom tests, so neither the exam-based prerequisite (criterion E) nor coverage of all main subjects is satisfied.
- "They were administered a pretest and a posttest to gauge their improvement in MWE knowledge at the level of form recognition and meaning recall." (p. 1)
Relevant Quotes:
1) "They were administered a pretest and a posttest to gauge their improvement in MWE knowledge at the level of form recognition and meaning recall." (p. 1, Abstract)
2) "A form-recognition test and a meaning-recall test were administered in the pretest and posttest..." (p. 6)
Detailed Analysis:
Criterion A requires standardised exam-based assessment of all main subjects, and per the instructions it cannot be met when criterion E is not met. Criterion E failed here because the outcome measures were custom-made MWE tests. Moreover, the study measured only one narrow facet of one subject area (English L2 multiword expressions); no other subjects were assessed. As a study with adult university volunteers, one might argue a specialised focus, but the exception cannot rescue the criterion because no standardised exams were used as outcomes at all. All quotes were re-checked against the PDF and are verbatim.
Criterion A is not met because criterion E is not met and only a single narrow domain (MWE knowledge) was assessed.
-
G
Graduation Tracking
- There was no delayed posttest or longer-term follow-up of any kind, let alone tracking of participants until graduation, the prerequisite criterion Y is unmet, and a fresh internet search found no follow-up publications on this cohort.
- "First, the retention of MWEs acquired under the different subtitling conditions was unknown due to the lack of a delayed posttest." (p. 13)
Relevant Quotes:
1) "First, the retention of MWEs acquired under the different subtitling conditions was unknown due to the lack of a delayed posttest." (p. 13)
2) "After the treatment, the participants completed a comprehension test and an unannounced MWE posttest." (p. 9)
Detailed Analysis:
Criterion G requires following participants until graduation from their educational stage. This study measured outcomes only immediately after the viewing treatment; the author explicitly acknowledges there was not even a delayed posttest. A dedicated internet search in July 2026 for subsequent publications by Wangyin Kenneth Li tracking this same participant cohort (e.g., delayed-retention or longitudinal follow-ups) found none; only the original 2025 article and unrelated works by other authors were located. The participants were adult university students recruited as anonymised volunteers for a single-session lab experiment, with no mechanism described for tracking them through to graduation. Additionally, per the instructions, criterion G cannot be met when criterion Y is not met, and Y failed here. All quotes were re-checked against the PDF and are verbatim.
Criterion G is not met because measurement stopped immediately after the intervention with no follow-up or graduation tracking, and no subsequent tracking papers were found online.
-
P
Pre-Registered
- The paper contains no mention of any pre-registration; direct inspection of the linked OSF project confirms it is an unregistered project hosting only video excerpts, an OASIS document, and appendices, not a pre-registered protocol.
Relevant Quotes:
1) "The R scripts used for the statistical analyses can be accessed via the Open Science Framework (https://doi.org/10.17605/OSF.IO/A2XHE)." (p. 9)
2) "Due to space constraints, the remaining appendices are available via the Open Science Framework (https://doi.org/10.17605/OSF.IO/A2XHE)." (p. 22)
Detailed Analysis:
Criterion P requires that the full study protocol, including hypotheses, methods, and planned analyses, be registered on a public registry before data collection begins, with verifiable timing. The paper's only references to an open-science platform concern sharing R analysis scripts and supplementary appendices on OSF, which is materials sharing, not pre-registration of a protocol. No registry entry, registration ID, or registration date is mentioned anywhere in the paper, and there is no statement that hypotheses or analysis plans were registered prior to data collection. Direct verification of the linked OSF project (osf.io/a2xhe) in July 2026 confirms via the OSF API that it is a standard, unregistered project ("registration": false, category "Project"), created April 29, 2025, containing video excerpt files, an "OASIS" document, and an appendices file rather than a frozen, timestamped pre-registration. No separate pre-registration entry for this study was found on OSF Registries or other registries (e.g., AsPredicted, ClinicalTrials.gov) in a further internet search. All quotes were re-checked against the PDF and are verbatim.
Criterion P is not met because no pre-registration of the study protocol is mentioned in the paper, and the linked OSF project itself confirms it is unregistered materials-sharing, not a pre-registered protocol.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.