CLIL for all? A randomised controlled field experiment with sixth-grade students on the effects of content and language integrated science learning

Nicole Piesche, Kathrin Jonkmann, Christiane Fiege, Jörg-U. Keßler

Published:
ERCT Check Date:
DOI: 10.1016/j.learninstruc.2016.04.001
  • science
  • K12
  • EU
0
  • C

    Randomisation was conducted at the class level (30 classes), which satisfies the class-level RCT requirement.

    "The randomization took part at the class level." (p. 3)

  • E

    The primary outcome was a custom-assembled "Floating and Sinking" test built largely from prior research items plus new author-written items, not a widely recognised standardised exam.

    "Altogether the test consisted of 36 items. ... four items were developed by us." (p. 4)

  • T

    Outcomes were measured immediately after a five-lesson intervention and again only 6 weeks later, far short of one academic term.

    "The lessons were followed by a 60-min posttest and a 30-min follow-up test 6 weeks after instruction ended." (p. 3)

  • D

    The comparison (monolingual) group's demographics, baseline achievement, and instructional exposure are documented in detail in Table 1 and the Procedure section.

    "There were 362 bilingually educated students (52.7% girls, ...) and 360 monolingually educated students (44.8% girls, ...)." (p. 3)

  • S

    Randomisation was explicitly conducted at the class level, not at the school level, so the stronger school-level requirement is not satisfied.

    "The randomization took part at the class level." (p. 3)

  • I

    The intervention was designed and personally taught by the paper's first author, with no independent or external evaluation team involved.

    "The whole unit of instruction was taught by the first author of this paper who is a formally trained secondary school teacher of both English and science. Hence, the teacher was not blind to the condition." (p. 3)

  • Y

    The study tracked outcomes for at most a few months total (intervention plus a 6-week follow-up), far short of 75% of an academic year, and criterion T was already not met.

    "... a 30-min follow-up test 6 weeks after instruction ended." (p. 3)

  • B

    Instructional time and teaching materials were held constant across both groups; the only difference was the language of instruction plus minor language scaffolding integral to the bilingual condition itself.

    "The teaching material and the instructional time were held constant between groups. The only variation was in the language of instruction, and additional language support ... was given to the bilingual group." (p. 3)

  • R

    No independent replication of this specific study by a different research team was found in the paper or via external search.

  • A

    Only science (physics) content knowledge was assessed; no other core subjects were measured, and criterion E was not met.

    "Second, the achievement measure that we used focused on one aspect of science competence, namely content knowledge." (p. 7, Limitations)

  • G

    Tracking stopped 6 weeks after the intervention with no further follow-up reported, and criterion Y was not met.

    "... a 30-min follow-up test 6 weeks after instruction ended." (p. 3)

  • P

    No pre-registration of the study protocol is mentioned anywhere in the paper, and none was found via external search.

Abstract

Content and language integrated learning (CLIL) has been widely implemented in Europe. This article presents a randomised controlled field experiment on the effects of CLIL on students' science learning. Thirty sixth-grade intermediate-track German secondary-school classes (722 students) were randomly assigned to learn (5 lessons, 90 min each) a physics topic taught either in German or in English and German. We expected that the monolingually taught students would outperform the bilingually taught ones immediately after the intervention. For the follow-up test 6 weeks later, the same or smaller differences between the groups were expected due to the potential for a deeper processing of the subject matter in the bilingual condition. The results showed that the bilingually educated students' learning gains were smaller than the monolingually educated ones' immediately after the intervention (d = -0.21) and at follow-up (d = -0.23). The expectation of more sustainable processing was not supported.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was conducted at the class level (30 classes), which satisfies the class-level RCT requirement.
      • "The randomization took part at the class level." (p. 3)
      • Relevant Quotes: 1) "722 students (48.7% girls) from 30 sixth-grade classes from 10 Realschulen located in middle-class neighbourhoods of small and midsize towns in Baden-Wuerttemberg, Germany were randomly assigned to receive either German or German/English instruction." (p. 3) 2) "The randomization took part at the class level." (p. 3) 3) "There were 362 bilingually educated students ... and 360 monolingually educated students ..." (p. 3) Detailed Analysis: The paper explicitly states that whole classes (not individual students within a shared classroom) were the unit of randomisation: 30 intact sixth-grade classes, drawn from 10 schools, were assigned wholly to either the monolingual or the bilingual condition. This is exactly the unit of randomisation the ERCT Standard's C criterion requires, and it rules out the within-class contamination problem the criterion is designed to prevent (students in the same class receiving different treatments). There is no indication that individual students within the same class were split across conditions. Final summary: Criterion C is met because the study explicitly reports that randomisation occurred at the whole-class level across 30 classes.
    • E

      Exam-based Assessment

      • The primary outcome was a custom-assembled "Floating and Sinking" test built largely from prior research items plus new author-written items, not a widely recognised standardised exam.
      • "Altogether the test consisted of 36 items. ... four items were developed by us." (p. 4)
      • Relevant Quotes: 1) "To measure students' learning gains on the topic 'Floating and Sinking', a test based on students' typical preconcepts ... was administered. Altogether the test consisted of 36 items." (p. 4) 2) "Twenty-eight items were published in Blumberg (2008); Hardy et al. (2006); Kleickmann (2008); Möller (2005); Möller et al. (2006); Stern, Möller, Hardy, and Jonen (2002). Three more items were translated from former Trends in International Mathematics and Science Studies (TIMSS ...), and four items were developed by us." (p. 4) 3) "For the IRT-based scaling of the test, a 2-parameter logistic model (Birnbaum model) was chosen ..." (p. 4) Detailed Analysis: The study's primary and only academic outcome measure is a bespoke, topic-specific instrument assembled by the research team from earlier studies' items (on the same "Floating and Sinking" topic), a handful of translated TIMSS items, and four items the authors wrote themselves, then custom-calibrated with an IRT model for this study. This is precisely the kind of researcher-constructed, intervention-aligned test the ERCT Standard's E criterion warns against: it is not a widely recognised, standardised exam (e.g., a national curriculum or state-wide test) administered as-is. Even the physics-preknowledge covariate (partly from TIMSS) is a secondary/control measure, not the primary outcome, and it too was adapted/supplemented by the authors. Final summary: Criterion E is not met because the outcome measure is a custom, author-assembled test rather than a standardised, widely recognised exam.
    • T

      Term Duration

      • Outcomes were measured immediately after a five-lesson intervention and again only 6 weeks later, far short of one academic term.
      • "The lessons were followed by a 60-min posttest and a 30-min follow-up test 6 weeks after instruction ended." (p. 3)
      • Relevant Quotes: 1) "Then, the intervention was implemented through a teaching unit on the topic 'Floating and Sinking' consisting of five lessons lasting 90 min each." (p. 3) 2) "The lessons were followed by a 60-min posttest and a 30-min follow-up test 6 weeks after instruction ended." (p. 3) Detailed Analysis: The intervention itself was very brief (five 90-minute lessons), and the final, longest-delayed measurement point (the follow-up test) occurred only 6 weeks after the intervention ended. Six weeks is well short of the 3-4 month academic term the ERCT Standard's T criterion requires between intervention start and outcome measurement. No later assessment point is reported anywhere in the paper. Final summary: Criterion T is not met because the longest follow-up interval reported is 6 weeks, far below the required one-term minimum.
    • D

      Documented Control Group

      • The comparison (monolingual) group's demographics, baseline achievement, and instructional exposure are documented in detail in Table 1 and the Procedure section.
      • "There were 362 bilingually educated students (52.7% girls, ...) and 360 monolingually educated students (44.8% girls, ...)." (p. 3)
      • Relevant Quotes: 1) "There were 362 bilingually educated students (52.7% girls, Mage = 11.5 years, SD = 0.61, 35.7% immigration background) and 360 monolingually educated students (44.8% girls, Mage = 11.5 years, SD = 0.57, 33.2% immigration background)." (p. 3) 2) "Table 1 provides an overview of the descriptive statistics. ... the subsamples differed significantly only on gender." (p. 5, Results 6.1) - Table 1 reports, separately for the monolingual and bilingual groups: gender, immigration background, parental education, general cognitive ability, English ability, science grade, physics preknowledge, science self-concept, interest in science, and floating-and-sinking pretest scores. 3) "The teaching material and the instructional time were held constant between groups." (p. 3) Detailed Analysis: Although this study compares two active instructional conditions rather than an intervention-vs-nothing design, the monolingually taught group functions as the comparison/control condition against which the bilingual condition is evaluated. Its size, gender split, immigration background, socioeconomic proxy (parental education), baseline cognitive ability, baseline physics knowledge, baseline English ability, motivation measures, and pretest performance are all explicitly reported and statistically compared against the treatment group in Table 1. This level of detail allows a reader to judge comparability at baseline. Final summary: Criterion D is met because the comparison group's characteristics and baseline data are thoroughly documented in Table 1 and the text.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was explicitly conducted at the class level, not at the school level, so the stronger school-level requirement is not satisfied.
      • "The randomization took part at the class level." (p. 3)
      • Relevant Quotes: 1) "The randomization took part at the class level." (p. 3) 2) "722 students (48.7% girls) from 30 sixth-grade classes from 10 Realschulen ... were randomly assigned to receive either German or German/English instruction." (p. 3) Detailed Analysis: The paper draws its sample from 10 schools but is explicit that the unit of random assignment was the individual class, not the school: multiple classes from the same school could, and likely did, end up in different conditions. The ERCT Standard's S criterion requires randomisation among whole schools (or equivalent higher-level units), which is not what happened here. Final summary: Criterion S is not met because the study explicitly describes class-level, not school-level, randomisation.
    • I

      Independent Conduct

      • The intervention was designed and personally taught by the paper's first author, with no independent or external evaluation team involved.
      • "The whole unit of instruction was taught by the first author of this paper who is a formally trained secondary school teacher of both English and science. Hence, the teacher was not blind to the condition." (p. 3)
      • Relevant Quotes: 1) "To create comparable conditions, the whole unit of instruction was taught by the first author of this paper who is a formally trained secondary school teacher of both English and science. Hence, the teacher was not blind to the condition." (p. 3) 2) "The study was financially supported by the Ministry of Science, Research, and Arts Baden-Wuerttemberg, Germany and the University of Education, Ludwigsburg, Germany. However, these institutions did not influence the research processes in any way." (Acknowledgements) Detailed Analysis: Far from being conducted independently of the intervention's designers, this study was delivered in-person by one of the paper's own authors, who designed and taught both the monolingual and bilingual versions of the unit herself and was explicitly not blind to condition. There is no mention anywhere in the paper of an external agency handling delivery, data collection, or analysis; the funding acknowledgement only concerns financial support, not independent conduct of the trial itself. Final summary: Criterion I is not met because the same author who designed the study also personally delivered the intervention and was aware of condition assignment, with no independent third party involved.
    • Y

      Year Duration

      • The study tracked outcomes for at most a few months total (intervention plus a 6-week follow-up), far short of 75% of an academic year, and criterion T was already not met.
      • "... a 30-min follow-up test 6 weeks after instruction ended." (p. 3)
      • Relevant Quotes: 1) "Then, the intervention was implemented through a teaching unit on the topic 'Floating and Sinking' consisting of five lessons lasting 90 min each." (p. 3) 2) "The lessons were followed by a 60-min posttest and a 30-min follow-up test 6 weeks after instruction ended." (p. 3) Detailed Analysis: Per the ERCT specification, Criterion Y automatically fails if Criterion T is not met, and T was not met above. Independently of that rule, the substantive facts also fail Y on their own terms: the entire measured period, from intervention start to the final follow-up assessment, spans at most a few months (five lessons plus a 6-week follow-up), nowhere near 75% of a 9-10 month academic year. Final summary: Criterion Y is not met, both because Criterion T was not met and because the actual tracking period is far shorter than an academic year.
    • B

      Balanced Control Group

      • Instructional time and teaching materials were held constant across both groups; the only difference was the language of instruction plus minor language scaffolding integral to the bilingual condition itself.
      • "The teaching material and the instructional time were held constant between groups. The only variation was in the language of instruction, and additional language support ... was given to the bilingual group." (p. 3)
      • Relevant Quotes: 1) "The teaching material and the instructional time were held constant between groups. The only variation was in the language of instruction, and additional language support (helpful words and phrases on word cards or on the worksheets) was given to the bilingual group." (p. 3) 2) "Following the idea of CLIL, the mother tongue was not abolished from the classroom completely; nevertheless, English was the dominant language for both instruction and classroom discourse in the bilingual group." (p. 3) 3) "In both groups students did not receive grades for their work during the intervention." (p. 3) Detailed Analysis: Applying the criterion B decision procedure: extra resources are present (word cards/phrases with helpful vocabulary given only to the bilingual group), but this difference is negligible in time/budget terms (simple printed vocabulary aids, not extra class time, staffing, or materials budget), and, more importantly, this language scaffolding is not a separable add-on but a necessary, integral component of implementing the CLIL treatment itself (students cannot receive content instruction in a foreign language without some lexical support). Under the ERCT B decision tree this falls under "resources are the treatment" / "negligible difference," both of which resolve to met. No additional instructional time or budget was granted to either arm: instructional time (5 x 90 min lessons) and core teaching material were explicitly held constant, and neither group received grades for their work. Final summary: Criterion B is met because instructional time and core materials were equal across conditions, and the only extra resource (language scaffolding) is integral to, and negligible relative to, the treatment variable under study (language of instruction).
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study by a different research team was found in the paper or via external search.
      • Relevant Quotes: 1) "These results are in line with the studies by Hartmannsgruber (2014) and Marsh et al. (2000)." (p. 7, Discussion) 2) "Hence, future studies should replicate these findings using a preparatory training on bilingual instruction and older students with higher ability levels." (p. 7, Discussion) 3) "Therefore, future studies should replicate our findings on small negative CLIL-effects on science learning in other settings of bilingual teaching ..." (p. 7-8, Discussion) Detailed Analysis: The paper itself explicitly frames replication as future work still to be done ("future studies should replicate ..."), not as something that has already occurred. The cited prior studies (Marsh et al., 2000; Hartmannsgruber, 2014) are earlier, independent works with different interventions/populations that happen to show broadly consistent (negative) CLIL effects, but they are not replications of this specific "Floating and Sinking" experiment conducted by a different team in a different context. An external search (OpenAlex citation graph, ~70 papers citing this study as of 2026, plus Unpaywall/DOI checks) did not identify any published attempt by another research team to independently reproduce this particular Piesche et al. (2016) experiment (same sixth-grade Realschule sample, same "Floating and Sinking" physics unit) and its results. Final summary: Criterion R is not met because no independent replication of this specific study was found either in the paper or through external search.
    • A

      All-subject Exams

      • Only science (physics) content knowledge was assessed; no other core subjects were measured, and criterion E was not met.
      • "Second, the achievement measure that we used focused on one aspect of science competence, namely content knowledge." (p. 7, Limitations)
      • Relevant Quotes: 1) "Second, the achievement measure that we used focused on one aspect of science competence, namely content knowledge." (p. 7, Limitations) 2) "We conducted a randomised controlled field experiment ... on the effects of CLIL on subject-matter learning in science (physics)." (p. 2) Detailed Analysis: Per the ERCT Standard, Criterion A requires Criterion E to be met as a prerequisite, and E was not met above (custom test, not a standardised exam), which alone means A cannot be met. Independently, the study's own Limitations section acknowledges that only one science sub-competence (content knowledge on "Floating and Sinking") was assessed; no other core subjects (e.g. mathematics, reading/language arts) were measured, and no justification is given for restricting assessment to this single subject in a way that would fit the standard's specialised-intervention exception. Final summary: Criterion A is not met, both because Criterion E was not met and because only one subject (science) was assessed with no all-subject coverage or justified exception.
    • G

      Graduation Tracking

      • Tracking stopped 6 weeks after the intervention with no further follow-up reported, and criterion Y was not met.
      • "... a 30-min follow-up test 6 weeks after instruction ended." (p. 3)
      • Relevant Quotes: 1) "The lessons were followed by a 60-min posttest and a 30-min follow-up test 6 weeks after instruction ended." (p. 3) 2) "Hence, future studies should replicate these findings using ... a longer follow-up interval." (p. 7, Discussion) - the authors themselves flag the short follow-up as a limitation to be addressed by later work. Detailed Analysis: Per the ERCT Standard, Criterion G automatically fails if Criterion Y is not met, and Y was not met above. Substantively, the paper's data collection ends 6 weeks after the intervention concluded; there is no mention of any further follow-up. An external search for later publications by these authors (Piesche, Jonkmann, Fiege, Keßler) tracking the same cohort found a related 2021 paper on test-language effects in bilingual education by overlapping authors, but it is a separate re-analysis of test-language effects, not a longitudinal follow-up of this cohort toward graduation; no graduation-tracking publication was found. Final summary: Criterion G is not met because tracking stopped 6 weeks post-intervention, Criterion Y was not met, and no subsequent graduation-tracking publication by the same authors was found.
    • P

      Pre-Registered

      • No pre-registration of the study protocol is mentioned anywhere in the paper, and none was found via external search.
      • Relevant Quotes: No quote could be found anywhere in the Methods, Procedure, Acknowledgements, or Discussion sections referencing a pre-registration platform, registry ID, or a published a-priori protocol/analysis plan for this study. Detailed Analysis: A full read of the paper, including the methods, statistical analysis, acknowledgements, and reference list, contains no reference to pre-registration (e.g., no OSF, AsPredicted, ISRCTN, or similar registry mention) and no statement that hypotheses or analyses were specified in advance of data collection. An external search (OSF registries, DOI/publisher record) did not locate a pre-registration record for this study, consistent with pre-registration not yet being standard practice in this subfield at the time of data collection (submitted 2014-2015). Final summary: Criterion P is not met because no evidence of pre-registration was found in the paper or through external search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.