Abstract
This report presents findings from a randomised controlled trial (RCT) of a Community-Based English Language (CBEL) intervention aimed at people with very low levels of functional English proficiency. The intervention consisted of 66 hours of guided learning and support delivered through 22 classes and 11 club sessions over an 11-week period. The programme sought to test whether the intervention was effective in supporting individuals from communities with very low levels of functional English to improve their ability to communicate in English, and integrate into their wider community and society. A two-armed randomised controlled trial was deployed, with a waiting list design for the control group. In total, 527 participants were recruited to the trial. The trial found a strong and sizable difference in overall English proficiency and the amount of progress made between the treatment and control groups across all the English proficiency domains (speaking and listening, reading and writing).
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was conducted at the individual level (with small friend/family clusters) within each delivery centre, not at the class level, and the intervention was group classes rather than one-to-one tutoring, so no exception applies.
- "For reasons of practicality and to ensure the intervention was viable, randomisation was clustered by groups of known learners registered on the course" (p. 11, footnote 3)
Relevant Quotes:
1) "For reasons of practicality and to ensure the intervention was viable, randomisation was clustered by groups of known learners registered on the course (discussed below)." (p. 11, footnote 3)
2) "It shows how the trial took place in multiple delivery centres and randomisation took place at each centre to create a treatment (T) and control group (C)." (p. 13)
3) "The trial therefore adopted a cluster randomisation approach. All participants were asked at registration if they knew anyone else attending the course in that or another centre... Participants known to each other were then randomised into the treatment and control conditions together." (p. 19)
4) "Around thirty per cent of trial participants belonged to a cluster larger than a single individual. The mean cluster size was 1.21 participants" (p. 19)
5) "Randomisation was further restricted by the need to ensure roughly equal numbers of participants were assigned to each of the trial conditions at each centre... Therefore, randomisation was stratified by centre." (p. 20)
6) "Learners were expected to attend three sessions per week - two sessions of the Talk English Together classes and one session of the Talk English Together club (both elements were compulsory)." (p. 15)
Detailed Analysis:
Criterion C requires randomisation of entire classes (or schools) to prevent contamination between treatment and control participants sharing a class. In this trial the unit of randomisation was the individual participant, or a small cluster of friends/family members who registered together (mean cluster size 1.21). Randomisation was stratified by delivery centre, meaning participants from the same centre and community were split between treatment and control conditions, then formed into classes after allocation. This is individual-level (small-cluster) randomisation within sites, not class-level randomisation. The intervention was delivered as group classes of up to 12-15 learners taught by an ESOL teacher, not personal one-to-one tutoring, so the tutoring exception does not apply. Although the friend-cluster approach was intended to reduce contamination risk, it does not satisfy the class-level (or stronger) randomisation requirement.
Criterion C is not met because randomisation occurred at the individual/friendship-cluster level within centres, not at the class or school level, and the group-class intervention does not qualify for the tutoring exception.
-
E
Exam-based Assessment
- English proficiency was measured with bespoke tests developed by the English Speaking Board specifically for this project, not a widely recognised standardised exam.
- "The primary outcome measure, English language proficiency, was assessed by bespoke tests (copies of which are available on request)." (p. 16)
Relevant Quotes:
1) "The primary outcome measure, English language proficiency, was assessed by bespoke tests (copies of which are available on request). The tests provide a ten-point scale on which English proficiency in speaking and listening, reading and writing can be measured" (p. 16)
2) "The scale and the tests were developed by ESB for this project based on their expert knowledge of assessment methods and were consistent with the ESOL core curriculum." (pp. 16-17)
3) "Thirdly, the trial employed bespoke measures of English proficiency which assessed ability from Pre-entry to Entry level 2... Further development of the tool (i.e. expansion to include assessment to at least Entry level 3) is recommended to enable assessment above this parameter." (p. 61)
4) "Some of the questions were adapted from existing surveys (including the Citizenship Survey and European Social Survey) while others were created specifically for this research." (p. 17)
Detailed Analysis:
Criterion E requires that outcomes be measured with standard, widely recognised standardised exams rather than instruments created for the study. Here the primary outcome (English proficiency across speaking and listening, reading and writing) was assessed using tests explicitly described as "bespoke" and "developed by ESB for this project". Although the English Speaking Board is a recognised awarding body and the tests were aligned with the national ESOL core curriculum, the instruments themselves were custom-built for this trial, are not a named, widely used standardised examination, and the authors themselves flag the bespoke nature and the ceiling limitations of the tool as a study limitation. The secondary social integration outcome was likewise measured with a survey partly created specifically for this research.
Criterion E is not met because the primary outcome was assessed with bespoke, project-specific tests rather than a widely recognised standardised exam.
-
T
Term Duration
- Outcomes were measured at the end of the 11-week intervention, an interval shorter than a full academic term of approximately 3-4 months.
- "The intervention consisted of 66 hours of guided learning and support delivered through 22 classes and 11 club sessions over an 11-week period." (p. 0, Abstract)
Relevant Quotes:
1) "The intervention consisted of 66 hours of guided learning and support delivered through 22 classes and 11 club sessions over an 11-week period." (p. 0, Abstract)
2) "Participants assigned to the treatment group received the intervention over an 11-week period between April - July 2016." (p. 7)
3) "Pre-intervention measures were collected as close to the commencement of the trial as possible (in most instances during the first week of treatment group classes). 'Post' measures were collected as close to the conclusion of the learning as possible (in most instances during the final week of treatment group classes)." (p. 16)
4) "The start of the intervention was staggered, with centres beginning classes between 25th April and 9th May 2016." (p. 16)
5) "Finally, the time-bound nature of the trial, which was constrained to the duration of the course, was unlikely to be sufficiently long to capture the full impacts of the intervention... Follow up work over a longer period is therefore recommended." (p. 61)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. In this trial, the intervention began between 25 April and 9 May 2016 and follow-up measures were taken during the final week of the 11-week course (by July 2016). The interval from intervention start to outcome measurement was therefore about 11 weeks (roughly 2.5 months), which falls short of a full academic term of 3-4 months. There was no later follow-up measurement; the authors explicitly acknowledge the time-bound nature of the trial as a limitation and recommend longer follow-up work.
Criterion T is not met because the interval from intervention start to outcome measurement was only about 11 weeks, shorter than a full academic term.
-
D
Documented Control Group
- The waiting-list control group is thoroughly documented, including its size, socio-demographic profile, baseline proficiency scores and the fact it received no provision during the trial period.
- "Following randomisation, 249 trial participants were assigned to the treatment condition and 278 to the control." (p. 21)
Relevant Quotes:
1) "A two-armed randomised controlled trial was deployed, with a waiting list design for the control group. All trial participants received the intervention, with those assigned to the treatment group receiving the intervention over an 11-week period between April and July 2016, and those assigned to the control group receiving the intervention from September 2016." (p. 0, Abstract)
2) "Following randomisation, 249 trial participants were assigned to the treatment condition and 278 to the control." (p. 21)
3) "Baseline scores for speaking and listening were 4.42 for the treatment group compared with 4.27 for the control group (t=0.768, p=.443). Baseline writing scores were 3.63 for the treatment group and 3.45 for the control group (t=0.893, p=.372)." (p. 30)
4) "Table 3: Demographic balance at randomisation, baseline and follow-up points" (p. 34) [reporting gender, age, language, area, education, years in the UK and children for treatment and control groups at three stages]
5) "there was little evidence of control group contamination insofar that the control group participants did not have access to course materials nor were they able to attend any classes during the intervention period." (pp. 58-59)
Detailed Analysis:
Criterion D requires detailed documentation of the control group, including demographics, baseline performance and the conditions it experienced. The report documents the control group extensively: its size (278 randomised), its socio-demographic composition (gender, age, language, local authority area, education, years in the UK, children) at randomisation, baseline and follow-up in Table 3, its baseline English proficiency scores in each domain (including the noted baseline imbalance in reading), and its condition (waiting-list, receiving the course from September 2016, with no access to materials or classes during the trial). Balance checks between the arms are reported and discussed in detail.
Criterion D is met because the control group's size, demographics, baseline scores and treatment conditions are clearly and comprehensively documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation was stratified within each of the 22 delivery centres, splitting participants inside every centre between arms, so no school/site-level randomisation took place.
- "Randomisation was further restricted by the need to ensure roughly equal numbers of participants were assigned to each of the trial conditions at each centre... Therefore, randomisation was stratified by centre." (p. 20)
Relevant Quotes:
1) "It shows how the trial took place in multiple delivery centres and randomisation took place at each centre to create a treatment (T) and control group (C)." (p. 13)
2) "Randomisation was further restricted by the need to ensure roughly equal numbers of participants were assigned to each of the trial conditions at each centre. This was important in order to minimise variation in class to tutor ratio (twelve learners to one tutor was considered the optimal balance). Therefore, randomisation was stratified by centre." (p. 20)
3) "Randomisation only occurred once a centre had registered a sufficient number of participants to run two courses (one immediately for those assigned to the treatment group and one in September for those in the control group)." (p. 20)
Detailed Analysis:
Criterion S requires randomisation among schools or equivalent implementing institutions (here, the 22 community delivery centres would be the analogous units). Instead, every centre contained both treatment and control participants: individuals (and small friend clusters) were randomised within each centre, with stratification by centre to keep the arms balanced at each site. No centres or institutions were randomly assigned as whole units to conditions.
Criterion S is not met because randomisation occurred within centres at the individual/cluster level rather than between whole centres or institutions.
-
I
Independent Conduct
- The intervention was developed and delivered by Manchester Talk English under MHCLG, while the trial procedures, randomisation, data collection and analysis were carried out by independent evaluators (Learning and Work Institute, BMG Research and ESB assessors) with external oversight.
- "The Learning and Work Institute, in partnership with BMG Research, were commissioned to implement RCT procedures, including the collection of measures and analysis." (p. 11)
Relevant Quotes:
1) "The programme was designed and overseen by the Ministry of Housing, Communities and Local Government (MHCLG) who were also responsible for the design of the trial with input from the Behavioural Insights Team. The Learning and Work Institute, in partnership with BMG Research, were commissioned to implement RCT procedures, including the collection of measures and analysis." (p. 11)
2) "Manchester Talk English were commissioned by MHCLG to develop and deliver a Community-Based English Language intervention, derived from their existing Talk English programme... They designed the overall programme, developed the intervention manual and teaching materials" (p. 13)
3) "Once registered, participants were randomised by L&W into either treatment or control groups" (p. 19)
4) "Outcome measures were obtained by a series of English language proficiency tests which were developed by the English Speaking Board (ESB) and administered by ESB assessors and CBEL tutors, and a paper based survey administered by BMG research (on behalf of L&W)." (p. 16)
5) "Oversight for the evaluation was provided by advisors from the Cross-Government Trial Advice Panel including Government and academic experts in the field of impact evaluation." (p. 11)
6) "All of the assessments were marked by ESB assessors." (p. 24)
Detailed Analysis:
Criterion I requires that the evaluation be conducted independently of the intervention's designers. The intervention itself (the Talk English Together course) was developed and delivered by the provider, Manchester Talk English, commissioned by the government funder MHCLG. The evaluation - randomisation, collection of outcome measures and analysis, and this report - was carried out by a separate research organisation, the Learning and Work Institute, working with BMG Research; proficiency assessments were developed and marked by the English Speaking Board, a third-party awarding organisation, and the whole evaluation was overseen by the Cross-Government Trial Advice Panel including academic experts. The report's authors are all L&W researchers, not intervention developers. This mirrors the standard's accepted example of a government-commissioned trial being independent of the programme developer, provided the evaluation itself (data collection, analysis, conclusions) is run by a separate party from the developer/deliverer. A caveat is that assessors and researchers could not be blinded to allocation at assessment events, and MHCLG (the programme funder/designer) designed the trial; nonetheless, the data collection, marking and analysis were performed by parties independent of the intervention developer and deliverer.
Criterion I is met because randomisation, outcome measurement and analysis were performed by independent evaluators (L&W, BMG Research, ESB) distinct from the intervention developer and deliverer, with external expert oversight.
-
Y
Year Duration
- The trial tracked outcomes only over the 11-week course, far short of 75 per cent of an academic year, and criterion T is also unmet.
- "Participants assigned to the treatment group received the intervention over an 11-week period between April - July 2016." (p. 7)
Relevant Quotes:
1) "Participants assigned to the treatment group received the intervention over an 11-week period between April - July 2016. The control group received the intervention from September 2016, once the trial had concluded. Measures of English proficiency and social integration were taken at the beginning and end of the intervention period." (p. 7)
2) "'Post' measures were collected as close to the conclusion of the learning as possible (in most instances during the final week of treatment group classes)." (p. 16)
3) "Finally, the time-bound nature of the trial, which was constrained to the duration of the course, was unlikely to be sufficiently long to capture the full impacts of the intervention... Follow up work over a longer period is therefore recommended." (p. 61)
Detailed Analysis:
Criterion Y requires outcome tracking covering at least 75 per cent of an academic year (roughly 9-10 months) from intervention start. Here the entire measurement window was the 11-week course (late April/early May to July 2016), about 2.5 months, with no later follow-up. This is far below the required duration, and because criterion T (term duration) is not met, criterion Y automatically fails as well.
Criterion Y is not met because outcomes were tracked for only about 11 weeks, well short of 75 per cent of an academic year.
-
B
Balanced Control Group
- The extra 66 hours of English provision is itself the treatment variable being tested against a waiting-list business-as-usual control, so the deliberate resource difference is integral to the research question.
- "Do individuals who attend an 11-week Community-Based English class have significantly better levels of English Proficiency than individuals who do not?" (p. 12)
Relevant Quotes:
1) "The trial therefore sought to address the following research questions: 1. Do individuals who attend an 11-week Community-Based English class have significantly better levels of English Proficiency than individuals who do not?" (p. 12)
2) "The intervention consisted of 66 hours of guided learning and support delivered through 22 classes and 11 club sessions over an 11 week period." (p. 7)
3) "A randomised controlled trial design was implemented using a waiting list design for the control group... The control group received the intervention from September 2016, once the trial had concluded." (p. 7)
4) "To strengthen the evidence base MHCLG commissioned this Randomised Controlled Trial to test the impact of a Community-Based English Language intervention." (p. 12)
5) "...the control group participants did not have access to course materials nor were they able to attend any classes during the intervention period." (pp. 58-59)
Detailed Analysis:
Following the criterion B decision tree: the intervention group clearly received substantial additional resources relative to the control group during the trial period - 66 guided learning hours of classes and clubs, qualified ESOL teachers, volunteers and teaching materials - while the waiting-list control group received nothing until September 2016. The difference is not negligible. The decisive question is whether these additional resources are the explicit treatment variable. They are: the trial was commissioned specifically "to test the impact of a Community-Based English Language intervention", and the research questions directly compare individuals who attend the 11-week course against individuals who do not. The provision of the course (its time, teaching and materials) is the intervention itself, tested against a business-as-usual baseline - exactly the situation the standard's exception covers (extra provision as the primary treatment variable versus business as usual). It should be clearly noted that the intervention group received 66 hours of additional educational input that the control group did not receive during the trial; this imbalance is intentional and integral to the design rather than a confounding add-on.
Criterion B is met because the additional 66 hours of CBEL provision is explicitly the treatment variable under test against a documented business-as-usual waiting-list control.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this CBEL trial by a different research team has been published; the report itself describes the trial as novel evidence-building, and a further internet search found no such replication.
Relevant Quotes:
1) "To strengthen the evidence base MHCLG commissioned this Randomised Controlled Trial to test the impact of a Community-Based English Language intervention." (p. 12)
2) "A RCT is commonly considered the strongest form of impact evaluation." (p. 12)
Detailed Analysis:
Criterion R requires that the study be independently replicated by a different research team in a different context, published in a peer-reviewed journal. The report contains no reference to any replication of this trial; it was commissioned precisely because rigorous evidence on community-based English language provision was lacking. An internet search (Google Scholar and general web search) for replications of the CBEL/Talk English Together trial and its authors identified only the original MHCLG report, its protocol annex (Annex A), the accompanying process evaluation and survey annexes, and earlier non-experimental CBEL pilot evaluations - no independent peer-reviewed replication of this specific trial by a different research team was found. No citing academic literature reproducing the CBEL design in a different context could be located.
Criterion R is not met because no independent, peer- reviewed replication of this trial could be identified in the paper or through external searching.
-
A
All-subject Exams
- Only English language proficiency was assessed, criterion E is unmet, and no other subjects were measured, so the all-subject exams requirement fails.
- "Assessments covered three domains of English proficiency: speaking and listening, reading and writing, each of which was assessed against a ten-point scale." (p. 23)
Relevant Quotes:
1) "Assessments covered three domains of English proficiency: speaking and listening, reading and writing, each of which was assessed against a ten-point scale." (p. 23)
2) "As the intervention was designed to specifically improve functional English language proficiency, the primary outcome measure for the trial was the impact on English language proficiency. Social integration measures were considered secondary outcomes." (p. 12)
Detailed Analysis:
Criterion A requires standardised exam-based assessment across all main subjects, and explicitly fails if criterion E fails. Here criterion E is not met (bespoke tests), which alone fails criterion A. Additionally, only English proficiency was assessed; no other subject was measured. For this adult ESOL population one could argue English is the sole relevant subject, but the prerequisite of a standardised exam is still unfulfilled, and the secondary outcomes (social integration) are non-academic survey measures.
Criterion A is not met because criterion E fails and the assessments covered only English proficiency using bespoke instruments.
-
G
Graduation Tracking
- Measurement stopped at the end of the 11-week course with no longer-term or completion/graduation tracking, prerequisite criterion Y is unmet, and no follow-up publications tracking this cohort were found online.
- "Finally, the time-bound nature of the trial, which was constrained to the duration of the course, was unlikely to be sufficiently long to capture the full impacts of the intervention." (p. 61)
Relevant Quotes:
1) "Finally, the time-bound nature of the trial, which was constrained to the duration of the course, was unlikely to be sufficiently long to capture the full impacts of the intervention. In particular, changes in social integration may further manifest over a longer period... Follow up work over a longer period is therefore recommended." (p. 61)
2) "'Post' measures were collected as close to the conclusion of the learning as possible (in most instances during the final week of treatment group classes)." (p. 16)
Detailed Analysis:
Criterion G requires tracking participants until graduation from their educational stage. In this adult community education context, the trial measured outcomes only at the end of the 11-week course, and the authors explicitly state the trial was constrained to the course duration and recommend future longer-term follow-up. An internet search for subsequent publications by the same authors (Patel, Hoya, Bivand, McCallum, Stevenson, Wilson) or MHCLG/Learning and Work Institute tracking this same CBEL/Talk English cohort further did not identify any follow-up publications. Additionally, criterion Y is not met, which under the ranking rules means criterion G cannot be met regardless of any later tracking.
Criterion G is not met because tracking ceased at the end of the 11-week course with no longer-term follow-up found, and prerequisite criterion Y is unmet.
-
P
Pre-Registered
- A trial protocol exists as an annex to the report, but there is no evidence of registration on a public trial registry before data collection began; no ISRCTN or similar registry entry could be found online.
- "In collaboration with the Behavioural Insights Team, MHCLG developed an RCT protocol (see Annex A) which underpinned the design and implementation of the trial." (p. 13)
Relevant Quotes:
1) "In collaboration with the Behavioural Insights Team, MHCLG developed an RCT protocol (see Annex A) which underpinned the design and implementation of the trial." (p. 13)
2) "Retention within the trial was high and in line with expectations (as outlined in the trial protocol)" (p. 8)
3) "Power calculations carried out prior to the implementation of the trial (detailed in the RCT protocol in Annex A) suggested that a final achieved sample of 400 participants equally split over both arms would be required" (p. 21)
Detailed Analysis:
Criterion P requires the full study protocol to be pre-registered on a public registry before data collection begins, with verifiable registration timing. The report shows that a written protocol with hypotheses, design and power calculations existed before implementation and was later published as Annex A alongside the report in March 2018 (the Annex A PDF's own file metadata records a modification date of 13 March 2018, i.e. published together with the final report, not before the trial began in April 2016). An internet search of the ISRCTN registry and general web/Google Scholar searches for "Talk English", "Community-Based English Language" and MHCLG trial registration found no registry entry, ID, or registration date on ISRCTN, ClinicalTrials.gov, the AEA RCT registry or any similar database predating data collection in 2016. An internal protocol published retrospectively with the final report does not satisfy the requirement for verifiable public pre-registration.
Criterion P is not met because no public pre-registration with a registry ID and date preceding data collection is documented in the paper or discoverable externally.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.