Abstract
The effects of focusing second/foreign language (L2) learners' attentions on phonological forms while communicating in meaningful discourse has recently attracted attention in L2 pronunciation research. One such treatment is focus-on-form (FonF) instruction wherein L2 learners practice and notice pronunciation features in communicative tasks rather than in decontextualized exercises and drills (i.e., focus-on-forms [FonFS]). Given this, the current study investigated the differential effects of FonF and FonFS instructions on improving Iranian English as a foreign language (EFL) learners' pronunciation of the most problematic English consonants. After identifying the problematic English consonants (i.e., /θ/, /ð/, /w/, /ŋ/) via remedial and expert judgment approaches, 45 pre-intermediate learners embarked on an 8-hour course. The experimental group received FonF, the comparison group received FonFS, and the control group had a free conversation class minus any feedback on the target consonants. Learners' pronunciations were measured in terms of phonemic accuracy and comprehensibility in controlled and spontaneous tasks. The results of immediate and delayed post-test for phonemic accuracy revealed that whereas both FonF and FonFS were equally effective in controlled tasks, only FonF instruction proved effective up to the delayed post-test in spontaneous tasks; no such improvements, however, were observed for the control group. Results also showed that improvements in phonemic accuracy led to overall comprehensibility enhancements in EFL learners' speech. The article concludes with some pedagogical implications of the findings.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was carried out at the individual student level within a single cohort at one university, not at the class or school level, and the study is not a personal tutoring intervention.
- "The learners were randomly assigned to three groups: the experimental group (n=15) received FonF instruction, the comparison group (n=15) received FonFS instruction, and the control group (n=15) had a free conversation class..." (p. 116)
Relevant Quotes:
1) "The learners were randomly assigned to three groups: the experimental group (n=15) received FonF instruction, the comparison group (n=15) received FonFS instruction, and the control group (n=15) had a free conversation class with feedback on various pronunciation features but not the target consonants of the study (Table 1)." (p. 116)
2) "The participants of the study were 45 Iranian EFL English-majors studying in a non-state university in Tehran, Iran. All learners were in their third semester of university study..." (p. 115)
3) "This study had a quasi-experimental design with experimental, comparison, and control groups performing in the pre-test, immediate and delayed post-tests." (p. 119)
Detailed Analysis:
The paper explicitly states that individual learners (n=45), all drawn from a single third-semester cohort at one non-state university, were randomly assigned to one of three groups. This is student-level randomisation within what functions as a single classroom-sized population, not randomisation at the class or school level. There is no mention of randomising intact classes or institutions. The ERCT tutoring exception does not apply because the intervention consists of group instruction (15 learners per condition, taught together in sessions) rather than one-to-one personal tutoring. The paper itself even labels the design "quasi-experimental" in one place, underscoring that the unit of assignment was the individual student rather than an intact class or school.
Final sentence: Criterion C is not met because randomisation occurred at the individual student level within a single university cohort, with no class- or school-level randomisation and no valid tutoring exception.
-
E
Exam-based Assessment
- Both the phonemic accuracy measures and the comprehensibility rating scale were custom-designed by the researchers for this study rather than being widely recognised standardised exams.
- "Two L1 English EFL teachers ... were recruited to phonemically transcribe the extracted target words produced by EFL learners." (p. 116)
Relevant Quotes:
1) "Two L1 English EFL teachers (one male and one female, with the mean age of 34.1), who were also experts in English phonology, were recruited to phonemically transcribe the extracted target words produced by EFL learners." (p. 116)
2) "As for the former elicitation task, for each target consonant of the study, two sentences were loaded with six words (i.e., eight sentences in total). As for the latter spontaneous task, one picture was presented to students for each target consonant (four pictures in total)." (p. 117)
3) "Following previous research ... a rating scale from 1 (very easy to understand) to 9 (very difficult to understand) was developed and given to the EFL teachers to rate the comprehensibility of learners' pronunciations." (p. 117)
Detailed Analysis:
Phonemic accuracy was measured via researcher-built sentence read-aloud and picture description tasks targeting four specific consonants, scored through phonemic transcription by recruited teacher-raters, not through any standardised, widely recognised exam instrument. Comprehensibility was measured with a custom 1-9 rating scale developed specifically for this line of research, not a standard, validated, widely used assessment. Neither instrument is a national, state-wide, or otherwise externally validated standardised exam; both were purpose-built for the target consonants and constructs under study.
Final sentence: Criterion E is not met because the study relies entirely on custom-built phonemic accuracy tasks and a custom comprehensibility rating scale rather than standardised, widely recognised exams.
-
T
Term Duration
- The interval from intervention start to the final (delayed) measurement spans only about ten weeks, well short of one academic term.
- "Learners in these two groups had two 1-hour sessions for each consonant (eight sessions in eight weeks, equalling with eight hours of instruction in total)." (p. 120)
Relevant Quotes:
1) "One week after the pre-test, the instructional treatments began for FonF and FonFS groups. Learners in these two groups had two 1-hour sessions for each consonant (eight sessions in eight weeks, equalling with eight hours of instruction in total)." (p. 120)
2) "The immediate post-tests were taken right after the last instruction session of each target consonant, and the delayed post-test was administered two weeks after the immediate post-test of each target consonant with the same procedure." (p. 120)
Detailed Analysis:
The intervention itself ran across eight weeks (eight one-hour sessions), and the final, delayed measurement point occurred two weeks after the immediate post-test, i.e., approximately ten weeks after the intervention began. Ten weeks (about two and a half months) is shorter than the one full academic term (roughly 3-4 months) required by the ERCT standard. There is no year-long tracking that would allow this weaker criterion to be satisfied by the stronger Y criterion either.
Final sentence: Criterion T is not met because the total interval from intervention start to final measurement is approximately ten weeks, short of a full academic term.
-
D
Documented Control Group
- The control group's size, demographics, baseline proficiency, and the nature of its "business as usual" activity are clearly documented.
- "the control group (n=15) had a free conversation class with feedback on various pronunciation features but not the target consonants of the study (Table 1)." (p. 116)
Relevant Quotes:
1) "the control group (n=15) had a free conversation class with feedback on various pronunciation features but not the target consonants of the study (Table 1)." (p. 116)
2) Table 1 (Demographical characteristics and proficiency scores) reports the control group's values as: Number 15; Gender (Female/Male) 8 / 7; Age (mean/SD) 27.3 / 4.3; OPT (mean/SD) 29.8 / 5.3. (p. 116; this is a paraphrase of tabular data, not a running-text quotation)
3) "the control group only had a free discussion class on various pre-determined themes and topics at every session and received different feedback on linguistic aspects minus the target consonants of the study (see Figure 1)." (p. 118)
4) "The control group took the immediate and delayed post-tests in the same time intervals as the treatment groups." (p. 120)
Detailed Analysis:
Table 1 provides detailed baseline demographic and proficiency data (gender split, mean age and SD, mean Oxford Placement Test score and SD) for the 15-member control group, showing it was comparable to the treatment groups at baseline. The paper also clearly describes what the control group actually did each session (free discussion with feedback on non-target linguistic features) and confirms they were tested on the same schedule as the treatment groups. This level of detail satisfies the documentation requirement.
Final sentence: Criterion D is met because the paper provides clear baseline demographic, proficiency, and activity documentation for the control group.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred among individual students at a single university, not among schools.
- "The participants of the study were 45 Iranian EFL English-majors studying in a non-state university in Tehran, Iran." (p. 115)
Relevant Quotes:
1) "The participants of the study were 45 Iranian EFL English-majors studying in a non-state university in Tehran, Iran." (p. 115)
2) "The learners were randomly assigned to three groups: the experimental group (n=15) received FonF instruction, the comparison group (n=15) received FonFS instruction, and the control group (n=15) had a free conversation class..." (p. 116)
Detailed Analysis:
All 45 participants came from a single institution (one non-state university in Tehran), and randomisation was performed on individual students within that single population, not across multiple schools or institutions. There is no description of school selection or of multiple sites being randomised. This does not meet the school-level RCT requirement.
Final sentence: Criterion S is not met because the study involved a single university with student-level randomisation, not randomisation across schools.
-
I
Independent Conduct
- The intervention was designed and personally delivered by one of the authors, and scoring of the outcome data was also performed by one of the authors, with no independent third-party conduct.
- "Learners were required to listen carefully to the teacher's (the first author) explanations and audio-based English-native pronunciations and produce consonants accordingly." (p. 117)
Relevant Quotes:
1) "Learners were required to listen carefully to the teacher's (the first author) explanations and audio-based English-native pronunciations and produce consonants accordingly." (p. 117)
2) "While performing the tasks, learners were also provided with explicit correction feedback by the teacher if they had articulatory mistakes." (pp. 117-118)
3) "Based on these transcriptions, one of the authors assigned scores to learners' productions." (p. 119)
Detailed Analysis:
The first author personally served as the classroom teacher delivering both the FonF and FonFS instruction, including giving explicit correction feedback to learners. Additionally, one of the authors (rather than an independent evaluator) converted the external teacher-transcribers' data into phonemic accuracy scores. There is no mention anywhere in the paper of an external, third-party agency independently conducting or overseeing the trial's implementation or core scoring/analysis decisions. This concentration of implementation and analysis roles in the study's own authors does not satisfy the independence requirement.
Final sentence: Criterion I is not met because the same author who designed the intervention also delivered instruction and scored outcome data, without independent third-party conduct.
-
Y
Year Duration
- Because criterion T (Term Duration) is not met, the stronger Year Duration criterion is automatically not met, and the actual tracking period is far shorter than a year in any case.
- "The immediate post-tests were taken right after the last instruction session of each target consonant, and the delayed post-test was administered two weeks after the immediate post-test..." (p. 120)
Relevant Quotes:
1) "One week after the pre-test, the instructional treatments began for FonF and FonFS groups. Learners in these two groups had two 1-hour sessions for each consonant (eight sessions in eight weeks, equalling with eight hours of instruction in total)." (p. 120)
2) "The immediate post-tests were taken right after the last instruction session of each target consonant, and the delayed post-test was administered two weeks after the immediate post-test of each target consonant with the same procedure." (p. 120)
Detailed Analysis:
As established for criterion T, the total interval from intervention start to the final delayed measurement is approximately ten weeks, which is far short of the roughly 9-10 month (75% of an academic year) requirement for Y. Per the ERCT specification, since T is not met, Y cannot be met either.
Final sentence: Criterion Y is not met because the study duration (about ten weeks) is far shorter than the required academic-year-length tracking, and T is also not met.
-
B
Balanced Control Group
- All three groups received identical total instructional/contact time (eight one-hour sessions), so no group was given extra time or resources beyond the others; the only difference is the instructional content/focus being tested.
- "Learners in these two groups had two 1-hour sessions for each consonant (eight sessions in eight weeks ... in total). Meanwhile, the control group had eight 1-hour free-discussion sessions." (p. 120)
Relevant Quotes:
1) "Learners in these two groups had two 1-hour sessions for each consonant (eight sessions in eight weeks, equalling with eight hours of instruction in total). ... Meanwhile, the control group had eight 1-hour free-discussion sessions." (p. 120; ellipsis marks three intervening sentences describing the 20-minute explicit phase and 100-minute practice phase for the FonF/FonFS groups)
2) "the control group (n=15) had a free conversation class with feedback on various pronunciation features but not the target consonants of the study." (p. 116)
3) "The control group took the immediate and delayed post-tests in the same time intervals as the treatment groups." (p. 120)
Detailed Analysis:
Applying the Criterion B decision procedure: the first question is whether the intervention groups received extra time or budget relative to the control. Here, the FonF and FonFS groups each received eight 1-hour sessions, and the control group likewise received eight 1-hour free-discussion sessions -- the same total contact time (8 hours) for all three arms, on the same testing schedule. No group received additional class time, materials budget, or teacher contact beyond the others. The only difference between conditions is the content and feedback focus of the sessions (focused communicative tasks vs. controlled drills vs. free discussion without target-consonant feedback), which is precisely the pedagogical variable under investigation, not an extraneous resource imbalance. Since no extra resources are present in any condition, the criterion is trivially satisfied.
Final sentence: Criterion B is met because all three groups received equal instructional time (eight 1-hour sessions each), with no group given extra time or resources beyond the others.
-
Level 3 Criteria
-
R
Reproduced
- No evidence of an independent replication of this specific study by a different research team was found in the paper, in citing literature, or through a dedicated internet search.
Detailed Analysis:
The paper does not reference any prior or subsequent independent replication of this specific FonF vs. FonFS pronunciation study by a different research team. An internet search (Semantic Scholar citation records for DOI 10.17576/3L-2018-2401-09) identified five citing works: Moeen, Nejadansari & Dabaghi (2019, implicit/explicit grammar teaching via scaffolding), Gonzalez Robaina & Diaz Larenas (2019, a theoretical model paper), Philip & Noyan (2018, metaphonological awareness instruction), Mustafa (2019, perception of /ɾ/ by Arabic speakers), and a same-authors follow-up, Tabandeh, Moinzadeh & Barati (2019), "Differential Effects of FonF and FonFS on Learning English Lax Vowels in an EFL Context" (Asia TEFL, DOI 10.18823/ASIATEFL.2019.16.2.4.499). None of these are independent replications of this specific study: the 2019 companion paper is by the same author team and targets a different linguistic feature (lax vowels, not the /θ/, /ð/, /w/, /ŋ/ consonants), so it does not count as independent reproduction, and the other citing works merely reference this paper rather than repeating its design in a new context.
Final sentence: Criterion R is not met because no independent replication of this specific study by a different research team was found in the paper or via internet search.
-
A
All-subject Exams
- Since criterion E (standardised exam-based assessment) is not met, criterion A cannot be met either; additionally, only pronunciation of four consonants was assessed, not other subjects.
- "the current study investigated the differential effects of FonF and FonFS instructions on improving Iranian ... learners' pronunciation of the most problematic English consonants." (p. 112, Abstract)
Relevant Quotes:
1) "the current study investigated the differential effects of FonF and FonFS instructions on improving Iranian English as a foreign language (EFL) learners' pronunciation of the most problematic English consonants." (p. 112, Abstract)
2) "four consonants (i.e., /θ/, /ð/, /w/, /ŋ/) achieved the highest scores on average and hence, were regarded as the target consonants of the study." (p. 116)
Detailed Analysis:
Per the ERCT specification, criterion E is a prerequisite for criterion A; since E was found not met (the assessments were custom, non-standardised instruments), A is automatically not met. Separately, the study exclusively measured pronunciation accuracy of four target consonants and overall speech comprehensibility, with no assessment of other subjects or broader academic outcomes.
Final sentence: Criterion A is not met because criterion E is not met and only a narrow pronunciation-focused outcome was assessed.
-
G
Graduation Tracking
- Tracking stopped two weeks after the immediate post-test, with no follow-up until graduation or any long-term tracking reported, and criterion Y is also not met.
- "the delayed post-test was administered two weeks after the immediate post-test of each target consonant with the same procedure." (p. 120)
Relevant Quotes:
1) "the delayed post-test was administered two weeks after the immediate post-test of each target consonant with the same procedure." (p. 120)
2) "Primarily, the EFL learners received only two hours of instruction for each target consonant. Although this amount of treatment time seems sufficient for experimental designs, longer treatments accompanied by more distant delayed post-tests reveals more valid results of long-term improvements (see Lee et al. 2015)." (p. 123, Discussion)
Detailed Analysis:
The study's final data collection point is the delayed post-test, occurring only two weeks after the immediate post-test and roughly ten weeks after the intervention began. The authors themselves note this as a limitation, calling for more distant delayed post-tests in future research. There is no mention of any follow-up study tracking these learners through to graduation from their degree program. An internet search for subsequent publications by the same authors found one related paper, Tabandeh, Moinzadeh & Barati (2019) on English lax vowels, but it studies a different linguistic feature and gives no indication of tracking the same 45-learner cohort toward graduation; no graduation-tracking follow-up was found. Per the ERCT specification, since criterion Y is not met, G cannot be met either.
Final sentence: Criterion G is not met because tracking ended shortly after the intervention with no graduation-level follow-up found in this paper or in subsequent publications, and Y is also not met.
-
P
Pre-Registered
- No statement of pre-registration, registry ID, or registration date is present anywhere in the paper, and no internet search evidence of registration was found.
Detailed Analysis:
A review of the methods, procedure, and acknowledgement sections of the paper reveals no mention of a pre-registration platform, registry ID, or a registration date prior to data collection. No protocol or statistical analysis plan is cited as having been published in advance of the study. An internet search did not surface any registry entry (e.g., OSF, AsPredicted) associated with this study or its authors for this design; pre-registration was also not a common practice in this subfield of applied linguistics at the time of data collection.
Final sentence: Criterion P is not met because there is no evidence of pre-registration anywhere in the paper or found through internet search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.