Abstract
English language teaching has newly been introduced to pre-school curriculum in Turkey. The purpose of this study was to investigate the effectiveness of teaching EFL vocabulary to pre-school children through Project Based Learning (PBL). For this purpose, an experimental design, consisted of observation checklists, exam scores and a short survey, was adopted. Firstly, through a short online survey, 150 kindergarten teachers were asked to specify which techniques they commonly used in their English classes. The primary aim here was to define traditional techniques and the rate of PBL use in Turkey. After defining common techniques, 28 children were randomly assigned to experimental (PBL instruction) and control groups (traditional instruction) equally and the data was collected in real time classroom setting for 8 weeks. The results showed that (1) PBL was rarely adopted in EFL classes in Turkey, (2) PBL instruction could increase EFL vocabulary learning gains when compared to common methods and (3) young learners were observed to have been more active in PBL classes. The effect of PBL instruction was discussed in local, cognitive and motivational perspectives in the light of previous related research. The potential benefits of further PBL use for young EFL learners and implications were also discussed.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was done at the individual student level rather than by class or school, and the tutoring exception does not apply, so this criterion is not met.
- "28 children were randomly assigned to experimental (PBL instruction) and control groups (traditional instruction) equally..." (Abstract, p. 426)
Relevant Quotes:
1) "After defining common techniques, 28 children were randomly assigned to experimental (PBL instruction) and control groups (traditional instruction) equally and the data was collected in real time classroom setting for 8 weeks." (Abstract, p. 426)
2) "The population of the study consisted of pre-school children with an age range of 5-6. The experimental group included 14 participants 6 boys and 8 girls) and the control group included 14 participants (8 boys and 6 girls)." (p. 430)
3) "The implementer of both groups is the same teacher." (p. 431)
Detailed Analysis:
Criterion C requires randomisation to occur at the class (or stronger, school) level, unless the intervention is a personal tutoring/one-to-one design, in which case student-level assignment is acceptable. Here, the paper describes randomly assigning 28 individual children to an experimental or control group, with no mention of pre-existing classes, sections, or schools being the unit of randomisation. Both groups were taught in group settings (PBL projects and traditional group techniques such as TPR, games, songs), not as one-to-one personal tutoring, so the tutoring exception does not apply. Since individual children drawn from what appears to be a single population were split into treatment/control at the student level, this creates a risk of the same contamination problem the criterion is meant to prevent (e.g., via shared exposure at the same kindergarten, potential interaction between the groups' children, or the same teacher implementing both conditions).
Final: Criterion C is not met because randomisation was conducted at the individual student level rather than at the class or school level, and no personal tutoring exception applies.
-
E
Exam-based Assessment
- The study used a custom, researcher-designed picture-matching test rather than a standardised, widely recognised exam.
- "Thus, the tests were kept simple and included the pictures and the art crafts of the vocabulary items as materials and real show of the students on the directed word." (p. 430)
Relevant Quotes:
1) "Both experimental and control groups took the same tests. After using appropriate methods and techniques for teaching during five days, the children were tested regularly on the fifth day (end of the week) after the instruction of the topic was completed. The exams were prepared and administered in accordance with the age range of the participants." (p. 430)
2) "Thus, the tests were kept simple and included the pictures and the art crafts of the vocabulary items as materials and real show of the students on the directed word. The pictures or the crafts of the related words (8 each week) were laid on the table and the students were asked to pin on the picture of the word they hear and see." (p. 430)
3) "The testing procedure was designed just like a game and the teacher gave 1 point for each correct answer. Total score that could be obtained was 8 points. Then these points were then arranged to 100-point grading system for more interpretable analysis and results." (p. 430)
Detailed Analysis:
Criterion E requires a standardised, widely recognised exam rather than a bespoke instrument designed specifically for the study. The paper explicitly describes a custom, researcher-designed weekly picture-pinning game created specifically to match the eight target vocabulary items selected for that week's topic. There is no reference to any externally validated, standardised vocabulary or language assessment (e.g., a national curriculum exam or a published/normed vocabulary test). The assessment was purpose-built to align tightly with the specific words taught, which is precisely the custom-assessment bias problem the criterion is designed to detect.
Final: Criterion E is not met because the outcome measure was a custom, researcher-designed weekly picture-matching test, not a standardised exam.
-
T
Term Duration
- The intervention and final measurement both occurred within an 8-week (about 2-month) window, which is shorter than the minimum one-term duration required.
- "For both groups, the study lasted 8 weeks in parallel fashion, in which a new topic was provided to children per week." (p. 431)
Relevant Quotes:
1) "The study was an experimental design including a control and an experimental group. Individuals in the experimental group received project based instruction while those in the control group received traditional instruction. Both instructional periods lasted for 8 weeks and each week, learner development was assessed by an exam and observer checklists." (p. 429)
2) "For both groups, the study lasted 8 weeks in parallel fashion, in which a new topic was provided to children per week." (p. 431)
3) "At the end of the 8 weeks, each groups took a total of 1400 minutes of instruction." (p. 431)
Detailed Analysis:
Criterion T requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. The intervention and its final measurement both occurred within an 8-week window (approximately 2 months), which is shorter than the minimum term-length duration required by the standard. There is no indication of a delayed or follow-up measurement extending the tracking period beyond the 8 weeks of instruction.
Final: Criterion T is not met because the total intervention-to-measurement interval was only about 8 weeks (roughly 2 months), short of a full academic term.
-
D
Documented Control Group
- The control group's size, demographics, prior exposure, and specific instructional content are all clearly documented in the paper.
- "For control group, traditional materials were chosen carefully depending on the most common techniques used in kindergarten for foreign language instruction in Turkey such as TPR, songs, games or flashcards." (p. 431)
Relevant Quotes:
1) "The experimental group included 14 participants 6 boys and 8 girls) and the control group included 14 participants (8 boys and 6 girls). All of the participants were Turkish and with L1 of Turkish. As the kindergarten did not have any English classes for children, it was the first time the children met with English as a foreign language. Moreover, they did not have contact of English at home." (p. 430)
2) "For control group, traditional materials were chosen carefully depending on the most common techniques used in kindergarten for foreign language instruction in Turkey such as TPR, songs, games or flashcards." (p. 431)
3) "Among these methods Total Physical Response was the most used method. As materials and activities; drawing and painting activities, printable worksheets, art craft activities, flashcards, pictures, real objects, games, cartoons, songs, and animations were also included." (p. 431)
4) "The courses took 35 minutes per a lesson each day and a total of 175 minutes per week for 8 weeks for each group. At the end of the 8 weeks, each groups took a total of 1400 minutes of instruction." (p. 431)
Detailed Analysis:
Criterion D requires clear documentation of the control group's demographics, baseline characteristics, and the conditions/treatment it received. The paper reports the control group's exact size (n=14), gender split (8 boys, 6 girls), L1 background (Turkish), prior exposure to English (none, at school or at home), and describes in detail the specific methods and materials used for "traditional instruction" (TPR, songs, games, flashcards, worksheets, etc.), which were themselves derived from a systematic survey of 150 teachers rather than left vague. The control group's instructional dosage (1400 minutes total) is also explicitly quantified.
Final: Criterion D is met because the control group's size, demographics, baseline exposure, and specific instructional content/dosage are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred among 28 individual children rather than among schools, so the stronger school-level criterion is not met.
- "After defining common techniques, 28 children were randomly assigned to experimental (PBL instruction) and control groups (traditional instruction) equally..." (Abstract, p. 426)
Relevant Quotes:
1) "After defining common techniques, 28 children were randomly assigned to experimental (PBL instruction) and control groups (traditional instruction) equally..." (Abstract, p. 426)
2) "The population of the study consisted of pre-school children with an age range of 5-6." (p. 430)
Detailed Analysis:
Criterion S requires randomisation at the level of whole schools (or equivalent institutions), not individual students. The paper gives no indication that multiple schools were involved or that randomisation occurred between schools; instead, 28 individual children were randomly assigned to conditions, apparently within a single population/ setting. There is no mention of school selection, number of participating schools, or a school-level randomisation procedure.
Final: Criterion S is not met because randomisation occurred at the individual student level, not at the school level.
-
I
Independent Conduct
- No independent, third-party evaluator is described; the same research team that designed the study (as a PhD thesis) also implemented and analysed it.
- "*This study was produced from PhD. Thesis by Fatma Kimsesiz supervised by Yavuz Konca in Atatürk University" (title page footnote, p. 426)
Relevant Quotes:
1) "*This study was produced from PhD. Thesis by Fatma Kimsesiz supervised by Yavuz Konca in Atatürk University" (title page footnote, p. 426)
2) "The implementer of both groups is the same teacher. The teacher introduced the topic and requested children to make suggestions about the content and design of the project." (p. 431)
3) "The children and the classroom environment was observed by using the checklists mentioned above." (p. 431)
Detailed Analysis:
Criterion I requires the evaluation to be conducted independently of those who designed the intervention, e.g., by an external, third-party evaluator, to reduce bias in implementation and analysis. The paper does not mention any external or independent evaluation team; instead, it states this study was produced from the first author's PhD thesis, and that a single teacher implemented both the experimental and control conditions, with the same authors apparently responsible for the observation, testing, and analysis. There is no statement of third-party oversight, blinded observers, or an independent data-collection agency.
Final: Criterion I is not met because there is no evidence of independent, third-party conduct; instead, the study was carried out by the same research team (as part of a PhD thesis) that designed the intervention.
-
Y
Year Duration
- The study covered only about 8 weeks, far short of 75% of an academic year, and criterion T is also not met.
- "For both groups, the study lasted 8 weeks in parallel fashion..." (p. 431)
Relevant Quotes:
1) "For both groups, the study lasted 8 weeks in parallel fashion, in which a new topic was provided to children per week." (p. 431)
2) "At the end of the 8 weeks, each groups took a total of 1400 minutes of instruction." (p. 431)
Detailed Analysis:
Criterion Y requires outcomes to be tracked over at least 75% of a full academic year (roughly 9-10 months). The intervention and measurement window in this study spanned only 8 weeks (about 2 months), which is far short of the required duration. Per the standard's own dependency rule, since the weaker Term Duration criterion (T) is not met, the Year Duration criterion (Y) cannot be met either.
Final: Criterion Y is not met because the study spanned only about 8 weeks, far short of 75% of an academic year, and the prerequisite T criterion is also not met.
-
B
Balanced Control Group
- Both groups received identical instructional time (1400 minutes total) and comparable materials, with only the teaching method differing, so the groups are balanced.
- "The courses took 35 minutes per a lesson each day and a total of 175 minutes per week for 8 weeks for each group. At the end of the 8 weeks, each groups took a total of 1400 minutes of instruction." (p. 431)
Relevant Quotes:
1) "For both groups, the study lasted 8 weeks in parallel fashion, in which a new topic was provided to children per week. The courses took 35 minutes per a lesson each day and a total of 175 minutes per week for 8 weeks for each group. At the end of the 8 weeks, each groups took a total of 1400 minutes of instruction." (p. 431)
2) "For experimental group, the implementation of the PBL included almost all the phases of the project based approach... The teacher prepared the required materials for the courses." (p. 431)
3) "For control group, traditional materials were chosen carefully depending on the most common techniques used in kindergarten for foreign language instruction in Turkey such as TPR, songs, games or flashcards." (p. 431)
Detailed Analysis:
Applying the Criterion B decision procedure: the paper explicitly states that both the experimental and control groups received identical amounts of instructional time (35 minutes/day, 175 minutes/week, 1400 minutes total over 8 weeks). Both groups also received teacher-prepared materials (project materials such as play-dough or paper for the experimental group; flashcards, worksheets, art craft materials, real objects, etc. for the control group). There is no indication that the experimental (PBL) group received extra class time, extra budget, or extra teacher attention beyond what the control group received; the two conditions differ only in pedagogical method/technique (PBL vs. TPR/games/ songs/flashcards), not in the quantity of time, money, or materials provided. Since no additional time or budget was given to either group beyond matched instructional dosage and comparable teacher- prepared materials, this criterion is trivially satisfied under the standard's decision tree (no extra resources present).
Final: Criterion B is met because both groups received an identical amount of instructional time and comparable teacher-prepared materials, with the only difference being the pedagogical method used.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study (pre-school EFL vocabulary via PBL) was found in the paper or elsewhere, including after an internet-based citation search.
Relevant Quotes:
1) "The study was conducted with the aim of exploring and evaluating the effect of using Project based approach in teaching English vocabulary to pre-school children compared to using traditional techniques in Turkish context." (p. 435)
2) No quotes in the paper reference any subsequent or independent replication of this specific study (pre-school EFL vocabulary via PBL, n=28, Turkey).
Detailed Analysis:
Criterion R requires that this specific study (its central experimental claim, design, and population) be independently replicated by a different research team, ideally published in a peer-reviewed outlet. The paper itself only cites earlier, related but distinct PBL studies with older learners (e.g., Çırak, 2006; Türker, 2007; Yıldız, 2009; Baş & Beyhan, 2010; Köroğlu, 2011), which are prior work the current study builds on, not replications of it.
Internet verification: the paper's citation record (39 citing works per OpenAlex, as of this check) was reviewed, including titles such as "An examination of vocabulary learning and retention levels of pre-school children using augmented reality technology in English language learning" (Yılmaz, Topu, & Takkuç Tulgar, 2022) and "Impact of Project-Based Learning on Critical Thinking Skills and Language Skills in EFL Context: A Review of Literature" (Song, Razali, Sulaiman, & Jeyaraj, 2024). None of the identified citing works are independent replications of this specific pre-school (age 5-6), n=28, Turkish-context PBL vocabulary study; they are either literature reviews/meta-analyses or studies of different populations, subjects, or designs. No evidence of an independent replication of this specific study was found in any available source.
Final: Criterion R is not met because no independent replication of this specific study has been found, including after an internet-based citation search.
-
A
All-subject Exams
- Only EFL vocabulary was assessed, no other subjects were measured, and Criterion E (a prerequisite) is not met.
- "Both experimental and control groups took the same tests... the children were tested regularly on the fifth day (end of the week)..." (p. 430)
Relevant Quotes:
1) "Both experimental and control groups took the same tests... the children were tested regularly on the fifth day (end of the week) after the instruction of the topic was completed." (p. 430)
2) "3. Is there an effect of PBL on vocabulary learning performance of young learners at the end of 8-week period?" (p. 429)
Detailed Analysis:
Criterion A requires assessment of all main subjects, not just the subject targeted by the intervention, and it presupposes that Criterion E (standardised exam) is met. Here, only EFL vocabulary (a single subject/skill area) was assessed via the weekly picture-matching tests; no other core subjects (e.g., mathematics, native-language literacy, science) were measured. In addition, since Criterion E was not met (the assessment was a custom, non-standardised instrument), Criterion A cannot be met either, per the standard's explicit dependency rule.
Final: Criterion A is not met because only vocabulary in the single subject of English was assessed, and because the prerequisite Criterion E is not met.
-
G
Graduation Tracking
- No follow-up tracking beyond the 8-week intervention is reported, and the prerequisite Criterion Y is not met; an internet search found no follow-up publication tracking this cohort.
Relevant Quotes:
1) "For both groups, the study lasted 8 weeks in parallel fashion, in which a new topic was provided to children per week." (p. 431)
2) No follow-up data collection beyond the 8-week instructional period is described anywhere in the paper; the Findings and Discussion, and Conclusion sections report only the results from the 8-week intervention window.
Detailed Analysis:
Criterion G requires participants to be tracked until graduation from their educational stage to assess long-term impact, and per the standard's dependency rule, it cannot be met if Criterion Y is not met. In this study, data collection ended with the final weekly test at the close of the 8-week intervention; there is no mention of any follow-up assessment after that point, nor of any subsequent publication tracking the same cohort of pre-school children through to graduation.
Internet verification: a search for subsequent publications by the same authors (Fatma Kimsesiz, Emrah Dolgunsöz, M. Yavuz Konca) was conducted via OpenAlex. No follow-up paper tracking the same n=28 pre-school cohort toward graduation was found; the only other located work by Kimsesiz is an unrelated 2019 study on EFL teacher burnout in Turkey, not a cohort follow-up.
Final: Criterion G is not met because tracking stopped at the end of the 8-week intervention with no follow-up to graduation, and Criterion Y is also not met, and no follow-up publication was found.
-
P
Pre-Registered
- The paper contains no mention of a pre-registered protocol, registry link, or registration date.
Relevant Quotes:
No statement anywhere in the paper (Method, Procedure, Data Analysis, or Discussion sections) mentions a trial registry, a pre-registration platform, or a publicly published protocol/analysis plan prior to data collection.
Detailed Analysis:
Criterion P requires quoted evidence of a pre-registered protocol (hypotheses, methods, planned analyses) published on a registry before data collection began. This paper, an experimental classroom study derived from a PhD thesis, contains no reference to any pre-registration, registry ID, or registration date. Absence of any such statement means the criterion cannot be considered met. As the paper itself gives no registry name or identifier to check, no further registry lookup was possible or warranted.
Final: Criterion P is not met because no pre-registration statement, registry link, or registration date is present anywhere in the paper.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.