Coding Science: The Role of Block-Based Programming Activities in Enhancing Computational Thinking and Science Achievement

Sukran Sungur, Gulbin Ozkan

Published:
ERCT Check Date:
DOI: 10.1007/s10956-026-10317-5
  • science
  • K12
  • EU
  • EdTech app
0
  • C

    Two intact sixth-grade classes were randomly assigned to conditions, making the class the unit of randomisation.

    Two classes were randomly assigned as either the experimental or control group.

  • E

    The study used research-developed and researcher-adapted instruments rather than recognised standardised exams.

    The pre- and post-test instruments for data collection in this research were the Computational Thinking Test (CTt; Roman-Gonzalez et al., 2017) and Force and Motion Knowledge test (FMKt; Cimentepe, 2019).

  • T

    The intervention and post-test spanned only six weeks, far shorter than one academic term, with no follow-up.

    The research lasted a total of six weeks. Both groups received an identical amount of instructional time (4 lesson hours per week for 6 weeks).

  • D

    The control group's size, demographics, instruction, and baseline scores are clearly documented.

    The control group followed the standard curriculum using materials (textbooks, worksheets, physical experiments) and did not engage in informal use of BBPEs or similar technologies during the instructional hours.

  • S

    Randomisation was at the class level within one school, not at the school level.

    potential information sharing between experimental and control groups within the same school remains a limitation. Future research could use separate schools to avoid this.

  • I

    The intervention designers also collected and analysed the data, with no independent or external evaluator.

    Material preparation, data collection and analysis were performed by Sukran Sungur and Gulbin Ozkan.

  • Y

    The study spanned only six weeks, far short of an academic year, and criterion T was not met.

    The research lasted a total of six weeks.

  • B

    Both groups received identical instructional time, and the only difference is the integral teaching medium being tested.

    Both groups received an identical amount of instructional time (4 lesson hours per week for 6 weeks).

  • R

    No independent replication of this specific study by a different research team could be found in any source.

    this research extends the ecological validity and generalizability of the original model, validating its effectiveness within a new cultural and institutional context.

  • A

    Only a single science topic and CT were assessed, and the prerequisite standardised-exam criterion E was not met.

    This study investigated the impact of a Scratch-integrated science intervention on sixth-grade students' CT and science achievement...

  • G

    There was no follow-up to graduation, and the prerequisite year-duration criterion Y was not met.

    Longitudinal studies are required to investigate potential sleeper effects...

  • P

    The paper reports ethics approval but no public pre-registration of the study protocol.

Abstract

In the digital era, fostering computational thinking (CT) skills is essential for scientific literacy and engaging students in authentic scientific practices. Block-based programming environments (BBPEs), such as Scratch, effectively integrate these skills into K-12 science by lowering the syntax barrier, allowing a focus on scientific logic. This study investigated the impact of a Scratch-integrated science intervention on sixth-grade students' CT and science achievement using a quasi-experimental design. Participants (n = 51) were assigned to either an experimental group receiving BBPEs or a control group following the standard curriculum. To ensure methodological rigor, CT test was first adapted and validated for the Turkish context (n = 353). Despite the benefits of BBPEs, implementation poses unique challenges for science teachers, necessitating new technical and pedagogical demands. Therefore, this research supported the quantitative data with the lived experiences of the teacher who underwent this integration process. Quantitative results indicate that the intervention significantly enhanced students' CT (p < .05). While both groups showed gains in science achievement, the experimental group achieved higher mean scores, though the difference was not statistically significant. Qualitative findings revealed that the instructor successfully navigated the liminal space of teacher re-novicing, fostering a collaborative environment that empowered lower-achieving students to emerge as programming experts. Furthermore, the integration promoted epistemic synergy and an iterative debugging mindset, transforming experimental failures into productive inquiry steps that strengthened student persistence and understanding. This research provides empirical evidence for using BBPEs to bridge computer science and science education while highlighting the support needed for teachers during this transition.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Two intact sixth-grade classes were randomly assigned to conditions, making the class the unit of randomisation.
      • Two classes were randomly assigned as either the experimental or control group.
      • Relevant Quotes: 1) "This study involved two sixth-grade classes from a public middle school in Türkiye... Two classes were randomly assigned as either the experimental or control group. The experimental group consisted of 23 sixth-grade students aged 11 to 13, while the control group included 28 sixth-grade students of the same age range." (p. 4) 2) "Specifically, a quasi-experimental research design was adopted, which is commonly employed when random assignment is not feasible, yet comparison between groups is necessary to evaluate the impact of an intervention (Cook & Campbell, 1979)." (p. 4) 3) "Participants (n=51) were assigned to either an experimental group receiving BBPEs or a control group following the standard curriculum." (Abstract, p. 1) Detailed Analysis: Criterion C requires that randomisation be conducted at the class level (or stronger) rather than assigning individual students within a single classroom. Here the authors describe two intact sixth-grade classes that were "randomly assigned as either the experimental or control group." Thus the unit of randomisation was the whole class, not individual students within one class, which is what the C criterion requires to avoid within-class contamination. Although the authors label the overall design "quasi-experimental" and note random assignment is often "not feasible," the specific allocation reported is a random assignment of the two intact classes to conditions. The number of randomised units is very small (only two classes in one school), so this is a weak randomisation, but the unit of randomisation reported is the class level, satisfying the literal requirement of this criterion. Criterion C is met because two intact classes were randomly assigned to the experimental and control conditions, making the class the unit of randomisation.
    • E

      Exam-based Assessment

      • The study used research-developed and researcher-adapted instruments rather than recognised standardised exams.
      • The pre- and post-test instruments for data collection in this research were the Computational Thinking Test (CTt; Roman-Gonzalez et al., 2017) and Force and Motion Knowledge test (FMKt; Cimentepe, 2019).
      • Relevant Quotes: 1) "The pre- and post-test instruments for data collection in this research were the Computational Thinking Test (CTt; Roman-Gonzalez et al., 2017) and Force and Motion Knowledge test (FMKt; Cimentepe, 2019)." (p. 7) 2) "For this study, the CTt was adapted into Turkish for use among middle school students across all grade levels. The adaptation process involved the collaboration of two science education faculty members, one expert in educational sciences, and a language specialist." (p. 7) 3) "The internal consistency of the FMKt, including 25 items, was determined to be.87, indicating a high level of reliability." (p. 7) 4) "At the end of each topic, students answered researcher-prepared questions identical to those given to the experimental group." (p. 6) Detailed Analysis: Criterion E requires the use of a standardised, externally recognised exam-based assessment (e.g., a state-wide or national curriculum exam) rather than an instrument created or assembled by the researchers for the study. The outcome measures here are the Computational Thinking Test (CTt) and the Force and Motion Knowledge test (FMKt). The CTt is a research instrument from Roman-Gonzalez et al. (2017) that the authors themselves adapted and re-validated into Turkish for this study, and the FMKt is drawn from an unpublished master's thesis (Cimentepe, 2019). Neither is a widely recognised national or state standardised examination; both are research-developed/adapted instruments. The fact that the authors had to conduct their own validity and reliability analyses (EFA, Cronbach's alpha) on the adapted CTt confirms it is not an established standardised exam used at scale in the education system. Likewise, the FMKt is a thesis-derived knowledge test rather than an official curriculum examination. Criterion E is not met because the study relied on research-developed and researcher-adapted instruments (CTt and FMKt) rather than widely recognised standardised exams.
    • T

      Term Duration

      • The intervention and post-test spanned only six weeks, far shorter than one academic term, with no follow-up.
      • The research lasted a total of six weeks. Both groups received an identical amount of instructional time (4 lesson hours per week for 6 weeks).
      • Relevant Quotes: 1) "The implementation of this study was conducted during the force and motion unit of the sixth-grade science curriculum in the spring semester of the 2022-2023 academic year. The research lasted a total of six weeks. Both groups received an identical amount of instructional time (4 lesson hours per week for 6 weeks)." (p. 4) 2) "In response to their call for more designs, we implemented a six-week quasi-experimental study featuring both experimental and control groups." (p. 4) 3) "The pre- and post-test instruments for data collection in this research were the Computational Thinking Test (CTt...) and Force and Motion Knowledge test (FMKt...)." (p. 7) Detailed Analysis: Criterion T requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins, or that there be at least term-long tracking from intervention start to outcome measurement. The intervention here ran for six weeks (4 lesson hours per week), and the post-tests were administered at the end of this six-week period. There is no delayed follow-up; the paper's limitations even call for future "longitudinal studies" to assess persistence. Six weeks is substantially shorter than one academic term and there is no term-length follow-up tracking from the start of the intervention to the outcome measurement. Criterion T is not met because the intervention and outcome measurement spanned only six weeks, far less than one academic term, with no term-long follow-up.
    • D

      Documented Control Group

      • The control group's size, demographics, instruction, and baseline scores are clearly documented.
      • The control group followed the standard curriculum using materials (textbooks, worksheets, physical experiments) and did not engage in informal use of BBPEs or similar technologies during the instructional hours.
      • Relevant Quotes: 1) "Two classes were randomly assigned as either the experimental or control group... the control group included 28 sixth-grade students of the same age range." (p. 4) 2) "No additional intervention was applied in the control group. The regular classroom teacher conducted the lessons in line with the existing curriculum framework. The control group followed the standard curriculum using materials (textbooks, worksheets, physical experiments) and did not engage in informal use of BBPEs or similar technologies during the instructional hours." (p. 5) 3) "According to the teacher, lessons followed the official Science 6 Course Book distributed by MoNE, (2018) ... including the completion of in-book activities and solution of relevant sample questions." (pp. 5-6) 4) "Table 5 Descriptive Statistics for FMKt and CTt Scores ... Control 28 ... CTt Pre-tests Control 11.85 ... FMKt Pre-tests Control 5.75" (Table 5, p. 9) Detailed Analysis: Criterion D requires clear documentation of the control group's composition, size, baseline performance, and the conditions it experienced. The paper documents the control group's size (28 students), age range (11-13), grade (sixth), and explicitly describes its instruction (standard MoNE curriculum with textbooks, worksheets, and physical experiments; no BBPEs). Baseline (pre-test) means and standard deviations for the control group on both the CTt and FMKt are reported in Table 5, allowing comparison with the experimental group. This level of detail confirms who the control group was, its baseline characteristics, and that it received only business-as-usual instruction. Criterion D is met because the control group's size, age, instruction, and baseline scores are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation was at the class level within one school, not at the school level.
      • potential information sharing between experimental and control groups within the same school remains a limitation. Future research could use separate schools to avoid this.
      • Relevant Quotes: 1) "This study involved two sixth-grade classes from a public middle school in Türkiye... Two classes were randomly assigned as either the experimental or control group." (p. 4) 2) "In addition, despite separate scheduling and confidentiality requests, potential information sharing between experimental and control groups within the same school remains a limitation. Future research could use separate schools to avoid this." (p. 12) Detailed Analysis: Criterion S requires randomisation at the school level, where entire schools (or equivalent implementing units) are randomly assigned to conditions. In this study, both the experimental and control classes are located within a single public middle school, and randomisation occurred at the class level within that one school. The authors themselves note that both groups were "within the same school," identifying potential contamination as a limitation. Because only one school was involved and randomisation was at the class (not school) level, this criterion cannot be satisfied. Criterion S is not met because randomisation occurred at the class level within a single school, not at the school level.
    • I

      Independent Conduct

      • The intervention designers also collected and analysed the data, with no independent or external evaluator.
      • Material preparation, data collection and analysis were performed by Sukran Sungur and Gulbin Ozkan.
      • Relevant Quotes: 1) "The current study involved the design and implementation of programming activities through Scratch embedded within the existing science curriculum." (p. 4) 2) "To ensure the integration of CT within the curriculum, researchers developed Scratch-based activities tailored to each of the five learning objectives related to force and motion..." (p. 5) 3) "Material preparation, data collection and analysis were performed by Sukran Sungur and Gulbin Ozkan." (Author's Contributions, p. 16) 4) "Prior to the implementation process, the researchers conducted a two-week (totaling 8 h) Scratch training program for the science teacher at the participating school." (p. 4) 5) "This study was produced from the Master's thesis of the first author Sukran Sungur completed at Yildiz Technical University." (Acknowledgements, p. 16) Detailed Analysis: Criterion I requires that the study be conducted independently from those who designed the intervention, typically via a third-party or external evaluation team. Here, the same authors designed the Scratch activities, trained the implementing teacher, and performed the data collection and analysis themselves (as stated in the Author's Contributions). There is no indication of any independent or external evaluator overseeing data collection or analysis; the work derives from the first author's master's thesis. Because the intervention designers also collected and analysed the data with no independent oversight, the independence requirement is not satisfied. Criterion I is not met because the same researchers designed the intervention and carried out the data collection and analysis without independent oversight.
    • Y

      Year Duration

      • The study spanned only six weeks, far short of an academic year, and criterion T was not met.
      • The research lasted a total of six weeks.
      • Relevant Quotes: 1) "The research lasted a total of six weeks. Both groups received an identical amount of instructional time (4 lesson hours per week for 6 weeks)." (p. 4) 2) "Longitudinal studies are required to investigate potential sleeper effects, determining if science achievement surpasses traditional groups once the initial cognitive load of coding diminishes." (p. 12) Detailed Analysis: Criterion Y requires outcomes to be measured at least 75% of one full academic year (~9-10 months) after the intervention begins. The intervention and measurement here spanned only six weeks, with no longer-term tracking; the authors explicitly note that longitudinal study is needed in future work. Per the standard, if the weaker Term Duration (T) criterion is not met, the stronger Year Duration (Y) criterion cannot be met. Criterion Y is not met because the study spanned only six weeks, far short of an academic year, and criterion T was also not met.
    • B

      Balanced Control Group

      • Both groups received identical instructional time, and the only difference is the integral teaching medium being tested.
      • Both groups received an identical amount of instructional time (4 lesson hours per week for 6 weeks).
      • Relevant Quotes: 1) "Both groups received an identical amount of instructional time (4 lesson hours per week for 6 weeks)." (p. 4) 2) "In the experimental group, the force and motion unit was taught using activities in BBPEs... these objectives were developed by integrating Scratch programming." (pp. 4-5) 3) "No additional intervention was applied in the control group. The regular classroom teacher conducted the lessons in line with the existing curriculum framework. The control group followed the standard curriculum using materials (textbooks, worksheets, physical experiments)..." (p. 5) 4) "The school was chosen due to its well-equipped computer lab, where each student had access to a fully functional computer." (p. 4) Detailed Analysis: Criterion B asks whether the intervention and control conditions were balanced in instructional time and resources, unless the additional resource is itself the treatment variable. The paper explicitly states both groups received "an identical amount of instructional time (4 lesson hours per week for 6 weeks)," so there is no imbalance in time-on-task. The intervention differs from the control only in the teaching method: the experimental group learns the force and motion unit through Scratch programming activities in the computer lab, while the control group learns the same unit via the standard curriculum (textbooks, worksheets, physical experiments). The Scratch tool and computer access are an integral part of the BBPE teaching method being tested, not an extra block of additional instructional time or a separable confounding resource. Following the decision tree: extra resources (computers, Scratch) are present, but they are integral to the treatment being tested (RESOURCES_ARE_TREATMENT), and instructional time is identical across groups. Therefore the groups are balanced for the purposes of this criterion. Criterion B is met because both groups received identical instructional time and the only difference is the integral instructional medium (Scratch and computer access) that constitutes the treatment being tested.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study by a different research team could be found in any source.
      • this research extends the ecological validity and generalizability of the original model, validating its effectiveness within a new cultural and institutional context.
      • Relevant Quotes: 1) "Inspired by the study of Aksit and Wiebe (2020) this study builds upon their foundational work by addressing identified limitations... this research extends the ecological validity and generalizability of the original model, validating its effectiveness within a new cultural and institutional context." (p. 4) 2) "These results align with the work of Aksit and Wiebe (2020), supporting the view that computational modeling is more than just learning to code..." (p. 10) Detailed Analysis: Criterion R requires that this specific study be independently replicated by a different research team in a different context, published in a peer-reviewed journal. The paper frames itself as inspired by and extending Aksit and Wiebe (2020), so the current study is a new study building on prior work, not a replication of the current study by others. Internet research was conducted across web search engines, ResearchGate, and Springer for any independent replication of this specific Scratch-integrated force-and-motion quasi-RCT. No such replication was found. The paper is very recent (received 29 December 2025, accepted 4 April 2026), so no subsequent independent reproduction by other authors could be located as of the review date (2026-06-01). The related works found (e.g., Aksit & Wiebe, 2020; Aytekin & Topçu, 2024; Koray & Bilgin, 2023) are prior or parallel studies, not replications of this specific study. Criterion R is not met because no independent replication of this specific study by another research team could be found in any available source.
    • A

      All-subject Exams

      • Only a single science topic and CT were assessed, and the prerequisite standardised-exam criterion E was not met.
      • This study investigated the impact of a Scratch-integrated science intervention on sixth-grade students' CT and science achievement...
      • Relevant Quotes: 1) "This study investigated the impact of a Scratch-integrated science intervention on sixth-grade students' CT and science achievement..." (Abstract, p. 1) 2) "The pre- and post-test instruments for data collection in this research were the Computational Thinking Test (CTt...) and Force and Motion Knowledge test (FMKt...)." (p. 7) Detailed Analysis: Criterion A requires measuring impact across all main subjects taught at that level, using standardised exam-based assessments. This study assessed only computational thinking (CTt) and a single science topic, force and motion (FMKt); no other core school subjects were measured. Moreover, criterion A is conditional on criterion E (standardised exam-based assessment) being met. Because the study used research-developed/adapted instruments rather than standardised exams, criterion E is not met, and therefore criterion A cannot be met. Criterion A is not met because only a single science topic and CT were assessed, and because the prerequisite criterion E (standardised exams) was not met.
    • G

      Graduation Tracking

      • There was no follow-up to graduation, and the prerequisite year-duration criterion Y was not met.
      • Longitudinal studies are required to investigate potential sleeper effects...
      • Relevant Quotes: 1) "Longitudinal studies are required to investigate potential sleeper effects, determining if science achievement surpasses traditional groups once the initial cognitive load of coding diminishes." (p. 12) 2) "The research lasted a total of six weeks." (p. 4) Detailed Analysis: Criterion G requires tracking participants through to their graduation from the relevant educational stage. This study measured outcomes immediately at the end of a six-week intervention with no follow-up, and the authors explicitly call for future longitudinal work to assess persistence. No graduation tracking is reported. Internet research was conducted for any follow-up publications by the same authors (Sukran Sungur, Gulbin Ozkan) that track this sixth-grade cohort through to graduation. No such follow-up paper could be found; the study derives from a recent master's thesis and no subsequent tracking study by these authors was located. In addition, criterion G is conditional on criterion Y (Year Duration) being met; since Y is not met, G cannot be met. Criterion G is not met because there was no follow-up to graduation (and none found in any subsequent publication), and the prerequisite criterion Y was not met.
    • P

      Pre-Registered

      • The paper reports ethics approval but no public pre-registration of the study protocol.
      • Relevant Quotes: 1) "This study was approved by the Yildiz Technical University Social and Humanities Research Ethics Board (Date: 03.10.2022, Number: 2022.10)." (Declarations, p. 16) 2) "This study was produced from the Master's thesis of the first author Sukran Sungur completed at Yildiz Technical University." (Acknowledgements, p. 16) Detailed Analysis: Criterion P requires that the full study protocol (hypotheses, methods, planned analyses) be pre-registered on a public registry before data collection began, with a verifiable date. The paper reports an institutional ethics approval (3 October 2022) but provides no reference to any public pre-registration platform (e.g., OSF, ClinicalTrials.gov, AEA registry, ISRCTN), no registration ID, and no pre-registration date. Ethics approval is not equivalent to pre-registration of the study protocol. Internet research found no pre-registration record for this study on any public registry. Only the ethics approval reported in the paper could be identified. Criterion P is not met because no public pre-registration of the study protocol is reported or could be found.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.