Enhancing Academic Performance in Tertiary Education Through Social Media: A Multi-arm Randomized Controlled Trial of the BE-Social Program

Aida Tarifa-Rodriguez, Javier Virues-Ortega, Ana Calero-Elvira

Published:
ERCT Check Date:
DOI: 10.1007/s10864-026-09629-8
  • higher education
  • EU
  • blended learning
  • EdTech platform
0
  • C

    Randomization was conducted at the individual student level (via Excel RAND), not at the class or school level, and the intervention is a group program rather than one-to-one tutoring.

    "We randomly assigned 141 participants to the five arms of the study. We used the RAND function in Microsoft Excel to assign participants into groups."

  • E

    Outcomes were measured with ad hoc weekly multiple-choice course-content tests created by the authors, not a recognized standardized exam.

    "Our outcome measures were limited to ad hoc course content tests and standard social media responses (i.e., views, likes, comments, posts)."

  • T

    The intervention phase lasted four weeks within a 12-week (about three-month) study, and outcomes were measured well under one full academic term after the intervention began.

    "We completed the study over 12 successive weeks (January 18 to April 11). The baseline phase comprised the first four weeks of the study. The intervention phase followed over the next four weeks."

  • D

    The control group's size, demographics, baseline performance, and business-as-usual conditions are documented in detail in Table 1 and the procedure section.

    "The control group remained under baseline conditions for the duration of the study (12 weeks), receiving daily multiple-choice discussion questions, weekly comments from the instructor, and weekly live video broadcasts, but without any of the educational components of the intervention groups."

  • S

    Randomization was at the individual student level within a single online course, not at the level of whole schools or institutions.

    "We randomly assigned 141 participants to the five arms of the study. We used the RAND function in Microsoft Excel to assign participants into groups."

  • I

    The same authors designed the BE-Social intervention and also conducted the trial, collected data, and analyzed the results, with no independent third-party evaluator.

    "Tarifa-Rodriguez et al. (2024) developed and evaluated the Behavioral Education and Social Media intervention package (the BE-Social Program) ..."

  • Y

    The intervention and tracking spanned only about 12 weeks, far short of 75% of a full academic year, and criterion T is not met.

    "We completed the study over 12 successive weeks (January 18 to April 11)."

  • B

    The control group received the same baseline inputs as all groups (daily questions, weekly feedback, weekly videos), and the educational components added to the treatment arms are the explicit treatment variables being tested.

    "Baseline activities were intended to maintain student engagement (thereby preventing participant dropout) and equate the instructor inputs across control and intervention groups."

  • R

    No independent replication by a different research team exists; this study is itself a same-team replication and expansion of Tarifa-Rodriguez et al. (2024).

    "The present study aimed to replicate and expand the findings by Tarifa-Rodriguez et al. (2024) ..."

  • A

    Only course-specific (applied psychology) outcomes were measured with custom tests, no standardized all-subject exams were used, and criterion E is not met.

    "Participants took a total of twelve weekly tests throughout the study ... Each test included 20 multiple-choice discussion questions ... covered aspects of the course content."

  • G

    Tracking ended after the 12-week study with no follow-up to graduation, and criterion Y is not met.

    "We completed the study over 12 successive weeks (January 18 to April 11)."

  • P

    The paper provides no pre-registration statement, registry identifier, or registration date predating data collection.

Abstract

Few randomized controlled trials have analyzed evidence-based educational practices delivered through a social media environment. This study used a multi-arm randomized controlled trial to evaluate the critical components of an educational intervention package: study self-management skills training delivered through video modeling, cooperative learning, and semi-immediate instructor feedback. We evaluated social media engagement and academic performance among 141 students in a postgraduate applied psychology program. Students were randomly assigned to five groups: control (n = 27); self-management (n = 27); cooperative learning (n = 33); self-management and cooperative learning (n = 27); and self-management, cooperative learning, and semi-immediate instructor feedback (n = 27). Results indicated that participants receiving the complete intervention showed numerically higher effect sizes, whereas all groups, except for the self-management group showed significantly higher academic performance than the control group. We discuss the conceptual, methodological, and practical implications of the study.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomization was conducted at the individual student level (via Excel RAND), not at the class or school level, and the intervention is a group program rather than one-to-one tutoring.
      • "We randomly assigned 141 participants to the five arms of the study. We used the RAND function in Microsoft Excel to assign participants into groups."
      • Relevant Quotes: 1) "We randomly assigned 141 participants to the five arms of the study. We used the RAND function in Microsoft Excel to assign participants into groups. Participants' names were listed in column A of a spreadsheet, and random numbers generated by the RAND function were pasted as values in column B." (p. 6) 2) "A 33-year-old male and a 32-year-old female, both doctoral students, created five closed Facebook groups with similar names ... and served as instructors and moderators within these social media groups." (p. 7) 3) "Students were randomly assigned to five groups: control (n = 27); self-management (n = 27); cooperative learning (n = 33) ..." (Abstract) Detailed Analysis: The ERCT 'C' criterion requires randomisation at the level of entire classes or schools, unless the intervention is a one-to-one tutoring / personal teaching intervention, in which case student-level randomisation is acceptable. Here the authors clearly randomised individual students (by name) into five arms using the Excel RAND function. The intervention was a group-based social-media program delivered through closed Facebook study groups, not a one-to-one tutoring intervention, so the tutoring exception does not apply. Because assignment was at the individual student level rather than by intact class or school, the design carries exactly the contamination risk the criterion seeks to avoid. Criterion C is not met because randomisation was performed at the individual student level and the intervention is not a personal/one-to-one tutoring exception.
    • E

      Exam-based Assessment

      • Outcomes were measured with ad hoc weekly multiple-choice course-content tests created by the authors, not a recognized standardized exam.
      • "Our outcome measures were limited to ad hoc course content tests and standard social media responses (i.e., views, likes, comments, posts)."
      • Relevant Quotes: 1) "Participants took a total of twelve weekly tests throughout the study. Weekly tests were conducted via the Moodle platform ... Each test included 20 multiple-choice discussion questions with three distractors and one correct answer." (p. 11) 2) "90% of the scenarios described practical applications of the course contents, while the remainder of the questions focused on conceptual information (e.g., definitions, characteristics, conceptual clarification, etc.)." (p. 11) 3) "Our outcome measures were limited to ad hoc course content tests and standard social media responses (i.e., views, likes, comments, posts)." (p. 20) 4) "future analyses would benefit from adding standardized measurements of student satisfaction and end-of-course assessments. The latter was not practical in the current study ..." (p. 21) Detailed Analysis: The 'E' criterion requires standardised, widely recognised exams rather than instruments built specifically for the study. The authors administered weekly 20-item multiple-choice tests built around the specific course content and explicitly describe these as "ad hoc course content tests." They further note that adding "standardized measurements" and "end-of-course assessments" was not done. No national, state-wide, or otherwise recognised standardised exam was used; the tests were custom-made for the intervention and aligned to the course material. Criterion E is not met because the study relied on custom, ad hoc course-content tests rather than a recognised standardised exam.
    • T

      Term Duration

      • The intervention phase lasted four weeks within a 12-week (about three-month) study, and outcomes were measured well under one full academic term after the intervention began.
      • "We completed the study over 12 successive weeks (January 18 to April 11). The baseline phase comprised the first four weeks of the study. The intervention phase followed over the next four weeks."
      • Relevant Quotes: 1) "We completed the study over 12 successive weeks (January 18 to April 11). The baseline phase comprised the first four weeks of the study. The intervention phase followed over the next four weeks ... Participants returned to baseline conditions during the last phase of the study for another four weeks (posttreatment)." (p. 6) 2) "The course began in September 2020 and ended in May 2021." (p. 6) 3) "Participants took a total of twelve weekly tests throughout the study." (p. 11) Detailed Analysis: The 'T' criterion requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. Although the surrounding course ran from September to May, the RCT itself spanned only 12 weeks (January 18 to April 11). The intervention phase began only after the four-week baseline (around mid-February) and ran four weeks, followed by a four-week post-treatment phase ending April 11. The interval from intervention start to the final outcome measurement is therefore roughly eight weeks, and even the entire study from its first day is under three months. This is shorter than one full academic term. Criterion T is not met because the interval from intervention onset to outcome measurement is well below one full academic term.
    • D

      Documented Control Group

      • The control group's size, demographics, baseline performance, and business-as-usual conditions are documented in detail in Table 1 and the procedure section.
      • "The control group remained under baseline conditions for the duration of the study (12 weeks), receiving daily multiple-choice discussion questions, weekly comments from the instructor, and weekly live video broadcasts, but without any of the educational components of the intervention groups."
      • Relevant Quotes: 1) "Individuals in the control group remained in baseline conditions for the complete duration of the study." (p. 7) 2) "The control group remained under baseline conditions for the duration of the study (12 weeks), receiving daily multiple-choice discussion questions, weekly comments from the instructor, and weekly live video broadcasts, but without any of the educational components of the intervention groups." (p. 8) 3) "Table 1 Sociodemographic Characteristics of Participants (n = 141) ... G5 (n = 27) ... Female 77.78 (21) ... Mean age, years (SD) 41.04 (8.52) ... Mean BL performance (SD) 11.67 (2.45) ..." (Table 1, p. 8) 4) "To ensure the effectiveness of the randomization process, we verified that age, gender, educational level, and socioeconomic standing were not statistically different across groups (p > .05)." (p. 6) Detailed Analysis: The 'D' criterion requires detailed documentation of the control group, including demographics, baseline performance, and the conditions it experienced. Table 1 provides a dedicated column (G5, n = 27) for the control group with gender, mean age, years of education, baseline performance, country, and socioeconomic status. The procedure section clearly explains that the control group remained in baseline conditions for the full 12 weeks, receiving the shared baseline inputs but none of the intervention components, and that baseline equivalence across groups was statistically verified. Criterion D is met because the control group's size, demographics, baseline performance, and conditions are clearly and quantitatively documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomization was at the individual student level within a single online course, not at the level of whole schools or institutions.
      • "We randomly assigned 141 participants to the five arms of the study. We used the RAND function in Microsoft Excel to assign participants into groups."
      • Relevant Quotes: 1) "We randomly assigned 141 participants to the five arms of the study. We used the RAND function in Microsoft Excel to assign participants into groups." (p. 6) 2) "Participants were students enrolled in an online college-level course in applied psychology." (p. 6) 3) "A ... male and a ... female, both doctoral students, created five closed Facebook groups ... Each group was assigned a different set of behavioral strategies ..." (p. 7) Detailed Analysis: The 'S' criterion requires randomisation at the level of whole schools or implementing institutions/units. In this study all 141 participants were drawn from a single online applied psychology course and were individually randomised into five Facebook-group arms. There is no school-, site-, or institution-level randomisation; the unit of randomisation was the individual student. The "groups" here are experimental arms (closed Facebook groups) created for the study, not pre-existing schools or institutional units. Criterion S is not met because randomisation occurred at the individual student level within one course rather than at the school or institution level.
    • I

      Independent Conduct

      • The same authors designed the BE-Social intervention and also conducted the trial, collected data, and analyzed the results, with no independent third-party evaluator.
      • "Tarifa-Rodriguez et al. (2024) developed and evaluated the Behavioral Education and Social Media intervention package (the BE-Social Program) ..."
      • Relevant Quotes: 1) "Tarifa-Rodriguez et al. (2024) developed and evaluated the Behavioral Education and Social Media intervention package (the BE-Social Program) ..." (p. 3) 2) "Author Contributions ATR ... Research design development. Data collection and data curation. Data analysis design. Manuscript writing and editing (first draft). JVO. Research design development. Data analysis design. Logistics and resources. Funding procurement. Manuscript writing and editing. Doctoral supervision of ATR. ACE. Research design development. Manuscript writing and editing." (p. 22) 3) "Regarding social media participation, two independent research assistants extracted data directly from Facebook groups." (p. 11) Detailed Analysis: The 'I' criterion requires that the evaluation be conducted independently of the intervention designers. The same author team that developed the BE-Social program (Tarifa-Rodriguez et al., 2024) also designed, ran, and analysed this trial, as shown by the author contribution statement (research design, data collection, data curation, and data analysis all performed by the authors). Although independent research assistants and a secondary observer assisted with extracting social-media engagement data and interobserver agreement, this is a data-coding support role, not an independent evaluation of the trial; the design, conduct, analysis, and conclusions remained with the intervention developers. There is no statement of an external/third-party evaluation team. Criterion I is not met because the intervention designers themselves conducted and analysed the trial without independent third-party evaluation.
    • Y

      Year Duration

      • The intervention and tracking spanned only about 12 weeks, far short of 75% of a full academic year, and criterion T is not met.
      • "We completed the study over 12 successive weeks (January 18 to April 11)."
      • Relevant Quotes: 1) "We completed the study over 12 successive weeks (January 18 to April 11). The baseline phase comprised the first four weeks of the study. The intervention phase followed over the next four weeks ... another four weeks (posttreatment)." (p. 6) 2) "the multi-component intervention may have optimal effects on achievement when delivered for the complete duration of a semester or year-long course." (p. 21, Conclusions) Detailed Analysis: The 'Y' criterion requires outcomes to be measured at least 75% of a full academic year (roughly 7+ months) after the intervention begins, and per the standard if 'T' is not met then 'Y' cannot be met. The entire RCT, from baseline through post-treatment, occupied only 12 weeks (about three months), with the intervention phase itself lasting only four weeks. This is far below 75% of an academic year. The authors themselves note that end-of-course assessment "was not practical in the current study owing to the multi-phase structure of the design, which had to fit within a one-year course," but the experimental tracking window was only the 12-week January-April period. Criterion Y is not met because the study tracked outcomes for only about 12 weeks, well under 75% of an academic year (and criterion T was also not met).
    • B

      Balanced Control Group

      • The control group received the same baseline inputs as all groups (daily questions, weekly feedback, weekly videos), and the educational components added to the treatment arms are the explicit treatment variables being tested.
      • "Baseline activities were intended to maintain student engagement (thereby preventing participant dropout) and equate the instructor inputs across control and intervention groups."
      • Relevant Quotes: 1) "The control group remained under baseline conditions for the duration of the study (12 weeks), receiving daily multiple-choice discussion questions, weekly comments from the instructor, and weekly live video broadcasts, but without any of the educational components of the intervention groups." (p. 8) 2) "Baseline activities were intended to maintain student engagement (thereby preventing participant dropout) and equate the instructor inputs across control and intervention groups." (p. 8) 3) "The present study aimed to replicate and expand the findings by Tarifa-Rodriguez et al. (2024) by incorporating additional methodological standards: ... (d) comparable exposure intensity to academic content across groups, and, (e) a multi-arm RCT to evaluate the the BE-Social intervention components combined and in isolation ..." (p. 5) 4) "Each weekday, the instructors created three posts using all program components." (Group 1, p. 9); for the baseline shared by the control group "The posts were identical for all groups." (p. 8) Detailed Analysis: Following the criterion B decision tree: the intervention arms add educational resources (self-management training via video-modeling, cooperative learning prompts, and semi-immediate feedback). However, these added components ARE the explicit treatment variables that the multi-arm RCT was designed to isolate and test (RESOURCES_ARE_TREATMENT = true). Moreover, the authors deliberately equated the underlying inputs: all groups, including the control, received the same daily multiple-choice questions, weekly instructor comments, and weekly one-hour live video broadcasts, and the authors explicitly designed baseline activities to "equate the instructor inputs across control and intervention groups" and ensure "comparable exposure intensity to academic content across groups." Thus the differences between arms are the specific educational components under test, layered on a balanced, business-as-usual background common to all groups. Criterion B is met because the shared baseline provided comparable time and instructor inputs across all groups, and the additional educational components are the explicit treatment variables being tested against that business-as-usual baseline.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication by a different research team exists; this study is itself a same-team replication and expansion of Tarifa-Rodriguez et al. (2024).
      • "The present study aimed to replicate and expand the findings by Tarifa-Rodriguez et al. (2024) ..."
      • Relevant Quotes: 1) "The present study aimed to replicate and expand the findings by Tarifa-Rodriguez et al. (2024) by incorporating additional methodological standards ..." (p. 5) 2) "Tarifa-Rodriguez et al. (2024) completed an initial evaluation of the model with a randomized controlled trial (RCT). Forty-six postgraduate students ... were randomly assigned to a control (n = 18) and a BE-Social intervention group (n = 28)." (p. 5) Detailed Analysis: The 'R' criterion requires that the study be independently replicated by a different research team in a different context and published in a peer-reviewed journal. This paper is itself a replication/expansion of the authors' own prior 2024 RCT (Tarifa-Rodriguez, A., Virues-Ortega, J., & Calero-Elvira, A. (2024). A behavioral education in social media program for postgraduate academic achievement: a randomized controlled trial. Journal of Behavioral Education) and is conducted by the same author team. An internet/literature search (Springer, ResearchGate, researcher profiles) for the very recently published (April 2026) BE-Social program found no independent replication of this specific program by an external team; the only related RCT located is the authors' own 2024 study. Same-team continuation does not satisfy the independence requirement. Criterion R is not met because no independent replication by a different research team has been identified; this study is a same-team replication of the authors' earlier work.
    • A

      All-subject Exams

      • Only course-specific (applied psychology) outcomes were measured with custom tests, no standardized all-subject exams were used, and criterion E is not met.
      • "Participants took a total of twelve weekly tests throughout the study ... Each test included 20 multiple-choice discussion questions ... covered aspects of the course content."
      • Relevant Quotes: 1) "All posted questions covered aspects of the course content taught during that week." (p. 8) 2) "We calculated test scores by dividing correct responses by 20 ... 90% of the scenarios described practical applications of the course contents ..." (p. 11) 3) "Our outcome measures were limited to ad hoc course content tests and standard social media responses ..." (p. 20) Detailed Analysis: The 'A' criterion requires that all main subjects be assessed using standardised exam-based assessments, and per the standard if criterion E (Exam-based Assessment) is not met then 'A' is automatically not met. Here the study only measured performance within the single applied psychology course using custom weekly tests, and no standardised exams were used. Both the single-subject scope and the failure of criterion E preclude meeting this criterion. Criterion A is not met because only one course's content was assessed with custom (non-standardised) tests, and criterion E was not met.
    • G

      Graduation Tracking

      • Tracking ended after the 12-week study with no follow-up to graduation, and criterion Y is not met.
      • "We completed the study over 12 successive weeks (January 18 to April 11)."
      • Relevant Quotes: 1) "We completed the study over 12 successive weeks (January 18 to April 11)." (p. 6) 2) "future analyses would benefit from adding ... end-of- course assessments. The latter was not practical in the current study owing to the multi-phase structure of the design ..." (p. 21) Detailed Analysis: The 'G' criterion requires tracking participants until graduation, and per the standard if criterion Y is not met then 'G' cannot be met. The study tracked participants only across the 12-week experimental window and explicitly did not conduct end-of-course assessments or any follow-up through program completion or graduation. An internet search for subsequent follow-up publications by the same authors tracking this cohort to graduation found none; the only related prior work is Tarifa-Rodriguez et al. (2024), which does not provide graduation tracking either. Criterion G is not met because the study did not track participants to graduation and criterion Y was also not met.
    • P

      Pre-Registered

      • The paper provides no pre-registration statement, registry identifier, or registration date predating data collection.
      • Relevant Quotes: 1) "The study was approved by the ethics committee of the Universidad Autonoma de Madrid (ethics approval number CEI 112-2204)." (p. 6) 2) "Data Availability The complete databased use for all analyses is available from the public repository with digital object identifier https://doi.org/10.6084/m9.figshare.31960806." (p. 22) Detailed Analysis: The 'P' criterion requires a pre-registered protocol with a registry reference and a registration date preceding data collection. The paper reports only ethics-committee approval and a post-hoc figshare data repository, neither of which constitutes pre-registration of hypotheses and analysis plans before data collection. No trial-registry identifier (e.g., ClinicalTrials.gov, ISRCTN, OSF) or pre-registration date is provided. An internet search of trial registries and OSF for the BE-Social program returned no pre-registration record for this study or its 2024 precursor. Criterion P is not met because no pre-registration reference or pre-data-collection registration date is reported.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.