Facilitating English Grammar Learning by a Personalized Mobile-Assisted System With a Self-Regulated Learning Mechanism

Xiao Wang, Jing Chen, Tingting Zhang

Published:
ERCT Check Date:
DOI: 10.3389/fpsyg.2021.624430
  • L2 languages
  • K12
  • China
  • online homework
  • EdTech app
  • mobile learning
  • formative assessment
0
  • C

    Randomisation was carried out at the individual student level within a single school rather than by class or school, and the intervention is a self-study mobile system, not one-to-one tutoring, so no exception applies.

    "A total of 598 EFL learners were recruited from a senior secondary school using the convenience sampling technique. They were randomly assigned into the experimental group (n = 278) or the control group (n = 320)." (p. 4)

  • E

    Outcomes were measured with 50-item multiple-choice grammar tests custom-designed by two EFL teachers for this study, not a widely recognised standardised exam.

    "Both the pre- and post- grammar tests comprised of 50 multiple choices questions, designed by two experienced EFL teachers, to examine the grammar structures that were introduced in the system, with each question worth two points." (p. 6)

  • T

    The intervention ran for 16 weeks (a full semester) from the pre-test at the beginning of the semester to the post-test after the treatment, which covers at least one academic term.

    "After that, the 16-week treatment started and the participants from the experimental group used the system to study English grammar for at least 15 min every day." (p. 7)

  • D

    The control group's size, gender proportion, year level, English-learning history, baseline test scores, and exact conditions (system access limited to weekly assignments) are clearly documented.

    "No statistical differences were found between the two groups concerning their average age, gender proportion, length of English learning (see Table 1)." (p. 4)

  • S

    The study took place in a single senior secondary school with randomisation of individual students, so there was no school-level randomisation.

    "A total of 598 EFL learners were recruited from a senior secondary school using the convenience sampling technique." (p. 4)

  • I

    The same authors designed the system, conceived the study, collected the data, and analysed the results, with no independent third-party evaluation or oversight reported.

    "XW conceived and designed the research project, collected the data, and contributed to the writing and editing of literature review, research methodology, and discussion." (p. 11)

  • Y

    The tracking period was only 16 weeks (about 4 months), well short of 75% of an academic year.

    "After that, the 16-week treatment started and the participants from the experimental group used the system to study English grammar for at least 15 min every day." (p. 7)

  • B

    Although the experimental group received substantially more system access and daily practice time (at least 15 min/day), this added access and usage is the personalised m-learning intervention itself being tested against a near business-as-usual control that used the same system only for weekly assignments, so the imbalance is integral to the treatment.

    "The experimental group, who had full access to the functions of the system (i.e., recommendations, feedback and an e-portfolio package), used the system to learn grammar, complete grammar exercise, and engage in for a semester, while the control group only use the system to submit weekly assignments." (p. 1)

  • R

    No independent replication of this specific study by a different research team is mentioned in the paper or found via an internet citation search, and the authors describe it as among the first of its kind.

    "The results obtained in the current study, which was among the first few to investigate EFL grammar instruction with middle school students using mobile devices..." (p. 9)

  • A

    Only English grammar was assessed with a custom test, so with criterion E unmet and no other core subjects measured, all-subject standardised assessment is absent.

    "The grammar tests focused exclusively on the grammar points and consisted of various discrete items to assess participants' knowledge of the target grammatical structures..." (p. 6)

  • G

    Measurement stopped at the post-test immediately after the 16-week treatment, criterion Y is not met (a prerequisite for G), and no follow-up publication tracking this cohort to graduation was found via internet search.

    "To measure the instructional effects of the system on students' performance in grammar texts, participants in both groups completed the post- grammar test after the treatment." (p. 7)

  • P

    There is no mention of any pre-registration, registry platform, or protocol registration anywhere in the paper, and no matching registration record was found via internet search.

Abstract

This study developed a personalized mobile-assisted system with a self-regulated learning (SRL) mechanism to facilitate English-as-a-foreign-language (EFL) students' learning of grammar. A quasi-experimental design, involving an experimental group (n = 278) and a control group (n = 320), was adopted to examine its effectiveness on students' grammar learning. The experimental group, who had full access to the functions of the system (i.e., recommendations, feedback and an e-portfolio package), used the system to learn grammar, complete grammar exercise, and engage in for a semester, while the control group only use the system to submit weekly assignments. The results suggest that the experimental group obtained significantly higher scores in English grammar tests than the control group and the system benefited both genders with no significant differences. The findings provide empirical evidence in support of the effectiveness of this system in improving students' grammar test scores, indicating its value as a supplementary tool to conventional classroom teaching of English grammar via supporting learners' SRL development in a m-learning context. The pedagogical implications for applying this system in classrooms and beyond were discussed.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was carried out at the individual student level within a single school rather than by class or school, and the intervention is a self-study mobile system, not one-to-one tutoring, so no exception applies.
      • "A total of 598 EFL learners were recruited from a senior secondary school using the convenience sampling technique. They were randomly assigned into the experimental group (n = 278) or the control group (n = 320)." (p. 4)
      • Relevant Quotes: 1) "A quasi-experimental design, involving an experimental group (n = 278) and a control group (n = 320), was adopted to examine its effectiveness on students' grammar learning." (p. 1, Abstract) 2) "A total of 598 EFL learners were recruited from a senior secondary school using the convenience sampling technique. They were randomly assigned into the experimental group (n = 278) or the control group (n = 320)." (p. 4) 3) "Although the participants were assigned into the conditions randomly, whether the two groups were equivalent in the use of SRL strategies had not been examined." (p. 10) Detailed Analysis: Criterion C requires randomisation of entire classes (or schools), unless the intervention is one-to-one personal teaching such as tutoring. The paper states that 598 individual students from one senior secondary school were "randomly assigned" to the experimental or control condition. There is no mention of classes or schools being the unit of randomisation; the unit is clearly the individual student. The intervention is a self-directed mobile learning system used by students as a supplement to their regular classroom instruction, not a personal tutoring or one-to-one teaching intervention, so the tutoring exception does not apply. Because treated and control students attend the same school (and plausibly the same classes), contamination between conditions cannot be ruled out. Note also that the authors themselves label the design "quasi-experimental" even though they describe random assignment of individuals. Criterion C is not met because randomisation was performed at the student level within a single school without a valid tutoring exception.
    • E

      Exam-based Assessment

      • Outcomes were measured with 50-item multiple-choice grammar tests custom-designed by two EFL teachers for this study, not a widely recognised standardised exam.
      • "Both the pre- and post- grammar tests comprised of 50 multiple choices questions, designed by two experienced EFL teachers, to examine the grammar structures that were introduced in the system, with each question worth two points." (p. 6)
      • Relevant Quotes: 1) "Each of the pre- and posttests was a cumulative review of the grammatical structures that were taught in class as well as in the system." (p. 6) 2) "Both the pre- and post- grammar tests comprised of 50 multiple choices questions, designed by two experienced EFL teachers, to examine the grammar structures that were introduced in the system, with each question worth two points." (p. 6) 3) "The internal consistency of the tests was acceptable for both the pre- and post-tests (above.80)." (p. 6) Detailed Analysis: Criterion E requires that outcomes be measured with a standard, widely recognised standardised exam rather than an instrument specially designed for the study. Here the pre- and post-tests were explicitly "designed by two experienced EFL teachers" to examine the grammar structures introduced in the system under study. Although the content was aligned with the nationally sanctioned Senior High School English Curriculum, the tests themselves are researcher/teacher-made instruments constructed for this study and closely aligned to the intervention content, which is exactly the situation the criterion is designed to guard against. No named national or state standardised examination (e.g., the College Entrance Examination) was used as the outcome measure. Criterion E is not met because the outcome measure was a custom-made grammar test rather than a recognised standardised exam.
    • T

      Term Duration

      • The intervention ran for 16 weeks (a full semester) from the pre-test at the beginning of the semester to the post-test after the treatment, which covers at least one academic term.
      • "After that, the 16-week treatment started and the participants from the experimental group used the system to study English grammar for at least 15 min every day." (p. 7)
      • Relevant Quotes: 1) "At the beginning of the semester, participants in both groups were invited to complete the pre-grammar test in classroom environments." (p. 6) 2) "After that, the 16-week treatment started and the participants from the experimental group used the system to study English grammar for at least 15 min every day." (p. 7) 3) "To measure the instructional effects of the system on students' performance in grammar texts, participants in both groups completed the post- grammar test after the treatment." (p. 7) 4) "The experimental group ... used the system to learn grammar, complete grammar exercise, and engage in for a semester, while the control group only use the system to submit weekly assignments." (p. 1, Abstract) Detailed Analysis: Criterion T requires the interval from intervention start to outcome measurement to be at least one full academic term (approximately 3-4 months). The intervention started shortly after the pre-test at the beginning of the semester, lasted 16 weeks (described as "a semester"), and the post-test was administered after the treatment ended. Sixteen weeks is approximately four months, which meets the definition of a full term/semester between intervention start and outcome measurement. Criterion T is met because outcomes were measured after a 16-week, semester-long intervention period, which is at least one full academic term after the intervention began.
    • D

      Documented Control Group

      • The control group's size, gender proportion, year level, English-learning history, baseline test scores, and exact conditions (system access limited to weekly assignments) are clearly documented.
      • "No statistical differences were found between the two groups concerning their average age, gender proportion, length of English learning (see Table 1)." (p. 4)
      • Relevant Quotes: 1) "No statistical differences were found between the two groups concerning their average age, gender proportion, length of English learning (see Table 1)." (p. 4) 2) "TABLE 1 | Participant demographic information. ... Control group 320 62.2% [male] 10 [year level] 6.14 (0.78) [average length of English learning]" (p. 4) 3) "Participants in the control group, however, had limited access to the functions of the system. They could only login the system to complete grammar exercises as assignments and submit their answers. Differing from the exercises provided for the experimental group, which were selected and recommended based on students' learning history and performance, the exercises used in the control group were chosen based on the year level without considering their learning progress." (p. 7) 4) "It can be seen from Figure 8 that the experimental group scored higher than the control group in the post-test, while the two groups obtained similar scores in the pre-test." (p. 9) Detailed Analysis: Criterion D requires detailed documentation of the control group, including demographics, baseline performance, and the treatment it received. The paper reports the control group's size (n = 320), gender proportion (62.2% male), year level (Grade 10), and average length of English learning (6.14 years, SD 0.78) in Table 1, confirms baseline comparability on demographics, and reports baseline (pre-test) grammar scores for both groups. It also describes precisely what the control group did during the study: they attended the same English courses and could only log into the system to complete and submit weekly year-level grammar assignments, without the personalised recommendation, feedback, and e-portfolio functions. This is sufficient to assess comparability and the control condition. Criterion D is met because the control group's composition, baseline characteristics, and conditions are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • The study took place in a single senior secondary school with randomisation of individual students, so there was no school-level randomisation.
      • "A total of 598 EFL learners were recruited from a senior secondary school using the convenience sampling technique." (p. 4)
      • Relevant Quotes: 1) "This research was conducted in a senior high school in mainland China, where all students were enrolled in English courses (240-300 min) per week..." (p. 4) 2) "A total of 598 EFL learners were recruited from a senior secondary school using the convenience sampling technique. They were randomly assigned into the experimental group (n = 278) or the control group (n = 320)." (p. 4) Detailed Analysis: Criterion S requires randomisation among schools (or equivalent implementing institutions). This study was conducted within one senior high school in mainland China, and the unit of randomisation was the individual student. There is no mention of multiple schools or sites, let alone random assignment of schools to conditions. Criterion S is not met because the study involved a single school and randomised individual students, not schools.
    • I

      Independent Conduct

      • The same authors designed the system, conceived the study, collected the data, and analysed the results, with no independent third-party evaluation or oversight reported.
      • "XW conceived and designed the research project, collected the data, and contributed to the writing and editing of literature review, research methodology, and discussion." (p. 11)
      • Relevant Quotes: 1) "To address these research gaps, a personalized mobile- assisted system with a self-regulated learning mechanism was developed in this study to help EFL students in the learning of English grammar." (p. 2) 2) "XW conceived and designed the research project, collected the data, and contributed to the writing and editing of literature review, research methodology, and discussion. JC performed the data analysis, produced figures and tables... TZ designed the research project and contributed to the design of research instruments..." (p. 11, Author Contributions) 3) "The present study adopted a quasi-experimental research design to evaluate the effectiveness of our system on students' performance in English grammar tests." (p. 4) Detailed Analysis: Criterion I requires the study to be conducted independently of the intervention's designers, or at least with documented third-party oversight of data collection and analysis. Here the authors developed the system themselves ("our system"), designed the research project, collected the data, designed the instruments, and performed the analysis, as stated in the Author Contributions section. There is no mention of an external evaluation team, independent enumerators, or any third-party oversight. Criterion I is not met because the intervention developers themselves conducted, administered, and analysed the study without any documented independent oversight.
    • Y

      Year Duration

      • The tracking period was only 16 weeks (about 4 months), well short of 75% of an academic year.
      • "After that, the 16-week treatment started and the participants from the experimental group used the system to study English grammar for at least 15 min every day." (p. 7)
      • Relevant Quotes: 1) "After that, the 16-week treatment started and the participants from the experimental group used the system to study English grammar for at least 15 min every day." (p. 7) 2) "To measure the instructional effects of the system on students' performance in grammar texts, participants in both groups completed the post- grammar test after the treatment." (p. 7) 3) "The experimental group ... used the system to learn grammar, complete grammar exercise, and engage in for a semester..." (p. 1, Abstract) Detailed Analysis: Criterion Y requires that outcomes be measured at least 75% of an academic year (roughly 6.75-7.5 months of a 9-10 month year) after the intervention begins. The intervention and outcome tracking here lasted 16 weeks (one semester, about 4 months), with the post-test administered immediately after the treatment. There was no longer follow-up measurement. Four months falls clearly short of 75% of an academic year. Criterion Y is not met because the interval from intervention start to final measurement was only about 4 months, far less than 75% of an academic year.
    • B

      Balanced Control Group

      • Although the experimental group received substantially more system access and daily practice time (at least 15 min/day), this added access and usage is the personalised m-learning intervention itself being tested against a near business-as-usual control that used the same system only for weekly assignments, so the imbalance is integral to the treatment.
      • "The experimental group, who had full access to the functions of the system (i.e., recommendations, feedback and an e-portfolio package), used the system to learn grammar, complete grammar exercise, and engage in for a semester, while the control group only use the system to submit weekly assignments." (p. 1)
      • Relevant Quotes: 1) "The experimental group, who had full access to the functions of the system (i.e., recommendations, feedback and an e-portfolio package), used the system to learn grammar, complete grammar exercise, and engage in for a semester, while the control group only use the system to submit weekly assignments." (p. 1, Abstract) 2) "After that, the 16-week treatment started and the participants from the experimental group used the system to study English grammar for at least 15 min every day. There was no time limitation for using the m-learning system and students were encouraged to use the system when needed." (p. 7) 3) "Participants in the control group, however, had limited access to the functions of the system. They could only login the system to complete grammar exercises as assignments and submit their answers. Differing from the exercises provided for the experimental group, which were selected and recommended based on students' learning history and performance, the exercises used in the control group were chosen based on the year level without considering their learning progress." (p. 7) 4) "This research was conducted in a senior high school in mainland China, where all students were enrolled in English courses (240-300 min) per week with explicit instructions of curriculum..." (p. 4) 5) "The findings provide empirical evidence in support of the effectiveness of this system in improving students' grammar test scores, indicating its value as a supplementary tool to conventional classroom teaching of English grammar..." (p. 1, Abstract) Detailed Analysis: Comparing inputs: both groups received identical regular classroom English instruction (240-300 min per week). The experimental group additionally had full access to the personalised system (recommendations, feedback, e-portfolio) and was asked to use it at least 15 minutes every day for 16 weeks, plus weekly exercises and weekly performance reports. The control group used the same system only to complete and submit non-personalised weekly grammar assignments. The experimental group therefore received substantially more supplementary study time and materials (roughly 15+ min/day versus one weekly assignment), which is a real resource and time imbalance and is documented here explicitly. Applying the decision tree: extra resources are present and not negligible. The next question is whether the extra resources are integral to the treatment being tested. The study's stated purpose is to evaluate the effectiveness of the personalised mobile-assisted system with an SRL mechanism "as a supplementary tool to conventional classroom teaching." The additional out-of-class system access and daily self-regulated practice time are not an optional add-on; they constitute the intervention itself (a personalised, self-regulated m-learning package layered on top of business-as-usual instruction). The control condition preserved the same core inputs (identical classroom teaching, and even a weekly system-based assignment), so the contrast is intervention-package versus business-as-usual, analogous to the ERCT exception examples (e.g., the DPL tool tested against schools that "continued to teach as usual"). It should be clearly noted, however, that this design cannot separate the effect of the personalisation/SRL functions from the effect of the additional practice time, since the experimental group simply practised much more; this confound is inherent to the package-versus-usual design. Criterion B is met because the additional system access and daily practice time are integral components of the m-learning intervention package being tested against a business-as-usual control that retained the same core classroom inputs.
  • Level 3 Criteria

    • R

      Reproduced

      • No independent replication of this specific study by a different research team is mentioned in the paper or found via an internet citation search, and the authors describe it as among the first of its kind.
      • "The results obtained in the current study, which was among the first few to investigate EFL grammar instruction with middle school students using mobile devices..." (p. 9)
      • Relevant Quotes: 1) "The results obtained in the current study, which was among the first few to investigate EFL grammar instruction with middle school students using mobile devices, also echo with empirical evidence in support of the positive effects of m-learning systems with an SRL mechanism..." (p. 9) 2) "Despite the prominent role SRL plays in EFL/ESL learning in mobile learning contexts, there were few systems with SRL mechanisms for the learning of English grammar. How to effectively improve EFL/ESL students' learning of English grammar via promoting their SRL abilities in mobile learning contexts remains largely unexplored." (p. 3) 3) "The findings could serve as initial evidence in support of the effectiveness of the personalized m-learning system with SRL an mechanism..." (p. 9) Internet Search Findings: A citation search (Semantic Scholar citation graph for DOI 10.3389/fpsyg.2021.624430) returned roughly two dozen citing works, including: Alqarni (2024) "Effect of Mobile Assisted Learning on English Language Vocabulary and Grammar: The Saudi Arabian Context"; Chernova, Litvinov, Telezhko, and Ermolova (2022) "Teaching science language grammar to would- be translators ... via m-learning"; Shadiev, Wang, Halubitskaya, and Huang (2022) "Enhancing Foreign Language Learning Outcomes ... in a Mobile-Assisted Language Learning Environment"; Murtiningsih, Wati, and Haryadi (2024) "The Effectiveness of Mobile Applications in Developing Students' English Grammar Skills"; and Sepia, Kalsum, Sunubi, and Munawir (n.d.) "Reinventing Grammar Pedagogy through Mobile- Assisted Learning: Busuu-Assisted Tense Mastery." None of these papers describe themselves as a replication of this specific study; each evaluates a different mobile system, in a different country/context, with a different design, and none references reproducing Wang, Chen, and Zhang's (2021) personalized SRL grammar system, procedure, or results. No paper by a different research team attempting to replicate this specific study's system, procedure, and findings in a peer-reviewed outlet was identified. Detailed Analysis: Criterion R requires an independent replication of the study by a different research team in a different context, published in a peer-reviewed journal. The paper positions itself as among the first studies of its kind and provides "initial evidence," which indicates no replication existed at publication. Related studies cited in the paper (e.g., Chen et al., 2019 on vocabulary; Chu et al., 2017; Kondo et al., 2012) involve different interventions, systems, skills, and populations and are prior work, not replications of this specific personalised grammar-learning system and trial. The internet citation search likewise identified only topically related but methodologically distinct mobile-grammar-learning studies, none of which replicate this specific system, study design, and cohort. Criterion R is not met because no independent, peer-reviewed replication of this specific study has been identified either in the paper or via internet search.
    • A

      All-subject Exams

      • Only English grammar was assessed with a custom test, so with criterion E unmet and no other core subjects measured, all-subject standardised assessment is absent.
      • "The grammar tests focused exclusively on the grammar points and consisted of various discrete items to assess participants' knowledge of the target grammatical structures..." (p. 6)
      • Relevant Quotes: 1) "The grammar tests focused exclusively on the grammar points and consisted of various discrete items to assess participants' knowledge of the target grammatical structures..." (p. 6) 2) "Both the pre- and post- grammar tests comprised of 50 multiple choices questions, designed by two experienced EFL teachers, to examine the grammar structures that were introduced in the system..." (p. 6) Detailed Analysis: Criterion A requires standardised exam-based assessment across all main subjects taught at the students' educational level, with criterion E as a prerequisite. Criterion E is not met (custom tests), so criterion A automatically fails. In addition, the study measured only English grammar knowledge; no other core senior high school subjects (e.g., mathematics, Chinese, sciences) were assessed, and no justification for a specialised exception is offered (this is general secondary education, not vocational training). Criterion A is not met because criterion E fails and only a single narrow outcome (English grammar) was measured.
    • G

      Graduation Tracking

      • Measurement stopped at the post-test immediately after the 16-week treatment, criterion Y is not met (a prerequisite for G), and no follow-up publication tracking this cohort to graduation was found via internet search.
      • "To measure the instructional effects of the system on students' performance in grammar texts, participants in both groups completed the post- grammar test after the treatment." (p. 7)
      • Relevant Quotes: 1) "To measure the instructional effects of the system on students' performance in grammar texts, participants in both groups completed the post- grammar test after the treatment." (p. 7) 2) "Whether students of various cognitive styles would progress similarly in their grammar performance after using this system could be the focus of further research." (p. 11) Internet Search Findings: A citation search (Semantic Scholar citation graph for DOI 10.3389/fpsyg.2021.624430) did not surface any subsequent paper by Xiao Wang, Jing Chen, or Tingting Zhang that follows up on this same senior-high-school cohort's grammar (or other) outcomes through to graduation. No such follow-up publication could be found; none is invented here. Detailed Analysis: Criterion G requires tracking participants until graduation from their educational stage, and per the instructions it cannot be met if criterion Y is not met (it is not, see above). Participants were Grade 10 students, and the last measurement was the post-test administered immediately after the 16-week treatment. No follow-up through the end of senior high school (Grade 12 graduation) or the College Entrance Examination is reported in the paper, and no follow-up publications tracking this cohort were located via internet search. Criterion G is not met because tracking ended immediately after the semester-long intervention, criterion Y (a prerequisite) is not met, and no follow-up graduation-tracking publication by the same authors was found.
    • P

      Pre-Registered

      • There is no mention of any pre-registration, registry platform, or protocol registration anywhere in the paper, and no matching registration record was found via internet search.
      • Relevant Quotes: 1) "Ethical review and approval was not required for the study on human participants in accordance with the local legislation and institutional requirements. Written informed consent to participate in this study was provided by the participants' legal guardian/next of kin." (p. 11) 2) "Received: 31 October 2020; Accepted: 10 September 2021; Published: 07 October 2021" (p. 1) Internet Search Findings: No pre-registration record for this study was located in the course of this verification (no reference to OSF, ClinicalTrials.gov, AEA RCT Registry, ISRCTN, or a similar registry was found in the paper or via search of the article's DOI landing page and citation record). The paper's own text contains no registration statement or identifier. Detailed Analysis: Criterion P requires the study protocol (hypotheses, methods, planned analyses) to be publicly pre-registered before data collection began. The paper contains no reference to any registry (e.g., ClinicalTrials.gov, OSF, AEA registry), no registration ID, and no statement about a pre-registered protocol or analysis plan. The ethics statement even notes that formal ethical review was not required, and no protocol publication is cited or found. Criterion P is not met because no pre-registration of the study protocol is mentioned in the paper or evidenced by internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.