Abstract
The researcher designed a smartphone app to help college students to learn English (L2) vocabulary. The app contained 3,402 English words that were compiled into an alphabetic wordlist with each word displayed on three features; namely: spelling, pronunciation and Chinese definitions. To test the effectiveness of the app, an experimental group (with app) was compared with a control group (without app) and knowledge of words was tested before and after the research. The study revealed that the students using the program significantly outperformed those in the control group in vocabulary acquisition. This paper introduced a research design method and set up a pedagogical paradigm which can be followed as a way to practice MALL.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Randomisation was performed at the class level, with one class of each matched pair randomly assigned to the experimental condition, satisfying the class-level RCT requirement.
- "Of each pair, one class was randomly chosen to be the experimental class; another class was assigned as the control class." (p. 6)
Relevant Quotes:
1) "The participant selection was to find two separate intact classes taught by one instructor, with the group of students who possessed Android smartphones in one class displayed no significant difference in knowledge of vocabulary recognition to the corps of students who also held Android smartphones in another class. The chosen two intact classes were tagged as a pair." (p. 4)
2) "Subsequently, English vocabulary recognition pretest was conducted and three pairs of classes were obtained." (p. 4)
3) "The selected three pairs of intact classes were taught by three colleagues. Of each pair, one class was randomly chosen to be the experimental class; another class was assigned as the control class." (p. 6)
4) "There were 47 male and 54 female students in the three randomly selected experimental classes, and 46 males and 52 females in the three control classes." (p. 4)
Detailed Analysis:
The ERCT 'C' criterion requires random assignment by entire class (or a stronger school-level unit), rather than by individual students within a single classroom. In this study, intact classes were matched into three pairs, and within each pair one whole class was randomly chosen as the experimental class and the other served as control. The unit of randomisation is therefore the class, not the individual student, which prevents within-class contamination. Although the paper does not detail the mechanism of the random choice, it explicitly states classes were "randomly chosen" and refers to "three randomly selected experimental classes."
Criterion C is met because randomisation was conducted at the class level, with entire classes assigned to experimental or control conditions.
-
E
Exam-based Assessment
- The outcome was measured with a researcher-generated vocabulary test drawn from the app's own word pool, not a recognised standardised exam.
- "By Sample Test function in Word Learning-CET4, random selections of words were performed twice to extract 100 words each time. These selected words were used to perform the pretest and posttest..." (p. 6)
Relevant Quotes:
1) "By Sample Test function in Word Learning-CET4, random selections of words were performed twice to extract 100 words each time. These selected words were used to perform the pretest and posttest in an effort to evaluate the vocabulary level of recipients in the research. Examinees were asked to write down the common Chinese meaning during the tests; one point was awarded if a correct Chinese meaning or interpretation was noted." (p. 6)
2) "In preparation, the 3,402 words were compiled into an applicable computer database... the researcher designed Word Learning-CET4..." (p. 4)
3) "The importance of passing College English Test-Band 4 (CET4) is not emphasized enough." (p. 3)
Detailed Analysis:
The ERCT 'E' criterion requires a standardised, widely recognised exam that was not created specifically for the study. Here, the assessment consisted of 100 words randomly drawn from the researcher's own 3,402-word app database using the app's Sample Test function, with students writing Chinese meanings scored by the researcher. This is a custom, study-specific instrument aligned directly to the intervention content, not an administered standardised exam. Although the national CET4 exam is discussed as motivation and the wordlist derives from CET4 vocabulary books, the students were never assessed with the actual CET4 or any other standardised test; outcomes were measured on the bespoke word-recognition test.
Criterion E is not met because a custom researcher-designed vocabulary test, rather than a standardised exam, was used to measure outcomes.
-
T
Term Duration
- The intervention and tracking spanned one full academic semester, from the first to the last class of the term, meeting the one-term duration requirement.
- "The third reason, and probably the most important, was the length of the time the research took. It was conducted over one full academic semester." (p. 8)
Relevant Quotes:
1) "The experiment was partially conducted with one pair of classes in the Spring semester of 2012–2013 academic year... the researcher easily obtained two more desired pairs of intact classes and compared them in the Spring semester of 2013–2014 academic year." (p. 6)
2) "The pretests were conducted in the first class of semester. The posttests were conducted in the last class of semester." (p. 6)
3) "Besides the treatment, time was also a variable since it took a semester to conduct the experiment." (p. 7)
4) "The third reason, and probably the most important, was the length of the time the research took. It was conducted over one full academic semester." (p. 8)
Detailed Analysis:
The ERCT 'T' criterion requires that outcomes be measured at least one full academic term (roughly 3-4 months) after the intervention begins. This study ran across a full spring semester, with the pretest administered in the first class of the semester and the posttest in the last class of the semester. The authors repeatedly emphasise that the study "took a semester" and was "conducted over one full academic semester." A semester corresponds to one academic term, satisfying the interval requirement between intervention start and outcome measurement.
Criterion T is met because the interval from intervention start to posttest covered one full academic semester (term).
-
D
Documented Control Group
- The control group's size, demographic composition, selection, baseline scores, and business-as-usual conditions are documented in the text and tables.
- "the control group consisted of 98 students who held Android smartphones in the three control classes." (p. 6)
Relevant Quotes:
1) "There were 47 male and 54 female students in the three randomly selected experimental classes, and 46 males and 52 females in the three control classes. They were sophomores and their age ranged between twenty to twenty-two." (p. 4)
2) "the control group consisted of 98 students who held Android smartphones in the three control classes. Word Learning-CET4 was only installed into smartphones in the experimental group..." (p. 6)
3) "All students in the three pairs of intact classes received printed copies of the 3,402 words in a wordlist to study. The university did not change textbook and the three instructors taught the three pairs of classes with the same curricular textbook, New Horizon College English: Reading and Writing, book 4, 2nd edition." (p. 6)
4) "Table 2. Descriptive group statistics... Control Group (n = 98) Pretest 49.80 ± 12.68 Posttest 51.08 ± 12.08" (p. 7)
Detailed Analysis:
The ERCT 'D' criterion requires the control group to be well documented, including demographics, baseline performance, size, and conditions received. The paper reports the control group's size (98 students), sex composition (46 male, 52 female), age range (20-22), academic level (sophomores in "Good" placement classes with statistically identical CEE scores), and its baseline and posttest scores in Table 2. It also states the control condition: control students did not receive the app but were given printed copies of the identical 3,402-word list and taught with the same textbook by the same paired instructor. This is sufficient documentation to confirm comparability.
Criterion D is met because the control group's demographics, size, baseline scores, and treatment conditions are clearly documented.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred among classes within a single university, not among whole schools, so the school-level RCT requirement is not satisfied.
- "Of each pair, one class was randomly chosen to be the experimental class; another class was assigned as the control class." (p. 6)
Relevant Quotes:
1) "For the selection, classes with similar number of students who possessed Android smartphones were screened from multiple sophomore 'Good' classes." (p. 4)
2) "The selected three pairs of intact classes were taught by three colleagues. Of each pair, one class was randomly chosen to be the experimental class; another class was assigned as the control class." (p. 6)
Detailed Analysis:
The ERCT 'S' criterion requires randomisation at the level of the school or equivalent implementing institution/unit. In this study, the entire experiment was carried out within a single institution (Jiujiang University), and the unit of randomisation was the class within that university, not the school. No multiple schools or sites were randomly assigned.
Criterion S is not met because randomisation was at the class level within one university, not at the school level.
-
I
Independent Conduct
- The same person designed the app and the study and conducted and analysed the experiment, with no independent or third-party evaluation.
- "Conceived and designed the experiments: QW. Performed the experiments: QW. Analyzed the data: QW... Wrote the paper: QW." (p. 11)
Relevant Quotes:
1) "I designed an application with broad practical usage to exploit MALL with smartphones in daily life..." (p. 2)
2) "the researcher designed Word Learning-CET4, a B4A application with touchscreen commands..." (p. 4)
3) "Author Contributions: Conceived and designed the experiments: QW. Performed the experiments: QW. Analyzed the data: QW. Contributed reagents/materials/analysis tools: QW. Wrote the paper: QW." (p. 11)
4) "The author thanks her colleagues Hui Jiang, Ling Zou and Huimin Liao, the three instructors to help to conduct this research." (p. 10)
Detailed Analysis:
The ERCT 'I' criterion requires the evaluation to be conducted independently from the intervention designer to reduce bias. Here, the single author (QW) designed the intervention (the app and wordlist), designed the study, performed the experiment, analysed the data, and wrote the paper. The three colleague instructors merely taught the classes; there is no statement of any independent or third-party evaluator overseeing data collection or analysis. Design, implementation, and analysis were all carried out by the intervention's creator.
Criterion I is not met because the intervention designer also conducted and analysed the study without independent oversight.
-
Y
Year Duration
- The study lasted only a single semester, which is far short of the 75% of a full academic year required.
- "It was conducted over one full academic semester." (p. 8)
Relevant Quotes:
1) "The pretests were conducted in the first class of semester. The posttests were conducted in the last class of semester." (p. 6)
2) "It was conducted over one full academic semester." (p. 8)
Detailed Analysis:
The ERCT 'Y' criterion requires outcomes to be tracked for at least 75% of a full academic year (roughly 7-8 months of a 9-10 month year). This study spanned only one spring semester (a single term of approximately 3-4 months) from the first to the last class. That is well below 75% of an academic year, so the stronger year-duration requirement is not satisfied even though the weaker term-duration criterion (T) is met.
Criterion Y is not met because tracking covered only one semester, far less than 75% of an academic year.
-
B
Balanced Control Group
- Both groups received the identical 3,402-word list and same textbook; the smartphone app itself was the explicit treatment variable, with the control given the equivalent printed materials as business as usual.
- "All students in the three pairs of intact classes received printed copies of the 3,402 words in a wordlist to study." (p. 6)
Relevant Quotes:
1) "The present study aimed to: 1. Explore the effectiveness of using smartphones as a tool for learning English vocabulary in a natural environment." (p. 3)
2) "Word Learning-CET4 was only installed into smartphones in the experimental group and access to it was protected by a password key." (p. 6)
3) "All students in the three pairs of intact classes received printed copies of the 3,402 words in a wordlist to study. The university did not change textbook and the three instructors taught the three pairs of classes with the same curricular textbook, New Horizon College English: Reading and Writing, book 4, 2nd edition." (p. 6)
4) "During the entire experiment, the researcher and instructors did not do any intervention, it was up to these participants to decide whether to use it or not." (p. 6)
5) "The design of the experiment that the participants in the experimental group received no additional treatment except with downloaded Word Learning-CET4 was also natural." (p. 9)
Detailed Analysis:
Applying the updated Criterion B decision tree: extra resources are present (the app itself, plus a one-off installation/teaching session "less than 30 minutes" given only to the experimental group), so the check proceeds. That one-off session is a negligible, purely administrative setup step for using the treatment, not an ongoing instructional or resource advantage, so it does not itself break balance. More importantly, the core resource being compared, the wordlist, is content-identical between groups: both received the same 3,402-word list and were taught the same textbook by the same paired instructor with no added formal class time or budget for either group. The paper's explicit aim is to test "the effectiveness of using smartphones as a tool for learning English vocabulary," i.e., the mode of delivery (app-based accessibility) is itself the treatment variable being evaluated, and the control's printed wordlist is the "business as usual" equivalent providing the same content without the accessibility advantage under test. The paper states explicitly that the experimental group "received no additional treatment except with downloaded Word Learning-CET4," confirming resources are integral to, and the subject of, the study's design rather than a separable confound. Under the decision tree, since the additional resource (the app) is the treatment variable itself, the criterion is met regardless of the control not receiving the app.
Criterion B is met because both groups received the same wordlist content and textbook, with the smartphone app itself being the explicit treatment variable under test and the control receiving the printed equivalent as business as usual.
-
Level 3 Criteria
-
R
Reproduced
- No independent replication of this specific study by another research team was found after internet searching.
Relevant Quotes:
1) "This study provides additional support to [30]'s conclusion that 'mobile technology can enhance learners' second language acquisition'." (p. 7)
2) "This research presents a neonatal idea to invite further scrutiny." (p. 10)
Detailed Analysis:
The ERCT 'R' criterion requires independent replication of the specific study by a different research team in a different context, published in a peer-reviewed journal. The paper cites prior MALL/SMS vocabulary studies for context but these are not replications of this particular Word Learning-CET4 app design and trial. The author describes the work as a "neonatal idea to invite further scrutiny," indicating no replication exists at the time of writing. An internet search for later papers citing or replicating this exact study (Wu 2015, PLoS ONE) did not surface any independent study that re-implemented the Word Learning-CET4 trial design with a different research team; subsequent MALL vocabulary papers found (e.g., on other apps, other populations) are related-field studies, not replications of this specific study, so no verbatim replication quotes could be provided.
Criterion R is not met because no independent replication of this specific study has been reported or found.
-
A
All-subject Exams
- Only English vocabulary was assessed, using a non- standardised custom test, so with criterion E unmet the all-subject requirement fails.
- "These selected words were used to perform the pretest and posttest in an effort to evaluate the vocabulary level of recipients in the research." (p. 6)
Relevant Quotes:
1) "By Sample Test function in Word Learning-CET4, random selections of words were performed twice to extract 100 words each time. These selected words were used to perform the pretest and posttest..." (p. 6)
2) "The first purpose of this study was to find a new way for Chinese students to improve their ability to retain English vocabulary." (p. 10)
Detailed Analysis:
The ERCT 'A' criterion requires that all main subjects be assessed with standardised exams, and it explicitly depends on criterion E being met. This study measured only English L2 vocabulary and used a custom, study-specific test rather than a standardised exam. No other subjects were assessed. Because criterion E is not met and only a single narrow outcome was measured, the all-subject requirement cannot be satisfied.
Criterion A is not met because only English vocabulary was measured with a non-standardised test and criterion E is unmet.
-
G
Graduation Tracking
- Outcomes were measured only at the end of one semester with no follow-up to graduation, and with criterion Y unmet this criterion fails.
- "The posttests were conducted in the last class of semester." (p. 6)
Relevant Quotes:
1) "The posttests were conducted in the last class of semester." (p. 6)
2) "Future research design should allow participants to study with either a list of words on paper or a smartphone for a set amount of time." (p. 10)
Detailed Analysis:
The ERCT 'G' criterion requires tracking participants until graduation and depends on criterion Y being met. This study collected its final data at the end of one semester with no longer-term or graduation follow-up of the sophomore participants. An internet search for subsequent papers by Qun Wu tracking the same cohort of students (e.g., toward their college graduation) did not surface any such follow-up publication; the only related work found by this author is the original PLoS ONE study and a similarly themed 2015 ScienceDirect paper on smartphone app design, neither of which reports graduation tracking. Because criterion Y is not met and no graduation tracking is reported anywhere, this criterion fails.
Criterion G is not met because measurement stopped at the end of one semester with no tracking through graduation, and no follow-up publications tracking the cohort were found.
-
P
Pre-Registered
- No pre-registration of the study protocol on any registry is mentioned anywhere in the paper, and none was found through internet searching.
Relevant Quotes:
1) "The Ethics Committee of Jiujiang University deemed the research was part of classroom teaching and the experiment could be conducted without its permission." (p. 3)
2) "Data Availability Statement: All relevant data are within the paper." (p. 1)
Detailed Analysis:
The ERCT 'P' criterion requires the study protocol, including hypotheses and planned analyses, to be publicly pre-registered before data collection begins. The paper contains no reference to any registry (e.g., ClinicalTrials.gov, ISRCTN, OSF) or a pre-registration ID or date. The only governance mentioned is an internal ethics committee note stating formal ethics approval was waived because the study was deemed part of classroom teaching. An internet search for a pre-registration record of this study (by title, author, and DOI) returned no matching registry entry. There is no evidence of a pre-registered protocol.
Criterion P is not met because the paper provides no indication of pre-registration of the study protocol, and none could be located online.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.