Abstract
This study was intended to compare processing instruction (VanPatten, 1993, 1996, 2000), an input-based focus on form technique, to dictogloss tasks, an output-oriented focus-on-form type of instruction to assess their effects in helping beginning-EFL (English as a Foreign Language) learners acquire the simple English passive voice. Two intact classes of Grade 7 beginning EFL learners (n = 110) in China were randomly assigned to each type of instruction: Group A (n = 55) to processing instruction (PI), and Group B (n = 55) to dictogloss tasks (DG). A pretest and posttest (immediate and delayed) design was used, where participants' ability to comprehend and produce the target feature was assessed. Results showed that the PI group performed significantly better than the DG group in comprehension, and as well as the DG group in production on the immediate posttest. One month later, the two groups' performances were similar in terms of both comprehension and production on the delayed posttest. Both groups improved significantly from the pretest to the two posttests in comprehension and production. One reasonable pedagogical implication is that both PI and DG are effective pedagogical tools to help beginning-EFL learners to acquire target grammatical forms.
Full
Article
ERCT Criteria Breakdown
-
Level 1 Criteria
-
C
Class-level RCT
- Whole intact classes, not individual students within a single class, were randomly assigned to the two instructional conditions, satisfying the unit-of- randomisation requirement, though the design used only two classroom clusters and is self-labelled quasi-experimental.
- "Two intact classes of Grade 7 beginning EFL learners (n = 110) in China were randomly assigned to each type of instruction: Group A (n = 55) to processing instruction (PI), and Group B (n = 55) to dictogloss tasks (DG)."
Relevant Quotes:
1) "Two intact classes of Grade 7 beginning EFL learners (n = 110) in China were randomly assigned to each type of instruction: Group A (n = 55) to processing instruction (PI), and Group B (n = 55) to dictogloss tasks (DG)." (Abstract, p. 61)
2) "The participants were two intact classes of the seventh-grade beginning-EFL learners (n = 115), ages 13 to 15." (p. 65)
3) "In the secondary school where the study was conducted, there were only two parallel classes at the seventh grade." (Note 2, p. 77)
4) "This study used a quasi-experimental pretest– posttest design." (p. 71)
Detailed Analysis:
The paper describes randomisation at the class level: one entire intact class received processing instruction (PI) and the other entire intact class received dictogloss tasks (DG). No individual students within the same classroom were split between conditions, so the contamination problem that Criterion C targets (mixing treatment and control students in one classroom) does not arise here. This satisfies the literal requirement of class-level (not within-class student-level) assignment.
However, this is a very weak instance of randomisation: the school had only two parallel seventh-grade classes in total, so the "randomisation" amounted to a single coin flip deciding which of the two existing classes got which treatment, with no replication across multiple classes per condition. The author explicitly labels the overall design "quasi-experimental," which is a meaningful caveat about the strength of causal inference even though the assignment mechanism is described as random and operates at the class level.
Criterion C is met on a technical basis because assignment was by whole class rather than by student within a class, although the very small number of clusters and the author's own "quasi-experimental" label warrant caution when interpreting the strength of this randomisation.
-
E
Exam-based Assessment
- The study used three parallel tests created by the researcher specifically for this study, not a widely recognised standardised exam.
- "Three parallel tests, Test A, B, C were used, with Test A as the pretest, Test B as the immediate posttest, and Test C as the delayed posttest ... These tests were created to assess the participants' ability to comprehend and produce English passive voice."
Relevant Quotes:
1) "Three parallel tests, Test A, B, C were used, with Test A as the pretest, Test B as the immediate posttest, and Test C as the delayed posttest. (Representative samples of the test appear in Appendix C.) These tests were created to assess the participants' ability to comprehend and produce English passive voice." (p. 70)
2) "Each test employed a variety of assessment tasks, and included seven tasks." (p. 70)
3) "Estimates of internal consistency, Cronbach's alpha values for the comprehension tasks in the immediate and the delayed posttests were the same, .74 ... the values of alpha for the production tasks ... were .60 and .61 respectively." (p. 71)
Detailed Analysis:
The outcome measures are three parallel researcher- designed tests (Test A/B/C) built specifically to probe comprehension and production of the single target grammatical feature (English passive voice) taught in this study. The paper reports internal-consistency reliability statistics for these custom instruments, but reliability statistics do not make an instrument a standardised, widely recognised exam; there is no reference to any national, state-wide, or otherwise externally validated standardised test being used. This is a bespoke, study-specific assessment tailored tightly to the intervention content (the same two stories used in instruction reappear as test items), which is exactly the kind of custom instrument the ERCT standard's E criterion is designed to flag.
Criterion E is not met because the assessments are custom-built tests created by the researcher for this study rather than standardised, widely recognised exams.
-
T
Term Duration
- The instructional treatment lasted only two 45-minute class periods and the final (delayed) measurement was taken about five weeks after the pretest, far short of a full academic term.
- "The instructional treatments lasted two class periods spread over two days in a row, with 90 minutes in total ... The delayed posttest, Test C was administered one month after the immediate posttest."
Relevant Quotes:
1) "The pretest, Test A was given to the participants one week before instruction. The instructional treatments lasted two class periods spread over two days in a row, with 90 minutes in total. The immediate posttest, Test B took place the following day after the completion of instruction. The delayed posttest, Test C was administered one month after the immediate posttest." (p. 71)
2) "Overall, the instruction lasted two class periods, with 45 minutes in each class." (p. 68, PI materials)
3) "The DG instruction lasted two class periods, with 45 minutes per class period." (p. 69)
Detailed Analysis:
The intervention itself is extremely brief: 90 minutes of instruction spread over two consecutive days. From intervention start to the final (delayed) outcome measurement, the total elapsed time is roughly one week (pretest) plus two days of treatment plus about one month to the delayed posttest, i.e. approximately five to six weeks overall, and the delayed measurement is only about one month after the treatment ended. This is well short of the one full academic term (roughly 3-4 months) that Criterion T requires between intervention start and outcome measurement.
Criterion T is not met because the interval from intervention start to the final measurement (about five weeks) is far shorter than one academic term.
-
D
Documented Control Group
- The study explicitly states it had no control group at all; both groups received an instructional treatment (PI or DG) and neither served as an untreated or business-as-usual comparison.
- "The study failed to have a control group due to both practical constraints and ethical considerations."
Relevant Quotes:
1) "The study failed to have a control group due to both practical constraints and ethical considerations." (p. 77)
2) "In the secondary school where the study was conducted, there were only two parallel classes at the seventh grade." (Note 2, p. 77)
3) "Therefore, future studies are suggested to include a control group to have a more valid design." (p. 77)
Detailed Analysis:
The paper directly and unambiguously states that it lacked a control group. Both experimental arms (PI and DG) received active instruction on the target feature; neither arm was a documented "business as usual" or no-treatment comparison group. The author attributes this to the school having only two parallel classes and to ethical concerns about withholding instruction, and explicitly recommends that future studies include a control group. Because there is no control group of any kind, there is nothing to document (demographics, baseline performance, or conditions of an untreated comparison group), so this criterion cannot be satisfied.
Criterion D is not met because the authors explicitly state the study had no control group.
-
Level 2 Criteria
-
S
School-level RCT
- Randomisation occurred between two classes within a single school, not between multiple schools, so the stronger school-level RCT requirement is not met.
- "Two intact classes of Grade 7 beginning EFL learners (n = 110) in China were randomly assigned to each type of instruction."
Relevant Quotes:
1) "The study was conducted in a secondary school on the east coast of China." (p. 65)
2) "Two intact classes of Grade 7 beginning EFL learners (n = 110) in China were randomly assigned to each type of instruction: Group A (n = 55) to processing instruction (PI), and Group B (n = 55) to dictogloss tasks (DG)." (Abstract, p. 61)
3) "In the secondary school where the study was conducted, there were only two parallel classes at the seventh grade." (Note 2, p. 77)
Detailed Analysis:
The entire study took place within a single secondary school, using its two seventh-grade classes as the two experimental arms. There is no allocation across multiple schools or educational institutions; the unit of randomisation is the classroom within one school, not the school itself. This is a single-site study, which is explicitly weaker than the school-level requirement of Criterion S.
Criterion S is not met because the study was conducted in a single school with randomisation only between its two classes, not across multiple schools.
-
I
Independent Conduct
- The same researcher who designed the instructional materials and study also administered the tests and conducted the analysis; there is no external or third-party evaluator.
- "The instructor was given detailed explanations from the researcher on how to carry out the treatments faithfully."
Relevant Quotes:
1) "The English instructor of the two parallel classes carried out the instruction in order to avoid any novelty effect that might be induced by bringing in a new person to the class. The instructor was given detailed explanations from the researcher on how to carry out the treatments faithfully." (p. 71)
2) "In addition, the instructor was asked to take notes on the activities that had been done during the treatments so that the researcher could check their fidelity. All the participants' worksheets and tests were collected after the treatments." (p. 71)
3) The paper is single-authored ("Jingjing Qin, Northern Arizona University, USA"), with no mention anywhere of an external evaluation team, agency, or blinded test administrators.
Detailed Analysis:
The classroom instructor (a member of school staff, not an independent evaluator) delivered the treatments, but did so under the direct, detailed guidance of the sole researcher/author, who designed both instructional materials and the test instruments, and who personally checked treatment fidelity and collected/analysed the resulting data. There is no external agency, independent data-collection team, or blinded assessor described anywhere in the paper; the same individual who designed the intervention and the outcome measures also controlled implementation fidelity and the analysis. This is the opposite of the independent, third-party evaluation Criterion I requires.
Criterion I is not met because the study's designer/ author retained full control over instruction fidelity, testing, and analysis, with no independent evaluator involved.
-
Y
Year Duration
- Because Criterion T (Term Duration) is not met, the stronger Year Duration criterion is automatically not met as well.
- "The delayed posttest, Test C was administered one month after the immediate posttest."
Relevant Quotes:
1) "The instructional treatments lasted two class periods spread over two days in a row, with 90 minutes in total ... The delayed posttest, Test C was administered one month after the immediate posttest." (p. 71)
Detailed Analysis:
Per the ERCT specification, if the weaker Term Duration criterion (T) is not met, the stronger Year Duration criterion (Y) cannot be met either. Here, tracking from intervention start to final measurement spanned only about five to six weeks in total, nowhere near the 75% of an academic year required by Y.
Criterion Y is not met because Criterion T is not met, and the actual follow-up period (about one month post- intervention) is far shorter than a year in any case.
-
B
Balanced Control Group
- Although there was no untreated control group, the two instructional arms (PI and DG) received explicitly matched amounts of class time and comparable materials, so no imbalance in resources existed between them.
- "In order to make sure that any differences in instructional results could only be attributed to treatments themselves, DG materials were kept similar to PI materials as much as possible."
Relevant Quotes:
1) "In order to make sure that any differences in instructional results could only be attributed to treatments themselves, DG materials were kept similar to PI materials as much as possible. The same metalinguistic explanation of the target grammatical feature used in PI materials was used in DG materials. The two stories used in PI materials were used as the reconstruction passages in DG ... Altogether two DG activities were used to keep the length of instruction similar to that of PI." (p. 68)
2) "Overall, the instruction lasted two class periods, with 45 minutes in each class." (PI, p. 68)
3) "The DG instruction lasted two class periods, with 45 minutes per class period." (DG, p. 69)
Detailed Analysis:
This study has no untreated "business as usual" control group (see Criterion D), so there is no comparison condition that received less time or fewer resources than an intervention condition in the classic sense. Instead, both experimental arms were active instructional treatments, and the author took explicit, documented steps to equalise the amount of instructional time (two 45-minute class periods, 90 minutes total, for both PI and DG) and to keep content (the same metalinguistic explanation and the same two source stories) as similar as possible between arms, differing only in the pedagogical technique (input-based vs output-based practice) being tested. Applying the Criterion B logic, no extra time or budget was given to one arm without a matching allocation to the other (i.e. no EXTRA_RESOURCES_PRESENT relative to each other); the resource allocation between the two conditions is balanced by design.
Criterion B is met because the two instructional conditions received matched instructional time and closely comparable materials, with no documented resource imbalance between them, even though the study lacked a separate untreated control group.
-
Level 3 Criteria
-
R
Reproduced
- No independent, peer-reviewed replication of this specific study (PI vs dictogloss on English passive voice with beginning Chinese EFL learners) could be confirmed with verbatim quotes, though a possibly related study was located by title during this check.
Relevant Quotes:
1) "This study is among the first few contributing to the PI literature by comparing a theoretically grounded output-based instruction to PI." (p. 76)
2) "This study has demonstrated the effect of two types of instruction on learners' acquisition of only one English grammatical feature within limited hours of instruction. A great contribution to the field would be made by investigating the effect of PI and DG on more target grammatical features..." (p. 77)
Detailed Analysis:
The paper itself frames this study as among the first to compare PI directly to dictogloss tasks, and closes by calling for future research extending this comparison to other grammatical features rather than citing any existing replication of its own specific design.
Internet research for this verification pass (via Semantic Scholar and OpenAlex citation records) found that this paper has been cited roughly 100-120 times. Among the citing works, one title stands out as a plausible conceptual replication: Tharamanit, A., & Kanprachar, N. (2017), "Effects of Processing Instruction and Dictogloss on the Acquisition of the English Passive Voice among Thai University Students," which appears from its title to compare the same two treatments (PI vs. dictogloss) on the same target feature (English passive voice) but with a different population (Thai university students rather than Chinese Grade 7 EFL learners) and by different authors. However, repeated attempts to retrieve this paper's full text, abstract, or venue during this check were blocked (search engines returned CAPTCHA challenges, academic APIs returned rate-limit errors or no abstract, and it is not indexed in OpenAlex or ERIC), so no verbatim quotes could be obtained from it, its peer-reviewed status could not be confirmed, and its relationship to this specific study (Qin, 2008) could not be verified. No other candidate replication was found.
Criterion R is not met because no independently verified, quoted evidence of a peer-reviewed replication of this specific study was found. The Tharamanit and Kanprachar (2017) title is flagged here for a future manual check rather than being counted toward this criterion.
-
A
All-subject Exams
- Because Criterion E (Exam-based Assessment) is not met, the stronger All-subject Exams criterion is automatically not met; in addition, only the single target grammatical feature (passive voice) was assessed, not all core subjects.
- "These tests were created to assess the participants' ability to comprehend and produce English passive voice."
Relevant Quotes:
1) "These tests were created to assess the participants' ability to comprehend and produce English passive voice." (p. 70)
Detailed Analysis:
Per the ERCT specification, Criterion A requires Criterion E to be met as a prerequisite; since E is not met (the tests are custom, non-standardised instruments), A cannot be met either. Independently, the outcome measures here assess only one narrow grammatical feature (the English passive voice) within a single subject (English/L2 acquisition), not performance across all main subjects taught to these students, so the substantive requirement of A is also not satisfied on its own terms.
Criterion A is not met both because Criterion E is not met and because only a single grammatical feature within one subject was assessed.
-
G
Graduation Tracking
- Because Criterion Y (Year Duration) is not met, the Graduation Tracking criterion is automatically not met; participants were only followed for about one month after the intervention, and no follow-up publications tracking this cohort were found.
- "The delayed posttest, Test C was administered one month after the immediate posttest."
Relevant Quotes:
1) "The delayed posttest, Test C was administered one month after the immediate posttest." (p. 71)
Detailed Analysis:
Per the ERCT specification, if Criterion Y is not met, Criterion G cannot be met either. In addition, the paper contains no mention of any follow-up tracking of participants toward graduation from secondary school; the last data collection point is the delayed posttest, one month after the intervention ended, with no indication of any subsequent follow-up study.
Internet research for this verification pass (searches of Semantic Scholar and OpenAlex records for author Jingjing Qin) found no subsequent publication by this author tracking the same seventh-grade cohort toward graduation or reporting any longer-term follow-up on their English passive voice acquisition.
Criterion G is not met because Criterion Y is not met, and no evidence of graduation tracking was found in the paper or in a search for follow-up publications.
-
P
Pre-Registered
- There is no mention anywhere in the paper of a pre-registered study protocol, hypotheses, or analysis plan published before data collection began, and no registry entry was found via external search.
Relevant Quotes:
No quotes were found anywhere in the paper referencing a study registry (e.g., a clinical-trials-style registry), a pre-registration date, or a publicly posted protocol predating data collection. The Method and Procedures sections describe the design, materials, and analysis plan only within the paper itself, without any reference to prior registration.
Detailed Analysis:
The ERCT standard requires quoted evidence of a pre-registration reference and a registration date preceding data collection. No such statement, registry name, or identifier appears anywhere in this paper (published in 2008, a period when pre-registration was not yet standard practice in applied linguistics/SLA research of this kind).
An internet search for this verification pass (for the paper title, author, and pre-registration/registry terms in connection with this study) found no registry entry or protocol associated with this study.
Criterion P is not met because no pre-registration statement or registry reference is found in the paper or via external search.
Request an Update or Contact Us
Are you the author of this study? Let us know if you have any questions or updates.