The effects of processing instruction, traditional instruction and meaning-output instruction on the acquisition of the English past simple tense

Alessandro Benati

Published:
ERCT Check Date:
DOI: 10.1191/1362168805lr154oa
  • L2 languages
  • K12
  • China
  • Asia
  • EU
0
  • C

    Students were randomly assigned individually to the PI, TI and MOI groups within each school rather than whole classes or schools being randomized, and no tutoring exception applies.

    "They were split into three groups using a random procedure from the beginning of the experiment: PI group (n = 15); TI group (n = 15); MOI group (n = 17)." (p. 74-75)

  • E

    Outcomes were measured with researcher-designed interpretation and production tasks piloted by the author, not with a widely recognised standardised exam.

    "The interpretation task (see Appendix D) consisted of 20 sentences (10 distractor items in the present regular form) of which 10 were in the English past simple tense (only regular forms)." (p. 79)

  • T

    Outcomes were measured immediately after a three-day, six-hour instructional period, with no delayed post-test, far short of one academic term.

    "In both studies, the three groups were taught for three consecutive days for a total of six hours of instruction (two hours per day) on the target feature." (p. 75)

  • D

    The study compares three active instructional treatments with no documented untreated or business-as-usual control group.

    "The first group received processing instruction; the second group was exposed to traditional instruction; the third group received meaning-based output instruction." (p. 67, abstract)

  • S

    Randomisation occurred among individual students within each school, not among whole schools, so the school-level RCT requirement is not satisfied.

    "They were split into three groups using a random procedure from the beginning of the experiment: PI group (n = 15); TI group (n = 15); MOI group (n = 17)." (p. 74-75)

  • I

    The same single author designed the PI framework, materials and study, and closely directed the classroom teachers who delivered it, with no independent evaluator for data collection or analysis.

    "In both classroom studies, the regular classroom teachers also served as the instructors for the present study. They were instructed on how to implement the instructional treatments and to act only as facilitators during the experiment. The instructors were observed in order to make sure that they followed very closely the content and procedures of the instructional packages." (p. 75)

  • Y

    Since the Term Duration criterion (T) is not met, the stronger Year Duration criterion is automatically not met as well; the whole study spanned three days with an immediate post-test.

    "In both studies, the three groups were taught for three consecutive days for a total of six hours of instruction (two hours per day) on the target feature." (p. 75)

  • B

    All three instructional groups (PI, TI, MOI) received an identical amount of instructional time and the same explicit grammatical information, differing only in the nature of the practice activities.

    "The three sets of materials developed for this experiment were balanced in terms of activity types, use of visuals and vocabulary which consisted in highly frequent items. An equal amount of processing practice in PI and output practice in TI and MOI was provided." (p. 76)

  • R

    No evidence was found that this specific study (PI vs. TI vs. MOI on English past simple tense with Chinese and Greek school-age learners) has been independently replicated by a different research team; a related but non-identical Hong Kong study on the same linguistic feature does not replicate this study's specific design.

    "Given that the present study has yielded findings quite different from those that have made comparisons between PI and MOI, there is need for further research in order to ascertain what factors are involved in the outcomes." (p. 85)

  • A

    Since criterion E (Exam-based Assessment) is not met, the stronger All-subject Exams criterion is automatically not met; only English past-tense interpretation and production were assessed.

    "The interpretation task (see Appendix D) consisted of 20 sentences... of which 10 were in the English past simple tense." (p. 79)

  • G

    Since criterion Y (Year Duration) is not met, the Graduation Tracking criterion is automatically not met; outcomes were measured only immediately after the three-day intervention with no follow-up, and no subsequent graduation-tracking publication was found.

    "IMMEDIATE POST-TESTS (Immediately after the end of the instructional period) Interpretation and production tasks" (Figure 1, p. 77)

  • P

    The paper contains no statement of a pre-registered protocol, registry link, or registration date prior to data collection, and no internet search located any pre-registration record for this 2005 study.

Abstract

This paper presents the results of a parallel classroom experiment investigating the effects of processing instruction, traditional instruction and meaning-based output instruction on the acquisition of the English past simple tense. The subjects involved in the present studies were Chinese and Greek school-age learners of English residing in their respective countries. The participants in both schools were divided into three groups. The first group received processing instruction; the second group was exposed to traditional instruction; the third group received meaning-based output instruction. One interpretation and one production measure were used in a pre-test and post-test design (immediate effect only). The results showed that processing instruction had positive effects on the processing and acquisition of the target feature. In both studies the processing instruction group performed better than the traditional instruction and meaning-based output instruction groups in the interpretation task and the three groups made equal gains in the production task.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Students were randomly assigned individually to the PI, TI and MOI groups within each school rather than whole classes or schools being randomized, and no tutoring exception applies.
      • "They were split into three groups using a random procedure from the beginning of the experiment: PI group (n = 15); TI group (n = 15); MOI group (n = 17)." (p. 74-75)
      • Relevant Quotes: 1) "A parallel study was carried out among first-semester school students in two different secondary schools in China and Greece." (p. 74) 2) "They were split into three groups using a random procedure from the beginning of the experiment: PI group (n = 15); TI group (n = 15); MOI group (n = 17)." (p. 74-75) 3) "In the second school, the participants were 30 in total. The subjects were studying English in a secondary Greek school... They were also split into three groups using the same random procedure: PI group (n = 10); TI group (n = 10): MOI group (n = 10)." (p. 75) 4) "In both classroom studies, the regular classroom teachers also served as the instructors for the present study." (p. 75) Detailed Analysis: The paper describes a parallel design carried out in two secondary schools (one in China, one in Greece), where within each school the pool of consenting students (47 in China, 30 in Greece) was split into three instructional groups (PI, TI, MOI) "using a random procedure." There is no statement that whole classes were randomized as intact units; instead, groups of 10-17 individual students were formed from the eligible pool at each school and taught by the regular classroom teachers acting as instructors. This matches the ERCT Standard's described problem of assigning treatment at the individual level within the same classroom/school population, which risks contamination between groups (e.g., students discussing material, or the same teacher inadvertently blending approaches). The intervention is not a personal tutoring intervention (each condition involves group classroom instruction of 10-17 students), so the tutoring exception in the ERCT Standard does not apply. No stronger school-level randomization is present either, since both schools implement all three conditions. Criterion C is not met because randomization occurred at the individual-student level within each school rather than at the class or school level, and no tutoring exception applies.
    • E

      Exam-based Assessment

      • Outcomes were measured with researcher-designed interpretation and production tasks piloted by the author, not with a widely recognised standardised exam.
      • "The interpretation task (see Appendix D) consisted of 20 sentences (10 distractor items in the present regular form) of which 10 were in the English past simple tense (only regular forms)." (p. 79)
      • Relevant Quotes: 1) "Two versions of the tests were designed, one interpretation task and one written production task, in order to form a split block design." (p. 79) 2) "The interpretation task (see Appendix D) consisted of 20 sentences (10 distractor items in the present regular form) of which 10 were in the English past simple tense (only regular forms). The participants had to listen to the sentences and indicate (interpret) whether the sentence they heard was related to a past or present action." (p. 79) 3) "The written production task (see Appendix E) was developed and used to measure learner's ability to produce correct sentences using the English past tense. The students were required to look at 10 pictures and use the verbs provided to produce a sentence for each of the picture." (p. 79) 4) "Tests were balanced in terms of difficulty and vocabulary (high frequency items) in a previous pilot experiment." (p. 79) Detailed Analysis: Both outcome measures (a listening interpretation task and a written production task) were purpose-built by the author for this experiment and piloted internally, rather than being widely recognised, externally validated standardised exams (e.g., a national English proficiency test). The paper explicitly describes constructing the items (20 sentences; 10 pictures), balancing them for difficulty/vocabulary in "a previous pilot experiment," and creating two parallel versions (A/B) for a split-block design. There is no reference to an established, standardised assessment instrument being used or adapted. Criterion E is not met because the assessments used are custom, researcher-designed interpretation and production tasks rather than standardised exams.
    • T

      Term Duration

      • Outcomes were measured immediately after a three-day, six-hour instructional period, with no delayed post-test, far short of one academic term.
      • "In both studies, the three groups were taught for three consecutive days for a total of six hours of instruction (two hours per day) on the target feature." (p. 75)
      • Relevant Quotes: 1) "In both studies, the three groups were taught for three consecutive days for a total of six hours of instruction (two hours per day) on the target feature." (p. 75) 2) "IMMEDIATE POST-TESTS (Immediately after the end of the instructional period) Interpretation and production tasks" (Figure 1, p. 77) 3) "Despite the outcomes of the present study, long-term effects of the variables under investigation should be re-examined as delayed post-tests were not available (the semester system in the two countries where data were collected did not allow for this)." (p. 85) Detailed Analysis: The intervention itself lasted only three consecutive days (six hours total), and the outcome measures were administered immediately after the instructional period ended, as shown in Figure 1's overview of the procedure. The authors explicitly acknowledge as a limitation that no delayed post-test was administered, meaning there is no tracking of outcomes even weeks later, let alone across a full academic term (approximately 3-4 months). This falls far short of the ERCT requirement of tracking outcomes at least one term after intervention start. Criterion T is not met because outcomes were measured immediately after a three-day intervention, with no term-long (or any delayed) follow-up.
    • D

      Documented Control Group

      • The study compares three active instructional treatments with no documented untreated or business-as-usual control group.
      • "The first group received processing instruction; the second group was exposed to traditional instruction; the third group received meaning-based output instruction." (p. 67, abstract)
      • Relevant Quotes: 1) "The first group received processing instruction; the second group was exposed to traditional instruction; the third group received meaning-based output instruction." (p. 67, abstract) 2) "The same instructional treatments were used in both experiments... The first group, received the PI treatment; the second group the TI treatment; and the third group the MOI treatment." (p. 76) 3) "All the three treatments were exposed to the same amount of explicit information... regarding the target feature so that the only difference between the three treatments was limited to the nature of the practice." (p. 76) Detailed Analysis: This study compares three active instructional treatments (PI, TI, MOI) against one another; there is no untreated, business-as-usual, or no-instruction control group in this experiment. All three groups received formal instruction targeting the identical grammatical feature over the same three-day period. While participant demographics and baseline scores for each of the three groups are documented in Tables 1-4, none of the three arms functions as a genuine control group receiving no intervention or standard instruction unrelated to the target feature; TI is itself an active treatment being compared, not a documented control condition. Criterion D is not met because the study contains no documented control group; all three groups received an active instructional intervention on the same target feature.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred among individual students within each school, not among whole schools, so the school-level RCT requirement is not satisfied.
      • "They were split into three groups using a random procedure from the beginning of the experiment: PI group (n = 15); TI group (n = 15); MOI group (n = 17)." (p. 74-75)
      • Relevant Quotes: 1) "A parallel study was carried out among first-semester school students in two different secondary schools in China and Greece." (p. 74) 2) "They were split into three groups using a random procedure from the beginning of the experiment: PI group (n = 15); TI group (n = 15); MOI group (n = 17)." (p. 74-75) 3) "In the second school, the participants were 30 in total... They were also split into three groups using the same random procedure: PI group (n = 10); TI group (n = 10): MOI group (n = 10)." (p. 75) Detailed Analysis: Both schools implemented all three instructional conditions internally, with individual students allocated to PI/TI/MOI groups within each school. No school was randomised as a whole unit to a single condition; instead, each school served as a separate replication in which students were split at the individual level. This does not meet the ERCT requirement that entire schools be randomised to condition. Criterion S is not met because randomisation was conducted among individual students within each school rather than among whole schools.
    • I

      Independent Conduct

      • The same single author designed the PI framework, materials and study, and closely directed the classroom teachers who delivered it, with no independent evaluator for data collection or analysis.
      • "In both classroom studies, the regular classroom teachers also served as the instructors for the present study. They were instructed on how to implement the instructional treatments and to act only as facilitators during the experiment. The instructors were observed in order to make sure that they followed very closely the content and procedures of the instructional packages." (p. 75)
      • Relevant Quotes: 1) "In both classroom studies, the regular classroom teachers also served as the instructors for the present study. They were instructed on how to implement the instructional treatments and to act only as facilitators during the experiment. The instructors were observed in order to make sure that they followed very closely the content and procedures of the instructional packages." (p. 75) 2) "I would like to thank my students for helping me in the collection of the data presented in this study. I also express my gratitude to the two anonymous reviewers for their valuable comments." (Acknowledgements, p. 86) 3) "Instructional materials and tests are available from the author." (Note 1, p. 86) Detailed Analysis: This is a single-author paper (Alessandro Benati), a prominent proponent of the Processing Instruction (PI) approach being evaluated. The author personally designed all three instructional packets (PI, TI, MOI) and the assessment instruments, and explicitly instructed and observed the classroom teachers "to make sure that they followed very closely the content and procedures of the instructional packages," indicating the author closely controlled implementation. The acknowledgements thank "my students" for helping collect data (i.e., the author's own students/research assistants) and anonymous peer reviewers, but there is no mention of an independent third-party team conducting data collection or analysis, nor of any organisational separation between the intervention designer and the evaluator. This does not meet the independence requirement. Criterion I is not met because the study was designed, implemented under close author supervision, and analysed by the same single author who developed the PI intervention, without independent third-party conduct.
    • Y

      Year Duration

      • Since the Term Duration criterion (T) is not met, the stronger Year Duration criterion is automatically not met as well; the whole study spanned three days with an immediate post-test.
      • "In both studies, the three groups were taught for three consecutive days for a total of six hours of instruction (two hours per day) on the target feature." (p. 75)
      • Relevant Quotes: 1) "In both studies, the three groups were taught for three consecutive days for a total of six hours of instruction (two hours per day) on the target feature." (p. 75) 2) "IMMEDIATE POST-TESTS (Immediately after the end of the instructional period) Interpretation and production tasks" (Figure 1, p. 77) 3) "Despite the outcomes of the present study, long-term effects of the variables under investigation should be re-examined as delayed post-tests were not available (the semester system in the two countries where data were collected did not allow for this)." (p. 85) Detailed Analysis: Per the criteria-specific instruction, if criterion T (Term Duration) is not met, criterion Y (Year Duration) is automatically not met. Independently of that rule, the study's own text confirms this: the entire instructional period was three consecutive days, and outcomes were measured immediately afterward with no delayed post-test, which is nowhere near the 75% of an academic year required by the ERCT Standard. Criterion Y is not met, both because criterion T is not met and because the study duration (three days, immediate post-test) is far shorter than one academic year.
    • B

      Balanced Control Group

      • All three instructional groups (PI, TI, MOI) received an identical amount of instructional time and the same explicit grammatical information, differing only in the nature of the practice activities.
      • "The three sets of materials developed for this experiment were balanced in terms of activity types, use of visuals and vocabulary which consisted in highly frequent items. An equal amount of processing practice in PI and output practice in TI and MOI was provided." (p. 76)
      • Relevant Quotes: 1) "The three sets of materials developed for this experiment were balanced in terms of activity types, use of visuals and vocabulary which consisted in highly frequent items. An equal amount of processing practice in PI and output practice in TI and MOI was provided. The vocabulary used in the activities was also roughly the same across the groups." (p. 76) 2) "All the three treatments were exposed to the same amount of explicit information (see Appendices F and G) regarding the target feature so that the only difference between the three treatments was limited to the nature of the practice." (p. 76) 3) "3 CONSECUTIVE DAYS / 6 HOURS INSTRUCTION" (Figure 1, applying identically to the PI, TI and MOI columns, p. 77) Detailed Analysis: Applying the Criterion B decision procedure: extra time/budget is not a relevant concern here because all three arms (PI, TI, MOI) received exactly the same amount of instructional time (three consecutive days, six hours total) and the same amount of explicit grammatical information, as confirmed both in the narrative text and in Figure 1's procedural overview. The only manipulated variable was the nature of the practice activities (input-processing practice vs. mechanical/communicative output practice vs. meaning-based output practice), which is precisely the intended contrast of the study. Since resources (time, explicit information, vocabulary, visuals) were explicitly balanced across all groups, there is no resource imbalance that could confound the comparison. Re-checked against the current ERCT B decision tree: EXTRA_RESOURCES_PRESENT is false (no group receives more time or budget than another), so the criterion is met trivially without needing to invoke the "integral-to-treatment" exception. Criterion B is met because the study explicitly balanced instructional time, explicit information and materials across all three groups, isolating the effect of the type of practice.
  • Level 3 Criteria

    • R

      Reproduced

      • No evidence was found that this specific study (PI vs. TI vs. MOI on English past simple tense with Chinese and Greek school-age learners) has been independently replicated by a different research team; a related but non-identical Hong Kong study on the same linguistic feature does not replicate this study's specific design.
      • "Given that the present study has yielded findings quite different from those that have made comparisons between PI and MOI, there is need for further research in order to ascertain what factors are involved in the outcomes." (p. 85)
      • Relevant Quotes (from the paper): 1) "Given that the present study has yielded findings quite different from those that have made comparisons between PI and MOI, there is need for further research in order to ascertain what factors are involved in the outcomes." (p. 85) 2) "It is interesting to note that the results from the present study differ from Farley's research (Farley, 2001a; 2001b) and Benati's (Benati, 2001) as it provides new evidence indicating that PI is better than output-oriented instruction." (p. 83) Internet Search Findings: A search for independent replications of this specific study (PI vs. TI vs. MOI, English past simple tense, Chinese and Greek secondary-school learners) did not locate any paper that reproduces this exact design with a different research team. The closest related work identified is: Chan, M. (2019) "The Role of Classroom Input: Processing Instruction, Traditional Instruction, and Implicit Instruction in the Acquisition of the English Simple Past by Cantonese ESL Learners in Hong Kong," System, 80, 246-256. According to its abstract and record, that single-authored study (a different author/team from Benati) compares Processing Instruction (PI), Traditional Instruction (TI) and Implicit Instruction (II) -- not MOI -- on the same target feature (English simple past) but with Primary 2 Cantonese-L1 learners in Hong Kong, a different population, age group and country than the Chinese secondary-school and Greek secondary-school samples in the present study. No verbatim quote from Chan (2019) referencing Benati (2005) could be independently confirmed from available sources, so no such quote is reproduced here to avoid fabrication. Detailed Analysis: The paper situates itself within a broader body of PI research (VanPatten and Cadierno, 1993; Cadierno, 1995; Benati, 2001; Farley, 2001a/b), but these are prior, conceptually related studies on different languages and features, not replications of this specific study (English past simple tense, Chinese and Greek school-age learners, PI vs. TI vs. MOI three-way design). The author explicitly calls for "further research" to verify these particular findings, which indicates no replication existed at time of writing. The Hong Kong study (Chan, 2019) targets the same grammatical feature but substitutes a different comparison arm (Implicit Instruction rather than MOI) and a markedly different population (young Cantonese primary-school learners vs. Chinese/Greek secondary learners), so it is a conceptually related follow-on study rather than an independent reproduction of this particular study's design and findings, consistent with the ERCT Standard's requirement that the replication reproduce the specific study rather than the general research programme. No other independent replication of this specific study by a different research team in a peer-reviewed outlet was identified. Criterion R is not met because no independent replication of this specific study was found in the text or in subsequent literature.
    • A

      All-subject Exams

      • Since criterion E (Exam-based Assessment) is not met, the stronger All-subject Exams criterion is automatically not met; only English past-tense interpretation and production were assessed.
      • "The interpretation task (see Appendix D) consisted of 20 sentences... of which 10 were in the English past simple tense." (p. 79)
      • Relevant Quotes: 1) "The interpretation task (see Appendix D) consisted of 20 sentences (10 distractor items in the present regular form) of which 10 were in the English past simple tense (only regular forms)." (p. 79) 2) "The written production task (see Appendix E) was developed and used to measure learner's ability to produce correct sentences using the English past tense." (p. 79) Detailed Analysis: Per the criteria-specific instruction, criterion A cannot be met if criterion E is not met, and E was judged not met because the assessments are researcher-designed, non-standardised tasks. In addition, on substance, the study measured only a single, narrow grammatical target (acquisition of the English past simple tense) via a listening interpretation task and a written production task; no other subjects or broader language domains (e.g., reading comprehension, other grammatical structures, or other school subjects such as mathematics or science) were assessed. Criterion A is not met, both because criterion E is not met and because only one narrow linguistic feature was assessed.
    • G

      Graduation Tracking

      • Since criterion Y (Year Duration) is not met, the Graduation Tracking criterion is automatically not met; outcomes were measured only immediately after the three-day intervention with no follow-up, and no subsequent graduation-tracking publication was found.
      • "IMMEDIATE POST-TESTS (Immediately after the end of the instructional period) Interpretation and production tasks" (Figure 1, p. 77)
      • Relevant Quotes: 1) "IMMEDIATE POST-TESTS (Immediately after the end of the instructional period) Interpretation and production tasks" (Figure 1, p. 77) 2) "Despite the outcomes of the present study, long-term effects of the variables under investigation should be re-examined as delayed post-tests were not available (the semester system in the two countries where data were collected did not allow for this). This is a clear limitation in this study..." (p. 85) Internet Search Findings: A search was conducted for later papers by Alessandro Benati that might track the same Chinese and Greek secondary-school cohorts from this 2005 study through to graduation. No such follow-up publication was found; Benati's subsequent papers (e.g., on Italian future tense and gender agreement, French causative replications, and later work with Angelovska and other co-authors) investigate different languages, features and participant samples rather than continuing to track the 47 Chinese and 30 Greek participants from the present study. No quotes from a graduation-tracking follow-up are provided because none could be located. Detailed Analysis: Per the criteria-specific instruction, since criterion Y is not met, criterion G is automatically not met. Substantively, the paper confirms that only an immediate post-test was administered, with the authors explicitly noting the absence of any delayed measurement as a limitation, let alone tracking of participants through to graduation from their educational stage. No follow-up publications tracking this cohort were referenced in the paper or located via internet search. Criterion G is not met because there was no follow-up tracking beyond the immediate post-test, criterion Y was not met, and no subsequent graduation-tracking publication by the author was found.
    • P

      Pre-Registered

      • The paper contains no statement of a pre-registered protocol, registry link, or registration date prior to data collection, and no internet search located any pre-registration record for this 2005 study.
      • Relevant Quotes: 1) No quotes referencing pre-registration, a trial registry, or a published protocol were found anywhere in the article text, method section, notes or references. Internet Search Findings: A search of pre-registration databases and general web sources found no record of a pre-registered protocol for this study. This is consistent with the paper's 2005 publication date, which predates the widespread adoption of pre-registration practice in educational and applied-linguistics RCTs; no pre-registration records for studies of this kind from that period were found in common registries. Detailed Analysis: A full review of the paper, including the Method section, Notes and References, reveals no mention of a pre-registered study protocol, hypotheses, or analysis plan published prior to data collection. This is unsurprising given the paper's 2005 publication date, predating widespread adoption of pre-registration practices in this field, but the criterion still requires explicit evidence of pre-registration, which is absent both from the paper and from internet searches. Criterion P is not met because no pre-registration statement, registry reference, or registration date is present in the paper or found via internet search.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.