The Effects of Perception- vs. Production-based Pronunciation Instruction

Bradford Lee, Luke Plonsky, Kazuya Saito

Published:
ERCT Check Date:
DOI: 10.1016/j.system.2019.102185
  • L2 languages
  • higher education
  • Asia
0
  • C

    Randomisation was done at the individual student level within one university, not at the class or school level, and the intervention was group classroom instruction, not one-to-one tutoring.

    "The participants were randomly divided into five groups: control group (CG; n = 23), syllabic perception instruction (SPe; n = 21), syllabic production instruction (SPr; n = 22), phonemic perception instruction (PPe; n = 24), and phonemic production instruction (PPr; n = 25)." (p. 5)

  • E

    Outcomes were measured with a custom researcher-designed elicitation instrument scored on a 9-point Likert scale by raters, not with any widely recognised standardised exam.

    "A PowerPoint slideshow of 30 slides was designed to test 10 English words three times each: in a free-response style question, a direct translation task from Japanese to English, and finally a read-aloud word list." (pp. 5-6)

  • T

    The entire study spanned only four weeks from pretest to delayed posttest, far short of the one academic term (roughly 3-4 months) required.

    "Although the duration of the study was relatively short (a total of 4 weeks from pretest to delayed posttest)..." (p. 5)

  • D

    The control group is clearly documented, including its size, baseline scores with confidence intervals, and confirmation that it received only normal coursework with no PI.

    "The control group received no treatment for the duration of the study, apart from their normal coursework at the university, and were only asked to convene for the pre- and posttests." (p. 5)

  • S

    Randomisation occurred among individual students at a single university; no schools or institutions were randomised.

    "This study was conducted in an EFL setting at a small university in rural Japan... The participants were randomly divided into five groups..." (p. 5)

  • I

    The lead researcher designed the study, personally delivered all treatments, and produced the sole rating dataset used in the analyses, with no independent third-party conduct or oversight.

    "...only the full dataset produced by the first author (who not only had the most teaching experience, but also conducted the experimental treatments for all groups) was used." (p. 10)

  • Y

    With only four weeks from pretest to delayed posttest, the study falls far short of tracking outcomes for 75% of an academic year; criterion T is also unmet, which entails Y is unmet.

    "Although the duration of the study was relatively short (a total of 4 weeks from pretest to delayed posttest)..." (p. 5)

  • B

    The four treatment groups received identical time and identically structured sessions, and the extra instruction relative to the no-PI control is the explicit treatment variable (provision of PI), with the control serving as a business-as-usual baseline to rule out practice effects.

    "Each experimental group received two treatment sessions of 30 minutes in duration from the lead researcher, with all groups following the same pattern of: explicit instruction lecture (10 minutes), teacher-led activities (10 minutes), pair work with a classmate (5 minutes) and finally worksheet completion (5 minutes)." (p. 7)

  • R

    An independent team (Alghazo, Jarrah & Al Salem, 2023, Frontiers in Education) conducted a peer-reviewed conceptual replication of this study's perception- vs. production-based PI comparison with Jordanian tertiary EFL learners, explicitly citing and using this study's own definitions and reporting a convergent conclusion for delayed-posttest gains.

    "Of particular importance is that by Lee et al. (2020) who examined the effect of the type of PI (perception-based vs. production-based) on pronunciation acquisition among 115 Japanese university students of English." (Alghazo, Jarrah & Al Salem, 2023, Frontiers in Education, 8:1182285, p. 3)

  • A

    Only pronunciation accuracy in English was assessed, using a custom instrument; no other school subjects were measured and criterion E is not met, so A automatically fails.

    "Pronunciation accuracy was therefore defined as the successful production of the Standard American English (SAE) features taught in the treatments." (p. 4)

  • G

    Measurement ended at a delayed posttest two weeks after instruction, with no tracking of participants to graduation and no follow-up publications found; Y is also unmet, which entails G is unmet.

    "The same three-part testing instrument was used for the collection at all three time points: pretest, posttest (immediately after the second treatment), and delayed posttest (exactly two weeks after the second treatment)." (p. 7)

  • P

    The paper contains no mention of any pre-registered protocol or registry entry; sharing materials on the IRIS database is not pre-registration, and no registration record was found externally.

Abstract

While research has shown that provision of explicit pronunciation instruction (PI) is facilitative of various aspects of second language (L2) speech learning (Thomson & Derwing, 2015), a growing number of scholars have begun to examine which type of instruction can best impact on acquisition. In the current study, we explored the effects of perception- vs. production-based methods of PI among tertiary-level Japanese students of English. Participants (N = 115) received two weeks of instruction on either segmental or suprasegmental features of English, using either a perception- or a production-based method, with progress assessed in a pre/post/delayed posttest study design. Although all four treatment groups demonstrated major gains in pronunciation accuracy, performance varied considerably across groups and over time. A close examination of our findings suggested that perception-based training may be the more effective training method across both segmental and suprasegmental features.

Full Article

ERCT Criteria Breakdown

  • Level 1 Criteria

    • C

      Class-level RCT

      • Randomisation was done at the individual student level within one university, not at the class or school level, and the intervention was group classroom instruction, not one-to-one tutoring.
      • "The participants were randomly divided into five groups: control group (CG; n = 23), syllabic perception instruction (SPe; n = 21), syllabic production instruction (SPr; n = 22), phonemic perception instruction (PPe; n = 24), and phonemic production instruction (PPr; n = 25)." (p. 5)
      • Relevant Quotes: 1) "Recruitment efforts resulted in a total of 119 participants. However, four members either chose to drop out or were unable to attend one or more of the training sessions. These individuals' data were excluded from the final analyses, resulting in a final sample of N = 115." (p. 5) 2) "The participants were randomly divided into five groups: control group (CG; n = 23), syllabic perception instruction (SPe; n = 21), syllabic production instruction (SPr; n = 22), phonemic perception instruction (PPe; n = 24), and phonemic production instruction (PPr; n = 25)." (p. 5) 3) "Each experimental group received two treatment sessions of 30 minutes in duration from the lead researcher, with all groups following the same pattern of: explicit instruction lecture (10 minutes), teacher-led activities (10 minutes), pair work with a classmate (5 minutes) and finally worksheet completion (5 minutes)." (p. 7) Detailed Analysis: The unit of randomisation was the individual student: participants recruited at a single small university in rural Japan were "randomly divided into five groups." There is no statement that intact classes or schools were assigned to conditions. The ERCT C criterion requires randomisation of entire classes (or schools) to prevent contamination, unless the intervention is one-to-one tutoring or personal teaching. Here the treatment was delivered as group classroom instruction (lectures, teacher-led activities, pair work, class-wide drills and performances in front of the class), so the tutoring exception does not apply. Students from the same institution were assigned to different conditions, creating exactly the contamination risk the criterion is designed to avoid. Criterion C is not met because randomisation was conducted at the individual student level for a group-based classroom intervention.
    • E

      Exam-based Assessment

      • Outcomes were measured with a custom researcher-designed elicitation instrument scored on a 9-point Likert scale by raters, not with any widely recognised standardised exam.
      • "A PowerPoint slideshow of 30 slides was designed to test 10 English words three times each: in a free-response style question, a direct translation task from Japanese to English, and finally a read-aloud word list." (pp. 5-6)
      • Relevant Quotes: 1) "A PowerPoint slideshow of 30 slides was designed to test 10 English words three times each: in a free-response style question, a direct translation task from Japanese to English, and finally a read-aloud word list." (pp. 5-6) 2) "Prior to the start of the experiment, a total of three piloting sessions were conducted by the lead researcher with the help of eight bilingual L1 Japanese individuals of similar age and ability as the target population." (p. 6) 3) "Each utterance of the target words was rated on a 9-point Likert scale, where only the ends of the scale were defined." (p. 10) 4) "The first author (a bilingual speaker of American English/Japanese, with a background in applied linguistics and over 20 years teaching experience in Japan) was the main expert coder." (p. 10) Detailed Analysis: The outcome measure was a bespoke three-part elicitation instrument (free response, translation, word list) created and piloted by the researchers specifically for this study, with pronunciation accuracy judged subjectively by expert raters on a 9-point Likert scale. This is precisely the situation the E criterion warns against: a custom test aligned to the taught content (the same 10 target words were used in the treatment sessions and the assessment). No national, state-wide, or otherwise widely recognised standardised exam (e.g., TOEIC, TOEFL, Eiken) was used to measure outcomes. Criterion E is not met because the study relied on a custom-made, researcher-designed assessment rather than a standardised exam.
    • T

      Term Duration

      • The entire study spanned only four weeks from pretest to delayed posttest, far short of the one academic term (roughly 3-4 months) required.
      • "Although the duration of the study was relatively short (a total of 4 weeks from pretest to delayed posttest)..." (p. 5)
      • Relevant Quotes: 1) "Participants (N = 115) received two weeks of instruction on either segmental or suprasegmental features of English, using either a perception- or a production-based method, with progress assessed in a pre/post/delayed posttest study design." (p. 1, Abstract) 2) "Although the duration of the study was relatively short (a total of 4 weeks from pretest to delayed posttest), and normal classwork at the university does not consist of any PI, the establishment of the control group was necessary..." (p. 5) 3) "The same three-part testing instrument was used for the collection at all three time points: pretest, posttest (immediately after the second treatment), and delayed posttest (exactly two weeks after the second treatment)." (p. 7) 4) "As the treatments only lasted two weeks, it is likely that the duration was not long enough for learners to fully form phonetic representations and they were still relying heavily on memory." (p. 16) Detailed Analysis: The T criterion requires that outcomes be measured at least one full academic term (approximately 3-4 months) after the intervention begins. Here the intervention consisted of two 30-minute sessions over two weeks, and the final (delayed) posttest was administered exactly two weeks after the second treatment, giving a total interval of only about four weeks from pretest to final measurement. The authors themselves acknowledge the short duration as a limitation. Four weeks is well below the minimum term-long follow-up window. Criterion T is not met because the interval from intervention start to final outcome measurement was only about four weeks, far less than one academic term.
    • D

      Documented Control Group

      • The control group is clearly documented, including its size, baseline scores with confidence intervals, and confirmation that it received only normal coursework with no PI.
      • "The control group received no treatment for the duration of the study, apart from their normal coursework at the university, and were only asked to convene for the pre- and posttests." (p. 5)
      • Relevant Quotes: 1) "The participants were randomly divided into five groups: control group (CG; n = 23)..." (p. 5) 2) "The control group received no treatment for the duration of the study, apart from their normal coursework at the university, and were only asked to convene for the pre- and posttests." (p. 5) 3) "...normal classwork at the university does not consist of any PI, the establishment of the control group was necessary to determine if any gains demonstrated by the experimental groups could be attributed solely to the treatments, or if there were other factors (e.g., test practice effects) that needed to be considered." (p. 5) 4) "CG 3.92 (.34) [3.77, 4.06] 4.05 (.46) [3.86, 4.25] 4.07 (.45) [3.87, 4.26]" (Table 3, p. 12) 5) "This study was conducted in an EFL setting at a small university in rural Japan... the students' oral proficiency level could be described as low-intermediate." (p. 5) Detailed Analysis: The D criterion requires the control group to be documented in terms of composition, baseline performance, and treatment received. The paper reports the control group's size (n = 23), explicitly states what it received (no treatment other than normal university coursework, which contains no pronunciation instruction), and provides its baseline pretest mean, standard deviation, and confidence interval in Table 3, alongside posttest and delayed posttest scores and gain scores (Table 4). The population from which all groups (including controls) were drawn is described (Japanese tertiary EFL students, low-intermediate oral proficiency, aged 18-20 per the plain language summary). Baseline comparability can be checked directly from the reported pretest scores. While demographic breakdown per group is limited, the documentation matches the level accepted under this criterion for a well-described no-treatment baseline group. Criterion D is met because the control group's size, baseline performance, and conditions are clearly documented.
  • Level 2 Criteria

    • S

      School-level RCT

      • Randomisation occurred among individual students at a single university; no schools or institutions were randomised.
      • "This study was conducted in an EFL setting at a small university in rural Japan... The participants were randomly divided into five groups..." (p. 5)
      • Relevant Quotes: 1) "This study was conducted in an EFL setting at a small university in rural Japan, where participants had little chance to use English in communicative contexts outside of the classroom." (p. 5) 2) "The participants were randomly divided into five groups: control group (CG; n = 23), syllabic perception instruction (SPe; n = 21), syllabic production instruction (SPr; n = 22), phonemic perception instruction (PPe; n = 24), and phonemic production instruction (PPr; n = 25)." (p. 5) Detailed Analysis: The S criterion requires randomisation among schools or equivalent implementing institutions. This study took place at one single university, and individual students within that university were randomly allocated to the five conditions. No quote indicates that multiple institutions were involved, let alone randomised. Since randomisation was at the weakest (student) level within a single site, the school-level requirement is clearly unmet. Criterion S is not met because the study randomised individual students within one university rather than randomising schools.
    • I

      Independent Conduct

      • The lead researcher designed the study, personally delivered all treatments, and produced the sole rating dataset used in the analyses, with no independent third-party conduct or oversight.
      • "...only the full dataset produced by the first author (who not only had the most teaching experience, but also conducted the experimental treatments for all groups) was used." (p. 10)
      • Relevant Quotes: 1) "Each experimental group received two treatment sessions of 30 minutes in duration from the lead researcher..." (p. 7) 2) "Prior to the start of the experiment, a total of three piloting sessions were conducted by the lead researcher..." (p. 6) 3) "The first author (a bilingual speaker of American English/Japanese, with a background in applied linguistics and over 20 years teaching experience in Japan) was the main expert coder." (p. 10) 4) "For the remainder of the discussion of the results and statistical analyses, only the full dataset produced by the first author (who not only had the most teaching experience, but also conducted the experimental treatments for all groups) was used." (p. 10) Detailed Analysis: The I criterion requires that the study be conducted independently of those who designed the intervention, or at minimum that data collection and analysis be handled by an external or clearly independent party. Here the same person (the first author) designed the instructional treatments, taught every treatment session, ran the piloting, edited the audio data, and served as the main rater whose ratings alone were used in the statistical analyses. The two additional raters only scored a one-third subset to check reliability, and one of them was still a co-investigator context rater trained by the first author; neither constitutes an external evaluation team. No statement of third-party oversight exists. This is a maximal case of non-independence: designer, implementer, and evaluator are the same individual. Criterion I is not met because the intervention designer also delivered all treatments and produced the outcome ratings used in the analysis, with no independent oversight.
    • Y

      Year Duration

      • With only four weeks from pretest to delayed posttest, the study falls far short of tracking outcomes for 75% of an academic year; criterion T is also unmet, which entails Y is unmet.
      • "Although the duration of the study was relatively short (a total of 4 weeks from pretest to delayed posttest)..." (p. 5)
      • Relevant Quotes: 1) "Participants (N = 115) received two weeks of instruction on either segmental or suprasegmental features of English..." (p. 1, Abstract) 2) "Although the duration of the study was relatively short (a total of 4 weeks from pretest to delayed posttest)..." (p. 5) 3) "The same three-part testing instrument was used for the collection at all three time points: pretest, posttest (immediately after the second treatment), and delayed posttest (exactly two weeks after the second treatment)." (p. 7) Detailed Analysis: The Y criterion requires outcome measurement at least 75% of an academic year (roughly 9-10 months) after the intervention begins. The entire study, from pretest to the final delayed posttest, lasted about four weeks. Additionally, per the ranking rules, because criterion T (Term Duration) is not met, criterion Y cannot be met either. Criterion Y is not met because the tracking period was approximately one month, vastly shorter than an academic year.
    • B

      Balanced Control Group

      • The four treatment groups received identical time and identically structured sessions, and the extra instruction relative to the no-PI control is the explicit treatment variable (provision of PI), with the control serving as a business-as-usual baseline to rule out practice effects.
      • "Each experimental group received two treatment sessions of 30 minutes in duration from the lead researcher, with all groups following the same pattern of: explicit instruction lecture (10 minutes), teacher-led activities (10 minutes), pair work with a classmate (5 minutes) and finally worksheet completion (5 minutes)." (p. 7)
      • Relevant Quotes: 1) "Each experimental group received two treatment sessions of 30 minutes in duration from the lead researcher, with all groups following the same pattern of: explicit instruction lecture (10 minutes), teacher-led activities (10 minutes), pair work with a classmate (5 minutes) and finally worksheet completion (5 minutes)." (p. 7) 2) "The control group received no treatment for the duration of the study, apart from their normal coursework at the university, and were only asked to convene for the pre- and posttests." (p. 5) 3) "...the establishment of the control group was necessary to determine if any gains demonstrated by the experimental groups could be attributed solely to the treatments, or if there were other factors (e.g., test practice effects) that needed to be considered." (p. 5) 4) "This study will test and compare the effects of such explicit PI using either perception- or production-based FFI." (p. 4) Detailed Analysis: Applying the criterion B decision tree: extra resources are present (the four treatment groups received 60 minutes of instruction the control group did not), so the analysis turns on whether that extra instruction is itself the treatment variable being tested. Two comparisons matter here. First, the study's primary research question compares the four treatment conditions against each other (perception vs. production, syllabic vs. phonemic). These four groups received exactly matched inputs: two 30-minute sessions each, delivered by the same instructor, with an identical internal structure (10-min lecture, 10-min teacher-led activities, 5-min pair work, 5-min worksheet). Time, materials, and instructor attention were therefore balanced across the conditions that drive the main findings. Second, relative to the no-PI control group, the treatment groups did receive an additional 60 minutes of instruction. Applying the decision tree: the extra instruction IS the intervention being tested - the study explicitly tests the effects of providing explicit pronunciation instruction (of two types), and normal classwork "does not consist of any PI," so the control represents the intended business-as-usual baseline whose stated purpose was to quantify test practice effects. The additional instructional time is integral to the treatment package rather than a separable confounding add-on, and the head-to-head contrasts of interest are between equally resourced active conditions. Criterion B is met because the active conditions were time-and-structure matched, and the extra 60 minutes of instruction relative to the business-as-usual control is the explicit treatment variable (provision of explicit PI) rather than a non-integral add-on.
  • Level 3 Criteria

    • R

      Reproduced

      • An independent team (Alghazo, Jarrah & Al Salem, 2023, Frontiers in Education) conducted a peer-reviewed conceptual replication of this study's perception- vs. production-based PI comparison with Jordanian tertiary EFL learners, explicitly citing and using this study's own definitions and reporting a convergent conclusion for delayed-posttest gains.
      • "Of particular importance is that by Lee et al. (2020) who examined the effect of the type of PI (perception-based vs. production-based) on pronunciation acquisition among 115 Japanese university students of English." (Alghazo, Jarrah & Al Salem, 2023, Frontiers in Education, 8:1182285, p. 3)
      • Relevant Quotes: 1) From the original paper: "Further research is therefore urgently needed which targets other segmental phonemes, suprasegmental features, L1 populations, and even other L2 target languages to establish the generalizability of our findings." (p. 16) - no replication is reported within the paper itself. 2) From Alghazo, Jarrah & Al Salem (2023), "The efficacy of the type of instruction on second language pronunciation acquisition," Frontiers in Education, 8:1182285 (verified directly against the published PDF): "Of particular importance is that by Lee et al. (2020) who examined the effect of the type of PI (perception-based vs. production-based) on pronunciation acquisition among 115 Japanese university students of English. PI lasted for 2 weeks, and improvement was assessed using pre-, post-, and delayed post-tests." (p. 3) 3) "In addition, we follow Lee et al. (2020) definitions of perception-based and production-based instruction which show that perception-based instruction aims 'at increasing the participants' identification or discrimination abilities' while production-based instruction aims at 'eliciting the correct articulation of the target features while making use of corrective feedback' (p. 1)." (p. 2) 4) "This result is consistent with Lee et al. (2020) that 'perception-based instruction would lead to greater improvement than production-based ones' (p. 8)." (p. 8) 5) Study design (verified from the Methods section of the 2023 paper): 64 tertiary-level Jordanian Arabic-speaking learners of English were recruited (60 retained after 4 dropouts), randomly divided into a perception-based group and a production-based group (no untreated control), each receiving 6 weeks (15 hours total) of instruction on segmental, syllabic, prosodic, global, and temporal aspects of pronunciation, with pre-, post- (week 6), and delayed post-tests (week 14) rated by an independent expert rater on a 10-point Likert scale. Perception-based instruction produced larger and more durable gains in segmental, syllabic, and prosodic accuracy, matching the direction of the original study's key finding. Detailed Analysis: The R criterion requires independent replication of the study's central experimental claim by a different team, in a different context, published in a peer-reviewed journal. The original paper (System, 2020) reports no replication, but an internet search identified Alghazo, Jarrah and Al Salem (2023), authored by researchers at the University of Jordan and University of Sharjah - a team fully independent of Lee, Plonsky, and Saito, with no overlapping authorship or stated affiliation. Their study explicitly adopts Lee et al.'s (2020) own operational definitions of perception- vs. production-based instruction, cites it as the key prior study on the effect of instruction type, and reproduces its central design contrast (perception- vs. production-based pronunciation instruction on segmental and suprasegmental targets, assessed via pre/post/ delayed-posttest design with tertiary EFL learners) in a different national and L1 context (Jordan, L1 Arabic). The 2023 study is not an exact procedural replication - it uses a longer 6-week treatment, drops the untreated control group in favour of comparing two active treatments directly, adds global and temporal outcome measures, and uses a single rater on a 10-point scale rather than three raters on a 9-point scale - but it explicitly and verbatim states its result is "consistent with Lee et al. (2020)" regarding the relative advantage of perception-based instruction, which is the original study's central claim. Frontiers in Education is a peer-reviewed journal. Taken together, this constitutes a credible independent conceptual replication of the study's central claim in a different context, notwithstanding the design differences noted above. Criterion R is met because an independent research team replicated the perception- vs. production-based PI comparison in a different context, using the original study's own definitions, and published convergent findings in a peer-reviewed journal.
    • A

      All-subject Exams

      • Only pronunciation accuracy in English was assessed, using a custom instrument; no other school subjects were measured and criterion E is not met, so A automatically fails.
      • "Pronunciation accuracy was therefore defined as the successful production of the Standard American English (SAE) features taught in the treatments." (p. 4)
      • Relevant Quotes: 1) "Pronunciation accuracy was therefore defined as the successful production of the Standard American English (SAE) features taught in the treatments." (p. 4) 2) "A PowerPoint slideshow of 30 slides was designed to test 10 English words three times each..." (pp. 5-6) 3) "Finally, the assessed segmental range of this study only included 10 phonemes... On the suprasegmental side, features included syllables, stress, and the lax syllable." (p. 16) Detailed Analysis: The A criterion requires standardised exam-based assessment across all main subjects taught at the educational level, and criterion E is a prerequisite. Criterion E is not met (the outcome measure was a custom instrument), so A automatically fails. Moreover, the study measured only one narrow outcome - English pronunciation accuracy on 10 target words - and no other university subjects were assessed. While a specialised focus can sometimes justify an exception in upper secondary/vocational contexts, the absence of any standardised assessment means the criterion cannot be satisfied regardless. Criterion A is not met because criterion E fails and only a single narrow outcome (pronunciation of 10 words) was measured.
    • G

      Graduation Tracking

      • Measurement ended at a delayed posttest two weeks after instruction, with no tracking of participants to graduation and no follow-up publications found; Y is also unmet, which entails G is unmet.
      • "The same three-part testing instrument was used for the collection at all three time points: pretest, posttest (immediately after the second treatment), and delayed posttest (exactly two weeks after the second treatment)." (p. 7)
      • Relevant Quotes: 1) "The same three-part testing instrument was used for the collection at all three time points: pretest, posttest (immediately after the second treatment), and delayed posttest (exactly two weeks after the second treatment)." (p. 7) 2) "Although the duration of the study was relatively short (a total of 4 weeks from pretest to delayed posttest)..." (p. 5) 3) "Future studies may wish to utilize a longer treatment duration." (p. 16) Detailed Analysis: The G criterion requires participants to be tracked until graduation from their educational stage. Data collection ended with a delayed posttest just two weeks after the final treatment session; the university students were not followed through to the end of their degree programmes. The paper proposes future research directions but mentions no planned or published follow-up of this cohort. A dedicated internet search for subsequent publications by Bradford Lee, Luke Plonsky, or Kazuya Saito tracking this same 115-student cohort (or any graduation-tracking follow-up to this study) found no such paper. Additionally, per the ranking rules, criterion Y is not met, which means G cannot be met regardless. Criterion G is not met because tracking ended two weeks after the intervention with no graduation follow-up, and no follow-up publication tracking this cohort could be found.
    • P

      Pre-Registered

      • The paper contains no mention of any pre-registered protocol or registry entry; sharing materials on the IRIS database is not pre-registration, and no registration record was found externally.
      • Relevant Quotes: 1) "The full instrument will be made available upon request and on the IRIS Database upon publication (see Marsden, Mackey, & Plonsky, 2016)." (p. 6, footnote 1) Detailed Analysis: The P criterion requires the full study protocol (hypotheses, methods, planned analyses) to be registered on a public registry before data collection began. The paper contains no reference to any registry (e.g., ClinicalTrials.gov, OSF, AsPredicted, AEA registry), no registration ID, and no registration date. The only openness statement concerns sharing the test instrument via the IRIS materials repository after publication, which is materials sharing, not protocol pre-registration. A dedicated internet search for a pre-registration record associated with this study (by title, authors, and DOI) found no OSF, AsPredicted, or other registry entry. Criterion P is not met because no pre-registration of the study protocol is reported or discoverable.

Request an Update or Contact Us

Are you the author of this study? Let us know if you have any questions or updates.

Have Questions
or Suggestions?

Get in Touch

Have a study you'd like to submit for ERCT evaluation? Found something that could be improved? If you're an author and need to update or correct information about your study, let us know.

  • Submit a Study for Evaluation

    Share your research with us for review

  • Suggest Improvements

    Provide feedback to help us make things better.

  • Update Your Study

    If you're the author, let us know about necessary updates or corrections.