EconBase
← All papers

Critical Values Robust to P-hacking

Adam McCloskey, Pascal Michaillat

arXiv 8 May 2020 · Econometrics · publishedThe Review of Economics and Statistics (2024) · 5 citations (OpenAlex)

arXiv:2005.04141 · PDF · DOI · OpenAlex · Extracted main text

Abstract

P-hacking is prevalent in reality but absent from classical hypothesis testing theory. As a consequence, significant results are much more common than they are supposed to be when the null hypothesis is in fact true. In this paper, we build a model of hypothesis testing with p-hacking. From the model, we construct critical values such that, if the values are used to determine significance, and if scientists' p-hacking behavior adjusts to the new significance standards, significant results occur with the desired frequency. Such robust critical values allow for p-hacking so they are larger than classical critical values. To illustrate the amount of correction that p-hacking might require, we calibrate the model using evidence from the medical sciences. In the calibrated model the robust critical value for any test statistic is the classical critical value for the same test statistic with one fifth of the significance level.

Citation extraction

78
references
126
in-text mentions
78
distinct cited
0
self-citations
8,106
main-text words

appendix boundary found by appendix_command · 49% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Dwan, Altman, Arnaiz, Bloom, Chan, Cronin, Decullier, Easterbrook, V… (2008) Systematic Review of the Empirical Evidence of Study Publication Bias and Outcome Reporting Bias0.98320395%
2Ferguson (2007) Optimal Stopping and Applications0.92843100%
3Lovell (1983) Data Mining0.7373367%
4Anscombe (1954) Fixed-Sample-Size Analysis of Sequential Observations0.73732100%
5Chen (2021) The Limits of P-hacking: Some Thought Experiments0.64422100%
6Christensen and Miguel (2018) Transparency, Reproducibility, and the Credibility of Economics Research0.64422100%
7Glaeser (2008) Researcher Incentives and Empirical Methods0.64422100%
8Christensen, Freese, and Miguel (2019) Transparent and Reproducible Social Science Research: How to Do Open Science0.5112250%
9Benjamin, Berger, Johannesson, Nosek, Wagenmakers, Berk, Bollen, Bre… (2018) Redefine Statistical Significance0.51121100%
10Cooper, DeNeve, and Charlton (1997) Finding The Missing Science: The Fate of Studies Submitted for Review By a Human Subjects Committee0.51121100%

Showing the top 10 of 78 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1The Power of Tests for Detecting $p$-Hacking0.64422
22412.164520.40511
3Publication Design with Incentives in Mind0.40511
42510.211780.40511