Adam McCloskey, Pascal Michaillat
arXiv 8 May 2020 · Econometrics · publishedThe Review of Economics and Statistics (2024) · 5 citations (OpenAlex)
arXiv:2005.04141 · PDF · DOI · OpenAlex · Extracted main text
P-hacking is prevalent in reality but absent from classical hypothesis testing theory. As a consequence, significant results are much more common than they are supposed to be when the null hypothesis is in fact true. In this paper, we build a model of hypothesis testing with p-hacking. From the model, we construct critical values such that, if the values are used to determine significance, and if scientists' p-hacking behavior adjusts to the new significance standards, significant results occur with the desired frequency. Such robust critical values allow for p-hacking so they are larger than classical critical values. To illustrate the amount of correction that p-hacking might require, we calibrate the model using evidence from the medical sciences. In the calibrated model the robust critical value for any test statistic is the classical critical value for the same test statistic with one fifth of the significance level.
appendix boundary found by appendix_command · 49% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Dwan, Altman, Arnaiz, Bloom, Chan, Cronin, Decullier, Easterbrook, V… (2008) Systematic Review of the Empirical Evidence of Study Publication Bias and Outcome Reporting Bias | 0.983 | 20 | 3 | 95% |
| 2 | Ferguson (2007) Optimal Stopping and Applications | 0.928 | 4 | 3 | 100% |
| 3 | Lovell (1983) Data Mining | 0.737 | 3 | 3 | 67% |
| 4 | Anscombe (1954) Fixed-Sample-Size Analysis of Sequential Observations | 0.737 | 3 | 2 | 100% |
| 5 | Chen (2021) The Limits of P-hacking: Some Thought Experiments | 0.644 | 2 | 2 | 100% |
| 6 | Christensen and Miguel (2018) Transparency, Reproducibility, and the Credibility of Economics Research | 0.644 | 2 | 2 | 100% |
| 7 | Glaeser (2008) Researcher Incentives and Empirical Methods | 0.644 | 2 | 2 | 100% |
| 8 | Christensen, Freese, and Miguel (2019) Transparent and Reproducible Social Science Research: How to Do Open Science | 0.511 | 2 | 2 | 50% |
| 9 | Benjamin, Berger, Johannesson, Nosek, Wagenmakers, Berk, Bollen, Bre… (2018) Redefine Statistical Significance | 0.511 | 2 | 1 | 100% |
| 10 | Cooper, DeNeve, and Charlton (1997) Finding The Missing Science: The Fate of Studies Submitted for Review By a Human Subjects Committee | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 78 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | The Power of Tests for Detecting $p$-Hacking | 0.644 | 2 | 2 |
| 2 | 2412.16452 | 0.405 | 1 | 1 |
| 3 | Publication Design with Incentives in Mind | 0.405 | 1 | 1 |
| 4 | 2510.21178 | 0.405 | 1 | 1 |