arXiv 24 Aug 2026 · Econometrics
arXiv:2608.23257 · PDF · Extracted main text
In the empirical sciences, significance thresholds often determine whether findings are treated as evidence of an effect. This paper studies how likely findings that just meet conventional significance thresholds are to remain significant in replications of the same sample size. To answer this question, we estimate the expected replication probability conditional on a given p-value among published studies for experimental economics, psychology, and social science. We validate this measure by showing it accurately predicts actual replication outcomes, outperforming prediction markets. A finding with a p-value of 0.05 has an expected replication probability ranging from 0.10 to 0.25 across fields. Low replicability reflects low power in original studies rather than publication bias. We then develop a nonparametric estimator and apply it to economics literatures that use larger samples, finding higher but still low replication probabilities. These results indicate that statistical significance in a single study provides only suggestive evidence of an effect. Stronger conclusions require cumulative evidence.
appendix boundary found by appendix_command · 59% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Brodeur, Abel and Cook, Nikolai and Heyes, Anthony (2020) Methods Matter: p-Hacking and Publication Bias in Causal Analysis in Economics | 1.000 | 7 | 3 | 100% |
| 2 | Andrews, Isaiah and Kasy, Maximilian (2019) Identification of and Correction for Publication Bias | 0.971 | 12 | 4 | 92% |
| 3 | Patrick Vu (2024) Why are replication rates so low? self | 0.950 | 7 | 4 | 86% |
| 4 | Camerer, Colin F. and Dreber, Anna and Forsell, Eskil and others (2016) Evaluating Replicability of Laboratory Experiments in Economics | 0.894 | 14 | 7 | 71% |
| 5 | Camerer, Colin F. and Dreber, Anna and Holzmeister, Felix and others (2018) Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015 | 0.894 | 14 | 7 | 71% |
| 6 | Open Science Collaboration (2015) Estimating the Reproducibility of Psychological Science | 0.894 | 14 | 7 | 71% |
| 7 | Stefan Faridani (2026) Testing for underpowered literatures self | 0.888 | 10 | 4 | 70% |
| 8 | Marine Carrasco and Jean-Pierre Florens (2011) A SPECTRAL METHOD FOR DECONVOLVING A DENSITY | 0.811 | 5 | 2 | 80% |
| 9 | Brodeur, Abel and Dreber, Anna and Hoces de la Guardia, Fernando and… (2023) Replication Games: How to Make Reproducibility Research More Systematic | 0.737 | 3 | 2 | 100% |
| 10 | Coffman, Lucas C. and Niederle, Muriel (2015) Pre-Analysis Plans Have Limited Upside, Especially Where Replications Are Feasible | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 44 scored citations.