arXiv 19 Jun 2024 · Econometrics
arXiv:2406.13122 · PDF · DOI · OpenAlex · Extracted main text
How many experimental studies would have come to different conclusions had they been run on larger samples? I show how to estimate the expected number of statistically significant results that a set of experiments would have reported had their sample sizes all been counterfactually increased. The proposed deconvolution estimator is asymptotically normal and adjusts for publication bias. Unlike related methods, this approach requires no assumptions of any kind about the distribution of true intervention treatment effects and allows for point masses. Simulations find good coverage even when the t-score is only approximately normal. An application to randomized trials (RCTs) published in economics journals finds that doubling every sample would increase the power of t-tests by 7.2 percentage points on average. This effect is smaller than for non-RCTs and comparable to systematic replications in laboratory psychology where previous studies enabled more accurate power calculations. This suggests that RCTs are on average relatively insensitive to sample size increases. Research funders who wish to raise power should generally consider sponsoring better-measured and higher quality experiments -- rather than only larger ones.
appendix boundary found by appendix_command · 72% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Andrews, I. and M. Kasy (2019, August) (2019) Identification of and correction for publication bias | 1.000 | 7 | 4 | 100% |
| 2 | Elliott, G., N. Kudrin, and K. Wüthrich (2022) Detecting p-hacking | 1.000 | 6 | 4 | 100% |
| 3 | Bartoš, F. and U. Schimmack (2022, 09) (2022) Z-curve 2.0: Estimating replication rates and discovery rates | 1.000 | 5 | 3 | 100% |
| 4 | Brodeur, A., N. Cook, and A. Heyes (2020, November) (2020) Methods matter: p-hacking and publication bias in causal analysis in economics | 0.979 | 16 | 6 | 94% |
| 5 | Carrasco, M. and J.-P. Florens (2011) A spectral method for deconvolving a density | 0.956 | 8 | 7 | 88% |
| 6 | Ioannidis, J. P. A., T. D. Stanley, and H. Doucouliagos (2017, Octob… (2017) The power of bias in economics research | 0.941 | 6 | 4 | 83% |
| 7 | Fan, J (1991) On the Optimal Rates of Convergence for Nonparametric Deconvolution Problems | 0.928 | 4 | 3 | 100% |
| 8 | McKenzie, D (2025) Designing and analysing powerful experiments: practical tips for applied researchers | 0.928 | 4 | 3 | 100% |
| 9 | Klein, R. A., K. A. Ratliff, M. Vianello, R. B. Adams, v. Bahnḱ, M.… (2014) Investigating variation in replicability | 0.843 | 5 | 3 | 60% |
| 10 | Brodeur, A., M. Lé, M. Sangnier, and Y. Zylberberg (2016, January) (2016) Star wars: The empirics strike back | 0.843 | 3 | 3 | 100% |
Showing the top 10 of 40 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | When is $p$-hacking detectable? | 0.405 | 1 | 1 |