EconBase
← Back to paper

How Replicable Are Statistically Significant Findings?

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

77,667 characters

How Replicable Are Statistically Significant Findings?



\makeatletter
\patchcmd{\@maketitle}{\LARGE \@title}{\fontsize{12}{12}\selectfont\@title}{}{}
\makeatother



\font\myfont=cmr12 at 16pt
\title{{\LARGE How Replicable Are Statistically Significant Findings?\thanks{
We thank Jonathan Roth and seminar participants at the University of Queensland and the University of Western Australia for helpful comments and suggestions.
}}}

\author[1]{\large Patrick Vu\thanks{Department of Economics, University of New South Wales. Email: \href{[[email removed]]}{[email removed]}} and
Stefan Faridani\thanks{Georgia Institute of Technology. Email: \href{[[email removed]]}{[email removed]}}.}

\date{\today}

\setcounter{footnote}{0}
\maketitle
\thispagestyle{empty}
\setcounter{footnote}{0}


\maketitle
\begin{abstract}
In the empirical sciences, significance thresholds often determine whether findings are treated as evidence of an effect. This paper studies how likely findings that just meet conventional significance thresholds are to remain significant in replications of the same sample size. To answer this question, we estimate the expected replication probability conditional on a given $p$-value among published studies for experimental economics, psychology, and social science. We validate this measure by showing it accurately predicts actual replication outcomes, outperforming prediction markets. A finding with a $p$-value of 0.05 has an expected replication probability ranging from 0.10 to 0.25 across fields. Low replicability reflects low power in original studies rather than publication bias. We then develop a nonparametric estimator and apply it to economics literatures that use larger samples, finding higher but still low replication probabilities. These results indicate that statistical significance in a single study provides only suggestive evidence of an effect. Stronger conclusions require cumulative evidence.
\end{abstract}


\newpage
\setcounter{page}{1}

\section{Introduction}
In the empirical sciences, conventional significance thresholds are widely used to establish evidence of an effect. In some cases, a statistically significant empirical claim is repeated using the same methodology and a new sample \citep{OpenScience2015, Camerer2016, Camerer2018}. A significant replication is deemed `successful' and strengthens the original claim, while an insignificant replication is deemed `unsuccessful' and weakens it.



Given the ubiquity of significance testing in applied research, this paper asks a simple but fundamental question: ``How likely are findings that just meet conventional significance thresholds to remain significant in replications of the same sample size?'' If significant findings are unlikely to remain significant in replication, then the initial result may provide weaker grounds for claiming a nonzero effect than is commonly supposed. This concern differs from critiques that significance testing obscures economic significance or encourages $p$-hacking. Instead, in this paper, we evaluate significance testing on its own terms.


To answer this question, we derive the predictive power curve in the \citet{Andrews2019} framework of selective publication. Predictive power is the probability of replication for a published finding conditional on a given $p$-value, where replication success is defined using the standard criterion of significance at the 5\% level in the same direction as the original estimate \citep{OpenScience2015, Camerer2016, Camerer2018}. Within a given empirical literature, it can be interpreted as the expected replication probability of a randomly chosen published finding with that $p$-value.\footnote{A $p$-value below 5\% does not, by itself, imply a high probability of replication at the 5\% level. A significant finding may reflect a large true effect, but it may also reflect a favorable sample realization from a small or null effect, in which case it is unlikely to replicate. More generally, for any $p$-value in $(0,1)$, the replication probability can in principle range from 0.025 to arbitrarily close to 1.} We define predictive power under the original design by considering a replication with the same sample size as the original study.\footnote{In practice, replication studies often use larger sample sizes than the original study, which tends to overstate replication rates relative to using the same sample size. Moreover, common power calculations that choose the replication sample size to detect the original estimate with a prespecified power yield expected replication power strictly below that target \citep{Vu2024}.} Importantly, the framework does not require that an actual replication is conducted, or is even feasible. Predictive power is closely related to statistical power and can be applied equally to observational studies, as shown in Section~\ref{section:nonparametric}.






Predictive power has two noteworthy properties. First, as the conditional mean of the replication outcome, it minimizes expected squared prediction error. This provides strong theoretical grounds for using it to estimate replicability at conventional significance thresholds. Second, predictive power is invariant to publication bias. This counterintuitive property arises because publication bias changes the distribution of observed findings \textit{across} $p$-values, but not the distribution of true effects \textit{within} a given published $p$-value, on which predictive power is conditioned. Hence, low predictive power at a given $p$-value cannot be attributed to publication bias.



We estimate the predictive power curve in three empirical literatures with systematic replication evidence: experimental economics \citep{Camerer2016}, psychology \citep{OpenScience2015}, and experimental social science \citep{Camerer2018}. Predictive power is completely determined by the distribution of true effects. We draw directly on estimates from \citet{Vu2024}, whose empirical applications use the `metastudy approach' from \citet{Andrews2019} to estimate this distribution under parametric assumptions.

The advantage of examining literatures with systematic replication evidence is that the estimated predictive power curves can be validated against realized replication outcomes. Predictive power forecasts actual replication outcomes with remarkable accuracy: we cannot reject perfect calibration, and it achieves lower average squared prediction error than prediction markets, despite using far less information. High predictive accuracy provides strong empirical justification for using the measure to evaluate replicability, and also suggests a simple, low-cost alternative to prediction markets for forecasting replication outcomes.


We next turn to the paper's main question: how replicable are findings that narrowly satisfy conventional significance thresholds when replicated using the original sample size? In all three applications, and across all standard conventional significance thresholds, the estimated predictive power curves point to the same conclusion: replication probabilities are generally quite low. For example, in experimental economics---which fares better than psychology and social science---a finding with a $p$-value of 0.05 has a one-in-four chance of being successfully replicated. In other words, a same-sized replication of a just-significant result is three times as likely to be insignificant as significant. At $p=0.01$, the expected replication probability remains relatively low at 0.40. Notably, moving from ``marginal significance'' at the 10\% level to statistical significance at the 5\% level increases expected replication probabilities from 0.19 to 0.25, implying little change in replicability. More generally, achieving high replication probabilities requires far stronger statistical evidence than demanded by conventional thresholds. In experimental economics, achieving a replication probability of at least 0.8 requires $p<0.00006$, a standard met by 16.7\% of findings in the sample. We interpret this as a diagnostic benchmark rather than a recommendation to impose stricter thresholds.

To broaden the applicability of our framework, we develop a nonparametric estimator of predictive power. Its principal advantages are that it imposes no parametric assumptions on the distribution of true effects and can still be used when each study reports multiple correlated $p$-values. Its main drawback is that it cannot be applied to the replication studies due to the small number of observations. Instead, we apply it to a large dataset of experimental and quasi-experimental findings in economics \citep{Brodeur2020}. For randomized control trials, the chance of replicating a finding with $p=0.05$ is 0.34. Quasi-experimental studies using observational data have notably higher replication probabilities at $p=0.05$, ranging from 0.43 to 0.50, although replicability remains low in absolute terms. Across all applications, differences in predictive power can largely be explained by differences in median sample size. Taken together, the parametric and nonparametric estimates reinforce the same conclusion: replicability at conventional significance thresholds is relatively low.





How should statistically significant findings be interpreted in light of these results? Conventional thresholds play a central role in determining whether a finding---often based on a single study---is viewed as evidence of an effect. Yet we find that an empirical result in experimental economics with $p=0.05$ is expected to replicate with probability $\frac{1}{4}$. This does not imply that three-quarters of just-significant findings are false positives. Nor do unsuccessful replications necessarily imply errors in original studies. Rather, low replicability reflects the winner's curse \citep{WinnersCurseQJE}: even if estimates are unbiased before conditioning on their $p$-values, the majority of just-significant estimates overstate their true effects because statistical significance frequently reflects favorable sampling variation around a smaller true effect.

To clarify what these results imply for evidence against the null, Subsection~\ref{subsection:interpreting_replication_outcomes} examines how a successful or unsuccessful replication updates the probability that the true effect is close to zero, beginning with an original finding with $p=0.05$. For experimental economics, we find that a successful replication can substantially strengthen evidence against a near-null effect, reducing its probability from 0.16 to 0.02. By contrast, an unsuccessful replication provides only limited evidence in its favor, increasing the probability from 0.16 to 0.21. This asymmetry arises because a successful replication is highly unlikely when the true effect is near zero, whereas an unsuccessful replication remains relatively likely even when the effect is not close to zero.


Overall, we recommend interpreting statistical significance in a single study as suggestive evidence against the null. Robust scientific conclusions require cumulative evidence across multiple studies. These findings support the longstanding view that replication is fundamental to scientific credibility \citep{Popper1934}, and reinforce more recent calls to strengthen incentives for conducting and publishing replications \citep{Nosek2012,Coffman2015,Brodeur2023}.

\bigskip
\textbf{Related Literature.} The paper contributes to two strands of the metascience literature. First, it contributes to research on replications and statistical power \citep{Cohen1988, Ioannidis2005, Loken2017, Stanley2017, Andrews2019, DellaVigna2022}. It is most closely related to two recent papers. First, \citet{Goodman2022} estimates predictive power in a parametric model without publication bias for medical research and finds low replicability. Second, \citet{Lang2025} estimates 65\% of narrowly rejected null hypotheses in economics are likely to be false rejections. This paper builds on the literature in several respects. First, to our knowledge, it is the first study to validate predictive power against actual replication outcomes. Second, it models publication bias, for which there is substantial empirical evidence across scientific fields \citep{Card1995, Franco2014, Andrews2019}. Finally, it develops a nonparametric estimator of predictive power under weak assumptions.

Second, the paper contributes to the literature on predicting replications \citep{Dreber2015, Altmejd2019, DellaVigna2019, DellaVigna2020, Gordon2021}. Existing approaches use surveys, prediction markets, and machine learning methods to forecast replication outcomes. Our simple model-based forecast uses $p$-values from original studies yet nevertheless outperforms prediction markets. Thus, it offers a simple, low-cost alternative to more resource-intensive forecasting methods.




The remainder of the article is organized as follows. Section \ref{section:theory} derives predictive power. Section \ref{section:empirical} discusses estimation and validation. Section \ref{section:results} presents the parametric results and Section \ref{section:nonparametric} the nonparametric results. Section \ref{section:recommendations} makes recommendations for researchers and replicators. Section~\ref{section:conclusion} concludes.



\section{Theory}\label{section:theory}
This section introduces the framework, derives the predictive power curve, and discusses its properties. Proofs are in Appendix \ref{appendix:proofs}.

\subsection{Setup}\label{subsection:setup}
Consider an empirical literature composed of research findings indexed by $i$. For instance, the literature of interest could be based on field (e.g. labor economics, electoral politics, epidemiology) or methodology (e.g. experimental, quasi-experimental). Research findings are characterized by an (unobserved) true effect, an estimate of the true effect, and a standard error, $(\beta_i, \hat{\beta}_i, \sigma_i)$.\footnote{Following \cite{Andrews2019} and \cite{Wuthrich2022}, we assume for simplicity that the researcher observes the true standard error. Under standard regularity conditions, estimation error in the standard error is typically much smaller than sampling variation in the effect estimate.} We focus on normalized studies, characterized by $(z_i, \hat{z}_i)$, where $z \equiv \beta/\sigma$, $\hat z \equiv \hat \beta/\sigma$. For expositional simplicity, we refer to $z$ as the \textit{true effect} and $\hat z$ as the \textit{estimated effect}.

The data-generating process follows the selective publication model of \citet{Andrews2019} presented in normalized units:

\begin{enumerate}
    \item \textbf{Draw true effect:} Draw $z$ from the distribution of true effects: $z \sim \pi$.
    \item \textbf{Estimate the effect:} Draw an estimate $\hat z$ from a normal distribution centered at the true effect $z$: $\hat z \mid z \sim N(z, 1)$.
    \item \textbf{Publication selection:} Selective publication is modeled by a function $s(\cdot)$, which gives the probability of publication for an estimate $\hat z$. Let $D$ be a Bernoulli random variable equal to 1 if a study is published and 0 otherwise. Then $\mathbbm{P}(D = 1 | \hat z) =   s\left( \hat z \right)$.
\end{enumerate}

We observe $|\hat z|$ for published findings, $D=1$. Focusing on the absolute value reflects common empirical settings in which only the magnitude of the test statistic -- or, equivalently, the two-sided $p$-value -- is reported.

To illustrate the model, consider the quasi-experimental labor economics literature. Papers in this literature address different questions, such as the impact of the minimum wage on employment or the effect of a job-training program on earnings. Each question has an associated true effect $z$, drawn from the distribution $\pi(z)$. For the job-training example, $z$ represents the program's true average treatment effect, normalized by the study's standard error. In the second step, a researcher applies an empirical method -- such as a difference-in-differences design -- to obtain the estimate $\hat z$. Under mild regularity conditions, the resulting estimator is approximately normally distributed with a consistently estimable variance, a standard assumption used in empirical practice. In the final step, the estimate $\hat z$ is published with probability $s(\hat z)$. This captures selective publication; for instance, statistically significant findings may be more likely to be published than null results.

Finally, note that while $\pi(z)$ can be interpreted as a `prior' in a Bayesian hierarchical model, it is not subjective. Instead, it should be interpreted as the \textit{objective} (unobserved) distribution of true effects across studies in the empirical literature of interest.


\subsection{Replication Under Original Design}
\label{subsection:replications_statisticalpower}

Our main objective is to estimate the probability that a finding with a given $p$-value successfully replicates. Formally, consider a replication of an original finding $(z,\hat z)$. To evaluate replicability under the original design, we consider a replication that uses the same sample size as the original study. Thus, the replication estimate satisfies $\hat z_r \mid z \sim N(z,1)$ and is conditionally independent of $\hat z$ given $z$.


Replication success is defined using the standard criterion from large-scale replication studies: statistical significance at the two-sided 5\% level with the same sign as the original estimate \citep{OpenScience2015, Camerer2016, Camerer2018}. Specifically, let the replication indicator be $R =
\mathbbm{1}
\left\{
|\hat z_r| \geq 1.96
\text{ and }
\operatorname{sign}(\hat z_r)
=
\operatorname{sign}(\hat z)
\right\}$, where $1.96$ is the critical value for statistical significance at the two-sided 5\% level. This measure of replication success is closely related to statistical power.\footnote{Conventional one-sided power is the probability that the replication estimate is significant in the direction of the true effect, whereas our measure uses the direction of the original estimate. The two coincide unless the original estimate is statistically significant with the wrong sign, and are therefore likely to be similar when such sign errors are rare.} As with statistical power, it is also well-defined for observational studies for which a direct replication may be infeasible.

We focus on replications using the same sample size as original studies. In practice, replication studies use various sample-size rules, most of which produce larger samples and therefore higher replication rates.\footnote{Appendix \ref{appendix:alternative_rep_sample_size_rules} estimates that the rules used in the experimental economics \citep{Camerer2016}, psychology \citep{OpenScience2015}, and experimental social science \citep{Camerer2018} replication projects raise estimated replication rates by 30--36\% relative to the same-sample-size benchmark. More generally, replication rates depend mechanically on the chosen sample-size rule: for any non-zero effect, statistical power approaches one as sample size increases without bound and approaches the size of the test as sample size approaches zero.} Our benchmark instead evaluates replicability under the original research design, which provides a common basis for comparison across literatures that is independent of replicators' sample-size choices.

\subsection{Predictive Power Curve}\label{subsection:replication_power_curve}
Define \textit{predictive power} as the probability of a successful replication conditional on the observed two-sided $p$-value and publication: $r(p)\equiv \Pr(R=1\mid p,\text{published})$. Tracing $r(p)$ over $p \in (0,1)$ yields the \textit{predictive power curve}. Predictive power is a function of the two-sided $p$-value to reflect the common empirical practice of using two-sided tests. We express predictive power in terms of the $p$-value because statistical significance is typically reported using thresholds such as $p<0.05$ or $p<0.01$. This is without loss of generality because the two-sided $p$-value maps one-to-one to $|\hat z|$.

It is useful to distinguish predictive power from other common definitions of power. \textit{Planned power} is the probability of obtaining a statistically significant result under a prespecified effect size. \textit{Actual power} is this probability under the true effect. Finally, \textit{post hoc power} substitutes the original estimate for the true effect. Because it is a one-to-one transformation of the $p$-value, it provides no additional information about the evidence contained in the original estimate \citep{Hoenig2001}.

The following theorem gives a convenient representation of predictive power:

\begin{theorem}[Predictive Power]\label{theorem:predictive_power}
Let $s(\cdot)$ be symmetric about zero and strictly positive. Then for every $p\in(0,1)$, predictive power is given by
\begin{align}
r(p)
=
\frac{
\int_{\mathbb R}
\bigl[1-\Phi(1.96-z)\bigr]\,
\varphi\!\left(cv(p)-z\right)\,
\widetilde{\pi}(z)\,dz
}{
\int_{\mathbb R}
\varphi\!\left(cv(p)-z\right)\,
\widetilde{\pi}(z)\,dz
}
\label{equation:posterior_replication_power}
\end{align}
where $cv(p)\equiv\Phi^{-1}(1-p/2)$ and $\widetilde{\pi}(z) \equiv \frac{\pi(z)+\pi(-z)}{2}$ is the symmetrized distribution of normalized true effects.
\end{theorem}

Predictive power in \eqref{equation:posterior_replication_power} can be interpreted as the probability that a randomly chosen published finding with a given $p$-value is statistically significant at the 5\% level in the same direction as the original estimate in a replication using the same sample size as the original study.

The assumptions that $s(\cdot)$ is symmetric about zero and strictly positive are relatively mild. Symmetry requires that publication depend on the magnitude rather than the sign of the test statistic. This is natural in the empirical literatures we examine, which address different research questions and therefore provide no general reason for positive estimates to be more publishable than negative ones. Strict positivity rules out regions of the test-statistic distribution in which publication is impossible. In practice, this is not restrictive, since meta-datasets contain published findings throughout the $p$-value distribution. The use of the symmetrized distribution is necessary because $\pi(\cdot)$ is not identified from absolute test statistics; only its symmetrization $\widetilde{\pi}(\cdot)$ is identified (Theorem~\ref{thm:identification_nonparm}). Nonetheless, the representation in Theorem~\ref{theorem:predictive_power} is exact because predictive power is invariant to replacing $\pi(\cdot)$ with $\widetilde{\pi}(\cdot)$.




Predictive power has three noteworthy features. First, it admits a decision-theoretic interpretation: since $R$ is binary, $r(p)$ is the conditional expectation of replication success given the original $p$-value, and therefore minimizes expected squared prediction error. This provides a theoretical justification for focusing on predictive power to evaluate expected replicability at conventional significance thresholds.

Second, predictive power does not depend on the publication bias function $s(\cdot)$. It is therefore invariant to arbitrary forms of selective publication. The reason for this counterintuitive result is that selection changes the relative frequency of different $p$-values but not the distribution of power conditional on a given $p$-value. In other words, since publication depends only on the observed $p$-value, it treats all studies with that $p$-value identically and therefore leaves the expected replication probability conditional on that $p$-value unchanged. Invariance to publication bias has substantive implications for interpreting $r(p)$. It implies that expected replicability at any fixed $p$-value is identical whether publication bias is highly prevalent or entirely absent.

Third, predictive power is completely characterized by the symmetrized distribution of normalized true effects, $\widetilde{\pi}(z)$. Consequently, differences in predictive power across literatures arise entirely from differences in $\widetilde{\pi}(z)$. For instance, in a literature containing mostly large true effects, predictive power will be relatively high for a given $p$-value. By contrast, if $\widetilde{\pi}(z)$ places substantial mass near zero, predictive power will be relatively low for the same $p$-value. Thus, the central empirical task is to estimate the symmetrized distribution of normalized true effects, which we turn to next.\footnote{Following \cite{Andrews2019}, we treat $\pi$ throughout as a density. This simplifies the exposition and the proofs for the nonparametric results. The analysis can be extended to discrete or mixed
distributions of true effects without changing the nonparametric estimator, similarly to \cite{faridani2026testingunderpoweredliteratures}.}


\section{Estimation and Validation}\label{section:empirical}
This section estimates the predictive power curve for three empirical literatures and validates its performance against observed replication outcomes. This section focuses on literatures with available replications, which permit direct validation. Section~\ref{section:nonparametric} develops a nonparametric estimator of $r(p)$ and applies it to a larger dataset without replication outcomes.


\subsection{Data}
We examine three large-scale replication projects that systematically replicated published studies. First, \citet{Camerer2016} replicate experimental economics results from all 18 between-subjects laboratory experiments published in \textit{American Economic Review} and \textit{Quarterly Journal of Economics} between 2011 and 2014. Second, \citet{OpenScience2015} replicate results from 100 psychology findings in 2008 from \textit{Psychological Science}, \textit{Journal of Personality and Social Psychology}, and \textit{Journal of Experimental Psychology: Learning, Memory, and Cognition}. Following \citet{Andrews2019}, we consider a subsample of 73 studies with test statistics that are well-approximated by $z$-statistics. Finally, \citet{Camerer2018} replicate 21 experimental studies in the social sciences published between 2010 and 2015 in \textit{Science} and \textit{Nature}.

\subsection{Parametric Estimation}
Predictive power in equation \eqref{equation:posterior_replication_power} is fully determined by the symmetrized distribution of true effects, $\widetilde{\pi}(z)$. Although publication bias does not enter the expression for $r(p)$, it must still be accounted for when estimating $\widetilde{\pi}(z)$ from published studies. Ignoring selective publication would generally yield inconsistent estimates of $\widetilde{\pi}(z)$, and hence of $r(p)$. To estimate this distribution, we use the metastudy approach of \citet{Andrews2019}, with parameter estimates taken from Table 1 in \citet{Vu2024}. Estimates from \citet{Vu2024} are used because \citet{Andrews2019} estimate the distribution of estimates $\beta$ but not of standard errors, so their estimates do not recover the distribution of $z$ required to calculate predictive power.

The metastudy approach corrects for publication bias and assumes that true effects and standard errors are independent, each following a gamma distribution. It treats findings within each empirical literature as i.i.d. draws from the model described in Subsection~\ref{subsection:setup}. It normalizes true effects to have positive support, so we symmetrize the estimated distribution around zero before calculating predictive power, as required by Theorem~\ref{theorem:predictive_power}. Since our target is the symmetrized $\widetilde{\pi}(z)$, this is a normalization rather than a substantive restriction.\footnote{Equivalently, we can retain the nonnegative distribution of true effects and calculate the probability that the replication estimate is significant with the same sign as the original estimate. This yields identical results.}

A key advantage of the metastudy approach is that it uses only data from original studies. It can therefore be applied to any suitable meta-sample, regardless of whether replications have been conducted. In our applications, systematic replication evidence is used only to validate the resulting predictive power estimates. The main limitation of the metastudy approach is it assumes that true effects and standard errors are independent. In Section~\ref{section:nonparametric}, we develop a nonparametric estimator that relaxes the independence assumption and the parametric distributional assumptions.

As a robustness check, Figure \ref{figure:posterior_power_curves_repest} in Appendix~\ref{appendix:empirical_results} repeats the analysis using the systematic-replication approach of \citet{Andrews2019}, which uses both original and replication studies and requires weaker identifying assumptions. The resulting predictive power curves are quantitatively similar to those obtained using the metastudy approach.

\subsection{Validation}\label{subsection:validation}
Do the estimated predictive power curves accurately predict whether original studies actually replicate? Answering this question is central to establishing the empirical validity of the measure. If predictive power predicts observed replication outcomes, then it validates the measure and justifies its use in evaluating replicability. Conversely, if it fails to predict observed outcomes, any conclusions drawn from it lack empirical support and should be interpreted with caution.

Unlike the same-sample-size benchmark used in our main analysis, the validation exercise uses the sample-size rules actually implemented in the replication studies. We therefore construct predictive power curves for each field using a modified version of equation \eqref{equation:posterior_replication_power} that replaces original-study standard errors with replication standard errors. This is feasible because, in each application, replication sample sizes were set using prespecified rules.\footnote{Specifically, we replace $1-\Phi(1.96-z)$ with $1-\Phi(1.96-z_r)$, where $z_r=\beta/\sigma_r$ and $\sigma_r$ is determined by the replicators' sample-size rule.} See Lemma \ref{lemma:posterior_power_fpr} in the Appendix for details.

To assess predictive accuracy, we estimate the following linear probability model:
\begin{equation}
R_{il}=\alpha+\beta\hat{R}_{il}+\epsilon_{il},
\label{equation:validation_regression}
\end{equation}
\noindent
where $R_{il}$ is an indicator equal to one if study $i$ in empirical literature $l$ (i.e.\ economics, psychology, or social science) successfully replicates with significance and the same sign as the original estimate, and zero otherwise; and $\hat{R}_{il}$ is its predicted replication probability. Estimation assigns equal weight to each literature.

We compare predictive power based on the metastudy method with prediction-market forecasts compiled by \citet{Gordon2021}. Before the replications were conducted, academics traded contracts whose final market prices represent aggregate predicted replication probabilities. Because prediction markets were conducted for only a subset of replicated findings, the validation exercise uses the common sample of 70 studies for which both predictions are available.

Figure \ref{figure:model_fit} presents the results graphically using binscatter plots with fitted lines. Visually, both predictors are highly predictive of realized replication outcomes. Remarkably, the fitted line for predictive power lies almost exactly on the 45-degree benchmark of perfect calibration, under which predicted replication probabilities equal realized replication rates. By contrast, the fitted prediction-market line lies below the benchmark, indicating that prediction markets are overly optimistic about replication success. Formally, a joint test of perfect calibration ($\alpha=0$, $\beta=1$) does not reject the null for predictive power ($p=0.762$), but marginally rejects it for prediction markets at the 10\% level ($p=0.084$).



\begin{figure} [t]
	\centering
 	\caption{Replication Prediction Accuracy}
	\label{figure:model_fit}
\includegraphics[width=0.625\textwidth]{graphs/two_panel_model_fit.pdf}
\caption*{\textit{Notes}: The figure presents binscatter plots corresponding to equation \eqref{equation:validation_regression} for predictions from predictive power and prediction markets. Solid lines show fitted values from estimating equation \eqref{equation:validation_regression}, and dashed lines show the 45-degree benchmark. Perfect calibration is assessed using a joint test of $\alpha=0$ and $\beta=1$. The corresponding $p$-values are 0.762 for predictive power and 0.084 for prediction markets. Prediction-market forecasts are from \citet{Gordon2021}.}
\end{figure}

Table \ref{table:predictive_accuracy} complements the calibration regressions by comparing predictors using average squared prediction error, also known as the Brier score.\footnote{The Brier score is strictly proper: a forecaster minimizes expected loss by reporting their true subjective probability. This makes it suitable for comparing probability forecasts because it rewards honest, well-calibrated probabilities rather than strategic overstatement or understatement.} Calibration in Figure~\ref{figure:model_fit} assesses whether predicted probabilities match realized replication rates on average, whereas the Brier score also rewards forecasts that distinguish between studies with different replication prospects. For example, assigning every study a probability of 0.30 is perfectly calibrated if 30\% replicate, but provides no information about which studies are more likely to replicate. The Brier score addresses this by penalizing prediction errors at the study level, thereby favoring forecasts that are both well calibrated and informative.

In the weighted pooled sample, predictive power slightly outperforms prediction markets, with Brier scores of 0.216 and 0.220, respectively. This is quite striking given its limited information set. It relies only on test statistics from original studies, combined with the simple model outlined in Section \ref{subsection:setup}. Since the distribution of true effects is estimated from the full distribution of test statistics in the literature, each predicted replication probability places an individual result in the context of the broader body of evidence. By contrast, prediction markets aggregate the judgments of many expert scientists, whose information sets may include detailed knowledge of the specific research question, broader field expertise, perceptions of author quality, and evidence from earlier replication studies. That a parsimonious statistical model can slightly outperform this aggregation of expert judgments suggests that much of the signal relevant for replication is already embedded in the reported statistical evidence itself.


\begin{table}[!t] \centering
 \scriptsize
\caption{Replication Forecasting Error}
  \label{table:predictive_accuracy}
\begin{tabular}{@{\extracolsep{5pt}} lcc}
\\[-1.8ex]\hline
\hline \\[-1.8ex]
 & Predictive Power & Prediction Market \\
\hline \\[-1.8ex]
Economics  & $0.188$ & $0.226$ \\
Psychology  & $0.243$ & $0.243$ \\
Social Science  & $0.217$ & $0.190$ \\
All (Weighted) & $0.216$ & $0.220$ \\
\hline \\[-1.8ex]
\end{tabular}
\caption*{\textit{Notes}: This table reports Brier scores for predictive power and prediction-market forecasts. For finding $i$, the Brier score is $(R_i-\hat R_i)^2$, averaged across findings. Lower scores indicate greater predictive accuracy.}
\end{table}

Overall, predictive accuracy supports using predictive power to evaluate replicability and offers a simple, low-cost alternative to prediction markets.


\section{Results}\label{section:results}

\subsection{Replicability at Conventional Significance Thresholds}
We turn now to the main question of interest: how likely are findings that just satisfy conventional significance thresholds to replicate when using the same sample size as the original study? Figure~\ref{figure:posterior_power_curves} plots the predictive power curve for experimental economics, psychology, and social science. Evaluating the curve at the 5\% significance threshold, we find that replicability ranges between 0.1 and 0.25 across fields. Thus, a randomly chosen finding that just meets the most commonly used threshold for statistical significance has, at most, about a one-in-four chance of successful replication. At the more stringent 1\% level, replication probabilities remain relatively low, ranging between 0.23 and 0.40 across fields. Thus, even at the most stringent threshold typically used in empirical research, the chance of replication remains less than one-half across fields.

\begin{figure} [!t]
\centering
\caption{Predictive Power Curves, by Replication Project}
\label{figure:posterior_power_curves}
\includegraphics[width=0.64\textwidth]{graphs/metaest_predictive_power_curves.pdf}
\caption*{\textit{Notes}: The figure plots estimated predictive power curves for experimental economics, psychology, and experimental social science, using the metastudy approach to estimate $\tilde{\pi}(z)$. Predictive power is the probability that a replication with the same sample size as the original study produces an estimate that is statistically significant at the two-sided 5\% level and has the same sign as the original estimate, conditional on the original study's two-sided $p$-value. Square brackets report the percentage of original studies with $p$-values at or below the indicated threshold. Underlying data are from the replication projects of \citet{Camerer2016}, \citet{OpenScience2015}, and \citet{Camerer2018}, respectively.}
\end{figure}


These results come from a model that abstracts from several commonly proposed causes of low replicability, including $p$-hacking, researcher manipulation, measurement error, and heterogeneity between original and replication studies \citep{Ioannidis2005, Brodeur2016, Loken2017, Wuthrich2022}. Moreover, as discussed in Subsection~\ref{subsection:replication_power_curve}, predictive power is invariant to arbitrary forms of publication bias. The model nevertheless accurately predicts observed replication outcomes (Figure~\ref{figure:model_fit} and Table~\ref{table:predictive_accuracy}), suggesting that these factors are not necessary for explaining the low rates of replicability observed in practice.

The mechanism driving these results is instead the substantial mass that the estimated distributions of true effects place on small effects, leaving original studies with low statistical power at conventional significance thresholds. Appendix Figure~\ref{figure:power_bins} shows the distribution of predictive power conditional on $p=0.05$ for each application. A substantial share of just-significant findings have extremely low replicability: the share with replication probabilities below 5\% is 20\% in economics, 30\% in psychology, and 77\% in the social sciences. Thus, very low replication probabilities are relatively common across all fields and particularly prevalent in the social sciences.

These results imply that many just-significant findings have true effects substantially smaller than their original estimates. Figure~\ref{figure:posterior_z} in Appendix~\ref{appendix:empirical_results} plots the distributions of the true effect $z$ conditional on $p=0.05$ for all three applications. The distributions place substantial mass below the observed estimate of $\hat z=1.96$, although their shapes differ across fields: the distribution is centered at a positive but considerably smaller effect in economics, while it is concentrated closer to zero in psychology and especially in social science. This bias reflects the ``winner's curse" \citep{WinnersCurseQJE}: even though the original estimators are unbiased before conditioning on the $p$-value, just-significant estimates often combine relatively small true effects with favorable sampling variation, causing them to overstate the underlying effects. Consequently, many findings with $p=0.05$ are also likely to fail replication criteria based on similarity in effect magnitude, such as the relative effect size, $\hat z_r/\hat z$.


\subsection{Interpreting Replication Outcomes}\label{subsection:interpreting_replication_outcomes}
How should replication outcomes be interpreted in light of these results? We begin by clarifying what the results do not imply. Low replicability at conventional significance thresholds does not mean that most just-significant findings are `false'. Nor does an unsuccessful replication refute the original finding or necessarily indicate errors in the original study. Instead, replication outcomes should be interpreted as additional evidence about the null hypothesis that contributes to the cumulative evidence on an empirical claim.

To formalize this idea, consider a researcher interested in whether the true effect is `close to zero'. This framing meets practitioners ``where they are,'' given the ubiquity of null-hypothesis significance testing in applied research.\footnote{For proposals to move beyond dichotomous significance testing, see \citet{McShane2019} and \citet{Amrhein2019}.} Define a \textit{near-null effect} as a true effect whose magnitude is small enough that a same-sized replication has less than a 5\% probability of achieving 5\% two-sided significance in the direction of the true effect.\footnote{We focus on near-null effects because the estimated distribution of true effects is continuous and therefore assigns zero probability to an effect of exactly zero.} Thus, a near-null corresponds to $\{|z|<0.315\}$.\footnote{For $z>0$, the probability that a same-sized replication is significant in the direction of the true effect is $1-\Phi(1.96-z)$. Equating this probability to 0.05 gives $z=1.96-\Phi^{-1}(0.95)\approx0.315$; symmetry gives the corresponding cutoff for $z<0$. Because the definition uses the normalized effect $z$, rather than the underlying treatment effect $\beta$, the cutoff does not represent a common substantive effect size across studies and should be viewed as an operational definition of a near-null effect.}



Let $q_0(p)\equiv\Pr(H_0\mid p)$ denote the probability of a near null conditional on the original $p$-value. Recall that $r(p)\equiv\Pr(R=1\mid p)$ is the probability of a successful replication, and define $r_0(p)\equiv\Pr(R=1\mid p,H_0)$ as the corresponding probability under a near null. By construction, $r_0(p)<0.05$. Then it follows that
\begin{equation}\label{equation:updating}
\begin{aligned}
\Pr(H_0\mid p,D=1,R=1)
&=
q_0(p)
\left(
\frac{r_0(p)}{r(p)}
\right)
\\
\Pr(H_0\mid p,D=1,R=0)
&=
q_0(p)
\left(
\frac{1-r_0(p)}{1-r(p)}
\right)
\end{aligned}
\end{equation}


The first term expresses the probability of a near null after a successful replication and the second after an unsuccessful replication. In both expressions, the first term is the probability of a near null based on the original evidence, and the second is the updating factor implied by the replication outcome.

\begin{figure}[!t]
\centering
\caption{Probability of Near Null Before and After Replication}
\label{figure:effective_null_updates}
\includegraphics[width=0.64\textwidth]{graphs/metaest_near_null_updates.pdf}
\caption*{\textit{Notes}: The figure plots the probability of a near null before replication, after a successful replication, and after an unsuccessful replication, as a function of the original two-sided $p$-value. A near null is defined by $\{|z|<0.315\}$, corresponding to a true effect whose magnitude implies less than a 5\% probability that a replication achieves 5\% two-sided significance in the direction of the true effect. A successful replication is statistically significant at the 5\% level and has the same sign as the original estimate. Replications are assumed to use the same sample size as the original study. Underlying data are from \citet{Camerer2016}, \citet{OpenScience2015}, and \citet{Camerer2018}, respectively.}
\end{figure}
Figure~\ref{figure:effective_null_updates} illustrates the updating in \eqref{equation:updating}, plotting the probability of a near null effect before replication (dotted line), after a successful replication (green line), and after an unsuccessful replication (orange line). Most strikingly, for most $p$-values, a successful replication sharply lowers the probability of a near null, whereas an unsuccessful replication raises it only modestly.

For example, in experimental economics, a finding with $p=0.05$ has a 0.16 probability of being a near null before replication. A successful replication reduces this probability sharply to 0.02, whereas an unsuccessful replication increases it only modestly to 0.21. This asymmetry arises because near nulls have very low replication probabilities, ranging from 0.025 to 0.05. A successful replication is therefore highly unlikely under $H_0$ and produces a large downward update in \eqref{equation:updating}, whereas an unsuccessful replication remains relatively likely even when the effect is not close to zero, and therefore produces only a modest upward update. This asymmetry also appears in psychology and experimental social science, although near nulls are generally more likely in these fields for the same original $p$-value and replication outcome.

This asymmetry has implications for interpreting replication outcomes in practice: a successful replication can substantially strengthen evidence against a near null, whereas an unsuccessful replication generally provides only limited evidence in its favor. Thus, replication outcomes are best interpreted as cumulative evidence rather than binary confirmation or refutation.



\begin{comment}
We emphasize that low replicability at conventional thresholds does not imply most just-significant findings are false. Similarly, unsuccessful replications do not refute original findings. Instead, original and replication results should be treated as cumulative evidence: a successful replication strengthens evidence against the null, while an unsuccessful one weakens it.

To formalize this point, consider a researcher who is interesting in learning about the null hypothesis. Let $q(p)\equiv\Pr(H_0\mid p)$ denote the probability of the null conditional on the original $p$-value, and recall that $r(p)\equiv\Pr(R=1\mid p)$ denotes the probability of a successful replication. If the replication uses significance level $\alpha$, it follows that $
\Pr(H_0\mid p,R=1)
=q(p) \cdot \left(
\frac{\alpha}{r_{\sigma_r}(p)} \right)$ and $\Pr(H_0\mid p,R=0)
=  q(p) \cdot \left(
\frac{1-\alpha}{1-r_{\sigma_r}(p)} \right)$. In both expressions, the first term in the product is the probability of the null based on the original evidence, while the second is the `updating' factor based on the observed the replication outcome.

These expressions show that the amount learned from a replication depends on how surprising its outcome is given the predictive power of the original finding. For instance, an insignificant replication increases the probability that the null is true by more when $r(p)$ is high because failure is then more informative. Conversely, a significant replication decreases the probability of the null by more when $r(p)$ is low.

As an illustration, consider a finding with $p=0.05$ in experimental economics, for which we estimate $r(0.05)=\frac{1}{3}$. Furthermore, suppose $\alpha=0.05$ and $q(0.05)=0.30$.\footnote{The estimated Gamma distribution of true effects places no point mass at exactly zero. As a back-of-the-envelope approximation, we proxy null findings by those with underlying power between 5\% and 10\%. Figure~\ref{figure:power_bins} shows that this group accounts for approximately 29.6\% of findings conditional on $p=0.05$.} Based on the expressions above, a successful replication reduces the probability of the null from $0.30$ to
$\Pr(H_0\mid p=0.05,R=1)=0.045$, whereas an unsuccessful replication increases it to
$\Pr(H_0\mid p=0.05,R=0)=0.428$. Thus, a successful replication provides very strong additional evidence against the null. On the other hand,  an unsuccessful replication -- which is twice as likely to be observed -- increases the probability of the null, but still leaves it less likely to be true than the alternative.
\end{comment}

\subsection{Additional Results}
We highlight two additional implications of the predictive power curves. The first concerns how strong the statistical evidence in the original study must be to achieve a high probability of replication. For any replication probability target $T$, inverting the predictive power curve yields the $p$-value cutoff $p^*$ satisfying $r(p^*)=T$. Lemma~\ref{lemma:uniqueness} proves that, under mild regularity conditions, this cutoff exists and is unique. Figure~\ref{figure:posterior_power_curves} shows that achieving predictive power of at least 80\% requires $p$-value cutoffs of 0.00006, 0.00005, and 0.0001 in economics, psychology, and the social sciences, respectively. Between 10\% and 33\% of findings in the respective samples meet these literature-specific benchmarks. Notably, these cutoffs are several orders of magnitude more stringent than conventional significance thresholds, and far below even the 0.005 threshold proposed in \citet{Benjamin2018}. Nevertheless, we do not interpret these cutoffs as a recommendation to mechanically adopt more stringent thresholds. Although stricter thresholds would increase the average replicability of findings that meet them, they might also discard informative evidence, disadvantage settings where very large samples are infeasible, and intensify specification searching.

The second additional result concerns the convexity of the predictive power curves. In current empirical practice, moving from a $p$-value of 0.10 to 0.05 is commonly treated as a material gain in credibility, marking a transition from ``marginal significance'' to genuine ``statistical significance.'' However, Figure~\ref{figure:posterior_power_curves} suggests that this transition does little to alter the underlying replicability of a finding: expected replication probabilities rise by only 3--6 percentage points, and remain at low levels. This implies that moving between conventional significance thresholds need not translate into meaningful improvements in replicability.

This observation also has implications for forms of $p$-hacking that misreport $p$-values just above a conventional threshold as falling just below it. Unlike selective reporting, which changes whether a result is observed but not its reported $p$-value, this form of $p$-hacking changes the reported $p$-value itself. Consider a finding with a true $p$-value of 0.06 that is misreported as $p=0.049$ in order to claim significance; its expected replicability remains $r(0.06)$. Figure~\ref{figure:posterior_power_curves} shows that $r(0.06)$ and $r(0.049)$ are nearly identical across all fields. Thus, forms of $p$-hacking that nudge reported $p$-values across the 5\% significance threshold are unlikely to account for much of the low replicability observed in practice.


\section{Nonparametric Predictive Power Curve}\label{section:nonparametric}

Thus far, the empirical analysis has imposed parametric assumptions about the distribution of true effects. In this subsection, we develop a nonparametric estimator of $r(p)$ that imposes no functional-form restrictions on $\tilde \pi(z)$. The nonparametric estimator requires large meta-samples that include many reported $p$-values and cannot be applied to the relatively small replication datasets analyzed above. Given these data requirements, we apply the estimator to the large dataset of experimental and quasi-experimental economics findings compiled by \citet{Brodeur2020}.

Nonparametric identification of $r(p)$ is established formally in Appendix \ref{app:nonparametrics}. Discontinuities in the density of published $t$-ratios identify changes in publication probability across significance thresholds. After adjusting for these changes, recovering the distribution of true effects from the distribution of estimated $t$-ratios is a standard deconvolution problem with normally distributed errors.\footnote{See e.g. \cite{Carroll,Fan_global,CarrascoFlorens,faridani2026testingunderpoweredliteratures}.} Below, we focus on the principal challenge of constructing a consistent estimator of $r(p)$. Proofs are in Appendix \ref{app:nonparametrics}.

\subsection{Nonparametric Estimation}

Suppose we observe a sample of absolute values of $n$ published $t$-ratios, $\{|\hat z_i|\}_{i=1}^n$.  Our aim is to estimate the symmetrized density $\tilde{\pi}(z) \equiv \frac{\pi(z)+\pi(-z)}{2}$, which determines the target parameter $r(p)$ and is identified from the distribution of $|\hat z|$. Following Example 1 of \citet{CarrascoFlorens} and \citet{faridani2026testingunderpoweredliteratures}, we can express  $\tilde\pi(z)$ as a linear combination of known polynomial basis functions, $\chi_j$, with unknown coefficients $b_j$:
\begin{equation}\label{eq:pi_series}
    \tilde{\pi}(z) = \sum_{j=0}^\infty b_j\chi_j(cv(p)-z),
\end{equation}
\noindent where $\chi_j(z) \equiv \text{\normalfont He}_{j}\!\left(z/\sqrt{2}\right)/\sqrt{j!}$ is the normalized $j$th probabilists' Hermite polynomial.

\begin{remark}
    Estimating the full density $\tilde \pi$ is severely ill-posed because many distributions of true effects can generate nearly indistinguishable distributions of observed $t$-ratios \citep{Carroll,Fan_global}. Distinguishing among them relies on higher-order coefficients $b_j$, which capture fine features of the data and are highly sensitive to sampling error. Despite this difficulty, the key insight underlying our estimator is that higher-order coefficients have diminishing influence on $r(p)$ because it averages over $\tilde{\pi}(z)$ and is therefore relatively insensitive to its finer features. Consequently, imprecise estimates of $\tilde{\pi}(z)$ can yield accurate estimates of $r(p)$. A similar principle appears in \citet{faridani2026testingunderpoweredliteratures}, but targeting $r(p)$ yields a new estimator and a distinct convergence rate.
\end{remark}

In practice, we truncate the infinite expansion and estimate only the first $J_n$ coefficients. The cutoff $J_n$ grows logarithmically with the sample size, allowing the estimator to capture finer features of $\tilde{\pi}(z)$ as more data become available. Formally, this amounts to regularizing the deconvolution problem using a spectral cutoff based on the singular value decomposition from \cite{CarrascoFlorens}.\footnote{See Example 1 of \cite{CarrascoFlorens} for the singular value decomposition. The Gibbs fluctuations incurred by spectral cutoff do not prevent consistency because they are integrated away when we calculate $r(p)$ from $\tilde{\pi}(z)$.} Substituting this truncated expansion into Equation~(\ref{equation:posterior_replication_power}) and simplifying yields the following estimator:
\begin{align}\label{equation:np_estimator}
   \widehat{r}^{np}(p)
   \equiv
   1-\frac{\sum_{j=0}^{J_n}a_j\widehat{b}_j}
   {\sum_{j=0}^{J_n}d_j\widehat{b}_j}
\end{align}

\noindent where $\widehat{b}_j$ estimates the coefficient $b_j$ in \eqref{eq:pi_series} times a common constant (that cancels in the ratio). The known constants $a_j$ and $d_j$ capture the contribution of the $j$th basis function to the numerator and denominator, respectively, and can be computed numerically.\footnote{\raggedright More specifically,
$a_j \equiv \int_{\mathbb{R}}
\varphi(u)e^{-u^2/2}
\Phi\left(c-cv(p)+\sqrt{2}u\right)\psi_j(u)\,du$,
and
$d_j \equiv \int_{\mathbb{R}}
\varphi(u)e^{-u^2/2}\psi_j(u)\,du$.\par}

To estimate the coefficients $b_j$ consistently, we must account for selective publication of the observed $t$-ratios. As before, let $D=1$ indicate that a finding is published, and suppose its publication probability depends on the observed $t$-ratio through the even selection function $s_{\theta_0}$:
\begin{equation}\label{eq:w_definition_main}
    \Pr(D=1\mid \hat z )=s_{\theta_0}(\hat z)
\end{equation}
\noindent where the functional form of $s_\theta$ is known but the parameter vector $\theta_0\in \Theta$ is unknown.

The properties of the Hermite basis allow each coefficient $b_j$ to be expressed as a moment of the unselected distribution of estimated $t$-ratios. Given a consistent estimator $\widehat{\theta}_n$ of the publication-bias parameters, we follow \citet{faridani2026testingunderpoweredliteratures} and estimate this moment by weighting each published $t$-ratio by the inverse of its estimated publication probability:
\begin{align}\label{equation:np_coefficients}
    \widehat{b}_j
    &\equiv
    \frac{2^{j/2}}{n}\sum_{i=1}^n
    \frac{
         \psi_j\left(cv(p)-|\hat z_i|\right)
        \varphi\left(cv(p)-|\hat z_i|\right) +\psi_j\left(cv(p)+|\hat z_i|\right)
        \varphi\left(cv(p)+|\hat z_i|\right)
    }{
        s_{\widehat{\theta}_n}\left(\hat z_i\right)
    }
\end{align}
where $\varphi(\cdot)$ is the standard normal density, $\psi_j(z)\equiv \text{\normalfont He}_{j}\!\left(z\right)/\sqrt{j!}$ and $cv(p)\equiv\Phi^{-1}(1-p/2)$ is the standard normal critical value corresponding to the two-sided $p$-value.

Next we formalize our assumptions about the sample. In particular, we allow test statistics to be dependent within an article, but not across articles.
\begin{assumption}\label{assum:sample}
    The following all hold:
    \begin{enumerate}
        \item Let $|\hat z_1|,\ldots,|\hat z_n|$ be an identically distributed sample from the distribution of $|\hat z|$ conditional on publication, i.e. $D=1$.
        \item Sets of $\hat z_i$ drawn from different articles are independent.
        \item There is a universal constant $L\geq 1$ such that each article contributes no more than $L$ observations to the sample.
    \end{enumerate}
\end{assumption}

Theorem \ref{thm:nonparametric_consistency_simplified} shows that, if the publication-bias parameters $\theta_0$ are known or consistently estimated by $\widehat{\theta}_n$, then the estimator $\widehat{r}^{np}(p)$ defined in \eqref{equation:np_estimator}, with coefficients given by \eqref{equation:np_coefficients}, is consistent for $r(p)$.\footnote{This result is a special case of the more general Theorem \ref{thm:nonparametric_consistency} in Appendix \ref{app:nonparametrics}, which allows the replication and original sample sizes to differ.}

\begin{theorem}\label{thm:nonparametric_consistency_simplified}
    Let the sample $|\hat z_1|,\ldots,|\hat z_n|$ satisfy Assumption \ref{assum:sample}. Suppose that the density $\pi(z)$ exists and has bounded height. Assume that $s_\theta$ is even and $\inf_{\theta\in\Theta,\,t\in\mathbb R}s_\theta(t)>0.$ Let $J_n=\left\lceil \frac{2}{3\log 2}\log(An)\right\rceil$ for any $A>0$. Assume further that an estimator $\widehat{\theta}_n$ satisfies: $$
        \sup_{t\in\mathbb{R}}
        n^{1/3}\left|
        \frac{1}{s_{\widehat{\theta}_n}(t)}
        -\frac{1}{s_{\theta_0}(t)}
        \right|
       =\mathcal{O}_p(1)$$ Then, for every $p\in(0,1)$,
    \[
        \widehat{r}^{np}(p)\xrightarrow{p}r(p)
    \]
\end{theorem}

We complete the estimation procedure by specifying and estimating the publication-bias function. To allow for flexible forms of publication bias, we model the publication probability as a step function with $M$ known critical values $\{v_m\}_{m=1}^M$ where the publication probability of the most-significant $\hat z$ is normalized to one:
\begin{align}
    s_{\theta}(\hat z)
    &= \theta_{(1)}\mathbf{1}\left\{|\hat z| \leq v_1\right\}
    + \sum_{m=2}^{M}\theta_{(m)}
    \mathbf{1}\left\{|\hat z| \in (v_{m-1},v_{m}]\right\}
    +
    \mathbf{1}\left\{|\hat z|>v_M\right\}
    \label{eq:stepfunction}
\end{align}

Given the known step locations, we estimate the publication-bias parameters using the caliper ratio estimator of \citet{faridani2026testingunderpoweredliteratures}. The estimator compares the numbers of published $t$-ratios within narrow intervals on either side of each critical value. Discontinuities in these counts identify changes in publication probability, from which the step heights can be recovered. The resulting estimates are uniformly consistent for $s_{\theta}$ and can be used in Theorem~\ref{thm:nonparametric_consistency_simplified}. For further details, see \citet{faridani2026testingunderpoweredliteratures}.


\subsection{Application}
We apply the nonparametric estimator to the dataset in \cite{Brodeur2020}, which consists of 21,740  $t$-ratios testing main hypotheses in 684 articles published in top economics journals. Each $t$-ratio corresponds to the test of a causal inference hypothesis where the study methodology is either a randomized controlled trial (RCT), Regression Discontinuity Design (RDD), Difference-in-Differences (DID), or Instrumental Variables (IV). Papers can contain multiple estimates. Following \cite{Kranz2022}, we de-round the $t$-ratios before estimation because rounding reported coefficients and standard errors can generate artificial bunching around conventional significance thresholds. The publication bias function, $s_{\theta}(\cdot)$, is specified to have steps at the three conventional two-sided critical values $1.645, 1.96,$ and $2.576$. Estimates for the publication probability function are presented in Appendix~\ref{appendix:publication_probability_ratios}.


\begin{figure}[!t]
    \centering
    \caption{Nonparametric Predictive Power Curves, by Research Design in Economics} \label{fig:brodeur-stacked-nonparametric}

 \includegraphics[width=\textwidth]{graphs/nonparametric_combined.pdf}
\caption*{\textit{Notes}: The figure plots nonparametric estimates of the predictive power curve by research design in economics. Predictive power is evaluated for a replication with the same sample size as the original study. Underlying data are from \citet{Brodeur2020}.}
\end{figure}

To make the replication exercise concrete for observational studies, consider an RDD using the Current Population Survey (CPS) to estimate the effect of Medicare eligibility at age 65 on health-insurance coverage. Replication is a hypothetical repeated-sampling exercise: imagine drawing a new independent sample of the same size from the same population and policy environment and applying the identical RDD specification. Replication is successful if the resulting estimate is significant at the 5\% level and has the same sign as the original estimate.

Figure \ref{fig:brodeur-stacked-nonparametric} presents the nonparametric predictive power curves for the four research designs. An RCT finding with a $p$-value of 0.05 has an estimated replication probability of 0.34, modestly exceeding the parametric estimate of 0.25 for experimental economics. Predictive power at $p=0.05$ is notably higher for the observational designs, ranging between 0.43 and 0.50, although, in absolute terms, replication probabilities remain relatively low. Similar to the parametric analysis in Figure~\ref{figure:posterior_power_curves}, extremely small $p$-values are required to achieve predictive power of at least 80\%.

Differences in replicability across the seven applications are closely related to differences in sample size. Figure~\ref{figure:median_sample_size_r0.05} plots median sample size against predictive power at $p=0.05$. It shows that predictive power is approximately linear in the logarithm of median sample size, with the applications lying close to the fitted line. Higher predictive power in observational studies is consistent with their much larger samples. Similarly, economics RCTs—which include relatively large field experiments—have higher replicability than smaller laboratory studies in experimental economics. These patterns are consistent with replicability being determined primarily by the distribution of normalized true effects, $\tilde \pi(z)$, as emphasized in Section~\ref{section:results}. For a given underlying effect, a larger sample reduces the standard error and increases the normalized true effect, shifting $\tilde \pi(z)$ away from zero and increasing predictive power.

These results offer a different perspective on the relationship between research practices and replicability. A literature with more selective reporting and $p$-hacking may nevertheless be more replicable. For example, \citet{Brodeur2020} use discontinuities at conventional significance thresholds to argue that IV studies—and, to a lesser extent, DID studies—exhibit substantially more $p$-hacking than RCTs. Yet Figure~\ref{figure:median_sample_size_r0.05} shows that just-significant RCT findings have much lower expected replicability than IV and DID findings, consistent with their smaller sample sizes.

This is because replicability depends primarily on the distribution of normalized true effects, which is strongly influenced by sample size. The distinction matters because low replicability is frequently attributed to $p$-hacking and publication bias \citep{OpenScience2015, Camerer2016, Camerer2018}. For example, in a 2016 \textit{Nature} survey, 90\% of researchers identified ``selective reporting'' as contributing to irreproducible research, more than any other factor \citep{Baker2016}. Against this background, our results suggest greater caution in attributing low replicability primarily to selective reporting and $p$-hacking.


\begin{figure} [!t]
\centering
\caption{Median Sample Size and Expected Replicability at $p=0.05$}
\label{figure:median_sample_size_r0.05}
\includegraphics[width=0.5\textwidth]{graphs/sample_size_against_r0.05.pdf}
\caption*{\textit{Notes}: The figure plots the expected replication probability at an original two-sided $p$-value of 0.05 against the median sample size for each application. Replication project data are from \citet{Camerer2016}, \citet{OpenScience2015}, and \citet{Camerer2018}. For RCTs, IV, RDD, and DiD, data are from \citet{Brodeur2020}, with median sample sizes calculated from the nonmissing values. The dotted line plots the OLS fitted values from $r_i(0.05)=\alpha+\beta\log_{10}(\operatorname{Median}N_i)+\varepsilon_i$.}
\end{figure}


Overall, the nonparametric analysis extends the earlier parametric results to a broader range of research designs while allowing for a flexible distribution of true effects. Expected replicability at conventional significance thresholds is slightly higher for observational designs than for experiments, but remains low in absolute terms across all applications.\footnote{The additional analyses in Figure \ref{figure:effective_null_updates} are not feasible with the nonparametric estimator because they require estimation of $\tilde{\pi}$ itself. While the estimator of $r(p)$ converges relatively quickly, the nonparametric estimator of $\tilde{\pi}$ converges slowly and is not guaranteed to yield a valid density.}

\section{Recommendations}\label{section:recommendations}
This section offers three recommendations for researchers and replicators. Recommendations are broadly applicable to replications evaluated using either binary significance criteria or effect sizes.


\begin{comment}
\begin{enumerate}
    \item \textbf{Cumulative evidence:} Treat evidence as cumulative, not binary. A single significant finding provides only partial evidence against the null, while evidence accumulates through replication. Replications are most informative when their outcomes are surprising---for example, when a finding with $p=0.05$ replicates or one with $p=0.005$ does not.
    \item \textbf{Replication design:} Use the largest feasible sample size in replications, as larger samples distinguish more sharply between the null and alternative hypotheses. When evidence informs a decision, choose the sample size that balances the marginal benefit of better-informed decisions against the marginal cost of additional observations. Avoid conventional power calculations that treat the original effect estimate as the true effect: these calculations yield expected power strictly below the nominal target \citep{Vu2024}.
    \item \textbf{Research reforms:} Encourage and reward replications. For example, top journals could publish short reports summarizing replication evidence for influential studies, and researchers could cite these replications alongside the originals \citep{Coffman2017}. Policies such as preregistration and pre-analysis plans remain valuable for limiting selective reporting and researcher discretion, but do not address the primary source of low replicability, which is low power.
\end{enumerate}
\end{comment}


\subsection{Cumulative Evidence}\label{subsection:cumulative_evidence}
\begin{recommendation}
Treat evidence as cumulative, not binary. A single significant finding provides only suggestive evidence against the null, while evidential strength accumulates through replication. Successful replications can provide strong evidence against the null, whereas insignificant replications do not, by themselves, refute the original finding or imply flaws in the study.
\end{recommendation}


A significant finding provides only partial evidence against the null hypothesis. It does not establish that the null is false. Similarly, a replication does not deliver a definitive conclusion about the original claim, confirming it if ``successful'' and refuting it otherwise.

For interpreting replication outcomes, the analysis in Subsection~\ref{subsection:interpreting_replication_outcomes} reveals an important asymmetry: a successful replication can substantially strengthen evidence against a near null result, whereas an unsuccessful replication generally provides only limited evidence in its favor. In this way, replications build the cumulative body of evidence required to distinguish persistent empirical relationships from isolated results driven by sampling variation, leading to more reliable scientific inferences. The ultimatum-game literature provides an instructive example, where a large number of replications have established robust empirical regularities \citep{Coffman2015}. Treating scientific evidence as cumulative aligns with arguments developed by \citet{Popper1934} in \textit{The Logic of Scientific Discovery}: ``We do not take even our own observations quite seriously, or accept them as scientific observations, until we have repeated and tested them. Only by such repetitions can we convince ourselves that we are not dealing with a mere isolated `coincidence', but with events which, on account of their regularity and reproducibility, are in principle inter-subjectively testable.''

\subsection{Replication Design}
\begin{recommendation}
    Use the largest feasible replication sample to better distinguish between the null and alternative hypotheses. Treat nominal power targets with caution when power calculations substitute the estimated effect for the true effect: expected replication rates will generally be below the nominal target, even without publication bias or $p$-hacking.
\end{recommendation}

The same-sample-size benchmark used in this paper evaluates replicability under the original research design; it is not intended as a recommendation for replication design. Instead, we recommend that replicators use the largest feasible sample, which maximizes predictive power. This is because larger samples make replication outcomes more informative: significance provides stronger evidence against the null, whereas insignificance provides stronger evidence in its favor.\footnote{Formally, consider the updating factors in equation~\eqref{equation:updating}. Increasing the replication sample size lowers $\left(\frac{r_0(p)}{r(p)}\right)$ following significance and raises $\left(\frac{1-r_0(p)}{1-r(p)}\right)$ following insignificance.}

When evidence informs a concrete decision, the replication sample size should balance the value of more precise evidence against the cost of collecting additional data. In principle, this requires weighing sampling costs against the consequences of adopting an ineffective intervention or rejecting a beneficial one.

Finally, we recommend caution when using the conventional approach of choosing a replication sample size to detect the original effect estimate with, for example, 80\% or 90\% power. Such calculations treat the original estimate as exactly equal to the true effect and ignore its sampling uncertainty. Due to the nonlinearity of the power function, \citet{Vu2024} shows that this approach delivers expected power strictly below the nominal target, even in the absence of $p$-hacking and publication bias.\footnote{The intuition is as follows. Let $T$ denote the nominal power target and $g(\cdot)$ denote the probability of replication. Suppose $\hat{\beta}\sim N(\beta,\sigma^2)$. Because $g(\cdot)$ is locally concave, Jensen's inequality implies that $\mathbb{E}[g(\hat{\beta})]<g(\mathbb{E}[\hat{\beta}])=g(\beta)=T$. Thus, even though the original estimates are unbiased, expected replication power is below the nominal target $T$.} Of course, outcomes from such replications can still be used to rationally update beliefs about the null. The concern is rather that stated replication-rate targets may create the mistaken impression that they are attainable absent research distortions, which is not the case.





\subsection{Encourage and Reward Replications}
\begin{recommendation}
Encourage the publication and citation of replications to establish empirical regularities. Preregistration and pre-analysis plans remain valuable for publishing null results and limiting $p$-hacking, but significant findings can still have low replicability even in the absence of $p$-hacking and publication bias.
\end{recommendation}

Our finding of low replicability at conventional thresholds reinforces recent calls for greater incentives to conduct, publish, and recognize replications \citep{Nosek2012, Coffman2015, Brodeur2023}. As shown in Subsection~\ref{subsection:interpreting_replication_outcomes}, successful replications can generate substantial information about the null hypothesis and hence provide considerable value for subsequent researchers and decision-makers.

Despite this, replications offer less novelty and professional reward than original studies. To address this imbalance, \citet{Coffman2017} make two proposals. First, leading journals could publish short reports summarizing the replication evidence for influential studies, including replications otherwise buried within broader papers. Second, researchers could cite replications alongside original studies, ensuring that credit and attention reflect the cumulative evidence rather than the first published result alone. Complementing these proposals, the Replication Games of \citet{Brodeur2023} use team-based workshops and joint meta-papers to coordinate replication efforts and reward participants.

Finally, it is important to emphasize that preregistration, pre-analysis plans, and registered reports can limit publication bias and $p$-hacking, but that low replicability can persist even without these distortions. Hence, improving replicability and limiting publication bias are distinct objectives. Three observations underscore this insight. First, the expected replicability for a finding with a given $p$-value is invariant to publication bias and instead determined by the underlying distribution of true effects (Subsection \ref{subsection:replication_power_curve}). Second, Section~\ref{section:nonparametric} finds higher replicability in observational studies than RCTs, despite stronger evidence of $p$-hacking and publication bias in the former \citep{Brodeur2020, Kranz2022}. This occurs due to larger samples in observational studies, which lead to higher replicability. Third, the impact of imposing more stringent publication thresholds would be to mechanically increase replicability, since this favors findings with higher statistical power. In other words, increasing selectivity \textit{improves} replicability. For these reasons, reforms to reduce selective reporting and promote systematic replication should be pursued as complementary objectives.


\section{Conclusion}\label{section:conclusion}
Null hypothesis significance testing has been central to scientific practice -- and the subject of sustained debate about its merits and limitations -- for nearly a century \citep{Fisher1925, NeymanPearson1933, Rozeboom1960, Cohen1994, Wasserstein2019}. Instead of comparing it to proposed alternatives, this paper evaluates the conventional practice of significance testing on its own terms. Using an empirically validated measure of predictive power across fields and research designs, we find that results meeting conventional significance thresholds have relatively low chances of remaining significant in replications of the same sample size. The mechanism is low power in original studies rather than publication bias or $p$-hacking.

These findings suggest that conventional significance thresholds provide a much weaker guarantee of replicability than is often assumed. A statistically significant result from a single study should instead be treated as suggestive rather than decisive evidence. Robust scientific conclusions require evidence that withstand repeated testing across independent studies.

\bigskip
\bibliographystyle{aer}
{\small
\bibliography{References}
}