Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
89,771 characters · 23 sections · 79 citation commands
The Power of Tests for Detecting $p$-Hacking
Researchers have a strong incentive to find and report significant results imbens2021statistical. simonsohn2014p used the term “$p$-hacking” to encompass decisions made by researchers in conducting their work to improve the publication prospects of their results. Their work has generated a flourishing literature that examines empirically the distribution of $p$-values across studies (the “$p$-curve”) to determine if $p$-hacking is prevalent or not.\footnote{See, e.g., masicampo2012peculiar,simonsohn2014p,lakens2015what,simonsohn2015better,head2015extent,ulrich2015p for early applications and further discussions, havranek2021does, brodeur2022we,malovana2022borrower, yang2022hedge, decker2023preregistration for recent applications, and christensen2018transparency for a review.}
Interpreting these empirical results and assessing the usefulness of existing tests for $p$-hacking requires knowing how powerful these tests are, and for this we need to understand how $p$-hacking impacts the $p$-curve.\footnote{We focus on the problem of detecting $p$-hacking based the $p$-curve and do not consider the popular Caliper tests gerber2008do,gerber2008publication,bruns2019reporting,vivalt2019specification,brodeur2020methods. Caliper tests aim to detect $p$-hacking based on excess mass in the distribution of $z$-statistics right above significance cutoffs. However, these tests do not control size in general kudrin2022robust.} Theoretically and using Monte Carlo analyses, this paper examines the power of tests for detecting likely forms of $p$-hacking, such as selecting controls in regression analyses, selecting among instruments in instrumental variables (IV) regression, and variance estimator selection in regression analyses. A careful study of power is important because the implications of $p$-hacking on the $p$-curve are not clear, and the magnitude of power depends on precisely how the $p$-curve is affected, which in turn depends on the empirical problem and how the $p$-hacking is undertaken.
We find that tests for $p$-hacking can have low power (above size, but not by much) and even no power (lower than or equal to size) in some cases, so failure to reject is not strong evidence of a lack of $p$-hacking; that different approaches to $p$-hacking have different effects resulting in different tests having power to detect such $p$-hacking; and that power depends on how easy it is to obtain statistically significant results. To illustrate, consider our first example below, where a researcher is interested in the coefficient on a variable in a linear regression but has a choice over which controls to include in the regression.
Multiple choices over controls results in the researcher having a choice over a number of $p$-values that can be reported. Selective reporting, or $p$-hacking, could include either reporting the smallest $p$-value (referred to as the “minimum” approach) or reporting results from a preferred specification if significant and if not then trying other specifications (referred to as the “threshold” approach). The minimum approach pushes the $p$-curve to the left (we get smaller $p$-values) without creating non-monotonicities or discontinuities, so tests that look for these have no power. The minimum approach can however violate bounds on the $p$-curve. Unfortunately, tests for this can have no or low power, even when more than half of the results are $p$-hacked.
The threshold approach creates discontinuities in the $p$-curve as well as violations of upper bounds, so tests for these features have power, and this approach is easier to detect. With half the researchers $p$-hacking in this way power is quite reasonable for the covariate selection problem. However, the threshold approach may not create humps in the $p$-curve, so tests for these may not have power. Publication bias can have similar impacts on the $p$-curve as $p$-hacking, and so contribute to rejections using the tests we examine; however, it can also make it more difficult to detect $p$-hacking.
If the true coefficient in the linear model to be tested for zero is very large, then $p$-values are likely small anyway, and without much motivation to $p$-hack, the impact of $p$-hacking is hard to detect. For hypotheses for which data is informative but the true effects are near popular significance cutoffs, the incentive to search amongst the controls is larger and so is the impact on the distribution of $p$-values, resulting in higher power of the tests we examine.
The results in this paper have important implications for meta-analysts choosing tests for detecting $p$-hacking. First, for both $p$-hacking approaches, combined tests for violations of upper bounds and monotonicity, such as the CS2B test elliott2022detecting, do best. Second, testing for violations of bounds on the $p$-curve is particularly important in settings where $p$-hacking is difficult to detect, for example, when researchers use the minimum approach. Third, tests for discontinuities in the $p$-curve, such as the discontinuity test of cattaneo2020simple, can be useful complements to tests based on monotonicity or bounds since the threshold approach can yield pronounced discontinuities. Finally, the classical Binomial test simonsohn2014p,head2015extent often has lower power, even in settings where other tests have high power for detecting $p$-hacking.
We consider empirical studies where individual researchers provide test results of a hypothesis by reporting a test statistic $T$ with distribution $F_{h}$, where $h \in \mathcal{H}$ indexes parameters of the distribution of $T$.\footnote{The setup and notation here follow elliott2022detecting.} Researchers are testing the null hypothesis that $h \in \mathcal{H}_0$ against $h \in \mathcal{H}_1$ with $\mathcal{H}_0 \cap \mathcal{H}_1$ empty. Suppose the test rejects when $T>cv(p)$, where $cv(p)$ is the level $p$ critical value. For any individual study the researcher tests a hypothesis at a particular $h$ (the “true” effect).\footnote{Here, $h$ is the local version of the difference between the true coefficient and its value under the null hypothesis, where `local' means it is scaled by the standard error of the estimator.} In what follows, we denote the power function of the test as $\beta\left(p,h \right)=\Pr\left(T>cv(p)\mid h\right)$.
In practice researchers can use different specifications that yield a set of test statistics $\{T_1, T_2, T_3, \dots \}$. This generates a set of $p$-values $\{P_1, P_2, P_3, \dots \}$ that they could report. We assume the researchers have a preferred specification, and in the absence of $p$-hacking would report that $p$-value. If they choose to $p$-hack, the approach they take would then comprise a method of choosing which $p$-value $P_r$ to report, i.e., a selection function $P_r=d(P_1, P_2,P_3, \dots).$ The joint distribution over possible $p$-values will depend on the testing situation and the distribution of the data, including the value for $h$. It is not the case that researchers can select any $p$-value they desire; the available $p$-values will be drawn from this joint distribution (at least, unless they directly fabricate or manipulate data). This relates to simonsohn2020blog's distinction of “slow $p$-hacking” and “fast $p$-hacking”: slow $p$-hacking corresponds to the case where $p$-values change little across analyses; fast $p$-hacking refers to settings where $p$-values change a lot across analyses. Slow $p$-hacking is then when the $p$-values are highly positively correlated, fast $p$-hacking when the $p$-values are less dependent.
In the absence of $p$-hacking, the researcher reports the $p$-value from the preferred specification. elliott2022detecting provided a theoretical characterization of the distribution of $p$-values across multiple studies in the absence of $p$-hacking for general distributions of true effects.\footnote{See, e.g., hung1997, simonsohn2014p, and ulrich2018some for numerical and analytical examples of $p$-curves for specific tests and/or effect distributions.} Across researchers, there is a distribution $\Pi$ of true effects $h$, which is to say that different researchers testing different hypotheses examine different problems that have different true effects. The resulting CDF of $p$-values across all these studies is then
Under mild regularity assumptions (differentiability of the null and alternative distributions, boundedness and support assumptions; see elliott2022detecting for details), we can write the $p$-curve (density of $p$-values) in the absence of $p$-hacking as
For the purposes of testing for $p$-hacking, the properties of $g$ describe the set of distributions contained in the null hypothesis of no $p$-hacking. Tests can be based on deviations from this set. elliott2022detecting provide general sufficient conditions for when the $p$-curve is non-increasing, $g'\le 0$, and continuous when there is no $p$-hacking, allowing for tests of these properties of the distribution to be interpreted as tests of the null hypothesis of no $p$-hacking. These conditions hold for many possible distributions $F_h$ that arise in research, for example, normal, folded normal (relevant for two-sided tests), and $\chi^2$ distributions.
When $T$ is (asymptotically) normally distributed (for example, tests on means or regression parameters when central limit theorems apply), elliott2022detecting show that in addition to being non-increasing, the $p$-curves are completely monotonic (i.e., have derivatives of alternating signs so that $g''\ge 0$, $g'''\le 0$, etc.) and there are testable upper bounds on the $p$-curve and its derivatives.
We study the power of tests for $p$-hacking that exploit different combinations of these testable restrictions. In our simulation study, we consider four different types of tests: (i) tests for non-increasingness of the $p$-curve, (ii) tests for continuity of the $p$-curve, (iii) tests for upper bounds on the $p$-curve and its derivatives, and (iv) tests for combinations of monotonicity restrictions and upper bounds. We describe the individual tests and how we implement them in detail in Section (ref) and Table (ref).
If researchers do $p$-hack, the distribution of the reported $p$-value will depend on the functional form of the selection function $d(\cdot)$ and the joint distribution of $\{P_1, P_2, P_3,\dots \}$ given $h$. To the extent that this differs from the set of distributions under the null, it is then possible that the resulting $p$-curve violates the properties listed above in one way or another, providing the opportunity to test for $p$-hacking. This means that the power of any particular test for detecting $p$-hacking will depend on the functional form of $d(\cdot)$ and the joint distribution of $\{P_1, P_2, P_3,\dots \}$, in the sense that some tests have power against the types of changes to the distribution caused by $p$-hacking but others might not. Understanding the power of tests for $p$-hacking will thus depend on the testing problem as well as how researchers $p$-hack. It is for this reason we provide analytical results in the next section for a variety of testing problems and approaches to $p$-hacking.
We consider two general approaches to $p$-hacking. The first is the threshold approach, where the researcher constructs a test from their preferred model resulting in a $p$-value $P_1$, accepting this test if $P_1$ is below a target value $\alpha$ (for example, $0.05$). If the $p$-value does not achieve this goal value, additional models are considered. This is representative of the “intuitive” approach to $p$-hacking that is discussed in much of the literature on testing for $p$-hacking, where humps or discontinuities in the $p$-curve around common critical levels are examined. This results in a choice of $d(\cdot)$ such that
Other thresholding approaches could be considered, such as those in Section (ref). Tests for humps or discontinuities will have some power for detecting $p$-hacking based on the threshold approach, as will tests for violations of the upper bounds on the $p$-curve.
The second is the minimum approach, where researchers take the smallest $p$-value from a set of models. Here the choice of $d(\cdot)$ results in
Intuitively, we would expect that for this approach the distribution of $p$-values would shift to the left, be monotonically non-increasing, and there would be no expected hump in the distribution of $p$-values near commonly reported significance levels. Tests for non-increasingness and discontinuities will not have power. Instead the bounds derived in elliott2022detecting may be violated.
Ultimately, carefully considering the functions $d(\cdot)$ and distributions of $p$-values allows us to examine power against empirically relevant alternatives.
Our focus is on the power of testing for various types of $p$-hacking. However, empirical studies of $p$-hacking often focus on published $p$-values, which may be subject to publication bias andrews2019identification. Publication bias can impact the distribution of $p$-values in ways similar to $p$-hacking.
To see this effect, let $S=1$ if the paper is selected for publication and $S=0$ otherwise. By Bayes' Law, the $p$-curve conditional on publication, $g_{S=1}(p):=g(p\mid S=1)$, is
Here $\Pr(S=1\mid p)$ is the publication probability given $p$-value $P=p$, and $g^d(p)$ refers to the potentially $p$-hacked distribution of $p$-values when there is no publication bias.
Without publication bias, the publication probability does not depend on $p$, $\Pr(S=1\mid p)=\Pr(S=1)$, so that $g_{S=1}=g^d$. With publication bias, $\Pr(S=1\mid p)$ depends on $p$ so that $\Pr(S=1\mid p)$ is not be equal to $\Pr(S=1)$ for some $p\in (0,1)$ and $g_{S=1}\ne g^d$. In this case, one can detect $p$-hacking and/or publication bias if $g_{S=1}$ violates the testable restrictions underlying the statistical tests, so we can regard the tests here as joint tests of the absence of both $p$-hacking and publication bias. It is not possible in general to distinguish $p$-hacking from publication bias without additional assumptions.
It is plausible to assume that papers with smaller $p$-values are more likely to get published so that $\Pr(S=1\mid p)$ is decreasing in $p$. In this case, $g_{S=1}$ is non-increasing in the absence of $p$-hacking. Selection through publication bias that favors smaller $p$-values can result in steeper $p$-curves violating the bounds derived under the null hypothesis of no $p$-hacking. Hence rejections of bounds tests may well be exacerbated by publication bias. Discontinuities in $\Pr(S=1\mid p)$ can generate discontinuities in the absence of $p$-hacking, generating power for discontinuity tests. We examine these effects via Monte Carlo analysis in Section (ref).
For a power evaluation to be informative, relevant choices for the joint distribution of the $p$-values and methods for $p$-hacking need to be considered. We deal with each of these through the following choices:
We study the shape of the distribution of $p$-values under four arguably prevalent empirical problems with the potential for $p$-hacking. We focus on selecting control variables, selecting instruments, and variance bandwidth selection in the main text. We discuss selecting across datasets, which formally is a special case of the examples in the main text, in Appendix (ref). The ability to $p$-hack depends on the distribution of the $p$-values conditional on $h$, and the four examples allow us to study power in these relevant testing situations.\footnote{While we focus on the impact of $p$-hacking on the shape of the $p$-curve and the power of tests for detecting $p$-hacking, explicit models of $p$-hacking are also useful in other contexts. For example, mccloskey2023critical use a model of $p$-hacking to construct critical values that are robust to $p$-hacking.}
For analytical tractability, we focus on the case where researchers use one-sided tests in this section, for a limited number of options for $p$-hacking. Appendix (ref) provides analogous numerical results for two-sided tests. The analytical results provide a clear understanding of the opportunities for tests to have power and what types of situations the tests will have power in. In the simulation study in Section (ref), we consider generalizations of these analytical examples for two-sided tests, and we also show results for one-sided tests in Appendix (ref).
In addition to analyzing the effects of $p$-hacking on the shape of the $p$-curve, we study its implications for the bias of the estimates and size distortions of the tests reported by researchers engaged in $p$-hacking. We present these results in Appendix (ref). Appendix (ref) presents the derivations underlying all analytical results.
Linear regression has been suggested to be particularly prone to $p$-hacking hendry1980econometrics,leamer1983lets,bruns2016p,bruns2017metaregression. Researchers usually have available a number of control variables that could be included in a regression along with the variable of interest. Selection of various configurations for the linear model allows multiple chances to obtain a small $p$-value, perhaps below a threshold such as $0.05$. The theoretical results in this section yield a careful understanding of the shape of the $p$-curve when researchers engage in this type of $p$-hacking.
We construct a stylized model and consider the two approaches to $p$-hacking discussed above in order to provide analytical results that capture the impact of $p$-hacking. Suppose the researchers estimate the impact of a scalar regressor $X_i$ on an outcome $Y_i$. The data are generated as $Y_i = X_i\beta+U_i$, $ i=1,\dots,N,$ where $U_i\overset{iid}\sim\mathcal{N}(0,1)$. For simplicity, we assume that $X_i$ is non-stochastic. The researchers test the hypothesis $H_0:\beta=0$ against $H_1:\beta>0$.
In addition to $X_i$, the researchers have access to two additional non-stochastic control variables, $Z_{1i}$ and $Z_{2i}$.\footnote{For simplicity, we consider a setting where $Z_{1i}$ and $Z_{2i}$ do not enter the true model so that their omission does not lead to omitted variable biases bruns2016p. It is straightforward to generalize our results to settings where $Z_{1i}$ and $Z_{2i}$ enter the model: $Y_i = X_i\beta_1+Z_{1i}\beta_2 +Z_{2i}\beta_3 + U_i$.} We assume that $(X_i,Z_{1i},Z_{2i})$ are scale normalized so that $N^{-1}\sum_{i=1}^{N}X^2_i = N^{-1}\sum_{i=1}^{N}Z^2_{1i}=N^{-1}\sum_{i=1}^{N}Z^2_{2i}=1$. To simplify the exposition, we further assume that $N^{-1}\sum_{i=1}^{N}Z_{1i}Z_{2i}=\gamma^2$ and that $N^{-1}\sum_{i=1}^{N}X_iZ_{1i}= N^{-1}\sum_{i=1}^{N}X_iZ_{2i}= \gamma$, where $|\gamma|\in (0, 1)$.\footnote{We omit $\gamma = 0$, i.e., adding control variables that are uncorrelated with $X_i$, because in this case the $t$-statistics and thus $p$-values for each regression are equivalent and hence there is no opportunity for $p$-hacking of this form.} These assumptions are not essential for our analysis and could be relaxed at the expense of a more complicated notation. We let $h := \sqrt{N}\beta \sqrt{1-\gamma^2}$.
In terms of $p$-hacking, researchers regress $Y_i$ on $X_i$ and $Z_{1i}$ and compute the resulting $p$-value, $P_1$. An additional $p$-value, $P_2$, arises from regressing $Y_i$ on $X_i$ and $Z_{2i}$ instead of $Z_{1i}$. The reported $p$-value $P_r$ is obtained from equation (ref) under the threshold approach and equation (ref) under the minimum approach. Each approach results in different distributions of $p$-values, and, consequently, tests for $p$-hacking will have different power properties.
In Appendix (ref), we show that for the threshold approach the resulting $p$-curve is
where $\rho = 1-\gamma^2$, $z_h(p) = \Phi^{-1}(1-p) - h$, $\Phi$ is the standard normal cumulative distribution function (CDF), and
In interpreting this result, note that when there is no $p$-hacking, then $\Upsilon_1(p; \alpha, h, \rho)=1$. It follows from the properties of $\Phi$ that the threshold $p$-curve lies above the curve without $p$-hacking for $p \le \alpha$. We can also see that, since $\Phi\left(\frac{z_{h}(\alpha) - \rho z_{h}(p)}{\sqrt{1-\rho^2}}\right)$ is decreasing in $h$, for larger $h$ the difference between the threshold $p$-curve and the curve without $p$-hacking becomes smaller. This is intuitive since for a larger $h$, the need to $p$-hack diminishes as most of the studies find an effect without resorting to manipulation. Note that a larger $h$, ceteris paribus, corresponds to more “evidential value” in the terminology of simonsohn2014p.
The distribution of $p$-values for the minimum approach is equal to
For $p$-hacking of this form, the entire distribution of $p$-values is shifted to the left. For some values of $p$ less than 0.5, the curve lies above the no-$p$-hacking curve. This distribution is monotonically decreasing for all $\Pi$, so does not have a hump and remains continuous. Because of this, only the tests based on upper bounds and higher-order monotonicity have any potential for detecting $p$-hacking. If $\Pi$ is a point mass distribution, there is a range over which $g_1^{m}(p)$ exceeds the upper bound $\exp(z_0(p)^2/2)$ derived in elliott2022detecting, the upper end (largest $p$) of which is at $p=1-\Phi(h)$.
Panels (a) and (b) of Figure (ref) show the theoretical $p$-curves for various $h$ and $\gamma$.\footnote{Note that the results depend on $\gamma$ via $\rho=1-\gamma^2$. Therefore, the results do not depend on the sign of $\gamma$, and we only show results for positive values of $\gamma$.} In terms of violating the condition that the $p$-curve is monotonically decreasing, violations for the threshold case can occur but only for $h$ small enough. For $p<\alpha$, the derivative is
where $\phi$ is the standard normal probability density function (PDF). Note that $\rho$ is always positive and, when all nulls are true (i.e., when $\Pi$ assigns probability one to $h=0$), ${g^t}'(p)$ is positive for all $p\in (0,\alpha)$.\footnote{For $p>\alpha$, ${g^{t}_1}'(p)$ is negative and equal to ${g^t_1}'(p)=-\int_{\mathcal{H}}\frac{\phi(z_h(p))}{\phi^2(z_0(p))}\left(h+ \sqrt{\frac{1-\rho}{1+\rho}}\phi\left(z_h(p)\sqrt{\frac{1-\rho}{1+\rho}}\right)\right)d\Pi(h).$} This can be seen for the dashed line in the Panel (a) of Figure (ref). However, at $h=1$, this effect no longer holds, and the $p$-curve is downward sloping. From Panel (b) in Figure (ref), we see that violations of monotonicity are larger for smaller $\gamma$. When $\gamma=0.1$, the $p$-curve even becomes bimodal (one mode at 0 and one mode at 0.05).
Figure (ref) indicates that the threshold approach to $p$-hacking implies a discontinuity at $\alpha$. The size of the discontinuity is larger for larger $h$ and remains for each $\gamma$, although how that translates to power of tests for discontinuity also depends on the shape of the rest of the curve. We examine this in Monte Carlo experiments in Section (ref). [FIGURE (ref) HERE.]
Panel (c) of Figure (ref) compares the threshold and minimum approach to $p$-hacking. Results are presented for $h=1$ and $\gamma=0.5$, with respect to the bounds under no $p$-hacking. We also report the no-$p$-hacking distribution. Simply taking the minimum $p$-value as a method of $p$-hacking results in a curve that remains downward sloping and has no discontinuity --- tests for these two features will have no power against such $p$-hacking. But as Panel (c) of Figure (ref) shows, the upper bounds on the $p$-curve are violated for both methods of $p$-hacking. The violation in the thresholding case is pronounced. By contrast, the violation in the minimum case is barely visible, so that the power of tests for detecting this type of $p$-hacking based on violations of the upper bounds will be low, except in large samples.
Suppose that the researchers use an IV regression to estimate the causal effect of a scalar regressor $X_i$ on an outcome $Y_i$. The data are generated as
where $(U_i,V_i)'\overset{iid}\sim\mathcal{N}(0,\Omega)$ with $\Omega_{12}\ne 0$. The instruments are generated as $Z_i\overset{iid}\sim \mathcal{N}(0, I_2)$ and independent of $(U_i, V_i)$. The researchers test the hypothesis $H_0:\beta=0$ against $H_1:\beta>0.$ To simplify the exposition, suppose that ${\Omega_{11}=\Omega_{22}=1}$ and $\gamma_1=\gamma_2=\gamma$. We let $h := \sqrt{N}\beta|\gamma|$.
For $p$-hacking, researchers start by using all the available information (both instruments) to compute $P_1$ but can also generate $p$-values by only using the first instrument to compute $P_2$ or only the second instrument for $P_3$. The threshold and minimum choices for $P_r$ follow from equations (ref) and (ref).\footnote{In practice, it is likely that researchers also consider the first stage $F$-statistic when selecting instruments andrews2019weak,brodeur2020methods. We explore this in our Monte Carlo simulations.}
For the threshold approach, the $p$-curve (see Appendix (ref) for derivations) is
where
and $\zeta(p) = 1 - 2\Phi((1-\sqrt{2})z_0(p))$ and $D_h(p)=\sqrt{2}z_0(p)-2h$.
In Figure (ref), the $p$-curves for $h\in \{0,1,2\}$ are shown for the threshold approach in Panel (a). As in the covariate selection example, it is only for small values of $h$ that we see upward sloping curves and a hump below size. For $h=1$ and $h=2$, no such violation of non-increasingness occurs, and tests aimed at detecting such a violation will have no power. The reason is similar to that of the covariate selection problem --- when $h$ becomes larger many tests reject anyway, so whilst there is still a possibility to $p$-hack the actual rejections overwhelm the “hump” building of the $p$-hacking. For all $h$, there is still a discontinuity in the $p$-curve arising from $p$-hacking, so tests for a discontinuity at $\alpha$ will still have power. [FIGURE (ref) HERE]
For the minimum approach in equation (ref), the distribution of $p$-values is equal to
where
Panel (b) of Figure (ref) displays the $p$-curves for $h\in \{0,1,2\}$ for the minimum approach. There is no hump, as expected, and all the curves are non-increasing. Only tests based on upper bounds for and higher-order monotonicity of the $p$-curve have the possibility of rejecting the null hypothesis of no $p$-hacking in this situation. The upper bound is violated around the 0.05 threshold for $h = 0$ and at lower $p$-values for larger $h$.
Panel (c) in Figure (ref) shows the comparable figure for the IV problem as Panel (c) of Figure (ref) shows for the covariate selection example. The results are qualitatively similar across these examples, although quantitatively the $p$-hacked curves in the IV problem are closer to and more likely to exceed the bounds than in the covariates problem.
Overall, as with the case of covariate selection, both the relevant tests for $p$-hacking and their power will depend strongly on the range of $h$ relevant to the studies underlying the data employed for the tests.
In time series regression, sums of random variables such as means or regression coefficients are standardized by an estimate of the spectral density of the relevant series at frequency zero. A number of estimators exist; the most popular in practice is a nonparametric estimator that takes a weighted average of covariances of the data. With this method, researchers are confronted with a choice of the bandwidth for estimation. Different bandwidth choices allow for multiple chances at constructing $p$-values, hence allowing for the potential for $p$-hacking.
To examine this analytically, consider the model $Y_{t}=\beta+U_{t}$, $t=1,\dots,N$, where we assume that $U_{t}\overset{iid}\sim \mathcal{N}(0,1)$. Researchers consider two statistics for testing $H_0:\beta=0$ against $H_1:\beta>0$. First, the usual $t$-statistic, $T_1=\sqrt{N}\bar{Y}_{N}$, which generates a $p$-value $P_1$. Secondly, $T_2=(\sqrt{N}\bar{Y}_{N})/\hat\omega$, which generates a $p$-value $P_2$, where $\bar{Y}_{N}=N^{-1} \sum_{t=1}^N Y_t$, $\hat\omega^2:={\omega}^2(\hat\rho):=1+2\kappa \hat{\rho}$ and $\hat{\rho} = (N-1)^{-1} \sum_{t=2}^N\hat{U}_t \hat{U}_{t-1}$. Here ${\kappa}$ is the weight in the spectral density estimator. For example, in the newey1987simple estimator with one lag, $\kappa=1/2$. We again consider both the threshold approach to $p$-hacking (ref) as well as the minimum approach (ref).\footnote{If $\hat\rho$ is such that $\hat\omega^2$ is negative, the researcher always reports the initial result.}
In Appendix (ref), we show that the distribution of $p$-values has the form
with $\Upsilon_3(p; \alpha, h, \kappa)$ taking different forms over different parts of the support of the distribution. Define $l(p)=(2\kappa)^{-1} \left( \left(\frac{z_0(\alpha)}{z_0(p)} \right)^2-1 \right)$ and let $H_N$ and $\eta_N$ be the CDF and PDF of $\hat\rho$, respectively. Then we have
Panel (a) in Figure (ref) presents the $p$-curves for the thresholding case. Notice that, unlike the earlier examples, thresholding creates the intuitive hump in the $p$-curve at the chosen size (here $0.05$) for all of the values for $h$. Thus tests that attempt to find such humps may have power. Discontinuities at the chosen size and violations of the upper bounds also occur. [FIGURE (ref) HERE]
When the minimum over the two $p$-values is chosen, the $p$-curve is given by
where
Panel (b) in Figure (ref) presents the $p$-curves for the minimum case. When $p$-hacking works through taking the minimum $p$-value, the impact is to move the distributions towards the left, making the $p$-curves fall more steeply. Of interest is what happens at $p=0.5$, where taking the minimum (this effect is also apparent in the thresholding case) results in a discontinuity. The reason for this is that choices over the denominator of the $t$-statistic used to test the hypothesis cannot change the sign of the $t$-test. Within each side, the effect is to push the distribution to the left, so this results in a (small) discontinuity at $p=0.5$. This effect will extend to all methods where $p$-hacking is based on searching over different choices of variance-covariance matrices --- for example, different choices in estimators, different choices in the number of clusters (as we consider in the Monte Carlo simulations), etc. Panel (b) of Figure (ref) shows that for $h=1,2$, the bound is not reached, and any discontinuity at $p=0.5$ is very small. For $h=0$, the bound is slightly below the $p$-curve after the discontinuity.
In this section, we discuss several statistical tests for the null hypothesis of no $p$-hacking based on a sample of $n$ $p$-values, $\{P_i\}_{i=1}^n$.
Histogram-based Tests for Combinations of Restrictions. Histogram-based tests elliott2022detecting provide a flexible framework for constructing tests for different combinations of testable restrictions. Let $0=x_0<x_1<\cdots< x_J=1$ be an equidistant partition of $[0,1]$ and define the population proportions $\pi_j = \int_{x_{j-1}}^{x_j}g(p)dp$, $j=1,\dots, J.$ The main idea of histogram-based tests is to express the testable implications of $p$-hacking in terms of restrictions on the population proportions $(\pi_{1},\dots,\pi_J)$. For instance, non-increasingness of the $p$-curve implies that $\pi_{j} - \pi_{j-1}\le 0$ for $j=2,\dots,J$. More generally, elliott2022detecting show that $K$-monotonicity (i.e., derivatives with alternating signs up to order $K$) and upper bounds on the $p$-curve and its derivatives can be expressed as $H_0:~ A\pi_{-J}\le b$, for a matrix $A$ and vector $b$, where $\pi_{-J} := (\pi_1,\dots, \pi_{J-1})'$.\footnote{Here we incorporate the adding up constraint $\sum_{j=1}^J\pi_j=1$ into the definition of $A$ and $b$ and express the testable implications in terms of the “core moments” $(\pi_1,\dots, \pi_{J-1})$ instead of $(\pi_1,\dots, \pi_{J})$.}
To test this hypothesis, we estimate $\pi_{-J}$ by the vector of sample proportions $\hat\pi_{-J}$. The estimator $\hat\pi_{-J}$ is asymptotically normal with mean $\pi_{-J}$ so that the testing problem can be recast as the problem of testing affine inequalities about the mean of a multivariate normal distribution kudo1963,wolak1987,cox2022simple. Following elliott2022detecting, we use the conditional chi-squared test of cox2022simple, which is easy to implement and remains computationally tractable when $J$ is moderate or large.
Tests for Non-Increasingness of the $p$-Curve. A popular test for non-increasingness of the $p$-curve is the Binomial test simonsohn2014p,head2015extent, where researchers compare the number of $p$-values in two adjacent bins right below significance cutoffs. Under the null of no $p$-hacking, the fraction of $p$-values in the bin closer to the cutoff should be weakly smaller than the fraction in the bin farther away. Implementation is typically based on an exact Binomial test. Binomial tests are “local” tests that ignore information about the shape of the $p$-curve farther away from the cutoff, which often leads to no or low power in our simulations.\footnote{A “global” alternative is Fisher's test simonsohn2014p. We do not report results for Fisher's test since we found that this test has essentially no power for detecting the types of $p$-hacking we consider.}
In addition to the classical Binomial test, we consider tests based on the least concave majorant (LCM) elliott2022detecting.\footnote{LCM tests have been successfully applied in many different contexts carolan2005,beare2015nonparametric,fang2019refinements.} LCM tests are based on the observation that non-increasingness of $g$ implies that the CDF $G$ is concave. Concavity can be assessed by comparing the empirical CDF of $p$-values, $\hat{G}$, to its LCM $\mathcal{M}\hat{G}$, where $\mathcal{M}$ is the LCM operator. We choose the test statistic $\sqrt{n}\|\hat{G}-\mathcal{M}\hat{G}\|_\infty$. The uniform distribution is least favorable for this test kulikov2008distribution,beare2021least, and critical values can be obtained via simulations.
Tests for Continuity of the $p$-Curve. Continuity of the $p$-curve at pre-specified cutoffs $p=\alpha$ can be assessed using standard density discontinuity tests mccrary2008manipulation,cattaneo2020simple. Following elliott2022detecting, we use the approach by cattaneo2020simple with the automatic bandwidth selection implemented in the R-package rddensity cattaneo2021rddensity.
In this section, we investigate the finite sample properties of the tests in Section (ref) using a Monte Carlo simulation study based on generalizations of the analytical examples of $p$-hacking in Section (ref).
In all examples, researchers test a hypothesis about a scalar parameter $\beta$:
The results for one-sided tests of $H_0:\beta=0$ against $H_{1}:\beta> 0$ are similar. See Figure (ref).
Researchers may $p$-hack their initial results by exploring additional model specifications or estimators and report a different result of their choice. Specifically, we consider the two general approaches to $p$-hacking discussed in Sections (ref) and (ref): the threshold and the minimum approach. In what follows, we discuss the generalized examples of $p$-hacking in more detail.
Researchers have access to a random sample with $N=200$ observations generated as $Y_i=X_{i}\beta+u_i$, where $X_i\sim \mathcal{N}(0,1)$ and $u_i\sim \mathcal{N}(0,1)$ are independent of each other. There are $K$ additional control variables, $Z_i:=(Z_{1i},\dots, Z_{Ki})'$, generated as $Z_{ki}= \gamma_kX_{i}+\sqrt{1-\gamma_k^2} \epsilon_{Z_k,i}$ for $k=1,\dots,K$, where $\epsilon_{Z_k,i} \sim \mathcal{N}(0,1)$ and $\gamma_k\sim U[-0.8, 0.8]$. We set $\beta=h/\sqrt{N}$ with $h\in \{0,1,2\}$, and we also show results for $h\sim \hat\Pi$, where $\hat\Pi$ is the Gamma distribution fitted to the RCT subsample of the brodeur2020methods data.\footnote{The normalization $\beta=h/\sqrt{N}$ is motivated by $h$ being a local effect in our notation and the asymptotic variance of the OLS estimator from a regression of $Y_i$ on $X_i$ being equal to $1$ in this simple model.}
Researchers use one of two threshold approaches or a minimum approach to $p$-hacking.
Threshold approach (general-to-specific): Researchers regress $Y_i$ on $X_i$ and $Z_i$, test (ref), and obtain $P$. If $P\le 0.05$, they report it. If $P>0.05$, they regress $Y_i$ on $X_i$, trying all $(K-1)\times 1$ subvectors of $Z_i$. They report the smallest $p$-value if it is smaller than $0.05$. If it is larger than $0.05$, they continue and explore all $(K-2)\times 1$ subvectors of $Z_i$ etc. If all results are insignificant, they report the smallest $p$-value.
\noindentThreshold approach (specific-to-general): Researchers start by regressing $Y_i$ on $X_i$ only, test (ref), and obtain $p$-value $P$. If $P\le 0.05$, they report it. If $P>0.05$, they regress $Y_i$ on $X_i$, trying every component of $Z_i$ as control variable and select the result with the smallest $p$-value. If the smallest $p$-value is larger than $0.05$, they continue and explore all $2\times 1$ subvectors of $Z_i$ etc. If all results are insignificant, they report the smallest $p$-value.
Minimum approach: Researchers run regressions of $Y_i$ on $X_i$ and each possible configuration of covariates $Z_i$ and report the minimum $p$-value.
Figure (ref) shows the null distributions ($p$-curves without $p$-hacking), including estimates of the average power of the underlying studies, and the alternative distributions ($p$-curves with $p$-hacking).\footnote{To generate these distributions, we run the algorithm one million times and collect $p$-hacked and non-$p$-hacked results.} The threshold approach leads to a discontinuity in the $p$-curve and may lead to non-increasing $p$-curves and humps below significance thresholds. By contrast, the minimum approach generally leads to continuous and non-increasing $p$-curves. The distribution of $h$ is an important determinant of the shape of the $p$-curve. It affects the joint distribution of $p$-values under both $p$-hacking approaches and the probability that the researchers find significant results earlier under the threshold approaches. Finally, as expected, the violations of the testable restrictions are more pronounced for large $K$ (i.e., when researchers have many degrees of freedom).
Researchers have access to a random sample with $N=200$ observations generated as
The instruments $Z_i:=(Z_{1i},\dots, Z_{Ki})'$ are generated as $Z_{ki}= \gamma_k\xi_{i}+\sqrt{1-\gamma_k^2} \epsilon_{Z_k,i}$ for $k=1,\dots,K$, where $\xi_i\sim \mathcal{N}(0,1)$, $\epsilon_{Z_k,i} \sim \mathcal{N}(0,1)$, and $\gamma_k\sim U[-0.8, 0.8]$. Here $\xi_i, \epsilon_{Z_k, i}$, and $\gamma_k$ are independent for all $k$. Also, $\pi_k\overset{iid}\sim U[1, 3]$ for $k=1,\dots, K$. We set $\beta=h/(3 \sqrt{N})$ with $h\in \{0,1,2\}$ and also show results for $h\sim \hat\Pi$.\footnote{We additionally scale $h$ by 3 to make the average power of studies more comparable to the other examples.}
Researchers use either a threshold or a minimum approach to $p$-hacking.
Threshold approach: Researchers estimate the model using all instruments $Z_i$, test (ref), and obtain the $p$-value $P$. If $P\le 0.05$, they report it. If $P>0.05$, they try all $(K-1)\times 1$ subvectors of $Z_i$ as instruments and select the result corresponding to the smallest $p$-value. If the smallest $p$-value is larger than $0.05$, they explore all $(K-2)\times 1$ subvectors of $Z_i$ etc. If all results are insignificant, they report the smallest $p$-value.
Minimum approach: The researchers run IV regressions of $Y_i$ on $X_i$ using each possible configuration of instruments and report the minimum $p$-value.
Figures (ref) and (ref) display the null distributions, including estimates of the average power of the underlying studies, and the $p$-hacked distributions.\footnote{Unlike in the covariate selection example, we do not show results for $K=7$ since there is a very high concentration of $p$-values at zero in this case.} We also show these distributions for a scenario where the researchers screen out specifications with first-stage $F$-statistics below 10. In this case, the researchers ignore such specifications while doing thresholding or minimum type searches described above. As with covariate selection, the threshold approach yields discontinuous $p$-curves and may lead to non-increasingness and humps, whereas reporting the minimum $p$-value leads to continuous and decreasing $p$-curves. The distribution of $h$ and $K$ are important determinants of the shape of the $p$-curve.
We consider two different types of standard error selection: lag length selection as in Section (ref) and selecting the level of clustering, given the prevalence of clustered standard errors in empirical research cameron2015practitioner,mackinnon2023cluster.
Lag length selection. Researchers have access to a random sample with $N=200$ observations from $Y_t=X_t\beta+U_t$, where $X_t\sim \mathcal{N}(0,1)$ and $U_t\sim \mathcal{N}(0,1)$ are independent. We set $\beta=h/\sqrt{N}$ with $h\in \{0,1,2\}$ and also show results for $h\sim \hat\Pi$.
Researchers use either a threshold or a minimum approach to $p$-hacking.
\noindentThreshold approach. Researchers first regress $Y_t$ on $X_t$ and calculate the standard error using the classical Newey-West estimator with the number of lags selected using the Bayesian Information Criterion (they only choose up to $4$ lags). They then use a $t$-test to test (ref) and calculate the $p$-value $P$. If $P\le 0.05$, they report it. If $P>0.05$, they try the Newey-West estimator with one extra lag. If the result is not significant, they try two extra lags etc. If all results are insignificant, they report the smallest $p$-value.
Minimum approach. Researchers regress $Y_t$ on $X_t$, calculate the standard error using Newey-West with $0$ to $4$ lags and report the minimum $p$-value.
The null distributions, including estimates of the average power of the underlying studies, and the $p$-hacked distributions are displayed in Figure (ref). The threshold approach induces a sharp spike right below 0.05. This is because $p$-hacking via lag selection does not lead to huge improvements in terms of $p$-value.
Cluster level selection. Consider the same model as for the lag length selection. The researchers calculate cluster-robust standard errors from grouping data into 20, 40, 50, and 100 clusters. They also calculate standard errors without clustering (or equivalently with 200 clusters). They then report either the minimum $p$-value across clustering levels (minimum approach) or the first significant $p$-value found while searching through clustering levels starting from 20 clusters (threshold approach). The null distributions, including estimates of the average power of the underlying studies, and the $p$-hacked distributions are displayed in Figure (ref).
We model the distribution of reported $p$-values as a mixture, $g^{o}(p) = \tau \cdot g^d(p)+(1-\tau)\cdot g^{np}(p)$. Here, $g^d$ is the distribution under the different $p$-hacking approaches described above; $g^{np}$ is the distribution in the absence of $p$-hacking (i.e., the distribution of the $p$-value from the preferred specification). The parameter $\tau\in [0,1]$ captures the fraction of researchers who engage in $p$-hacking.
To generate the data, we first simulate the $p$-hacking algorithms one million times to obtain samples corresponding to $g^d$ and $g^{np}$. Then, to construct samples in every Monte Carlo iteration, we draw $n=5000$ $p$-values with replacement from a mixture of those samples. Appendix (ref) presents results for smaller sample sizes. Following elliott2022detecting, we apply the tests to the subinterval $(0, 0.15]$. Therefore, the effective sample size depends on the $p$-hacking strategy, the distribution of $h$, and the fraction of $p$-hackers $\tau$.\footnote{Note that the data-generating processes in Sections (ref)--(ref) imply that, for any $p > \alpha$, the number of observations in the interval $(0, p]$ is identical under both $p$-hacking approaches for covariate, IV, and cluster selection.}
We compare the finite sample performance of the tests described in Section (ref). See Table (ref) for more details.\footnote{For CS1, CSUB, and CS2B, the optimization routine fails to converge for some realizations of the data due to nearly singular covariance matrix estimates. We count these cases as non-rejections of the null in our Monte Carlo simulations.} The simulations are implemented using MATLAB MATLAB:2023 and R R2023. [TABLE (ref) HERE]
In this section, we present figures showing how power varies with the fraction of $p$-hackers, $\tau$, referred to as power curves, for the different data generating processes (DGPs). For covariate and instrument selection, we focus on the results for $K=3$ in the main text and present the results for larger values of $K$ in Appendix (ref). We focus on two-sided tests in the main text. Figure (ref) in Appendix (ref) presents the results for covariate selection with one-sided tests. The nominal level is 5%. All results are based on 5000 simulation draws. Figures (ref)--(ref) present the results. [FIGURES (ref)--(ref) HERE.]
The power for detecting $p$-hacking crucially depends on whether the researchers use a thresholding or a minimum approach to $p$-hacking, the econometric method, the fraction of $p$-hackers, $\tau$, and the distribution of $h$. When researchers $p$-hack using a threshold approach, the $p$-curves are discontinuous at the threshold, may violate the upper bounds, and may be non-monotonic. Thus, tests exploiting these testable restrictions may have power when the fraction of $p$-hackers is large enough.
CS2B, which exploits monotonicity restrictions and bounds, has the highest power overall. Among the tests that exploit monotonicity of the entire $p$-curve, CS1 typically exhibits higher power than LCM. LCM can exhibit non-monotonic power curves because the test statistic converges to zero in probability for strictly decreasing $p$-curves beare2015nonparametric.
The Binomial test often exhibits no or low power. The reason is that the $p$-hacking approaches we consider do not lead to isolated humps or spikes near $0.05$, even if researchers use a threshold $p$-hacking approach. There is one notable exception. When researchers engage in lag length selection, $p$-hacking based on the threshold approach can yield isolated humps right below the cutoff. By construction, the Binomial test is well-suited for detecting this type of $p$-hacking and is among the most powerful tests in this case. Our results for the Binomial test demonstrate the inherent disadvantage of using tests that only exploit testable implications locally. Such tests only have power against very specific forms of $p$-hacking, which limits their usefulness in practice.
Discontinuity tests are a useful complement to tests based on monotonicity and upper bounds because $p$-hacking based on threshold approaches often yields pronounced discontinuities. These tests are particularly powerful for detecting $p$-hacking based on lag length selection, which leads to spikes and pronounced discontinuities at $0.05$, as discussed above.
When researchers always report the minimum $p$-value, the power of the tests is much lower than when they use a threshold approach. The minimum approach to $p$-hacking does not lead to violations of monotonicity and continuity over $p\in (0,0.15]$. Therefore, by construction, tests based on these restrictions have no power, irrespective of the fraction of researchers who are $p$-hacking.
The minimum approach may yield violations of the upper bounds. The range over which the upper bounds are violated and the extent of these violation depend on the distribution of $h$ and the econometric method used. The simulations show power only for the tests based on upper bounds (CSUB and CS2B) and for covariate and IV selection when a sufficiently large fraction of researchers $p$-hacks. It is noteworthy that no test has power under the minimum approach when $h\sim \hat\Pi$. This is because $\hat\Pi$ puts mass on large values of $h$, which as explained below, can lead to no or low power.
Under the minimum approach, the power curves of CSUB and CS2B are quite similar, suggesting that the power of CS2B comes mainly from using upper bounds. This finding demonstrates the importance of exploiting upper bounds in addition to monotonicity and continuity restrictions in practice.
The relationship between the power of the tests and the value of $h$ need not be monotonic. The value of $h$ affects the shape of the $p$-curve in complicated ways, so that the power can be non-monotonic depending on the setting, testable restrictions, and specific test. Under the minimum approach, large values of $h$ lead to $p$-values close to zero, where the upper bounds are more difficult to violate. This is the reason why tests based on those upper bounds have very low or no power for $h=2$ (and $h\sim \hat\Pi$).
Finally, the results in Appendix (ref) show that the larger $K$ (i.e., the more degrees of freedom the researchers have when $p$-hacking) the higher the power of CSUB and CS2B.
Overall, the tests' ability to detect $p$-hacking is highly context-specific and can be low in some cases or lacking entirely. This is because $p$-hacking may not lead to violations of the testable restrictions used by the statistical tests for $p$-hacking. Moreover, even if $p$-hacking leads to violations of the testable restrictions, these violations may be small and can thus only be detected based on large samples of $p$-values. Regarding the choice of testable restrictions, the simulations demonstrate the importance of exploiting upper bounds in addition to monotonicity and continuity for constructing powerful tests against plausible $p$-hacking alternatives.
Here we investigate the impact of publication bias on the power of the tests for testing the joint null hypothesis of no $p$-hacking and no publication bias. We generate a sample of $n=5000$ $p$-values as in Section (ref) and keep each $p$-value, $P_i$, with probability $\Pr(S=1\mid P_i)$.
We consider two types of publication bias that differ with respect to how $\Pr(S=1\mid p)$ varies with $p$: sharp publication bias and smooth publication bias. Under sharp publication bias, $\Pr(S=1\mid p)$ is a step function: $\Pr(S=1\mid p)=1_{\{p\le 0.05\}}+0.1\times 1_{\{p> 0.05\}}$. Hence, significant results are 10 times more likely to be published than insignificant ones. Under smooth publication bias, we set $\Pr(S=1\mid p)=\exp(-A\cdot p)$, where we choose $A=8.45$ to make results comparable across both types of publication bias.\footnote{When $A=8.45$, the ratio between $\int_{0}^{0.05}\Pr(S=1\mid p)dp$ and $\int_{0.05}^{1}\Pr(S=1\mid p)dp$ is the same for both types of publication bias.}
Table (ref) shows the power of the tests for three levels of $p$-hacking ($\tau \in \{0,0.5,1\}$) with no publication bias, sharp publication bias, and smooth publication bias for covariate selection with $K=3$ and $h=0$. We show results for $h\in \{1,2\}$ and $h\sim \hat{\Pi}$ in Appendix (ref). The impact of publication bias on power depends on the testable restrictions that the tests exploit. The CSUB and CS2B tests have high power for detecting publication bias in the absence of $p$-hacking (when $\tau=0$), and publication bias substantially increases their power for rejecting the joint null when $p$-hacking alone is difficult to detect. This is expected since both forms of publication bias favor small $p$-values, which leads to $p$-curves that are more likely to violate the upper bounds, as discussed in Section (ref). [TABLE (ref) HERE.]
For the tests based on monotonicity of the entire $p$-curve (CS1 and LCM), the results depend on the type of publication bias. Sharp publication bias tends to increase power, whereas smooth publication bias can substantially lower power. Due to the local nature of the Binomial test, sharp publication bias does not increase its power. This again demonstrates the advantages of using “global” tests.
Sharp publication bias accentuates existing discontinuities and leads to discontinuities in otherwise smooth $p$-curves and thus increases the power of the discontinuity test. By contrast, smooth publication bias can decrease the power of the discontinuity test.
Overall, our results suggest that the presence of publication bias leads to higher power for tests based on upper bounds, especially when $p$-hacking is difficult to detect. For the other tests, whether publication bias increases or decreases the power depends on the exact nature of the publication bias: sharp publication bias tends to increase power, whereas smooth publication bias can decrease power.
In this section, we illustrate how our results can help guide the interpretation of empirical results from tests for $p$-hacking by reanalyzing the data in brodeur2020methods.\footnote{The data are form brodeur2022data. kudrin2022robust,kudrin2024jmp analyze some of the subsamples in brodeur2020methods using overlapping sets of tests. The differences between our results and theirs based on the same subsamples and tests are due to differences between V1 and V2 of the replication package.} brodeur2020methods study how $p$-hacking and publication bias vary by causal inference methods (difference-in-differences, RCT, regression discontinuity design, IV), strength of the instrument (whether the first-stage $F$-statistic is below or above 30), journal rank, and over time. We apply the tests for $p$-hacking in Table (ref) to 13 different subsamples (incl.\ the overall sample) analyzed by brodeur2020methods.\footnote{We do not consider the comparison between working and published papers in Section IV.B because the data on working papers only contain the $t$-statistics, so that we cannot apply our de-rounding strategy.} See Table (ref) for a list of the different subsamples.
An important practical issue is that the test results reported in research papers are typically rounded elliott2022detecting,kranz2022methods. Failure to de-round leads the tests to overreject. The reason is that rounding induces violations of the testable restrictions, such as discontinuities and non-monotonicities. The issue is particularly pronounced for the Binomial and the discontinuity test because it induces a mass point at $t=2$ ($p=0.046$) elliott2022detecting,kranz2022methods. For example, in the full sample, there are 258 observations with $t=2$, making up more than 37% of the 693 observations with $p\in [0.04,0.05]$ used by the Binomial test.
To document and correct for the impact of rounding, we present results based on the rounded original data in Panel (a) of Table (ref) and based on de-rounded data in Panel (b) of Table (ref). The de-rounded data are obtained by adding uniformly distributed noise from the interval $[-0.5, 0.5] \cdot 10^{dp}$ to each statistic (coefficient estimate, standard error, $t$-statistic, or $p$-value) recorded with $dp$ decimal places.\footnote{This de-rounding strategy is similar to brodeur2016. De-rounding the $p$-values preserves the monotonicity of the $p$-curve elliott2020detecting but could lead to violations of higher-order monotonicity that induce some power for tests targeting these restrictions. kranz2022methods propose an improved de-rounding method designed for Caliper tests. We leave an extension of their method to the tests we analyze here for future research.} To mitigate the impact of the added randomness from de-rounding, we present the results in terms of the average rejection rates over 1,000 draws of the de-rounding. In addition to reporting the results for each subsample and the overall sample, we also compute an overall rejection rate across all samples to get an estimate of the “empirical power” of the tests.\footnote{The dataset contains multiple tests per paper. We therefore adjust for clustering for the tests for which cluster-robust versions are available (i.e., CS2B, CSUB, and CS1).} [TABLE (ref) HERE.]
The results based on the rounded data (Panel (a)) suggest that there is substantial heterogeneity across the different subsamples. For example, all tests except CS1 reject at the 5% level in the DID subsample, whereas no test rejects in the $F\ge 30$ subsample. Overall, the Binomial test rejects in 77% of the subsamples, followed by CS2B, which rejects in 54% of the subsamples, whereas the other tests reject in less than 50% of the subsamples. De-rounding drastically changes the results and empirical conclusions. The overall rejection rates drop substantially for all tests. The overall rejection rate of the Binomial test drops from 75% to 0% after de-rounding, and only CSUB and CS2B have overall rejection rates larger than 10%.
The results in this paper can help explain these findings and clarify their interpretation. First, test results based on rounded data are misleading because the tests overreject. This is because the tests are often not able to distinguish a mass point at $p=0.046$ due to rounding only from $p$-hacking based on thresholding. As we show in this paper, many tests have some power to detect this type of $p$-hacking, which helps explain the rejections based on the rounded data. Second, once we focus on the de-rounded data, there is only limited empirical evidence for $p$-hacking: the average rejection rates are relatively low, even for the best tests. However, given the limited ability of tests to detect $p$-hacking we document in this paper, such findings do not imply that there is only limited $p$-hacking. Finally, there are substantial differences between the results of the different tests. These results are broadly consistent with our simulation results. For example, the most powerful tests (CS2B and CSUB) have the highest average rejection rates after de-rounding.
Concerns about $p$-hacking have motivated a fast-growing literature testing for $p$-hacking based on $p$-curves. Interpreting empirical work based on these tests requires a careful understanding of their ability to detect $p$-hacking. We examine how well existing tests are able to detect $p$-hacking in practice. Threshold approaches to $p$-hacking (where a predetermined significance level is targeted) result in $p$-curves that typically have discontinuities, $p$-curves that exceed upper bounds under no $p$-hacking, and less often violations of monotonicity restrictions. Many tests have some power to find such $p$-hacking, and the best tests are those exploiting both monotonicity and upper bounds and those based on testing for discontinuities. $p$-Hacking based on reporting the minimum $p$-value does not result in $p$-curves exhibiting discontinuities or monotonicity violations. While tests based on bound violations have some power, $p$-hacking based on mininum approaches is generally much harder to detect than $p$-hacking based on thresholding. The presence of publication bias (in addition to $p$-hacking) often increases but can also decrease the power of the tests for rejecting the joint null hypothesis of no $p$-hacking and no publication bias.
Overall, we find that tests for $p$-hacking can have low power in many settings and even no power in some cases. Therefore, failure to reject using tests for $p$-hacking based on $p$-values does not really indicate the absence of $p$-hacking. The lack of power we document has already motivated the development of new tests kudrin2024jmp, and we believe that this is a promising area for future research.
In this paper, we focus on tests of the null of no $p$-hacking (and no publication bias). From an empirical perspective, testing this null hypothesis is a useful starting point, and rejecting it can then help motivate further analyses and attempts to correct for selective reporting.\footnote{We view the role of this null hypothesis as similar to the role of sharp nulls in experiments. See, for example, the discussion in Chapter 5.11 of imbens2015causal.} These tests can also be useful to evaluate the impact of interventions to curb $p$-hacking and publication bias, such as mandating pre-registration, editorial statements blancoperez2020publication, or data-sharing policies brodeur2024phacking. From an econometric perspective, the advantage of focusing on the null of no $p$-hacking is that it allows for developing tests that control size absent any restrictions on how researchers $p$-hack. Methods that go beyond testing for the existence of $p$-hacking typically rely on specific models of $p$-hacking mccloskey2023critical or additional assumptions andrews2019identification,brodeur2020methods. A key finding of this paper is that even the power of the best tests can be low, despite the null of no $p$-hacking being strong and potentially unrealistic in some applications. This suggests that considering less stringent nulls could lead to tests with even lower power, absent strong additional assumptions.\footnote{An example would be the null of there being less than a certain fraction of $p$-hacking kudrin2024jmp.}
We end with two final remarks. First, there are important practical issues when testing $p$-hacking that our theoretical and Monte Carlo results do not directly speak to. Examples include how to optimally deal with rounding, how to define and select tests of interest, and how to ensure quality control when collecting large samples of $p$-values simonsohn2015blog.\footnote{In Appendix (ref), we show that mistakes in selecting the $p$-values that researchers target when $p$-hacking can substantially lower the power of tests for detecting $p$-hacking.} Second, we examine situations where the model is correctly specified, so estimators are consistent for their true values. For poorly specified models, for example, the omission of important variables that leads to omitted variables (confounding) effects, it is possible to generate a larger variation in $p$-values. Such problems with empirical studies are well understood and perhaps best found through theory and replication than meta-studies.