Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
60,414 characters · 13 sections · 40 citation commands
Detecting $p$-hacking
A researcher's ability to explore various ways of analyzing and manipulating data and then selectively report the ones that yield better-looking results, commonly referred to as $p$-hacking, compromises the reliability of research and undermines the scientific credibility of reported results. Absent systematic replication studies or meta analyses, a popular approach for assessing the extent of $p$-hacking is to examine distributions of $p$-values across studies, referred to as $p$-curves simonsohn2014p; see Section 2 in christensen2018transparency for a review.\footnote{Examples include: masicampo2012peculiar, leggett2013life, simonsohn2014p,simonsohn2015better, head2015extent, dewinter2015surge, and snyder2018sniff. Another strand of the literature uses the distribution of $t$-statistics to test for $p$-hacking gerber2008do,brodeur2016,brodeur2020methods,bruns2019reporting,vivalt2019specification.}
We consider the problem of testing the null hypothesis of no $p$-hacking against the alternative hypothesis of $p$-hacking and provide theoretical foundations for developing tests for $p$-hacking. We characterize analytically under general assumptions the null set of distributions of $p$-values implied in the absence of $p$-hacking and provide general sufficient conditions under which, for any distribution of the true effects, the $p$-curve is non-increasing and continuous in the absence of $p$-hacking. These conditions are shown to hold for many, but not all popular approaches to testing for effects.
For the leading case where $p$-curves are based on $t$-tests, we derive additional previously unknown testable restrictions. Specifically, the $p$-curves based on $t$-tests are completely monotone in the absence of $p$-hacking, and their magnitude and the magnitude of their derivatives are restricted by upper bounds. These restrictions are particularly useful when $p$-hacking fails to induce an increasing $p$-curve---for example when researchers engage in specification search across independent tests. In such cases tests based on non-increasingness have no power.
Our theoretical results allow us to develop more powerful statistical tests for $p$-hacking, which we apply to two large datasets of $p$-values. We find evidence for $p$-hacking in settings where the existing tests do not reject the null of no $p$-hacking.
When there is publication bias, our results characterize the $p$-curve under the null hypothesis of no $p$-hacking and no publication bias. Our tests become joint tests for $p$-hacking and publication bias, complementing available methods for identifying publication bias andrews2019identification.
Here we provide general sufficient conditions under which the $p$-curve is non-increasing under the null hypothesis of no $p$-hacking. These results are useful because tests for $p$-hacking often assume non-increasingness of the $p$-curve simonsohn2014p,simonsohn2015better,head2015extent. This assumption has been justified through analytical and numerical examples, which rely on specific choices of tests and distributions of true effects being tested hung1997,simonsohn2014p,ulrich2018some. However, such analyses are not sufficient for guaranteeing size control of statistical tests for $p$-hacking since the true effect distribution is never known. Instead, what is required for size control in a wide range of applications is a characterization of the shape of the $p$-curve for general tests and effect distributions.
Consider a test statistic $T$ that is distributed according to a distribution with cumulative distribution function (CDF) $F_{h}$, where $h$ indexes parameters of either the exact or asymptotic distribution of the test. We assume that the parameters $h$ only contain the parameters of interest. This is suitable for settings with large enough samples and asymptotically pivotal test statistics, which are prevalent in applied research.
Suppose researchers are testing the hypothesis
where $ \mathcal{H}_0 \cap \mathcal{H}_1 =\emptyset$. Let $ \mathcal{H} = \mathcal{H}_0 \cup \mathcal{H}_1 $. Denote as $F$ the CDF of the chosen null distribution from which critical values are determined. We assume that the test rejects for large values of the test statistic and denote the critical value for a level $p$ test as $cv(p)$. We will focus on settings with a continuous and strictly increasing $F$ (see Assumption (ref) below) and set $cv(p)=F^{-1}(1-p)$. For any $h$, we denote by $\beta\left(p,h \right) = \Pr\left(T>cv(p)\mid h\right)=1-F_{h}\left(cv(p)\right)$ the rejection rate of a level $p$ test with parameters $h$. For $ h \in \mathcal{H}_1$, this is the power of the test, and we refer to $\beta(p,h)$ as the power function.
For the remainder of the paper, we focus on settings where the tests generating the $p$-values satisfy Assumption (ref). This allows us to work with a well-defined density function and provide general results.
Assumption (ref) holds for many tests with parametric $F$ and $F_h$, including $t$-tests and Wald-tests. A necessary condition for Assumption (ref) is the absolute continuity of $F$ and $F_h$. This is not too restrictive since, in many cases, $F$ and $F_h$ are the asymptotic distributions of test statistics, which typically satisfy this condition. Further, in cases where the test statistics have a discrete distribution, size does not typically equal level, which could lead to $p$-curves that violate non-increasingness.
Consider the distribution of the $p$-values across studies, where we compute $p$-values from a distribution of $T$ given values of $h$, which themselves are drawn from a probability distribution $\Pi$. We refer to $\Pi$ as the distribution of true effects. The CDF of the $p$-values is
Under Assumption (ref), define the $p$-curve as follows.
In Section (ref), we analyze the shape of $g$ for general tests and distributions $\Pi$.
Here we derive conditions under which the $p$-curve is non-increasing in the absence of $p$-hacking for any distribution of true effects. We show that this property holds for most but not all popular statistical tests.
Under Assumption (ref), the curvature of the $p$-curve follows from
The sign of $g'(p)$ is determined by the second derivative of the rejection probability, $\partial^2 \beta\left(p,h\right)/\partial p^2$. As we will show in the proof of Theorem (ref) below, the following condition implies that $\partial^2 \beta\left(p,h\right)/\partial p^2$ is non-positive for all $h\in \mathcal{H}$.
Assumption (ref) is a restriction on how the power function changes when the critical value changes, which is governed by the shape of the density. When $\mathcal{H}_0=\{0\}$ and $F=F_0$ (as, for example, for one-sided $t$-tests), Assumption (ref) is of the form of a monotone likelihood ratio property, which relates the shape of the density of $T$ under the null to the shape of the density of $T$ under alternative $h$. The next lemma shows that this condition holds for many popular tests. Let $\Phi$ denote the CDF of the standard normal distribution.
The following theorem shows that the $p$-curve is non-increasing and continuously differentiable under the maintained assumptions for any distribution of true effects.
The result in Theorem (ref) holds for many commonly-used statistical tests such that, in many empirically relevant settings, the $p$-curve will be non-increasing in the absence of $p$-hacking. To our knowledge, Theorem (ref) provides the first general formal justification for the existing tests for $p$-hacking that exploit non-increasingness of the $p$-curve. Theorem (ref) further motivates the use of density discontinuity tests as an alternative to tests based on non-increasingness of the $p$-curve.
The results can be extended to settings with nuisance parameters. In such settings, $h$ contains both the parameters of interest, $h_1$, as well as additional nuisance parameters, $h_2$, such that $h=(h_1,h_2)$. Let $\mathcal{H}^1$ and $\mathcal{H}^2$ denote the supports of $h_1$ and $h_2$. Allow the null distribution to depend on $h_2$ with CDF $F_{h_2}$. The CDF of $p$-values becomes
where $\beta(p,h_1,h_2)=1-F_{h}\left(cv_{h_2}(p)\right)$ and $cv_{h_2}(p)=F_{h_2}^{-1}(1-p)$. The results of Theorem (ref) extend to the $p$-curve generated from this distribution after changing the notation to include the dependence on $h_2$. For $h_2\in \mathcal{H}^2$, $F_{h_2}$, $f_{h_2}, f_{h_2}'$ have the same properties as $F$, $f, f'$ in Assumption (ref), and the assumptions on $F_{h}$, $f_{h}, f_{h}'$ hold for $h=(h_1,h_2)$. Assumption (ref) becomes $f'_{h}(cv_{h_2}(p))f_{h_2}(cv_{h_2}(p))\ge f_{h_2}'(cv_{h_2}(p))f_{h}(cv_{h_2}(p))$ for $(h_1,h_2)\in \mathcal{H}^1\times \mathcal{H}^2$. The proof then follows directly from that of Theorem (ref).
In applications, often only a part of the $p$-curve is examined. The $p$-curve over subintervals $\mathcal{I}\subset (0,1)$ is given by $g_{\mathcal{I}}(p)=g(p)/\int_{\mathcal{I}}g(p)dp$ for $p\in \mathcal{I}$. Therefore, the results extend directly to this situation. Moreover, the $p$-curve constructed from a finite aggregation of different tests satisfying the assumptions of Theorem (ref) is continuously differentiable and non-increasing.
The assumptions of Theorem (ref) directly suggest $p$-curves for which the results of Theorem (ref) fail. For example, when the tests are non-similar, the $p$-curve can be non-monotonic in the absence of $p$-hacking, which arises through a violation of Assumption (ref). To illustrate, consider testing $H_0:h\le 0$ against $H_1:h> 0$ using a (non-similar) one-sided $t$-test, where $f$ is the density of the $\mathcal{N}(0,1)$ distribution and $f_{h}$ is the density of the $\mathcal{N}(h,1)$ distribution. It follows that $f'(x)/f(x)=-x$ and $f'_{h}(x)/f_{h}(x)=-(x-h)$, such that Assumption (ref) holds when $h \ge 0$ but is violated when $h<0$. Thus, when the weight in $\Pi$ on $h<0$ is large enough, the $p$-curve can be non-monotonic or increasing. For example, suppose that $\Pi$ is a normal distribution with mean $\mu$ and variance $1$, which places some mass on $h<0$, mixing increasing and decreasing $p$-curves. Figure (ref) shows that the resulting $p$-curve is non-increasing when $\mu=0$ and non-monotonic when $\mu=-2.5$.
We now show that for the leading case where $p$-curves are generated from $t$-tests with exact or asymptotic normal distributions, there are additional previously unknown testable restrictions. These restrictions allow us to develop more powerful statistical tests for $p$-hacking (see Section (ref)). In particular, these tests have power in situations where $p$-hacking does not lead to a violation of non-increasingness.
Consider first the problem of testing a one-sided hypothesis
where $h$ is a scalar, $\mathcal{H}_0=\{0\}$, and $\mathcal{H}_1=(0,\infty)$. We assume that $T\sim \mathcal{N}(h,1)$. This holds when using one-sided $t$-tests to test a hypothesis concerning a scalar parameter $\theta$: $H_0: \theta=\theta_0$ against $H_1: \theta > \theta_0.$ Let $\sqrt{N}\left(\hat\theta - \theta \right) \sim \mathcal{N}\left(0,\sigma^2\right)$, where $\hat\theta$ is an estimator of $\theta$ based on $N$ observations and $\sigma^2$ is assumed to be known. Denote the usual $t$-statistic as $\hat{t}$ and set $T=\hat{t}$. Defining $h:=\sqrt{N}\left((\theta-\theta_0)/\sigma\right)$ this fits (ref). More generally, testing problems with limiting normal experiments employed to test hypotheses of the form (ref) are common in empirical work (e.g., a one-sided test of a regression parameter using normal critical values).
The chosen null distribution is the standard normal distribution, $F=\Phi$. A level $p$ test rejects the null hypothesis when $T$ is larger than $cv_1(p):=\Phi^{-1}\left(1-p \right)$. Note that $cv_1(p)\ge 0$ for $p\in\left(0,1/2\right]$. Then $\beta\left(p,h \right) = 1-\Phi \left(cv_1(p)-h \right)$ and the CDF of $p$-values is
We also consider the two-sided version of this test. Here the hypothesis is
with $\mathcal{H}_0=\{0\}$ and $\mathcal{H}_1= \mathbb{R}\backslash \{0\}$. The two-sided test statistic $T$ is assumed to have a folded normal distribution. This holds when using a two-sided $t$-test with $T=|\hat{t}|$ for testing a two-sided hypothesis about $\theta$ : $H_0: \theta=\theta_0$ against $H_1: \theta \ne \theta_0$. More generally, testing problems with limiting normal experiments employed to test hypotheses of the form (ref) are also common in empirical work.
The chosen null distribution is the half normal distribution with scale parameter $1$. A level $p$ test rejects the null hypothesis when $T$ is larger than $cv_2(p):=\Phi^{-1}\left(1-\frac{p}{2}\right)$. The CDF of the $p$-values is
In addition to the results of Section (ref), previously unknown testable restrictions for $p$-curves based on $t$-tests follow from the shape of the power functions for these tests. These additional restrictions enable us to better pin down the space of potential $p$-curves when there is no $p$-hacking, allowing us to construct more powerful statistical tests for $p$-hacking. They also enable distinguishing non-increasing $p$-curves, which can arise from certain types of $p$-hacking, from curves where there is no $p$-hacking.
The $p$-curve based on one-sided $t$-tests testing hypothesis (ref) is
For two-sided $t$-tests testing hypothesis (ref), the $p$-curve is
Our next theorem shows that the $p$-curves (ref) and (ref) are completely monotone. A function $\xi$ is completely monotone on an interval $\mathcal{I}$ if $0\le (-1)^{k}\xi^{(k)}(x)$ for every $x\in \mathcal{I}$ and all $k=0,1,2,\dots$, where $\xi^{(k)}$ is the k$^{th}$ derivative of $\xi$.
Complete monotonicity yields additional restrictions that can be exploited to improve the power of statistical tests for $p$-hacking. Whilst available for one- and two-sided $t$-tests, not all tests yield completely monotonic $p$-curves. For example, a direct calculation shows that complete monotonicity may fail for tests based on $\chi^2$ distributions with more than two degrees of freedom (e.g., Wald tests).
The next theorem presents additional testable restrictions in the form of upper bounds on the $p$-curves and their derivatives.
As with the results in Theorem (ref), the results in Theorem (ref) yield additional restrictions, allowing more powerful tests for $p$-hacking.\footnote{One can use similar arguments as in Theorem (ref) to derive bounds for $p$-curves based on other specific tests such as Wald tests.} The bounds in Theorem (ref) do not only rule out large humps around significance cutoffs such as 0.01, 0.05, and 0.1 but also restrict the magnitude of the $p$-curves near zero. For the two-sided test, tests for $p$-hacking can be either constructed using the sharper (but not explicit) bound $\tilde{\mathcal{B}}^{(0)}_{2}(p)$ or the simpler explicit bound $\exp\left(\frac{cv_2(p)^2}{2}\right)$.
The bounds of Theorem (ref) are particularly useful when $p$-hacking fails to induce an increasing $p$-curve, a situation where tests based on non-increasingness of the $p$-curve have no power. Intuitively we might suspect this happens when all researchers $p$-hack but this simply shifts mass of the $p$-curve to the left, rather than inducing humps. A concrete example is when researchers run a finite number of $M>1$ independent analyses and report the smallest $p$-value, for example, when engaging in specification search across independent subsamples or data sets. The resulting $p$-curve under $p$-hacking is $g^{p}(p;M) = M(1-G^{np}(p))^{M-1}g^{np}(p)$, where $G^{np}$ and $g^{np}$ are the CDF and density of $p$-values in the absence of $p$-hacking.\footnote{This generalizes the example in ulrich2015p, who studied the special case where all null hypotheses are true such that $G(p)=p$.} Note that $g^{p}$ is non-increasing (completely monotone) whenever $g^{np}$ is non-increasing (completely monotone).\footnote{Since the products of completely monotone functions are completely monotone, complete monotonicity of $g^{p}(p;M)$ follows from complete monotonicity of $1-G^{np}(p)$ and $g^{np}(p)$.} Thus, $g^p$ will not violate the testable implications of Theorems (ref)--(ref), so tests based on these restrictions do not have power. However, $g^{p}$ can violate the bounds in Theorem (ref) whenever $M(1-G^{np}(p))^{M-1}>1$. For example, consider the one-sided case and let $\Pi$ be a half-normal distribution with scale parameter 1. Figure (ref) shows that $g^{p}$ violates the upper bound in Theorem (ref) to an extent that depends on $M$.
Upper bounds also help with testing for $p$-hacking with non-similar tests. In Section (ref), we show that non-increasingness may fail for non-similar one-sided $t$-tests, in which case tests of $p$-hacking based on non-increasingness may well reject because of non-similarity rather than $p$-hacking. Since upper bounds can also be derived for non-similar tests, we can still use bounds on the $p$-curve and its derivatives to test for $p$-hacking.\footnote{For instance, for $p\le 1/2$, the upper bound on the $p$-curve for non-similar one-sided $t$-tests coincides with that in Part (i) of Theorem (ref).}
Finally, the characterizations in Theorems (ref)--(ref) imply related characterizations of $p$-curves over subintervals $\mathcal{I}\subset (0,1)$, $g_{s,\mathcal{I}}(p)=g_s(p)/\int_{\mathcal{I}}g_s(p)dp$. In particular, complete monotonicity of $g_s$ implies the complete monotonicity of $g_{s,\mathcal{I}}$, because the sign of $g^{(k)}_{s,\mathcal{I}}$ equals the sign of $g_s^{(k)}$ for $k=0,1,2\dots$. Moreover, (conservative) upper bounds on $g_{s,\mathcal{I}}(p)$ for $\mathcal{I}=(0,\alpha]$ are given by the upper bounds in Theorem (ref), re-scaled by $\alpha$ since $G_s(\alpha)\ge \alpha$ for $s=1,2$.
Here we consider tests for $p$-hacking based on a sample of $n$ $p$-values. We consider three types of tests that differ with respect to the specification of the null hypothesis (the null space of $p$-curves). As a result, the different tests will differ with respect to the violations of the null of no $p$-hacking that they are able to detect.
In the absence of publication bias, our tests are tests for $p$-hacking; when there is also publication bias, they are joint tests for $p$-hacking and publication bias in general.
Theorem (ref) shows that, under general conditions, the $p$-curve is non-increasing. Consider the following testing problem
Popular tests based on hypothesis testing problem (ref) include the Binomial test simonsohn2014p,head2015extent and Fisher's test simonsohn2014p. Here we describe two alternative and more powerful tests.
Histogram-based tests. Let $0=x_0<x_1<\cdots< x_J=1$ be an equidistant partition of the unit interval. Define the population proportions as $\pi_j := \int_{x_{j-1}}^{x_j}g(p)dp, ~ j=1,\dots, J$. When $g$ is non-increasing, $\Delta_j :=\pi_{j+1} - \pi_j$ is non-positive for all $j=1,\dots,J-1$. Thus, the null hypothesis in testing problem (ref) can be reformulated as $H_0: \Delta_j\le 0$ for all $j=1,\dots,J-1$. To test this hypothesis, we apply the conditional chi-squared test of cox2020simple. We describe the implementation of this test in Section (ref) and Appendix (ref), where we propose more general tests that nest the histogram-based test for non-increasingness.
LCM test based on concavity of the CDF of $p$-values. Under the null hypothesis (ref), the CDF of $p$-values is concave. This observation allows us to apply tests based on the least concave majorant (LCM) carolan2005,beare2015nonparametric,fang2019refinements. LCM-based tests assess concavity of the CDF based on the distance between the empirical CDF of $p$-values, $\hat{G}$, and its LCM, $\mathcal{M}\hat{G}$, where $\mathcal{M}$ is the LCM operator.\footnote{For a function $f$, the LCM operator is defined as $\mathcal{M}f=\inf \{g: g \text{ is concave and } f\le g\}$ beare2015nonparametric.} We consider the test statistic $T = \sqrt{n}\| \mathcal{M}\hat{G} - \hat{G} \|_\infty$. The uniform distribution is least favorable for LCM tests kulikov2008distribution,beare2021least, in which case $T$ converges weakly to $\| \mathcal{M}B - B \|_\infty$, where $B$ is a standard Brownian Bridge on $[0,1]$.
Theorem (ref) shows that the $p$-curve is continuous in the absence of $p$-hacking. Tests for continuity of the $p$-curve at significance thresholds $\alpha$ such as $\alpha=0.05$, thus, provide an alternative to the tests based on non-increasingness of the $p$-curve. Consider the following testing problem:
Testing (ref) requires estimating two densities at the boundary point $\alpha$. Traditional kernel density estimators are not suitable for this task because they suffer from boundary bias karunamuni2005boundary. A popular approach to overcome this problem is to use local linear density estimators that rely on prebinning the data mccrary2008manipulation. We apply the density discontinuity test of cattaneo2020simple with data-driven bandwidth selection cattaneo2020rddensity, which is based on boundary adaptive local polynomial density estimators and avoids prebinning.
Theorem (ref) shows that $p$-curves based on $t$-tests are completely monotone, and Theorem (ref) establishes upper bounds on the $p$-curves and their derivatives. Here we develop tests based on these testable restrictions.
We say a function $\xi$ is $K$-monotone on some interval $\mathcal{I}$ if $0\le (-1)^{k}\xi^{(k)}(x)$ for every $x\in \mathcal{I}$ and all $k=0,1,\dots,K$, where $\xi^{(k)}$ is the k$^{th}$ derivative of $\xi$. By definition, a completely monotone function is $K$-monotone. Consider the null hypothesis
where $s=1$ for one-sided $t$-tests, $s=2$ for two-sided $t$-tests, and $\mathcal{B}^{(k)}_s$ is defined in Theorem (ref). Hypothesis (ref) implies restrictions on the population proportions $\boldsymbol{\pi}:=(\pi_1,\dots,\pi_J)'$, which can be expressed as $H_0: A\boldsymbol{\pi}_{-J}\le b$, where $\boldsymbol{\pi}_{-J} := (\pi_1,\dots, \pi_{J-1})'$.\footnote{The upper bounds on $\boldsymbol{\pi}$ implied by hypothesis (ref) are not sharp in general. Sharp bounds can be obtained by directly extremizing the proportions and their differences; see Appendix (ref).} The matrix $A$ and vector $b$ are defined in Appendix (ref).\footnote{We use $\boldsymbol{\pi}_{-J}$ because the variance matrix of the estimator of $\boldsymbol{\pi}$ is singular by construction and we want to express the left-hand side of our moment inequalities as a combination of “core” moments.}
We estimate $\boldsymbol{\pi}_{-J}$ using the sample proportions $\hat{\boldsymbol{\pi}}_{-J}$.\footnote{Given a sample of $n$ $p$-values, $\{P_i\}_{i=1}^n$, the sample proportions are defined as $\hat\pi_i=\frac{1}{n}\sum_{i=1}^n1\{x_{i-1}< P_i \le x_i\}$, $i=1,\dots,J$.} This estimator is $\sqrt{n}$-consistent and asymptotically normal with mean $\boldsymbol{\pi}_{-J}$ and non-singular (if all proportions are positive) covariance matrix $\Omega = \text{diag}\{{\pi}_1, \dots, \pi_{J-1}\} - \boldsymbol{\pi}_{-J}\boldsymbol{\pi}_{-J}'$. Following cox2020simple, we test the null by comparing $T = \inf_{q:\: Aq\le b}n (\hat{\boldsymbol{\pi}}_{-J} - q)'\hat{\Omega}^{-1} (\hat{\boldsymbol{\pi}}_{-J} - q) \label{eq: CS_stat}$ to the critical value from a $\chi^2$ distribution with $rank(\hat{A})$ degrees of freedom, where $\hat{A}$ is the matrix formed by the rows of $A$ corresponding to active inequalities.
The analyses were done using R R20 and Stata stata2019.
Here we reanalyze the data collected by brodeur2016, which contain information about 50,078 $t$-tests from 641 papers published in the AER, QJE, and JPE 2005--2011 brodeur2016data. We convert $t$-statistics into $p$-values associated with two-sided $t$-tests based on the standard normal distribution.\footnote{The original data contain $p$-values for less than 10% of observations. Where available, we work with the reported $p$-values.} After excluding observations with missing information, there are 49,838 tests from 640 papers.
Because the $p$-values may be correlated within papers, we use cluster-robust estimators of the variance of the sample proportions for the cox2020simple tests. In addition, we apply all tests to random subsamples with one $p$-value per paper, allowing us to use exact tests in the presence of within-paper correlation. To test for $p$-hacking, we focus on $p$-values smaller than $0.15$. We consider a Binomial test on $[0.04,0.05]$, Fisher's test, a histogram-based test for non-increasingness (CS1), a histogram-based test for $2$-monotonicity and bounds on the $p$-curve and the first two derivatives (CS2B), the LCM test, and a density discontinuity test at 0.05.\footnote{For the Binomial test, we split $[0.04,0.05]$ into two subintervals $[0.04,0.045]$ and $(0.045,0.05]$. Under the null of no $p$-hacking, the fraction of $p$-values in $(0.045,0.05]$ should be smaller than or equal to 0.5, which we assess using an exact Binomial test. For CS1 and CS2B, we use 30 bins when testing based on all $p$-values and 15 bins when testing based on random subsamples of $p$-values.}
Figure (ref) shows the results before and after de-rounding and based on the full sample and random subsamples. There is a large number of very small $p$-values, which is sometimes interpreted as indicative of evidential value (e.g., simonsohn2014p; in our notation, this is a large mass of $\Pi$ away from zero). The data exhibit a noticeable mass point at $\hat{t}=2$ (there are 427 such observations), which translates into a mass point in the $p$-curve at $p=0.046$.\footnote{This mass point could be due to low precision reporting brodeur2016, but also due to $p$-hacking, publication bias, or a combination thereof.} To analyze the impact of rounding, we also apply the tests to the de-rounded data provided by brodeur2016.\footnote{The de-rounded data were constructed by randomly redrawing estimates and standard errors; see Section II in brodeur2016 for a detailed description.}
In what follows, we say that a test rejects the null of no $p$-hacking if its $p$-value is smaller than $0.1$. Based on the original raw (rounded) data on all $p$-values, all tests reject the null except Fisher's test and the density discontinuity test. There are no rejections based on the random subsample, suggesting that the tests may be underpowered in small samples.
We find different results based on the de-rounded data.\footnote{Note that the (sub)sample sizes for the rounded and de-rounded data differ due to de-rounding.} There are no rejections based on the full sample of $p$-values. This finding suggests that the rejections based on the raw data are mainly due to the mass point just below $0.05$ and shows that de-rounding may substantially affect empirical conclusions.
Based on the random subsample of de-rounded $p$-values, only the CS2B test rejects the null of no $p$-hacking. The CS1 test comes close to rejecting ($p=0.11$). These two tests yield the smallest $p$-values across all four samples.
Here we reanalyze the data collected by head2015extent, which contain $p$-values obtained from text-mining open access papers in the PubMed database head2016data. There are $p$-values from 21 different disciplines. We focus on biology, chemistry, education, engineering, medical and health sciences, and psychology and cognitive science. The data contain $p$-values from the abstracts and the results sections in the main text. We use $p$-values from the results sections, allowing us to work with larger samples and present results for $p$-values smaller than $0.15$.
Since the data do not only contain $t$-tests, we consider tests based on non-increasingness and continuity of the $p$-curve (Theorem (ref)): a Binomial test on $[0.04,0.05]$, Fisher's test, a histogram-based test for non-increasingness (CS1), the LCM test, and a density discontinuity test at 0.05.\footnote{For CS1, we use 60 bins (all data) and 30 bins (random subsamples) for biological and medical and health sciences given the large sample sizes, and 30 and 15 bins for the other disciplines.} To account for within-paper dependence of $p$-values, we use a cluster-robust variance estimator for the CS1 test, and also present results based on random subsamples with one $p$-value per paper.
The left panel of Figure (ref) shows a histogram of the raw data on all $p$-values for the medical and health sciences (the largest subsample). A substantial fraction of $p$-values is rounded to two decimal places, which results in sizable mass points at $0.01, 0.02, \dots, 0.15$. Rounding makes the $p$-curve non-monotonic and discontinuous even in the absence of $p$-hacking and, thus, invalidates the testable restrictions in Theorem (ref). Therefore, we also show results based on de-rounded data.\footnote{We de-round the data as follows. To each observed $p$-value rounded up to the k$^{th}$ decimal point we add a random number generated from the uniform distribution supported on the interval $[\underline{u}, 0.5] \cdot 10^{-k}$, where $\underline{u}=0$ for zero $p$-values and $\underline{u}=-0.5$ for non-zero $p$-values.} In an earlier version of this paper elliott2020detecting, we show that de-rounding restores the non-increasingness but not the continuity of the $p$-curve. The right panel of Figure (ref) shows the impact of de-rounding on the shape of the $p$-curve. We note that density discontinuity tests are poorly suited here because rounding induces substantial discontinuities, which remain even after de-rounding. This means that rejections of the null can be either due to rounding or due to $p$-hacking.
In what follows, define a rejection of the null of no $p$-hacking for $p$-values smaller than $0.1$. Table (ref) presents the results for the full sample of $p$-values. For the original (rounded) data, the CS1 and the LCM test reject the null for all disciplines. De-rounding leads to fewer rejections. The CS1 test only rejects for biological sciences, engineering, and medical and health sciences; the LCM test rejects for medical and health sciences. This shows that rounding and de-rounding can substantially affect empirical results. The Binomial and Fisher's test do not reject the null for any discipline, which demonstrates the importance of using our more powerful tests.
Table (ref) shows the results based on random samples with one $p$-value per paper. We find that the CS1 test (biological sciences, engineering, medical and health sciences) and the LCM test (all disciplines except chemical sciences) reject the null based on the rounded data. None of the tests based on non-increasingness rejects the null based on the de-rounded data. A comparison to the results based on all $p$-values shows that the sample sizes required for detecting $p$-hacking may be quite large.
Finally, the density discontinuity test rejects for at least three disciplines based on the full sample and the random subsamples. After de-rounding, it only rejects for biological sciences (full sample) and chemical sciences (random subsample). These rejections are expected because of the prevalence of rounding-induced discontinuities.
We provide theoretical foundations for testing for $p$-hacking based on the distribution of $p$-values across scientific studies. We establish general results on the $p$-curve, providing conditions under which a null set of $p$-curves can be shown to be non-increasing. For $p$-values based on $t$-tests, we derive previously unknown additional restrictions on the $p$-curve when there is no $p$-hacking. These restrictions lead to the suggestion of more powerful tests that can be used to test the absence of $p$-hacking. A reanalysis of two datasets from the literature shows that the new tests based on additional restrictions are useful in testing for $p$-hacking.