Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
55,545 characters · 12 sections · 33 citation commands
When is $p$-hacking detectable?
\onehalfspacing
\noindentKeywords: $p$-Hacking, Deconvolution
\noindentJEL Codes: C12, C14
\pagenumbering{arabic}
Publication bias and $p$-hacking threaten the validity of empirical research. When researchers report only the most surprising statistical results, they are likely to over-claim. Scientific findings are therefore more trustworthy when they are drawn from literatures where selective reporting is not evident.
Learning what kinds of studies are $p$-hacked helps evaluate empirical claims and target reform. Much recent work has tested for selective reporting in a given literature by checking for distortions in the empirical distribution of reported $t$-scores.\footnote{See for example: Simonsohn,Head,Brodeur,havranek_housing,bb,Elliott,elliott2024powertestsdetectingphacking,Kudrinjmp,kudrinjmp24,Havranek24. The $p$-curve is distorted if and only if the $t$-curve is distorted. } All of these methods share a common shortcoming: They can be evaded by selective reporting strategies that do not violate the specific restrictions that the meta-analyst tested. For example, the Caliper test can be evaded by $p$-hacking that does not produce a discontinuity in the histogram of $t$-scores at the critical value. The bounds tested in Elliott can be evaded by $p$-hacking that distorts the overall shape of the $p$-curve but does not make its derivatives extreme at any particular point.
This paper proposes a new test that can only be evaded by forms of $p$-hacking that are totally undetectable by any examination of the $t$-curve. To our knowledge this is the first test that comes accompanied by such a guarantee. Our proposed test statistic is the distance between the observed $t$-curve and the set of all feasible $t$-curves under no selective reporting. The testable restriction that this distance equals zero is {\it sharp} because if a p-hacking process did not violate it, then the resulting $t$-curve could be explained without selective reporting by some distribution of true treatment effects.
Our test can only be evaded in large meta-samples by $p$-hacking processes that are undetectable in principle. We illustrate with a simple example: when a researcher reports the maximum of two correlated one-sided $t$-scores and the distribution of true effects is smooth, then the resulting $t$-curve can be explained either with or without $p$-hacking. This is an example of a kind of selective reporting that is undetectable by any examination of the $t$-curve. Outside of such undetectable cases, our test is consistent against all forms of selective reporting.
We call our test a projection test. The key technical step is to compute the test statistic using the Singular Value Decomposition (SVD). The SVD allows us to rewrite the projection of one smoothed density onto a set of other densities as simply the orthogonal projection of a vector of sample means onto a convex set of vectors in Euclidean space. Critical values are computed using the bootstrap.
Simulations show that the projection test is as powerful as the most sensitive of the six tests studied by Elliott. Further simulations show an example where $p$-hacking is nearly undetectable, the projection test is consistent against it while the six tests studied by Elliott all have power no greater than size in a large meta-sample. Therefore the projection test makes a meaningful contribution to the existing body of tests for $p$-hacking.
The second major advantage of the projection test is that the test statistic can still be interpreted when the $t$-scores reported by individual researchers are not exactly normally distributed. Recent nonparametric tests for $p$-hacking that examine every part of the $t$-curve like kudrinjmp24 and Elliott are highly sensitive but will reject the null hypothesis of no selective reporting whenever the empirical $t$-curve is not smooth enough to have been generated by a convolution with a normal kernel. Yet, non-smoothness in the $t$-curve can potentially arise for other reasons not related to selective reporting. For example, the $t$-score is more accurately described as following a student-$t$ distribution with $\nu$ degrees of freedom, where $\nu$ is the effective sample size of the experiment. In addition, the $t$-score is often reported after rounding. Even when de-rounded, this process introduces non-smoothness in the distribution of the $t$-score that could theoretically lead to spurious rejections of the null of no p-hacking.
We account for finite-sample deviations from the normal distribution in the following way. We first calculate the {\it Breakdown Statistic} $\widehat{B}$ defined as how severe the deviation from normality must be to overturn the meta-analyst's detection of $p$-hacking. The breakdown statistic is computed from the observed $t$-curve and is equal to the difference between the test statistic and the critical value. We then compare the breakdown statistic to the maximum possible deviations from normality that could arise from either the student-$t$ distribution or from (de)rounding. If the breakdown statistic is larger than could be explained by either of these sources of non-smoothness, then the rejection stands.
We apply our method to the meta-sample of 21,740 two-sided $t$-tests reported by 684 top economics journals collected by bb. We find that the $t$-curves for RCTs, DIDs, and IVs are significantly distorted and that these distortions are too large to be explained either by (1) rounding or de-rounding of the reported $t$-scores or by the limited degrees of freedom of the studies themselves. We conclude that significant selective reporting is likely present in all three subgroups of studies. We cannot say the same of RDDs.
This paper is organized as follows. Section (ref) sets up the problem. Section (ref) defines the meta-analyst's null hypothesis and provides examples of $p$-hacking processes that are detectable and undetectable by any test. Section (ref) presents the projection test and shows its validity and consistency when the $t$-score is exactly normal. Section (ref) extends to situations where the $t$-score is only approximately normal even in the absence of selective reporting. Section (ref) presents simulations and Section (ref) shows an empirical application. Proofs are in Appendix (ref).
First define several key pieces of notation. Let $\varphi(z)$ denote the probability density function of the standard normal distribution. Adding a subscript $\varphi_{\sigma^2}(z)$ denotes the density of the normal distribution with variance $\sigma^2$. No subscript means that the variance is unity. Let $\mathcal{F}_\infty$ denote the set of valid PDFs of continuous real-valued random variables with bounded height.
Consider a population of t-tests. Each researcher first draws the unobserved latent true effect $h\in\mathbb{R}$ from an unknown distribution $\Pi_0$. We call $\Pi_0$ the {\it distribution of true effects}. The researcher then reports the $t$-statistic, which is $h$ plus normal noise $ T\: |\: h \sim N(h,1) $. The conditional probability density $f_{T|h}\left(t\:|\:h\right)$ of $T$ is:
The unconditional density of $T$ is the expectation over the conditional density.
Selective reporting means that the meta-analyst does not observe a random sample of $T$ because some $t$-scores are truncated. Let $R$ be the event that $T$ is reported. The meta-analyst instead observes draws from the conditional distribution of the reported $t$-scores: $T|R$. Reporting is called {\it selective} if $R$ depends on the realization of $T$. For example, small values of $T$ may go unreported (e.g. by publication bias) or the researcher may repeatedly draw correlated $T$ with the same $h$ until they draw a large one and then stop. The form of selective reporting can be almost anything. In this paper, we will focus on two basic examples of selective reporting described below.
In principle, selective reporting can take many forms and is not limited to the three illustrative examples above. The meta-analyst does not know the form of selective reporting or the distribution of true effects and will not attempt to learn these. Instead, they wish only to determine whether the distribution of reported $t$-scores $f_{T|R}$ can be explained without selective reporting. We now state the meta-analyst's null hypothesis ${\mathbf{H}_0}$ in Equation ((ref)) below.
The meta-analyst's null hypothesis ${\mathbf{H}_0}$ states that there is some distribution of true effects $\Pi$ that explains the population $t$-curve without any selective reporting. This is true if and only if the PDF $f_{T|R}$ is sufficiently smooth.\footnote{To be precise, ${\mathbf{H}_0}$ is true when the deconvolved function $e^{t^2/2}\phi_{T|R}$ is a valid characteristic function where $\phi_{T|R}$ is the characteristic function of the reported $t$-score. This can only occur when $\phi_{T|R}$ has thin tails, i.e. the PDF $f_{T|R}$ is smooth.} If $f_{T|R}$ has a discontinuity or oscillates rapidly, then it could not have arisen “naturally" and ${\mathbf{H}_0}$ is violated. While some violations of ${\mathbf{H}_0}$ are apparent to the naked eye (e.g. jumps or kinks), others are not.
When ${\mathbf{H}_0}$ is false, then something has distorted the $t$-curve. However, the converse does not hold. If ${\mathbf{H}_0}$ is true, then either there is no selective reporting, or selective reporting is {\it undetectable}. We formally define the term “undetectable" below.
Undetectable forms of selective reporting change the distribution of $T$, but keep it smooth. Thus, undetectable $p$-hacking yields a $t$-curve that is still so smooth that it could have been produced “naturally" without any selective reporting by some other distribution of true effects. We stress that it is impossible to design a valid test that has power against undetectable forms of selective reporting. This is because when $\mathbf{H}_0$ is true, the population distribution of $T|R$ can be explained without selective reporting. Below we return to the three illustrative examples from the previous section and show that most forms of publication bias and threshold $p$-hacking are detectable, but whether maximization $p$-hacking is detectable depends on the smoothness of $\Pi_0$.
\vskip 0.07in {\bf Example (ref). Maximization $p$-hacking (continued)}
Consider the Maximization $p$-hacking setup from Example (ref). This kind of $p$-hacking produces a continuous $t$-curve with no discontinuities or kinks. Is it detectable? Surprisingly, this turns out to depend on the distribution of true effects $\Pi_0$! We show in the Appendix (ref) that if $h$ is any mixture of normals all with variance at least $2(1-\rho)$, then this form of $p$-hacking is {\bf undetectable.} If instead the distribution of $h$ places any amount of probability mass at zero, then this form of $p$-hacking is {\bf detectable}. The intuition for the result is that when $\Pi_0$ is smooth, then the $t$-curve becomes so smooth that reporting only the larger $t$-score doesn't take away enough smoothness to reveal itself. The implication for theory is that detectability depends jointly on the distribution of true effects and the form $p$-hacking takes. The implication for practice is that many forms of $p$-hacking are totally undetectable by any examination of the $t$-curve or $p$-curve.
\vskip 0.07in {\bf Example (ref). Threshold $p$-hacking (continued)}\\ Consider the Threshold $p$-hacking setup from Example (ref). Threshold $p$-hacking is non-smooth and will produce a jump discontinuity in the $t$-curve. Since any $f_T$ must be smooth regardless of the distribution of true effects, $f_{T|R}$ under threshold $p$-hacking cannot be explained without selective reporting. So ${\mathbf{H_0}}$ is false and threshold $p$-hacking is {\bf detectable}. \vskip 0.07in
$p$-hacking is detectable when it removes enough smoothness from the $t$-curve to reveal itself. Figure (ref) below visualizes detectability in three examples. All three simulated $t$-curves were subject to severe $p$-hacking but the detectability of that hacking varies widely. The red solid line is the $t$-curve after threshold-type $p$-hacking (described in Example (ref)). This is easy to detect because of the clear kink in the curve caused by the threshold. The solid blue line is the $t$-curve after maximization-type $p$-hacking where the researcher reports the largest of eight independent $t$-scores (Example (ref)) and every $h=0$. The non-smoothness is subtle, diffuse throughout the $t$-curve, and not visually apparent. Nevertheless $p$-hacking is still detectable here (and the test proposed later in this paper is consistent against it). Finally, the dashed blue line is an identical DGP to the solid blue line except $h\sim N(0.5,1)$. Now the smoothness of the distribution of true effects fully masks the selective reporting and the $p$-hacking is undetectable despite its severity.
In this section we rewrite $\mathbf{H}_0$ in a testable form and then propose a consistent test for it. There are three main steps. First, we express $\mathbf{H}_0$ as a zero-distance restriction. Second, we rewrite the zero-distance restriction in the spectral domain as the sum of sample moments. Finally, we specify the test statistic and critical values.
The first step is to express the null hypothesis as a restriction on the density $f_T$. To express this restriction we first need to specify a notion of distance between distributions. We choose the following norm because it allows us to naturally take the Singular Value Decomposition later on. Fix $\sigma_Y^2>0$; throughout we set $\sigma_Y^2=1$. Define the norm $||\cdot||_Y\: : \: \mathcal{F}_{\infty}\to \mathbb{R}^+$ as:
We now state the testable restriction in Equation ((ref)) below. The restriction says that the distance between $f_{T|R}$ and the set of all possible densities of $f_T$ for which $h$ has a continuous distribution $\pi \in \mathcal{F}_\infty$ is zero. This essentially means that the observed $t$-curve cannot be separated from the possible set of undistorted $t$-curves.
Theorem (ref) below shows that Restriction ((ref)) is sharp, i.e. identical to the null hypothesis. This means that selective reporting is detectable if and only if it separates $f_{T|R}$ from the set of undistorted $t$-curves in the $||\cdot||_Y$ norm.
Restriction ((ref)) is expressed in the PDF domain, which makes it easy to interpret but hard to see how it can be tested. The next step is to rewrite the restriction in the discrete spectral domain. Specifically, we use the Singular Value Decomposition to express the null hypothesis in terms of expectations. The meta-analyst can easily test restrictions on a set of expectations by plugging in sample means.
The Singular Value Decomposition expresses the mapping from a density of true effects $\pi\in \mathcal{F}_{\infty}$ into an undistorted $t$-curve $f_T$ in terms of a weighted sum of orthonormal basis polynomials where the weights depend on expectations over $\pi$. Below we take the Singular Value Decomposition from CarrascoPaper and recently repurposed by faridani2025testingunderpoweredliteratures.
The singular functions are scalings of the Hermite polynomials: $\chi_j(t) = \frac{1}{\sqrt{j!}} He_j\left(\frac{t}{\sqrt{1+\sigma_Y^2}}\right)$, $\psi_j(t) =\frac{1}{\sqrt{j!}} He_j\left(\frac{t}{\sigma_Y}\right)$.\footnote{The Hermite Polynomials are defined as: $ He_j(t)=\sum_{l=0}^{[j/2]}(-1)^l \frac{(2l)!}{2^ll!}\binom{j}{2l}t^{j-2l}$. } It is useful to know that the functions $\chi_j(h)\varphi_{\sigma_Y^2+1}(h)$ and $\psi_j(t)\varphi_{\sigma_Y^2}(t)$ are uniformly bounded over all $t,h,j$ and that the Hermite polynomials form orthonormal bases of of a weighted $\mathcal{L}^2$ space that contains $\mathcal{F}_\infty$ in the inner products $\langle \cdot,\cdot \rangle_Y$ and $\langle \cdot,\cdot \rangle_X$ defined in Section (ref).
The Singular Value Decomposition allows us to express the difference between two distributions either in the PDF domain (where differences are easy to interpret) or the spectral domain (where differences are easy to calculate). Specifically Lemma (ref) rewrites the distance between two probability densities in terms of the Euclidean distance between vectors of expectations.
Now that we can express the distance between two PDFs as the sum of squared differences in expectations, we can rewrite the projection problem in terms of vectors in Euclidean space. Specifically, we will show that projecting $f_{T|R}$ onto the set of all possible $t$-curves is equivalent to projecting a vector onto a certain convex set. Define $\theta$ as the vector of expectations with countable elements below. Here the element $\theta_j$ is the $j^{\text{th}}$ weighted Hermite coefficient of the observed $t$-curve. This vector fully characterizes $f_{T|R}$ in the spectral domain.
Define $\mathbf{b}_x$ as the vector with elements: $b_{x,j}=\eta_j\mathbb{E}\left[\chi_j(x)\varphi_{\sigma_Y^2+1}(x)\right] $. Define the set $U$ as the convex hull of all $\mathbf{b}_x$ over all $x\in \mathbb{R}$. Since the functions $\psi_j(t)\varphi_{\sigma_Y^2}(t)$ are continuous and uniformly bounded, $U$ is compact.
We would like to project $\theta$ onto $U$. But, since $\theta$ and $U$ are infinite-dimensional, the projection is computationally infeasible. To make it feasible, we regularize the projection by spectral cutoff, i.e. we only consider the first $J$ elements of the vectors involved. Define $ d_J({U},\mathbf{v})$ as the residual of the $||\cdot||_2$ projection of any vector $\mathbf{v}$ onto $U$ after truncation to the first $J$ elements:
We can now state Restriction ((ref)) which (we will show) is the regularized version of ((ref)).
Testing ((ref)) seems much easier than testing ((ref)) directly because the objects involved are vectors in $\mathbb{R}^J$. We now show that this is as good as testing $\mathbf{H}_0$ when $J$ is large enough. First we show validity. Theorem (ref) below says that $\theta$ is no farther from $U$ than $f_{T|R}$ is from the set of all possible PDFs under the null. A valid test of $d_J(U,\theta) = 0$ will therefore also always be a valid test of $\mathbf{H}_0 $.
A consistent test of $d_J(U,\theta)=0$ will also be consistent against all alternatives when $J$ is large enough. Theorem (ref) below shows that when $\mathbf{H}_0$ is false and $J$ is large enough, a consistent test of $d_J(U,\theta) = 0$ is also a consistent test of $\mathbf{H}_0$ for detectable selective reporting.
Theorem (ref) and (ref) prove that for large enough $J$, a valid and consistent test of ((ref)) is also a valid and consistent test of any detectable violation of $\mathbf{H}_0$.
Next we define the test statistic explicitly. Projecting a vector onto $U$ is computationally challenging because the basis of the convex hull $U$ is uncountable. But it is feasible to project onto $\widetilde{U}$, the convex hull of the $J\times 1$ vectors $\mathbf{u}_x=\eta_j\mathbb{E}\left[\chi_j(x)\varphi_{\sigma_Y^2+1}(x)\right] $ over a finite set of $x\in \mathcal{X}\subset \mathbb{R}$. This incurs some approximation error that can be forced to be as small as we want by using a grid $\mathcal{X}$ wide and fine enough. Lemma (ref) tells us how wide and fine the grid must be in order to control the approximation error. Notice that the bound holds uniformly across all $J\in \mathbb{N}$.
The meta-analyst has only a sample of $n$ $t$-scores $\{t_i\}_{i=1}^n$. They do not observe the vector of expectations $\theta$ but rather the analogous $J\times 1$ vector of sample means $\widehat{\theta}$:
The $n$ $t$-scores are reported by $m$ articles. While $t$-scores are dependent within an article, they must be independent across articles and the number of $t$-scores per article must be uniformly bounded.
This paper proposes $d_J(\widetilde{U},\widehat{\theta})$ as the test statistic for $\mathbf{H}_0$. The computational task of computing this projection distance is well-conditioned even though the problem of actually identifying the closest member of $\widetilde{U}$ is ill-conditioned. We compute projections with the \verb|osqp| package in R and set $J=30$ and $J=20$ in our empirical application and simulations.
The last task is to specify the critical value. First notice that the $J\times 1$ vector $\sqrt{n}\left(\widehat{\mathbf{\theta}}-\mathbf{\theta}\right)$ is asymptotically normal and its quantiles can be estimated with the bootstrap. This is immediate because $\widehat{\theta}$ is a vector sample means, the functions $\varphi_{\sigma^2_Y}(t_i)\psi_j(t_i)$ are uniformly bounded and everywhere infinitely differentiable, and Assumption (ref) guarantees weak dependence across observations.
The critical value for $d_J(\widetilde{U},\widehat{\theta})$ can be computed with a particular bootstrap. Specifically, the critical value for a test of size $\alpha$ will be the discretization error $\epsilon$ plus the $1-\alpha$ quantile of the distribution of $d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e})$ where $\mathbf{P}_{\widetilde{U}}\widehat{\theta}$ is the closest member of $\widetilde{U}$ to $\widehat{\theta}$ and $\sqrt{n}\mathbf{e}$ has the same asymptotic distribution as $\sqrt{n}\left(\widehat{\mathbf{\theta}}-\mathbf{\theta}\right)$.\footnote{This will work if $\mathbf{P}_{\widetilde{U}}\widehat{\theta}$ is any point that converges in probability at rate $n^{-1/2}$ to the member of $\widetilde{U}$ closest to $\theta$.} We compute this quantile by using the bootstrap to draw $\mathbf{e}$'s. The following paragraphs show the validity and consistency of this procedure in large samples.
Theorem (ref) says that if we know the distribution of $ d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e}) $ then we can construct critical values for $ d_J(\widetilde{U},\widehat{\theta} )$. We can learn the distribution of $d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e})$ by the bootstrap. We simply resample article-by-article to draw $\mathbf{e}^* = \widehat{\theta}^*-\widehat{\theta}$. The bootstrap is valid here because $\widehat{\theta}$ is a sample mean of uniformly bounded random variables and articles are assumed to have a bounded number of t-scores each. Resample many times to estimate the distribution of $d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e})$. By Theorem (ref), this distribution is sufficient to construct critical values. If $F$ is the CDF of the bootstrapped $d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e}^*)$, then the critical value is $\text{cv}(\alpha) = F^{-1}\left(1-\alpha\right)+\epsilon$. Our testing procedure is: $$\text{Reject } \mathbf{H}_0 \text{ if } d_J(\widetilde{U},\widehat{\theta}) > \text{cv}(\alpha) $$
Size control is Corollary (ref) which is an immediate consequence of Theorem (ref).
Finally we show consistency. Theorem (ref) says that comparing $d_J(\widetilde{U},\widehat{\theta})$ to $cv(\alpha)$ is a consistent test given a sufficiently large $J$ and a sufficiently wide and dense grid for $\widetilde{U}$.
A key consequence of Theorem (ref) is that by increasing $J,L,\delta^{-1}$ with $n$, the test can be guaranteed to be consistent against any selective reporting process that is identified. In particular, we can always increase $J,L,\delta^{-1}$ slowly enough that all the results that assume these values are fixed still hold. Then $\liminf_{n\to \infty}\mathbb{P}\left[d_{J_n}(\widetilde{U}_n,\widehat{\theta} ) > cv(\alpha)\right]=1$ no matter how slowly $J,L,\delta^{-1}$ increase.
So far we have assumed that $T|h \sim N(h,1)$ exactly in the absence of selective reporting. In practice, this normality assumption can be violated. We wish to reject ${\mathbf{H}_0}$ only when the distortions in the $t$-curve were likely due to selective reporting and to avoid spurious rejections originating from non-normality of $T$. We consider two main causes of violations: (1) the $t$-score is actually distributed student-$t$ instead of exactly normal and (2) the $t$-score was reported after rounding and then de-rounded by the meta-analyst. By augmenting the critical value, the meta-analyst can avoid spurious rejections from either of these two sources. The ease of this augmentation is one of the main advantages of the projection test.
From the outset, we have set up this test so that it is simple to account for misspecification of the exact distribution of $T|h$. We now generalize the setup from Section (ref) by allowing the noise term in the $t$-score to have a distribution $g$ that is close—--but not equal—--to the standard normal. Suppose that the meta-analyst knows that the total difference between the two densities is upper bounded by some constant $\Delta \geq 0$. Assumption (ref) makes this situation explicit.
Under misspecification, the null hypothesis no longer guarantees the restriction in ((ref)), but instead implies the potentially untestable condition in ((ref)).
Since the meta-analyst does not know $g$, they cannot test ((ref)). Instead, they run the misspecified projection onto the Gaussian convolutions and upper bound the error using the triangle inequality. The inequality below shows that even though the projection is misspecified, the resulting residual can still be bounded.
The null hypothesis $\mathbf{H}_0^{\text{miss}}$ for the non-normal model is that the residual of the misspecified projection is small enough that it can be explained by a deviation from $N(0,1)$ of $\Delta$.
We test $\mathbf{H}_0^{\text{\normalfont miss}}$ using the same test statistic as before: $d_J(\widetilde{U},\widehat{\theta})$. The difference is that we now add $\Delta$ to the critical value.
The main limitation of this strategy is that it requires a known upper bound on $\Delta$. When the researcher is unwilling to commit to a specific value of $\Delta$, they can instead report the {\it breakdown statistic}, or the smallest value $\widehat{B}$ of $\Delta$ that would overturn their detection of selective reporting.
$\widehat{B}$ is the smallest deviation from normality that would be sufficient to explain away the observed lack of smoothness in the $t$-curve. In other words, it is the minimal value of $\Delta$ such that the test would no longer reject $\mathbf{H}_0$. The next two sections discuss how large $\widehat{B}$ needs to be for the meta-analyst to be confident that the rejection was probably not the spurious result of common distortions in the $t$-curve that do not come from selective reporting.
In practice, the researcher does not know the true standard error of the point estimate with certainty. This means that when they construct the $t$-score, they divide the point estimate by an estimate of the standard error and the result is that $T-h$ is distributed student-$t$ with $\nu$ degrees of freedom, where $\nu$ is the effective sample size. We can calculate value of $\Delta$ associated with approximating the student-$t$ with the normal using numerical integration. For instance, for a study with effective sample size of $\nu=120$:
Therefore, if we had a set of studies with effective samples sizes predominantly larger than 120 and $\widehat{B}>.0009$, we could conclude that the detection of $p$-hacking was not just a spurious outcome of approximating the normal distribution with the student-$t$ distribution.
In practice, $t$-scores are always reported after some rounding in practice. To avoid spurious detections of $p$-hacking resulting from rounding discontinuities, the meta-analyst must de-round the meta-data by adding random uniformly distributed noise commensurate to the number of significant digits elliott2024powertestsdetectingphacking. This process does not leave the $t$-curve completely undistorted. It is therefore prudent to check whether $\widehat{B}$ is small enough to be explained by derounding distortions to the $t$-curve.
We can model the derounded $t$-score as $T=h+Z+V$ where $V$ is mean zero noise independent of $h,Z$ that comes from derounding. Since $V$ is small and of expectation zero, we can bound $\Delta$ by bounding the sup-norm of the difference between the density $f_{Z+V}$ of $f_Z = \varphi$. Proposition (ref) bounds the distortion from normality $\Delta$ that can be incurred by rounding and derounding.
If the $t$-score is observed up to two decimal places (i.e. we can distinguish $T=1.96$ from $T=1.95$), then $|V|< 0.005$ with probability 1. Thus, when $\Delta \leq \sqrt{\frac{1}{2\pi}} \frac{.05^2}{2} \leq 0.00049$. This means that whenever $\widehat{B}> .00049$, we can conclude that the detection of $p$-hacking was not a spurious consequence of de-rounding. This is true of all of the breakdown statistics reported in our empirical application later on in Section (ref). We conclude that the projection test can be made robust to non-normality induced by derounding and that this robustness does not cost us enough power to matter in application.
This section presents two sets of simulations that compare the power and size control of this paper's projection test with the six tests studied by Elliott. The takeaway message of these Monte Carlo exercises is that the projection test is as sensitive as the best of these six and controls size much better when $T|h$ is not exactly normal.
Figure (ref) displays power curves showing that the projection test is as sensitive as the most sensitive of the six Elliott tests. Here the DGP is as follows. The true effect $h$ equals 2 almost surely and $T$ is exactly normal centered on $h$. The researcher $p$-hacks with probability $q$ on the $x$-axis. So on the left part of the graph there is no $p$-hacking and on the right part of the graph everyone $p$-hacks. The $p$-hacking process matches Example (ref) where the researcher draws $T_1,T_2$ which share an $h$ and are otherwise iid. The researcher reports $T_1$ if it is greater than 1.96 in absolute value and otherwise reports $\max\{|T_1|,|T_2|\}$. Since the power curves for the projection test and the maximum of the six Elliott tests track each other almost exactly, we conclude that for this DGP no power is lost by using the projection test.
In a second simulation we show that the projection test in this paper can detect $p$-h
This section applies the projection test to the data from bb. This data consists of 21,740 two-sided t-statistics drawn from 684 articles. These articles were the universe of articles published in 25 top economics journals between 2015 and 2018 that used one of four causal methods. Each $t$-score corresponds to a two-sided $t$-test of a main hypothesis of interest. The coefficients in the numerators of the $t$-scores were estimated using Randomized Controlled Trials (RCT), Difference in Differences (DID), Regression Discontinuity Designs (RDD), or Instrumental Variables (IV). All “regression controls, constant terms, balance and robustness checks, heterogeneity of effects, and placebo tests" were excluded. Some tests were reported as $p$-values and bb converted these to $t$-scores using the standard normal distribution.
We run our tests on each of the four causal inference methods separately. The results are in Table (ref). We report results for $J=30$ and $J=20$ to demonstrate that our conclusions are not very sensitive to the choice of smoothing parameter $J$. The $p$-values in columns 4 and 7 mostly confirm the findings from kudrinjmp24. We detect selective reporting for RCTs, IVs, and DIDs.
Are these rejections spurious outcomes coming from (de)rounding or the limited degrees of freedom of the underlying studies? To determine this, we compute the breakdown statistic $\widehat{B}$ which is the smallest distortion in the $t$-curve from non-$p$-hacking sources that it would take to explain the observed $t$-curve and overturn our rejection. The smallest $\widehat{B}$ in Table (ref) is 0.0010. Section (ref) shows that when the effective sample size of the underlying studies are greater than $\nu=120$, the student-$t$ distribution cannot explain a breakdown statistic of 0.0010 or larger. Are the effective sample sizes in this data small enough to explain this? To find out we took a random sample of 100 RCT papers, 100 IV papers, and 100 DID papers from bb. We calculated their effective sample sizes as the number of clusters where the cluster was the level at which the standard error was clustered. The median effective sample sizes are reported in Table (ref). We calculate $\Delta_\nu$ as the average over all the effective sample sizes $\nu_j$ of $\frac{1}{100}\sum_{j=1}^{100}||\varphi - t_{\nu_j}||_Y$, which can be computed directly by numerical integration.\footnote{We were unable to determine even a lower bound on effective sample size for 20/100 DID papers and 15/100 IV papers.} For all three methods, $\Delta_\nu < \widehat{B}$. This means that the effective sample sizes were overall too large to explain the breakdown statistic and thus the rejection of $\mathbf{H}_0$ stands.
In Section (ref), we showed that if the $t$-score is observed up to two decimal places, (de)rounding can add at most $\sqrt{\frac{1}{2\pi}} \frac{.05^2}{2} \leq 0.00049$ to the test statistic. Since $\widehat{B} > 0.00049$ everywhere in Table (ref) we conclude that (de)rounding cannot explain the non-smoothness of the $t$-curves in the bb data.
These empirical results add a novel insight to existing knowledge. The $t$-curves for RCTs, IVs, and DIDs are significantly distorted. These distortions are larger than can be explained by random chance, (de)rounding, or the student-$t$ distribution. We conclude that the most likely culprit is selective reporting of $t$-scores. The projection test proposed in this paper is the first (to our knowledge) that is simultaneously sensitive enough to detect these distortions in bb's sample and also robust to basic concerns related to (de)rounding, the unknown distribution of true effects, and finite degrees of freedom.