EconBase
← Back to paper

A Sharp and Robust Test for Selective Reporting

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

55,545 characters · 12 sections · 33 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

When is $p$-hacking detectable?

\onehalfspacing

abstractSome forms of $p$-hacking cannot be detected by examining the $t$-curve (or $p$-curve). Standard tests may also fail to find even detectable forms of selective reporting. We propose a novel test that is consistent against {\it every} detectable form of $p$-hacking and remains interpretable even when the $t$-scores are not exactly normal. The test statistic is the distance between the smoothed empirical $t$-curve and the set of all distributions that would be possible in the absence of any selective reporting. This novel {\it projection test} can only be evaded in large meta-samples by selective reporting that also evades all other valid tests of restrictions on the $t$-curve. A second benefit of the projection test is that under the null hypothesis of no $p$-hacking we can check whether the projection residual could have been produced by other distortions not related to selective reporting, e.g. rounding and de-rounding. Applying the test to the bb meta-data, we find that the $t$-curves for RCTs, IVs, and DIDs are more distorted than could arise by chance. We confirm that these distortions cannot be explained by (de)rounding of $t$-scores or by the limited degrees of freedom of the underlying studies.

\noindentKeywords: $p$-Hacking, Deconvolution

\noindentJEL Codes: C12, C14

\pagenumbering{arabic}

Introduction

Publication bias and $p$-hacking threaten the validity of empirical research. When researchers report only the most surprising statistical results, they are likely to over-claim. Scientific findings are therefore more trustworthy when they are drawn from literatures where selective reporting is not evident.

Learning what kinds of studies are $p$-hacked helps evaluate empirical claims and target reform. Much recent work has tested for selective reporting in a given literature by checking for distortions in the empirical distribution of reported $t$-scores.\footnote{See for example: Simonsohn,Head,Brodeur,havranek_housing,bb,Elliott,elliott2024powertestsdetectingphacking,Kudrinjmp,kudrinjmp24,Havranek24. The $p$-curve is distorted if and only if the $t$-curve is distorted. } All of these methods share a common shortcoming: They can be evaded by selective reporting strategies that do not violate the specific restrictions that the meta-analyst tested. For example, the Caliper test can be evaded by $p$-hacking that does not produce a discontinuity in the histogram of $t$-scores at the critical value. The bounds tested in Elliott can be evaded by $p$-hacking that distorts the overall shape of the $p$-curve but does not make its derivatives extreme at any particular point.

This paper proposes a new test that can only be evaded by forms of $p$-hacking that are totally undetectable by any examination of the $t$-curve. To our knowledge this is the first test that comes accompanied by such a guarantee. Our proposed test statistic is the distance between the observed $t$-curve and the set of all feasible $t$-curves under no selective reporting. The testable restriction that this distance equals zero is {\it sharp} because if a p-hacking process did not violate it, then the resulting $t$-curve could be explained without selective reporting by some distribution of true treatment effects.

Our test can only be evaded in large meta-samples by $p$-hacking processes that are undetectable in principle. We illustrate with a simple example: when a researcher reports the maximum of two correlated one-sided $t$-scores and the distribution of true effects is smooth, then the resulting $t$-curve can be explained either with or without $p$-hacking. This is an example of a kind of selective reporting that is undetectable by any examination of the $t$-curve. Outside of such undetectable cases, our test is consistent against all forms of selective reporting.

We call our test a projection test. The key technical step is to compute the test statistic using the Singular Value Decomposition (SVD). The SVD allows us to rewrite the projection of one smoothed density onto a set of other densities as simply the orthogonal projection of a vector of sample means onto a convex set of vectors in Euclidean space. Critical values are computed using the bootstrap.

Simulations show that the projection test is as powerful as the most sensitive of the six tests studied by Elliott. Further simulations show an example where $p$-hacking is nearly undetectable, the projection test is consistent against it while the six tests studied by Elliott all have power no greater than size in a large meta-sample. Therefore the projection test makes a meaningful contribution to the existing body of tests for $p$-hacking.

The second major advantage of the projection test is that the test statistic can still be interpreted when the $t$-scores reported by individual researchers are not exactly normally distributed. Recent nonparametric tests for $p$-hacking that examine every part of the $t$-curve like kudrinjmp24 and Elliott are highly sensitive but will reject the null hypothesis of no selective reporting whenever the empirical $t$-curve is not smooth enough to have been generated by a convolution with a normal kernel. Yet, non-smoothness in the $t$-curve can potentially arise for other reasons not related to selective reporting. For example, the $t$-score is more accurately described as following a student-$t$ distribution with $\nu$ degrees of freedom, where $\nu$ is the effective sample size of the experiment. In addition, the $t$-score is often reported after rounding. Even when de-rounded, this process introduces non-smoothness in the distribution of the $t$-score that could theoretically lead to spurious rejections of the null of no p-hacking.

We account for finite-sample deviations from the normal distribution in the following way. We first calculate the {\it Breakdown Statistic} $\widehat{B}$ defined as how severe the deviation from normality must be to overturn the meta-analyst's detection of $p$-hacking. The breakdown statistic is computed from the observed $t$-curve and is equal to the difference between the test statistic and the critical value. We then compare the breakdown statistic to the maximum possible deviations from normality that could arise from either the student-$t$ distribution or from (de)rounding. If the breakdown statistic is larger than could be explained by either of these sources of non-smoothness, then the rejection stands.

We apply our method to the meta-sample of 21,740 two-sided $t$-tests reported by 684 top economics journals collected by bb. We find that the $t$-curves for RCTs, DIDs, and IVs are significantly distorted and that these distortions are too large to be explained either by (1) rounding or de-rounding of the reported $t$-scores or by the limited degrees of freedom of the studies themselves. We conclude that significant selective reporting is likely present in all three subgroups of studies. We cannot say the same of RDDs.

This paper is organized as follows. Section (ref) sets up the problem. Section (ref) defines the meta-analyst's null hypothesis and provides examples of $p$-hacking processes that are detectable and undetectable by any test. Section (ref) presents the projection test and shows its validity and consistency when the $t$-score is exactly normal. Section (ref) extends to situations where the $t$-score is only approximately normal even in the absence of selective reporting. Section (ref) presents simulations and Section (ref) shows an empirical application. Proofs are in Appendix (ref).

Setup

First define several key pieces of notation. Let $\varphi(z)$ denote the probability density function of the standard normal distribution. Adding a subscript $\varphi_{\sigma^2}(z)$ denotes the density of the normal distribution with variance $\sigma^2$. No subscript means that the variance is unity. Let $\mathcal{F}_\infty$ denote the set of valid PDFs of continuous real-valued random variables with bounded height.

equation[equation omitted — 160 chars of source]

Consider a population of t-tests. Each researcher first draws the unobserved latent true effect $h\in\mathbb{R}$ from an unknown distribution $\Pi_0$. We call $\Pi_0$ the {\it distribution of true effects}. The researcher then reports the $t$-statistic, which is $h$ plus normal noise $ T\: |\: h \sim N(h,1) $. The conditional probability density $f_{T|h}\left(t\:|\:h\right)$ of $T$ is:

equation[equation omitted — 61 chars of source]

The unconditional density of $T$ is the expectation over the conditional density.

equation[equation omitted — 121 chars of source]

Selective reporting means that the meta-analyst does not observe a random sample of $T$ because some $t$-scores are truncated. Let $R$ be the event that $T$ is reported. The meta-analyst instead observes draws from the conditional distribution of the reported $t$-scores: $T|R$. Reporting is called {\it selective} if $R$ depends on the realization of $T$. For example, small values of $T$ may go unreported (e.g. by publication bias) or the researcher may repeatedly draw correlated $T$ with the same $h$ until they draw a large one and then stop. The form of selective reporting can be almost anything. In this paper, we will focus on two basic examples of selective reporting described below.

comment\begin{example}{\bf Publication Bias}\normalfont Consider a process where some $t$-scores are truncated from the meta-analyst's sample and reporting $R$ is independent across $t$-scores (even those in the same manuscript). Under publication bias, the meta-analyst only observes a random sample of conditional draws of $T|R$ and not an unconditional random sample of $T$. Applying Bayes' Theorem and assuming that $\mathbb{P}[R]>0$ guarantees that $T|R$ has the following conditional density: \begin{equation} f_{T|R}(t) = \frac{ \mathbb{P}[R|T=t]f_T(t)}{\mathbb{P}[R]} \end{equation} This includes the “classical" publication bias where $t$-scores are reported with probability $p$ if they are significant and with probability $q$ if not Andrews. In this case $p(t)$ would be a step function. \end{example}
example{\bf Maximization $p$-hacking} \normalfont Consider a $p$-hacking process where a researcher first draws a single unobserved true effect $h$. Then they draw two correlated $t$-scores $\{T_1,T_2\}$ both with marginal distributions $N(h,1)$ and covariance $\rho \in (-1,1)$. They report the larger of the two $t$-scores (where “larger" means rightmost on the number line). This type of $p$-hacking produces no jumps in the $t$-curve and is known to be hard to detect with existing tests elliott2024powertestsdetectingphacking, kudrinjmp24. In Section (ref), we show that when the distribution of $h$ is smooth, this kind of $p$-hacking is undetectable in principle by {\it any} test.
example{\bf Threshold $p$-hacking}\normalfont Suppose that the researcher $p$-hacks only if their initial result is insignificant. Consider a similar setup to Example (ref) where a researcher draws two correlated $t$-scores with the same $h$. Now the researcher reports $T_1$ if $|T_1|>1.96$. Otherwise, they report $\max\{T_1,T_2\}$. This kind of $p$-hacking does produce jumps in the $t$-curve $f_{T|R}$ and sensitive tests have been proposed elliott2024powertestsdetectingphacking, kudrinjmp24.

Null Hypothesis and Detectability

In principle, selective reporting can take many forms and is not limited to the three illustrative examples above. The meta-analyst does not know the form of selective reporting or the distribution of true effects and will not attempt to learn these. Instead, they wish only to determine whether the distribution of reported $t$-scores $f_{T|R}$ can be explained without selective reporting. We now state the meta-analyst's null hypothesis ${\mathbf{H}_0}$ in Equation ((ref)) below.

equation[equation omitted — 168 chars of source]

The meta-analyst's null hypothesis ${\mathbf{H}_0}$ states that there is some distribution of true effects $\Pi$ that explains the population $t$-curve without any selective reporting. This is true if and only if the PDF $f_{T|R}$ is sufficiently smooth.\footnote{To be precise, ${\mathbf{H}_0}$ is true when the deconvolved function $e^{t^2/2}\phi_{T|R}$ is a valid characteristic function where $\phi_{T|R}$ is the characteristic function of the reported $t$-score. This can only occur when $\phi_{T|R}$ has thin tails, i.e. the PDF $f_{T|R}$ is smooth.} If $f_{T|R}$ has a discontinuity or oscillates rapidly, then it could not have arisen “naturally" and ${\mathbf{H}_0}$ is violated. While some violations of ${\mathbf{H}_0}$ are apparent to the naked eye (e.g. jumps or kinks), others are not.

When ${\mathbf{H}_0}$ is false, then something has distorted the $t$-curve. However, the converse does not hold. If ${\mathbf{H}_0}$ is true, then either there is no selective reporting, or selective reporting is {\it undetectable}. We formally define the term “undetectable" below.

definitionWe say that selective reporting is {\bf undetectable} if $T \stackrel{d}{\neq }T|R$ but ${\mathbf{H}_0}$ is true. We say that selective reporting is {\bf detectable} if ${\mathbf{H}_0}$ is false and $T\sim N(h,1)$.

Undetectable forms of selective reporting change the distribution of $T$, but keep it smooth. Thus, undetectable $p$-hacking yields a $t$-curve that is still so smooth that it could have been produced “naturally" without any selective reporting by some other distribution of true effects. We stress that it is impossible to design a valid test that has power against undetectable forms of selective reporting. This is because when $\mathbf{H}_0$ is true, the population distribution of $T|R$ can be explained without selective reporting. Below we return to the three illustrative examples from the previous section and show that most forms of publication bias and threshold $p$-hacking are detectable, but whether maximization $p$-hacking is detectable depends on the smoothness of $\Pi_0$.

comment\vskip 0.07in {\bf Example (ref). Publication Bias (continued)} Consider the publication bias setup from Example (ref). Suppose that the conditional probability of publication is nonsmooth---meaning that either $\mathbb{P}[R|T=t]$ or any of its derivatives in $t$ has a jump discontinuity somewhere. For instance, the probability of publication might increase sharply at the critical value $1.96$ as in Andrews. Then $f_{T|R}$ cannot be smooth. Since all convolutions with the normal distributions are smooth, $f_{T|R}$ could not be explained without selective reporting and $\mathbf{H}_0$ is false. So non-smooth publication bias is {\bf detectable}.

\vskip 0.07in {\bf Example (ref). Maximization $p$-hacking (continued)}

Consider the Maximization $p$-hacking setup from Example (ref). This kind of $p$-hacking produces a continuous $t$-curve with no discontinuities or kinks. Is it detectable? Surprisingly, this turns out to depend on the distribution of true effects $\Pi_0$! We show in the Appendix (ref) that if $h$ is any mixture of normals all with variance at least $2(1-\rho)$, then this form of $p$-hacking is {\bf undetectable.} If instead the distribution of $h$ places any amount of probability mass at zero, then this form of $p$-hacking is {\bf detectable}. The intuition for the result is that when $\Pi_0$ is smooth, then the $t$-curve becomes so smooth that reporting only the larger $t$-score doesn't take away enough smoothness to reveal itself. The implication for theory is that detectability depends jointly on the distribution of true effects and the form $p$-hacking takes. The implication for practice is that many forms of $p$-hacking are totally undetectable by any examination of the $t$-curve or $p$-curve.

\vskip 0.07in {\bf Example (ref). Threshold $p$-hacking (continued)}\\ Consider the Threshold $p$-hacking setup from Example (ref). Threshold $p$-hacking is non-smooth and will produce a jump discontinuity in the $t$-curve. Since any $f_T$ must be smooth regardless of the distribution of true effects, $f_{T|R}$ under threshold $p$-hacking cannot be explained without selective reporting. So ${\mathbf{H_0}}$ is false and threshold $p$-hacking is {\bf detectable}. \vskip 0.07in

$p$-hacking is detectable when it removes enough smoothness from the $t$-curve to reveal itself. Figure (ref) below visualizes detectability in three examples. All three simulated $t$-curves were subject to severe $p$-hacking but the detectability of that hacking varies widely. The red solid line is the $t$-curve after threshold-type $p$-hacking (described in Example (ref)). This is easy to detect because of the clear kink in the curve caused by the threshold. The solid blue line is the $t$-curve after maximization-type $p$-hacking where the researcher reports the largest of eight independent $t$-scores (Example (ref)) and every $h=0$. The non-smoothness is subtle, diffuse throughout the $t$-curve, and not visually apparent. Nevertheless $p$-hacking is still detectable here (and the test proposed later in this paper is consistent against it). Finally, the dashed blue line is an identical DGP to the solid blue line except $h\sim N(0.5,1)$. Now the smoothness of the distribution of true effects fully masks the selective reporting and the $p$-hacking is undetectable despite its severity.

figure[figure omitted — 206 chars of source]

A Consistent Test

In this section we rewrite $\mathbf{H}_0$ in a testable form and then propose a consistent test for it. There are three main steps. First, we express $\mathbf{H}_0$ as a zero-distance restriction. Second, we rewrite the zero-distance restriction in the spectral domain as the sum of sample moments. Finally, we specify the test statistic and critical values.

The first step is to express the null hypothesis as a restriction on the density $f_T$. To express this restriction we first need to specify a notion of distance between distributions. We choose the following norm because it allows us to naturally take the Singular Value Decomposition later on. Fix $\sigma_Y^2>0$; throughout we set $\sigma_Y^2=1$. Define the norm $||\cdot||_Y\: : \: \mathcal{F}_{\infty}\to \mathbb{R}^+$ as:

equation[equation omitted — 95 chars of source]

We now state the testable restriction in Equation ((ref)) below. The restriction says that the distance between $f_{T|R}$ and the set of all possible densities of $f_T$ for which $h$ has a continuous distribution $\pi \in \mathcal{F}_\infty$ is zero. This essentially means that the observed $t$-curve cannot be separated from the possible set of undistorted $t$-curves.

equation[equation omitted — 171 chars of source]
remark\normalfont Importantly, even though the minimization in ((ref)) is over all continuous distributions, we do not assume that $\Pi_0$ is continuous. Notice that $\mathcal F_\infty$ is dense in the space of bounded densities under the $||\cdot||_Y$ norm. So the non-continuous $\Pi$ lie in the boundary of the set of continuous $\Pi$. We will show that excluding discontinuous distributions from the minimization will not change the infimum.

Theorem (ref) below shows that Restriction ((ref)) is sharp, i.e. identical to the null hypothesis. This means that selective reporting is detectable if and only if it separates $f_{T|R}$ from the set of undistorted $t$-curves in the $||\cdot||_Y$ norm.

theoremIf $T|h \sim N(h,1)$, then $$\mathbf{H}_0 \iff (\ref{eq:testable_restriction})$$ Proof: Section (ref)

Restriction ((ref)) is expressed in the PDF domain, which makes it easy to interpret but hard to see how it can be tested. The next step is to rewrite the restriction in the discrete spectral domain. Specifically, we use the Singular Value Decomposition to express the null hypothesis in terms of expectations. The meta-analyst can easily test restrictions on a set of expectations by plugging in sample means.

The Singular Value Decomposition expresses the mapping from a density of true effects $\pi\in \mathcal{F}_{\infty}$ into an undistorted $t$-curve $f_T$ in terms of a weighted sum of orthonormal basis polynomials where the weights depend on expectations over $\pi$. Below we take the Singular Value Decomposition from CarrascoPaper and recently repurposed by faridani2025testingunderpoweredliteratures.

align*[align* omitted — 232 chars of source]

The singular functions are scalings of the Hermite polynomials: $\chi_j(t) = \frac{1}{\sqrt{j!}} He_j\left(\frac{t}{\sqrt{1+\sigma_Y^2}}\right)$, $\psi_j(t) =\frac{1}{\sqrt{j!}} He_j\left(\frac{t}{\sigma_Y}\right)$.\footnote{The Hermite Polynomials are defined as: $ He_j(t)=\sum_{l=0}^{[j/2]}(-1)^l \frac{(2l)!}{2^ll!}\binom{j}{2l}t^{j-2l}$. } It is useful to know that the functions $\chi_j(h)\varphi_{\sigma_Y^2+1}(h)$ and $\psi_j(t)\varphi_{\sigma_Y^2}(t)$ are uniformly bounded over all $t,h,j$ and that the Hermite polynomials form orthonormal bases of of a weighted $\mathcal{L}^2$ space that contains $\mathcal{F}_\infty$ in the inner products $\langle \cdot,\cdot \rangle_Y$ and $\langle \cdot,\cdot \rangle_X$ defined in Section (ref).

The Singular Value Decomposition allows us to express the difference between two distributions either in the PDF domain (where differences are easy to interpret) or the spectral domain (where differences are easy to calculate). Specifically Lemma (ref) rewrites the distance between two probability densities in terms of the Euclidean distance between vectors of expectations.

lemma$$ \left|\left|f_{T|R}- \int_{-\infty}^\infty \varphi(t-h)\pi(h)dh \right|\right|_Y =\sqrt{ \sum_{j=0}^\infty \left(\mathbb{E}\left[\psi_j(T)\varphi_{\sigma_Y^2}(T)|R\right] - \eta_j\mathbb{E}_\pi\left[\chi_j(h)\varphi_{\sigma_Y^2+1}(h)\right]\right)^2 }$$ Proof: Section (ref)

Now that we can express the distance between two PDFs as the sum of squared differences in expectations, we can rewrite the projection problem in terms of vectors in Euclidean space. Specifically, we will show that projecting $f_{T|R}$ onto the set of all possible $t$-curves is equivalent to projecting a vector onto a certain convex set. Define $\theta$ as the vector of expectations with countable elements below. Here the element $\theta_j$ is the $j^{\text{th}}$ weighted Hermite coefficient of the observed $t$-curve. This vector fully characterizes $f_{T|R}$ in the spectral domain.

equation[equation omitted — 87 chars of source]

Define $\mathbf{b}_x$ as the vector with elements: $b_{x,j}=\eta_j\mathbb{E}\left[\chi_j(x)\varphi_{\sigma_Y^2+1}(x)\right] $. Define the set $U$ as the convex hull of all $\mathbf{b}_x$ over all $x\in \mathbb{R}$. Since the functions $\psi_j(t)\varphi_{\sigma_Y^2}(t)$ are continuous and uniformly bounded, $U$ is compact.

We would like to project $\theta$ onto $U$. But, since $\theta$ and $U$ are infinite-dimensional, the projection is computationally infeasible. To make it feasible, we regularize the projection by spectral cutoff, i.e. we only consider the first $J$ elements of the vectors involved. Define $ d_J({U},\mathbf{v})$ as the residual of the $||\cdot||_2$ projection of any vector $\mathbf{v}$ onto $U$ after truncation to the first $J$ elements:

equation[equation omitted — 103 chars of source]

We can now state Restriction ((ref)) which (we will show) is the regularized version of ((ref)).

equation[equation omitted — 59 chars of source]

Testing ((ref)) seems much easier than testing ((ref)) directly because the objects involved are vectors in $\mathbb{R}^J$. We now show that this is as good as testing $\mathbf{H}_0$ when $J$ is large enough. First we show validity. Theorem (ref) below says that $\theta$ is no farther from $U$ than $f_{T|R}$ is from the set of all possible PDFs under the null. A valid test of $d_J(U,\theta) = 0$ will therefore also always be a valid test of $\mathbf{H}_0 $.

theoremFor any $J\in\mathbb{N}$, $$ d_J(U,\theta)\leq \lim_{K\to \infty}d_{K}(U,\theta) = \inf_{\pi \in \mathcal{F}_\infty }\left|\left|f_{T|R}- \int_{-\infty}^\infty \varphi(t-h)\pi(h)dh \right|\right|_Y$$ Proof: Section (ref)

A consistent test of $d_J(U,\theta)=0$ will also be consistent against all alternatives when $J$ is large enough. Theorem (ref) below shows that when $\mathbf{H}_0$ is false and $J$ is large enough, a consistent test of $d_J(U,\theta) = 0$ is also a consistent test of $\mathbf{H}_0$ for detectable selective reporting.

theoremIf $\mathbf{H}_0$ is false and selective reporting is detectable, $\exists J^*$ such that for all $J>J^*$ $$d_J(U,\theta)>0$$ Proof: Section (ref)

Theorem (ref) and (ref) prove that for large enough $J$, a valid and consistent test of ((ref)) is also a valid and consistent test of any detectable violation of $\mathbf{H}_0$.

Test Statistic and Critical Values

Next we define the test statistic explicitly. Projecting a vector onto $U$ is computationally challenging because the basis of the convex hull $U$ is uncountable. But it is feasible to project onto $\widetilde{U}$, the convex hull of the $J\times 1$ vectors $\mathbf{u}_x=\eta_j\mathbb{E}\left[\chi_j(x)\varphi_{\sigma_Y^2+1}(x)\right] $ over a finite set of $x\in \mathcal{X}\subset \mathbb{R}$. This incurs some approximation error that can be forced to be as small as we want by using a grid $\mathcal{X}$ wide and fine enough. Lemma (ref) tells us how wide and fine the grid must be in order to control the approximation error. Notice that the bound holds uniformly across all $J\in \mathbb{N}$.

lemmaLet $\mathcal{X}$ be a grid of evenly spaced points between $-L$ and $L$ with spacing $\delta$. For any $\mathbf{v}\in \mathbb{R}^J$,\begin{align*} \left|d_J({U},\mathbf{v})-d_J(\widetilde{U},\mathbf{v})\right| &< \max\left\{\sqrt{\int_{-\infty}^\infty \varphi_{\sigma_Y^2}(t)\varphi(t-L)^2dt },\sqrt{\int_{-\infty}^\infty \varphi_{\sigma_Y^2}(t)\left(\varphi(t-1)-\varphi(t-1+\delta/2)\right)^2dt } \right\} \end{align*} Proof: Section (ref).
remark\normalfont In our simulations and empirical application we set $\mathcal{X}$ to be a grid of 3000 evenly spaced points between $-6.5$ and $6.5$. This yields $\epsilon = 0.0003$ which is small compared to the critical values and test statistics in the application.

The meta-analyst has only a sample of $n$ $t$-scores $\{t_i\}_{i=1}^n$. They do not observe the vector of expectations $\theta$ but rather the analogous $J\times 1$ vector of sample means $\widehat{\theta}$:

equation[equation omitted — 96 chars of source]

The $n$ $t$-scores are reported by $m$ articles. While $t$-scores are dependent within an article, they must be independent across articles and the number of $t$-scores per article must be uniformly bounded.

assumptionIf $t_i,t_j$ are reported by different articles, then $t_i \protect\mathpalette{\protect\independenT}{\perp} t_j$. Furthermore there exists a universal constant $c_0>0$ such that no article reports more than $c_0$ $t$-scores.

This paper proposes $d_J(\widetilde{U},\widehat{\theta})$ as the test statistic for $\mathbf{H}_0$. The computational task of computing this projection distance is well-conditioned even though the problem of actually identifying the closest member of $\widetilde{U}$ is ill-conditioned. We compute projections with the \verb|osqp| package in R and set $J=30$ and $J=20$ in our empirical application and simulations.

The last task is to specify the critical value. First notice that the $J\times 1$ vector $\sqrt{n}\left(\widehat{\mathbf{\theta}}-\mathbf{\theta}\right)$ is asymptotically normal and its quantiles can be estimated with the bootstrap. This is immediate because $\widehat{\theta}$ is a vector sample means, the functions $\varphi_{\sigma^2_Y}(t_i)\psi_j(t_i)$ are uniformly bounded and everywhere infinitely differentiable, and Assumption (ref) guarantees weak dependence across observations.

The critical value for $d_J(\widetilde{U},\widehat{\theta})$ can be computed with a particular bootstrap. Specifically, the critical value for a test of size $\alpha$ will be the discretization error $\epsilon$ plus the $1-\alpha$ quantile of the distribution of $d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e})$ where $\mathbf{P}_{\widetilde{U}}\widehat{\theta}$ is the closest member of $\widetilde{U}$ to $\widehat{\theta}$ and $\sqrt{n}\mathbf{e}$ has the same asymptotic distribution as $\sqrt{n}\left(\widehat{\mathbf{\theta}}-\mathbf{\theta}\right)$.\footnote{This will work if $\mathbf{P}_{\widetilde{U}}\widehat{\theta}$ is any point that converges in probability at rate $n^{-1/2}$ to the member of $\widetilde{U}$ closest to $\theta$.} We compute this quantile by using the bootstrap to draw $\mathbf{e}$'s. The following paragraphs show the validity and consistency of this procedure in large samples.

theoremLet Assumption (ref) hold. Let $\sqrt{n}\mathbf{e}\to_d\sqrt{n}\left(\widehat{\mathbf{\theta}}-\mathbf{\theta}\right)$. For each $J \in \mathbb{N}$ there is some $N\in \mathbb{N}$ such that if $n>N$, then: $$ \mathbb{P}\left[ d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e}) > x+\epsilon|\widehat{\theta}\right] \leq \mathbb{P}\left[d_J(\widetilde{U},\widehat{\theta} )> x\right]\qquad \forall x\in \mathbb{R} $$ Proof: Section (ref).

Theorem (ref) says that if we know the distribution of $ d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e}) $ then we can construct critical values for $ d_J(\widetilde{U},\widehat{\theta} )$. We can learn the distribution of $d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e})$ by the bootstrap. We simply resample article-by-article to draw $\mathbf{e}^* = \widehat{\theta}^*-\widehat{\theta}$. The bootstrap is valid here because $\widehat{\theta}$ is a sample mean of uniformly bounded random variables and articles are assumed to have a bounded number of t-scores each. Resample many times to estimate the distribution of $d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e})$. By Theorem (ref), this distribution is sufficient to construct critical values. If $F$ is the CDF of the bootstrapped $d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e}^*)$, then the critical value is $\text{cv}(\alpha) = F^{-1}\left(1-\alpha\right)+\epsilon$. Our testing procedure is: $$\text{Reject } \mathbf{H}_0 \text{ if } d_J(\widetilde{U},\widehat{\theta}) > \text{cv}(\alpha) $$

Validity and Consistency

Size control is Corollary (ref) which is an immediate consequence of Theorem (ref).

corollaryLet Assumption (ref) hold. If $\mathbf{H}_0$ is true and $T|h\sim N(h,1)$: $$\limsup_{n\to \infty}\mathbb{P}\left[d_J(\widetilde{U},\widehat{\theta} ) > cv(\alpha)\right]\leq \alpha$$

Finally we show consistency. Theorem (ref) says that comparing $d_J(\widetilde{U},\widehat{\theta})$ to $cv(\alpha)$ is a consistent test given a sufficiently large $J$ and a sufficiently wide and dense grid for $\widetilde{U}$.

theoremLet Assumption (ref) hold. For every $\Pi_0$, if $\mathbf{H}_0$ is false, $T|h\sim N(h,1)$, and $J,L,\delta^{-1}$ are sufficiently large: $$\liminf_{n\to \infty}\mathbb{P}\left[d_J(\widetilde{U},\widehat{\theta} ) > cv(\alpha)\right]=1$$
proofBy Theorem (ref), $d_J(U,\theta)>0$. Set $L$ large enough and $\delta$ small enough that $d_J(\widetilde{U},\theta) >\epsilon $. Then $d_J(\widetilde{U},\widehat{\theta}) \to_p d_J(\widetilde{U}_n,\theta) > \epsilon $. Since $\limsup_{n\to \infty}cv(\alpha)\leq \epsilon$, the test must reject with probability approaching 1.

A key consequence of Theorem (ref) is that by increasing $J,L,\delta^{-1}$ with $n$, the test can be guaranteed to be consistent against any selective reporting process that is identified. In particular, we can always increase $J,L,\delta^{-1}$ slowly enough that all the results that assume these values are fixed still hold. Then $\liminf_{n\to \infty}\mathbb{P}\left[d_{J_n}(\widetilde{U}_n,\widehat{\theta} ) > cv(\alpha)\right]=1$ no matter how slowly $J,L,\delta^{-1}$ increase.

remark\normalfont In practice it is usually necessary to symmetrize the distribution of $t$-scores. $t$-scores of two-sided t-tests are usually reported as an absolute value. To account for this, take the sample of t-scores $\{t_i\}_{i=1}^n$ and replace it with $\{|t_i|,-|t_i|\}_{i=1}^n$. Under the null hypothesis this is equivalent to replacing $\Pi$ with its symmetrized version. When resampling with the bootstrap, first resample from the original dataset $\{t_i\}_{i=1}^n$ and then symmetrize. Symmetrization is smooth.
remark\normalfont In practice it is helpful for power to shift all the $t$-scores by 1.96 or another relevant critical value. This is because the norm $||\cdot||_Y$ weights down distortions in $f_{T|R}$ far from zero. But we expect distortions to be near $1.96$. Therefore, power may increase if the meta-analyst simply subtracts $1.96$ from each $t$-score in the dataset (after symmetrization).

Avoiding Spurious Rejections

So far we have assumed that $T|h \sim N(h,1)$ exactly in the absence of selective reporting. In practice, this normality assumption can be violated. We wish to reject ${\mathbf{H}_0}$ only when the distortions in the $t$-curve were likely due to selective reporting and to avoid spurious rejections originating from non-normality of $T$. We consider two main causes of violations: (1) the $t$-score is actually distributed student-$t$ instead of exactly normal and (2) the $t$-score was reported after rounding and then de-rounded by the meta-analyst. By augmenting the critical value, the meta-analyst can avoid spurious rejections from either of these two sources. The ease of this augmentation is one of the main advantages of the projection test.

The Approximately Normal Model

From the outset, we have set up this test so that it is simple to account for misspecification of the exact distribution of $T|h$. We now generalize the setup from Section (ref) by allowing the noise term in the $t$-score to have a distribution $g$ that is close—--but not equal—--to the standard normal. Suppose that the meta-analyst knows that the total difference between the two densities is upper bounded by some constant $\Delta \geq 0$. Assumption (ref) makes this situation explicit.

assumption$ T= h+U $ where $U \protect\mathpalette{\protect\independenT}{\perp} h$ and $U$ is continuous with unknown density $g(t)$ that satisfies: $$ \sup_{h\in\mathbb{R}}\left|\left|g(t-h) - \varphi(t-h)\right|\right|_Y \leq \Delta,\qquad ||g||_\infty <\infty $$ Where $\Delta \geq 0$ is a constant.

Under misspecification, the null hypothesis no longer guarantees the restriction in ((ref)), but instead implies the potentially untestable condition in ((ref)).

equation[equation omitted — 167 chars of source]

Since the meta-analyst does not know $g$, they cannot test ((ref)). Instead, they run the misspecified projection onto the Gaussian convolutions and upper bound the error using the triangle inequality. The inequality below shows that even though the projection is misspecified, the resulting residual can still be bounded.

align*[align* omitted — 189 chars of source]

The null hypothesis $\mathbf{H}_0^{\text{miss}}$ for the non-normal model is that the residual of the misspecified projection is small enough that it can be explained by a deviation from $N(0,1)$ of $\Delta$.

equation[equation omitted — 195 chars of source]

We test $\mathbf{H}_0^{\text{\normalfont miss}}$ using the same test statistic as before: $d_J(\widetilde{U},\widehat{\theta})$. The difference is that we now add $\Delta$ to the critical value.

The main limitation of this strategy is that it requires a known upper bound on $\Delta$. When the researcher is unwilling to commit to a specific value of $\Delta$, they can instead report the {\it breakdown statistic}, or the smallest value $\widehat{B}$ of $\Delta$ that would overturn their detection of selective reporting.

equation[equation omitted — 86 chars of source]

$\widehat{B}$ is the smallest deviation from normality that would be sufficient to explain away the observed lack of smoothness in the $t$-curve. In other words, it is the minimal value of $\Delta$ such that the test would no longer reject $\mathbf{H}_0$. The next two sections discuss how large $\widehat{B}$ needs to be for the meta-analyst to be confident that the rejection was probably not the spurious result of common distortions in the $t$-curve that do not come from selective reporting.

The student-$t$ distribution

In practice, the researcher does not know the true standard error of the point estimate with certainty. This means that when they construct the $t$-score, they divide the point estimate by an estimate of the standard error and the result is that $T-h$ is distributed student-$t$ with $\nu$ degrees of freedom, where $\nu$ is the effective sample size. We can calculate value of $\Delta$ associated with approximating the student-$t$ with the normal using numerical integration. For instance, for a study with effective sample size of $\nu=120$:

align*[align* omitted — 157 chars of source]

Therefore, if we had a set of studies with effective samples sizes predominantly larger than 120 and $\widehat{B}>.0009$, we could conclude that the detection of $p$-hacking was not just a spurious outcome of approximating the normal distribution with the student-$t$ distribution.

Derounding

In practice, $t$-scores are always reported after some rounding in practice. To avoid spurious detections of $p$-hacking resulting from rounding discontinuities, the meta-analyst must de-round the meta-data by adding random uniformly distributed noise commensurate to the number of significant digits elliott2024powertestsdetectingphacking. This process does not leave the $t$-curve completely undistorted. It is therefore prudent to check whether $\widehat{B}$ is small enough to be explained by derounding distortions to the $t$-curve.

We can model the derounded $t$-score as $T=h+Z+V$ where $V$ is mean zero noise independent of $h,Z$ that comes from derounding. Since $V$ is small and of expectation zero, we can bound $\Delta$ by bounding the sup-norm of the difference between the density $f_{Z+V}$ of $f_Z = \varphi$. Proposition (ref) bounds the distortion from normality $\Delta$ that can be incurred by rounding and derounding.

propositionSuppose that $Z\sim N(0,1)$, but instead of observing $T=h+Z$ we observe $T=h+Z+V$ where $V$ is continuous, $\mathbb{E}[V]=0$, $V\protect\mathpalette{\protect\independenT}{\perp} h+Z$, and $|V|<1$ almost surely. Then, $$\Delta < \sqrt{\frac{1}{2\pi}} \frac{\mathbb{E}[V^2]}{2}$$ Proof: Section (ref).

If the $t$-score is observed up to two decimal places (i.e. we can distinguish $T=1.96$ from $T=1.95$), then $|V|< 0.005$ with probability 1. Thus, when $\Delta \leq \sqrt{\frac{1}{2\pi}} \frac{.05^2}{2} \leq 0.00049$. This means that whenever $\widehat{B}> .00049$, we can conclude that the detection of $p$-hacking was not a spurious consequence of de-rounding. This is true of all of the breakdown statistics reported in our empirical application later on in Section (ref). We conclude that the projection test can be made robust to non-normality induced by derounding and that this robustness does not cost us enough power to matter in application.

Simulations

This section presents two sets of simulations that compare the power and size control of this paper's projection test with the six tests studied by Elliott. The takeaway message of these Monte Carlo exercises is that the projection test is as sensitive as the best of these six and controls size much better when $T|h$ is not exactly normal.

Figure (ref) displays power curves showing that the projection test is as sensitive as the most sensitive of the six Elliott tests. Here the DGP is as follows. The true effect $h$ equals 2 almost surely and $T$ is exactly normal centered on $h$. The researcher $p$-hacks with probability $q$ on the $x$-axis. So on the left part of the graph there is no $p$-hacking and on the right part of the graph everyone $p$-hacks. The $p$-hacking process matches Example (ref) where the researcher draws $T_1,T_2$ which share an $h$ and are otherwise iid. The researcher reports $T_1$ if it is greater than 1.96 in absolute value and otherwise reports $\max\{|T_1|,|T_2|\}$. Since the power curves for the projection test and the maximum of the six Elliott tests track each other almost exactly, we conclude that for this DGP no power is lost by using the projection test.

figure[figure omitted — 1,323 chars of source]

In a second simulation we show that the projection test in this paper can detect $p$-h

table[table omitted — 1,350 chars of source]
commentFigure (ref) displays size control as the number of studies in the meta-sample $n$ increases. Here there is no selective reporting at all and $h=0$. In addition $T-h$ is not exactly normal---instead it is the normalized sample mean of fifty iid uniforms. Since the support of $T$ is compact, the $p$-curve must slope upward near zero and the Cox-Shi tests of Elliott must reject with probability approaching one as $n$ grows even though there is no $p$-hacking. The graph in Figure (ref) shows that the size of the CS2B test and the LCM test grow as $n$ increases while the size of the projection test remains controlled--whether we adjust for $\Delta=.0014$ or not.\footnote{Adjusting for $\Delta$ makes our test conservative (undersized). Fortunately, we know that this adjustment does not make the test too conservative because is not enough to overturn the rejections in the empirical application later on.} Since the CS2B test is almost always equal to the upper envelope of power in Figure (ref), we conclude that in order for the tests of Elliott to match the sensitivity of our projection test, they must also accept some sensitivity to small deviations from normality of $T$. \begin{figure}[h!] \caption{ Simulated Size Curves. The $y$-axis is rejection rate (size) and the $x$-axis is the number of studies in the meta-sample $n$. We set $h=0$ almost surely and there is no $p$-hacking. Here $T$ is the normalized sample mean of 50 iid uniforms. The solid lines are the size of the projection test adjusting for $\Delta$ (blue) and unadjusted (orange). The adjustment of the critical values of $.0014$ is calculated in Example (ref). The dashed red line is the CS2B test from Elliott and the dashed purple line is their LCM test. Other tests from their study (not shown) were properly sized. Otherwise the simulation parameterization matches Figure (ref)} \end{figure}

Empirical Application

This section applies the projection test to the data from bb. This data consists of 21,740 two-sided t-statistics drawn from 684 articles. These articles were the universe of articles published in 25 top economics journals between 2015 and 2018 that used one of four causal methods. Each $t$-score corresponds to a two-sided $t$-test of a main hypothesis of interest. The coefficients in the numerators of the $t$-scores were estimated using Randomized Controlled Trials (RCT), Difference in Differences (DID), Regression Discontinuity Designs (RDD), or Instrumental Variables (IV). All “regression controls, constant terms, balance and robustness checks, heterogeneity of effects, and placebo tests" were excluded. Some tests were reported as $p$-values and bb converted these to $t$-scores using the standard normal distribution.

We run our tests on each of the four causal inference methods separately. The results are in Table (ref). We report results for $J=30$ and $J=20$ to demonstrate that our conclusions are not very sensitive to the choice of smoothing parameter $J$. The $p$-values in columns 4 and 7 mostly confirm the findings from kudrinjmp24. We detect selective reporting for RCTs, IVs, and DIDs.

Are these rejections spurious outcomes coming from (de)rounding or the limited degrees of freedom of the underlying studies? To determine this, we compute the breakdown statistic $\widehat{B}$ which is the smallest distortion in the $t$-curve from non-$p$-hacking sources that it would take to explain the observed $t$-curve and overturn our rejection. The smallest $\widehat{B}$ in Table (ref) is 0.0010. Section (ref) shows that when the effective sample size of the underlying studies are greater than $\nu=120$, the student-$t$ distribution cannot explain a breakdown statistic of 0.0010 or larger. Are the effective sample sizes in this data small enough to explain this? To find out we took a random sample of 100 RCT papers, 100 IV papers, and 100 DID papers from bb. We calculated their effective sample sizes as the number of clusters where the cluster was the level at which the standard error was clustered. The median effective sample sizes are reported in Table (ref). We calculate $\Delta_\nu$ as the average over all the effective sample sizes $\nu_j$ of $\frac{1}{100}\sum_{j=1}^{100}||\varphi - t_{\nu_j}||_Y$, which can be computed directly by numerical integration.\footnote{We were unable to determine even a lower bound on effective sample size for 20/100 DID papers and 15/100 IV papers.} For all three methods, $\Delta_\nu < \widehat{B}$. This means that the effective sample sizes were overall too large to explain the breakdown statistic and thus the rejection of $\mathbf{H}_0$ stands.

In Section (ref), we showed that if the $t$-score is observed up to two decimal places, (de)rounding can add at most $\sqrt{\frac{1}{2\pi}} \frac{.05^2}{2} \leq 0.00049$ to the test statistic. Since $\widehat{B} > 0.00049$ everywhere in Table (ref) we conclude that (de)rounding cannot explain the non-smoothness of the $t$-curves in the bb data.

These empirical results add a novel insight to existing knowledge. The $t$-curves for RCTs, IVs, and DIDs are significantly distorted. These distortions are larger than can be explained by random chance, (de)rounding, or the student-$t$ distribution. We conclude that the most likely culprit is selective reporting of $t$-scores. The projection test proposed in this paper is the first (to our knowledge) that is simultaneously sensitive enough to detect these distortions in bb's sample and also robust to basic concerns related to (de)rounding, the unknown distribution of true effects, and finite degrees of freedom.

table[table omitted — 1,496 chars of source]