EconBase
← Back to paper

A Sharp and Robust Test for Selective Reporting

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

55,545 characters

When is $p$-hacking detectable?



\maketitle
\onehalfspacing
\begin{abstract}
    Some forms of $p$-hacking cannot be detected by examining the $t$-curve (or $p$-curve). Standard tests may also fail to find even detectable forms of selective reporting. We propose a novel test that is consistent against {\it every} detectable form of $p$-hacking and remains interpretable even when the $t$-scores are not exactly normal. The test statistic is the distance between the smoothed empirical $t$-curve and the set of all distributions that would be possible in the absence of any selective reporting. This novel {\it projection test}  can only be evaded in large meta-samples by selective reporting that also evades all other valid tests of restrictions on the $t$-curve. A second benefit of the projection test is that under the null hypothesis of no $p$-hacking we can check whether the projection residual could have been produced by other distortions not related to selective reporting, e.g. rounding and de-rounding. Applying the test to the \cite{bb} meta-data, we find that the $t$-curves for RCTs, IVs, and DIDs are more distorted than could arise by chance. We confirm that these distortions cannot be explained by (de)rounding of $t$-scores or by the limited degrees of freedom of the underlying studies.
\end{abstract}

\vspace{0.1in}
\noindent\textbf{Keywords:} $p$-Hacking, Deconvolution

\vspace{0.1in}
\noindent\textbf{JEL Codes:} C12, C14


\pagenumbering{arabic}



\section{Introduction}

Publication bias and $p$-hacking threaten the validity of empirical research. When researchers report only the most surprising statistical results, they are likely to over-claim. Scientific findings are therefore more trustworthy when they are drawn from literatures where selective reporting is not evident.

 Learning what kinds of studies are $p$-hacked helps evaluate empirical claims and target reform. Much recent work has tested for selective reporting in a given literature by checking for distortions in the empirical distribution of reported $t$-scores.\footnote{See for example: \cite{Simonsohn,Head,Brodeur,havranek_housing,bb,Elliott,elliott2024powertestsdetectingphacking,Kudrinjmp,kudrinjmp24,Havranek24}. The $p$-curve is distorted if and only if the $t$-curve is distorted. } All of these methods share a common shortcoming: They can be evaded by selective reporting strategies that do not violate the specific restrictions that the meta-analyst tested. For example, the Caliper test can be evaded by $p$-hacking that does not produce a discontinuity in the histogram of $t$-scores at the critical value. The bounds tested in \cite{Elliott} can be evaded by $p$-hacking that distorts the overall shape of the $p$-curve but does not make its derivatives extreme at any particular point.

 This paper proposes a new test that can only be evaded by forms of $p$-hacking that are totally undetectable by any examination of the  $t$-curve. To our knowledge this is the first test that comes accompanied by such a guarantee. Our proposed test statistic is the distance between the observed $t$-curve and the set of all feasible  $t$-curves under no selective reporting. The testable restriction that this distance equals zero is {\it sharp} because if a p-hacking process did not violate it, then the resulting $t$-curve could be explained without selective reporting by some distribution of true treatment effects.

 Our test can only be evaded in large meta-samples by $p$-hacking processes that are undetectable in principle. We illustrate with a simple example: when a researcher reports the maximum of two correlated one-sided $t$-scores and the distribution of true effects is smooth, then the resulting $t$-curve can be explained either with or without $p$-hacking. This is an example of a kind of selective reporting that is undetectable by any examination of the $t$-curve. Outside of such undetectable cases, our test is consistent against all forms of selective reporting.


 We call our test a \textit{projection test}. The key technical step is to compute the test statistic using the Singular Value Decomposition (SVD). The SVD allows us to rewrite the projection of one smoothed density onto a set of other densities as simply the orthogonal projection of a vector of sample means onto a convex set of vectors in Euclidean space. Critical values are computed using the bootstrap.

Simulations show that  the projection test is as powerful as the most sensitive of the six tests studied by \cite{Elliott}. Further simulations show an example where $p$-hacking is nearly undetectable, the projection test is consistent against it while the six tests studied by \cite{Elliott} all have power no greater than size in a large meta-sample. Therefore the projection test makes a meaningful contribution to the existing body of tests for $p$-hacking.

 The second major advantage of the projection test is that the test statistic can still be interpreted when the $t$-scores reported by individual researchers are not exactly normally distributed. Recent nonparametric tests for $p$-hacking that examine every part of the $t$-curve like \cite{kudrinjmp24} and \cite{Elliott} are highly sensitive but will reject the null hypothesis of no selective reporting whenever the empirical $t$-curve is not smooth enough to have been generated by a convolution with a normal kernel. Yet, non-smoothness in the $t$-curve can potentially arise for other reasons not related to selective reporting. For example, the $t$-score is more accurately described as following a student-$t$ distribution with $\nu$ degrees of freedom, where $\nu$ is the effective sample size of the experiment. In addition, the $t$-score is often reported after rounding. Even when de-rounded, this process introduces non-smoothness in the distribution of the $t$-score that could theoretically lead to spurious rejections of the null of no p-hacking.

 We account for finite-sample deviations from the normal distribution in the following way. We first calculate the {\it Breakdown Statistic} $\widehat{B}$ defined as how severe the deviation from normality must be to overturn the meta-analyst's detection of $p$-hacking.  The breakdown statistic is computed from the observed $t$-curve and is equal to the difference between the test statistic and the critical value. We then compare the breakdown statistic to the maximum possible deviations from normality that could arise from either the student-$t$ distribution or from (de)rounding. If the breakdown statistic is larger than could be explained by either of these sources of non-smoothness, then the rejection stands.

 We apply our method to the meta-sample of 21,740 two-sided $t$-tests reported by 684 top economics journals collected by \cite{bb}. We find that the $t$-curves for RCTs, DIDs, and IVs are significantly distorted and that these distortions are too large to be explained either by (1) rounding or de-rounding of the reported $t$-scores or by the limited degrees of freedom of the studies themselves. We conclude that significant selective reporting is likely present in all three subgroups of studies. We cannot say the same of RDDs.

This paper is organized as follows. Section \ref{sec:setup} sets up the problem. Section \ref{sec:h0} defines the meta-analyst's null hypothesis and provides examples of $p$-hacking processes that are detectable and undetectable by any test. Section \ref{sec:correct_spec} presents the projection test and shows its validity and consistency when the $t$-score is exactly normal. Section \ref{sec:misspecification} extends to situations where the $t$-score is only approximately normal even in the absence of selective reporting. Section \ref{sec:sims} presents simulations and Section \ref{sec:application} shows an empirical application. Proofs are in Appendix \ref{sec:proofs}.

\section{Setup}\label{sec:setup}

First define several key pieces of notation. Let $\varphi(z)$ denote the probability density function of the standard normal distribution. Adding a subscript $\varphi_{\sigma^2}(z)$ denotes the density of the normal distribution with variance $\sigma^2$. No subscript means that the variance is unity.   Let  $\mathcal{F}_\infty$ denote the set of valid PDFs of continuous real-valued random variables with bounded height.
 \begin{equation}
\mathcal{F}_\infty \equiv  \left\{ g : \mathbb{R} \to [0, \infty) \,\middle|\, \int_{\mathbb{R}} g(h)\,dh = 1,\ \|g\|_\infty < \infty \right\}.
 \end{equation}



Consider a population of t-tests. Each researcher first draws the unobserved latent true effect $h\in\mathbb{R}$ from an unknown distribution $\Pi_0$. We call $\Pi_0$ the {\it distribution of true effects}. The researcher then reports the $t$-statistic, which is $h$ plus normal noise $ T\: |\: h \sim N(h,1) $.   The conditional probability density $f_{T|h}\left(t\:|\:h\right)$ of $T$  is:
\begin{equation}
   f_{T|h}\left(t\:|\:h\right) = \varphi(t-h)
\end{equation}


The unconditional density of $T$ is the expectation over the conditional density.
\begin{equation}
    f_{T}(t)  =\mathbb{E}_{\Pi_0}\left[\varphi(t-h)\right]= \int_{-\infty}^\infty \varphi(t-h)  d\Pi_0(h)
\end{equation}




Selective reporting means that the meta-analyst does not observe a random sample of $T$ because some $t$-scores are truncated. Let $R$ be the event that $T$ is reported.  The meta-analyst instead observes draws from the conditional distribution of the reported $t$-scores: $T|R$. Reporting is called {\it selective} if $R$ depends on the realization of $T$. For example, small values of $T$ may go unreported (e.g. by publication bias) or the researcher may repeatedly draw correlated $T$ with the same $h$ until they draw a large one and then stop. The form of selective reporting can be almost anything. In this paper, we will focus on two basic examples of selective reporting described below.

\begin{comment}
    \begin{example}\label{ex:pb}{\bf Publication Bias}\normalfont

   \noindent Consider a process where some $t$-scores are truncated from the meta-analyst's sample and reporting $R$ is independent across $t$-scores (even those in the same manuscript).
Under publication bias, the meta-analyst only observes a random sample of conditional draws of $T|R$ and not an unconditional random sample of $T$. Applying Bayes' Theorem and assuming that $\mathbb{P}[R]>0$ guarantees that $T|R$ has the following conditional density:
\begin{equation}
    f_{T|R}(t) = \frac{ \mathbb{P}[R|T=t]f_T(t)}{\mathbb{P}[R]}
\end{equation}

\noindent This includes the ``classical" publication bias where $t$-scores are reported with probability $p$ if they are significant and with probability $q$ if not \citep{Andrews}. In this case $p(t)$ would be a step function.
\end{example}

\end{comment}


\begin{example}\label{ex:minimiztion}{\bf Maximization $p$-hacking} \normalfont

    \noindent  Consider a $p$-hacking process where a researcher first draws a single unobserved true effect $h$. Then they draw two correlated $t$-scores $\{T_1,T_2\}$ both with marginal distributions $N(h,1)$ and covariance $\rho \in (-1,1)$. They report the larger of the two $t$-scores (where ``larger" means rightmost on the number line). This type of $p$-hacking produces no jumps in the $t$-curve and is known to be hard to detect with existing tests \citep{elliott2024powertestsdetectingphacking, kudrinjmp24}. In Section \ref{sec:h0}, we show that when the distribution of $h$ is smooth, this kind of $p$-hacking is undetectable in principle by {\it any} test.
\end{example}

\begin{example}\label{ex:threshold}{\bf Threshold $p$-hacking}\normalfont

\noindent Suppose that the researcher $p$-hacks only if their initial result is insignificant. Consider a similar setup to Example \ref{ex:minimiztion} where a researcher draws two correlated $t$-scores with the same $h$. Now the researcher reports $T_1$ if $|T_1|>1.96$. Otherwise, they report $\max\{T_1,T_2\}$.  This kind of $p$-hacking
does produce jumps in the $t$-curve $f_{T|R}$ and sensitive tests have been proposed \citep{elliott2024powertestsdetectingphacking, kudrinjmp24}.
\end{example}

\section{Null Hypothesis and Detectability} \label{sec:h0}
In principle, selective reporting can take many forms and is not limited to the three illustrative examples above. The meta-analyst does not know the form of selective reporting or the distribution of true effects and will not attempt to learn these. Instead, they wish only to determine whether the distribution of reported $t$-scores $f_{T|R}$ can be explained without selective reporting. We now state the meta-analyst's null hypothesis ${\mathbf{H}_0}$ in Equation (\ref{eq:h0}) below.
\begin{equation}\label{eq:h0}
     {\mathbf{H}_0:}\quad \exists \Pi \in \mathcal{P} \quad \text{ s.t. } \quad T|R  \stackrel{d}{=}  h'+Z\qquad h'\sim \Pi,\: Z\sim N(0,1)
\end{equation}

The meta-analyst's null hypothesis ${\mathbf{H}_0}$ states that there is some distribution of true effects $\Pi$ that explains the population $t$-curve without any selective reporting. This is true if and only if the PDF $f_{T|R}$ is sufficiently smooth.\footnote{To be precise, ${\mathbf{H}_0}$ is true when the deconvolved function $e^{t^2/2}\phi_{T|R}$ is a valid characteristic function where $\phi_{T|R}$ is the characteristic function of the reported $t$-score. This can only occur when $\phi_{T|R}$ has thin tails, i.e. the PDF $f_{T|R}$ is smooth.} If $f_{T|R}$ has a discontinuity or oscillates rapidly, then it could not have arisen ``naturally" and ${\mathbf{H}_0}$ is violated. While some violations of ${\mathbf{H}_0}$ are apparent to the naked eye (e.g. jumps or kinks), others are not.

When ${\mathbf{H}_0}$ is false, then something has distorted the $t$-curve. However, the converse does not hold. If ${\mathbf{H}_0}$ is true, then either there is no selective reporting, or selective reporting is {\it undetectable}. We formally define the term  ``undetectable" below.

\begin{definition}
    We say that selective reporting is {\bf undetectable} if $T \stackrel{d}{\neq }T|R$ but ${\mathbf{H}_0}$ is true.  We say that selective reporting is {\bf detectable} if ${\mathbf{H}_0}$ is false and $T\sim N(h,1)$.
\end{definition}

Undetectable forms of selective reporting change the distribution of $T$, but keep it smooth. Thus, undetectable $p$-hacking yields a $t$-curve that is still so smooth that it could have been produced ``naturally" without any selective reporting by some other distribution of true effects. We stress that it is impossible to design a valid test that has power against undetectable forms of selective reporting. This is because when $\mathbf{H}_0$ is true, the population distribution of $T|R$ can be explained without selective reporting. Below we return to the three illustrative examples from the previous section and show that most forms of publication bias and threshold $p$-hacking are detectable, but whether maximization $p$-hacking is detectable depends on the smoothness of $\Pi_0$.

\begin{comment}


\vskip 0.07in
\noindent {\bf Example \ref{ex:pb}. Publication Bias (continued)}

\noindent Consider the publication bias setup from Example \ref{ex:pb}. Suppose that the conditional probability of publication is nonsmooth---meaning that either $\mathbb{P}[R|T=t]$ or any of its derivatives in $t$ has a jump discontinuity somewhere. For instance, the probability of publication might increase sharply at the critical value $1.96$ as in \cite{Andrews}. Then $f_{T|R}$ cannot be smooth. Since all convolutions with the normal distributions are smooth, $f_{T|R}$ could not be explained without selective reporting and $\mathbf{H}_0$ is false. So non-smooth publication bias is {\bf detectable}.

\end{comment}

\vskip 0.07in
\noindent {\bf Example \ref{ex:minimiztion}. Maximization $p$-hacking (continued)}

\noindent Consider the Maximization $p$-hacking  setup from Example \ref{ex:minimiztion}. This kind of $p$-hacking produces a continuous $t$-curve with no discontinuities or kinks. Is it detectable? Surprisingly, this turns out to depend on the distribution of true effects $\Pi_0$! We show in the Appendix \ref{proof:ex:minimization} that if $h$ is any mixture of normals all with variance at least $2(1-\rho)$, then this form of $p$-hacking is {\bf undetectable.} If instead the distribution of $h$ places any amount of probability mass at zero, then this form of $p$-hacking is {\bf detectable}. The intuition for the result is that when $\Pi_0$ is smooth, then the $t$-curve becomes so smooth that reporting only the larger $t$-score doesn't take away enough smoothness to reveal itself. The implication for theory is that detectability depends jointly on the distribution of true effects and the form $p$-hacking takes. The implication for practice is that many forms of $p$-hacking are totally undetectable by any examination of the $t$-curve or $p$-curve.

\vskip 0.07in
\noindent {\bf Example \ref{ex:threshold}. Threshold $p$-hacking (continued)}\\
\noindent Consider the Threshold $p$-hacking setup from Example \ref{ex:threshold}. Threshold $p$-hacking is non-smooth and will produce a jump discontinuity in the $t$-curve. Since any $f_T$ must be smooth regardless of the distribution of true effects, $f_{T|R}$ under threshold $p$-hacking cannot be explained without selective reporting. So ${\mathbf{H_0}}$ is false and threshold $p$-hacking is {\bf detectable}.
\vskip 0.07in

$p$-hacking is detectable when it removes enough smoothness from the $t$-curve to reveal itself. Figure \ref{fig:detectability} below visualizes detectability in three examples. All three simulated $t$-curves were subject to severe $p$-hacking but the detectability of that hacking varies widely. The red solid line is the $t$-curve after threshold-type $p$-hacking (described in Example \ref{ex:threshold}). This is easy to detect because of the clear kink in the curve caused by the threshold. The solid blue line is the $t$-curve after maximization-type $p$-hacking where the researcher reports the largest of eight independent $t$-scores (Example \ref{ex:minimiztion}) and every $h=0$. The non-smoothness is subtle, diffuse throughout the $t$-curve, and not visually apparent. Nevertheless $p$-hacking is still detectable here (and the test proposed later in this paper is consistent against it). Finally, the dashed blue line is an identical DGP to the solid blue line except $h\sim N(0.5,1)$. Now the smoothness of the distribution of true effects fully masks the selective reporting and the $p$-hacking is undetectable despite its severity.

\begin{figure}
    \centering
    \includegraphics[width=0.7\linewidth]{figures/Detectability.pdf}
    \caption{Three examples of severe $p$-hacking with varying detectability.}
    \label{fig:detectability}
\end{figure}

\section{A Consistent Test}\label{sec:correct_spec}

In this section we rewrite $\mathbf{H}_0$ in a testable form and then propose a consistent test for it. There are three main steps. First, we express $\mathbf{H}_0$ as a zero-distance restriction. Second, we rewrite the zero-distance restriction in the spectral domain as the sum of sample moments. Finally, we specify the test statistic and critical values.

The first step is to express the null hypothesis as a restriction on the density $f_T$. To express this restriction we first need to specify a notion of distance between distributions. We choose the following norm because it allows us to naturally take the Singular Value Decomposition later on. Fix $\sigma_Y^2>0$; throughout we set $\sigma_Y^2=1$. Define the norm $||\cdot||_Y\: : \: \mathcal{F}_{\infty}\to \mathbb{R}^+$ as:\begin{equation}
    ||g||_Y \equiv \sqrt{\int_{-\infty}^\infty g(t)^2\varphi_{\sigma_Y^2}(t)dt}
\end{equation}

 We now state the testable restriction in Equation (\ref{eq:testable_restriction}) below. The restriction says that the distance between  $f_{T|R}$ and the set of all possible densities of $f_T$ for which $h$ has a continuous distribution $\pi \in \mathcal{F}_\infty$ is zero. This essentially means that the observed $t$-curve cannot be separated from the possible set of undistorted $t$-curves.
 \begin{equation}{\label{eq:testable_restriction}}
    \inf_{\pi \in \mathcal{F}_\infty } \left|\left|f_{T|R}- \int_{-\infty}^\infty \varphi(t-h)\pi(h)dh \right|\right|_Y =0
\end{equation}

\begin{remark}
    \normalfont Importantly, even though the minimization in  (\ref{eq:testable_restriction}) is over all continuous distributions, we do not assume that $\Pi_0$ is continuous. Notice that $\mathcal F_\infty$ is dense in the space of bounded densities under the $||\cdot||_Y$ norm. So the non-continuous $\Pi$ lie in the boundary of the set of continuous $\Pi$. We will show that excluding discontinuous distributions from the minimization will not change the infimum.
\end{remark}

Theorem \ref{thm:testable_implication} below shows that Restriction  (\ref{eq:testable_restriction}) is sharp, i.e. identical to the null hypothesis. This means that selective reporting is detectable if and only if it separates $f_{T|R}$ from the set of undistorted $t$-curves in the $||\cdot||_Y$ norm.
\begin{theorem}\label{thm:testable_implication}
  If $T|h \sim N(h,1)$, then  $$\mathbf{H}_0 \iff  (\ref{eq:testable_restriction})$$
  Proof: Section \ref{proof:thm:testable_implication}
\end{theorem}



Restriction (\ref{eq:testable_restriction}) is expressed in the PDF domain, which makes it easy to interpret but hard to see how it can be tested. The next step is to rewrite the restriction in the discrete spectral domain.  Specifically, we use the Singular Value Decomposition to express the null hypothesis in terms of expectations. The meta-analyst can easily test restrictions on a set of expectations by plugging in sample means.




The Singular Value Decomposition expresses the mapping from a density of true effects $\pi\in \mathcal{F}_{\infty}$ into an undistorted $t$-curve $f_T$ in terms of a weighted sum of orthonormal basis polynomials where the weights depend on expectations over $\pi$. Below we take the Singular Value Decomposition from \cite{CarrascoPaper} and recently repurposed by \cite{faridani2025testingunderpoweredliteratures}.
\begin{align*}
   \int_{-\infty}^\infty \varphi(t-h)\pi(h)dh &=  \sum_{j=0}^\infty \eta_j \mathbb{E}_\pi\left[\varphi_{1+\sigma_Y^2}(h)\chi_j(h)\right] \psi_j(t),\qquad \eta_j \equiv \left(\frac{\sigma_Y^2}{1+\sigma_Y^2}\right)^{j/2}
\end{align*} The singular functions are scalings of the  Hermite polynomials: $\chi_j(t) = \frac{1}{\sqrt{j!}} He_j\left(\frac{t}{\sqrt{1+\sigma_Y^2}}\right)$, $\psi_j(t) =\frac{1}{\sqrt{j!}} He_j\left(\frac{t}{\sigma_Y}\right)$.\footnote{The  Hermite Polynomials are defined as: $
He_j(t)=\sum_{l=0}^{[j/2]}(-1)^l \frac{(2l)!}{2^ll!}\binom{j}{2l}t^{j-2l}$. } It is useful to know that the functions $\chi_j(h)\varphi_{\sigma_Y^2+1}(h)$ and $\psi_j(t)\varphi_{\sigma_Y^2}(t)$ are uniformly bounded over all $t,h,j$ and that the Hermite polynomials form orthonormal bases of of a weighted $\mathcal{L}^2$ space that contains $\mathcal{F}_\infty$ in the inner products $\langle \cdot,\cdot \rangle_Y$ and $\langle \cdot,\cdot \rangle_X$ defined in Section \ref{proof:lem:bound_coeffs}.



The Singular Value Decomposition allows us to express the difference between two distributions either in the PDF domain (where differences are easy to interpret) or the spectral domain (where differences are easy to calculate). Specifically Lemma \ref{lem:pdf_spectral} rewrites the distance between two probability densities in terms of the Euclidean distance between vectors of expectations.

\begin{lemma}\label{lem:pdf_spectral}
    $$ \left|\left|f_{T|R}- \int_{-\infty}^\infty \varphi(t-h)\pi(h)dh \right|\right|_Y =\sqrt{  \sum_{j=0}^\infty    \left(\mathbb{E}\left[\psi_j(T)\varphi_{\sigma_Y^2}(T)|R\right] - \eta_j\mathbb{E}_\pi\left[\chi_j(h)\varphi_{\sigma_Y^2+1}(h)\right]\right)^2 }$$

    Proof: Section \ref{proof:lem:pdf_spectral}
\end{lemma}

Now that we can express the distance between two PDFs as the sum of squared differences in expectations, we can rewrite the projection problem in terms of vectors in Euclidean space. Specifically, we will show that projecting $f_{T|R}$ onto the set of all possible $t$-curves is equivalent to projecting a vector onto a certain convex set. Define $\theta$ as the vector of expectations with countable elements below. Here the element $\theta_j$ is the $j^{\text{th}}$ weighted Hermite coefficient of the observed $t$-curve. This vector fully characterizes $f_{T|R}$ in the spectral domain.
\begin{equation}
    \theta_j= \mathbb{E}\left[\psi_j(T)\varphi_{\sigma_Y^2}(T)|R\right]
\end{equation}
 Define $\mathbf{b}_x$ as the vector  with elements: $b_{x,j}=\eta_j\mathbb{E}\left[\chi_j(x)\varphi_{\sigma_Y^2+1}(x)\right] $. Define the set $U$ as the convex hull of all $\mathbf{b}_x$ over all $x\in \mathbb{R}$. Since the functions $\psi_j(t)\varphi_{\sigma_Y^2}(t)$ are continuous and uniformly bounded, $U$ is compact.

 We would like to project $\theta$ onto $U$. But, since $\theta$ and $U$ are infinite-dimensional, the projection is computationally infeasible. To make it feasible, we regularize the projection by spectral cutoff, i.e. we only consider the first $J$ elements of the vectors involved. Define $ d_J({U},\mathbf{v})$ as the residual of the $||\cdot||_2$ projection of any vector $\mathbf{v}$ onto $U$ after truncation to the first $J$ elements:
 \begin{equation}
     d_J({U},\mathbf{v}) \equiv \inf_{\mathbf{u}\in U}\sqrt{ \sum_{j=1}^J (v_j-u_j)^2}
 \end{equation}

We can now state Restriction (\ref{eq:restriction_J}) which (we will show) is the regularized version of (\ref{eq:testable_restriction}).
\begin{equation}\label{eq:restriction_J}
    d_J(U,\theta)=0
\end{equation}

Testing (\ref{eq:restriction_J}) seems much easier than testing (\ref{eq:testable_restriction}) directly because the objects involved are vectors in $\mathbb{R}^J$. We now show that this is as good as testing $\mathbf{H}_0$ when $J$ is large enough. First we show validity. Theorem \ref{thm:validity_dj} below says that $\theta$ is no farther from $U$ than $f_{T|R}$ is from the set of all possible PDFs under the null. A valid test of $d_J(U,\theta) = 0$ will therefore also always be a valid test of $\mathbf{H}_0 $.


\begin{theorem}\label{thm:validity_dj} For any $J\in\mathbb{N}$,

$$ d_J(U,\theta)\leq \lim_{K\to \infty}d_{K}(U,\theta) =  \inf_{\pi \in \mathcal{F}_\infty }\left|\left|f_{T|R}- \int_{-\infty}^\infty \varphi(t-h)\pi(h)dh \right|\right|_Y$$

Proof: Section \ref{proof:thm:validity_dj}

\end{theorem}

A consistent test of $d_J(U,\theta)=0$ will also be consistent against all alternatives when $J$ is large enough. Theorem \ref{thm:consistency_dj_large} below shows that when $\mathbf{H}_0$ is false and $J$ is large enough, a consistent test of $d_J(U,\theta) = 0$  is also a consistent test of $\mathbf{H}_0$ for detectable selective reporting.

\begin{theorem}\label{thm:consistency_dj_large}
  If $\mathbf{H}_0$ is false and selective reporting is detectable, $\exists J^*$ such that for all $J>J^*$
    $$d_J(U,\theta)>0$$
    Proof: Section \ref{proof:thm:consistency_dj_large}
\end{theorem}

Theorem \ref{thm:validity_dj} and \ref{thm:consistency_dj_large} prove that for large enough $J$, a valid and consistent test of (\ref{eq:restriction_J}) is also a valid and consistent test of any detectable violation of $\mathbf{H}_0$.

\subsection{Test Statistic and Critical Values}

Next we define the test statistic explicitly. Projecting a vector onto $U$ is computationally challenging  because the basis of the convex hull $U$ is uncountable. But it is feasible to project onto $\widetilde{U}$, the convex hull of the $J\times 1$ vectors $\mathbf{u}_x=\eta_j\mathbb{E}\left[\chi_j(x)\varphi_{\sigma_Y^2+1}(x)\right] $ over a finite set of $x\in \mathcal{X}\subset \mathbb{R}$. This incurs some approximation error that can be forced to be as small as we want by using a grid $\mathcal{X}$ wide and fine enough. Lemma \ref{lem:max_grid_approx_error} tells us how wide and fine the grid must be in order to control the approximation error. Notice that the bound holds uniformly across all $J\in \mathbb{N}$.

\begin{lemma}\label{lem:max_grid_approx_error}
Let $\mathcal{X}$ be a grid of evenly spaced points between $-L$ and $L$ with spacing $\delta$. For any $\mathbf{v}\in \mathbb{R}^J$,\begin{align*}
    \left|d_J({U},\mathbf{v})-d_J(\widetilde{U},\mathbf{v})\right| &< \max\left\{\sqrt{\int_{-\infty}^\infty \varphi_{\sigma_Y^2}(t)\varphi(t-L)^2dt },\sqrt{\int_{-\infty}^\infty \varphi_{\sigma_Y^2}(t)\left(\varphi(t-1)-\varphi(t-1+\delta/2)\right)^2dt } \right\}
\end{align*}

Proof: Section \ref{proof:lem:max_grid_approx_error}.
\end{lemma}


\begin{remark} \normalfont
    In our simulations and empirical application we set $\mathcal{X}$ to be a grid of 3000 evenly spaced points between $-6.5$ and $6.5$. This yields $\epsilon = 0.0003$ which is small compared to the critical values and test statistics in the application.
\end{remark}

The meta-analyst has only a sample of $n$ $t$-scores $\{t_i\}_{i=1}^n$. They do not observe the vector of expectations $\theta$ but rather the analogous $J\times 1$ vector of sample means $\widehat{\theta}$:
\begin{equation}
 \widehat{\theta}_j =\frac{1}{n}\sum_{i=1}^n\varphi_{\sigma_Y^2}(t_i)\psi_j(t_i)
\end{equation}

\noindent The $n$ $t$-scores are reported by $m$ articles. While $t$-scores are dependent within an article, they must be independent across articles and the number of $t$-scores per article must be uniformly bounded.
\begin{assumption}\label{assum:articles}
    If $t_i,t_j$ are reported by different articles, then $t_i \protect\mathpalette{\protect\independenT}{\perp} t_j$. Furthermore there exists a universal constant $c_0>0$ such that no article reports more than $c_0$ $t$-scores.
\end{assumption}

This paper proposes $d_J(\widetilde{U},\widehat{\theta})$ as the test statistic for $\mathbf{H}_0$. The computational task of computing this projection distance is well-conditioned even though the problem of actually identifying the closest member of $\widetilde{U}$ is ill-conditioned. We compute projections with the \verb|osqp| package in R and set $J=30$ and $J=20$ in our empirical application and simulations.

The last task is to specify the critical value. First notice that the $J\times 1$ vector $\sqrt{n}\left(\widehat{\mathbf{\theta}}-\mathbf{\theta}\right)$ is asymptotically normal and its quantiles can be estimated with the bootstrap. This is immediate because $\widehat{\theta}$ is a vector sample means, the functions $\varphi_{\sigma^2_Y}(t_i)\psi_j(t_i)$ are uniformly bounded and everywhere infinitely differentiable, and Assumption \ref{assum:articles} guarantees weak dependence across observations.


The critical value for $d_J(\widetilde{U},\widehat{\theta})$  can be computed with a particular bootstrap. Specifically, the critical value for a test of size $\alpha$ will be the discretization error $\epsilon$ plus the $1-\alpha$ quantile of the distribution of $d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e})$ where $\mathbf{P}_{\widetilde{U}}\widehat{\theta}$ is the closest member of $\widetilde{U}$ to $\widehat{\theta}$ and $\sqrt{n}\mathbf{e}$ has the same asymptotic distribution as  $\sqrt{n}\left(\widehat{\mathbf{\theta}}-\mathbf{\theta}\right)$.\footnote{This will work if $\mathbf{P}_{\widetilde{U}}\widehat{\theta}$ is any point that converges in probability at rate $n^{-1/2}$ to the member of $\widetilde{U}$ closest to $\theta$.} We compute this quantile by using the bootstrap to draw $\mathbf{e}$'s. The following paragraphs show the validity and consistency of this procedure in large samples.


\begin{theorem}\label{thm:Ru_thetahat}
     Let Assumption \ref{assum:articles} hold. Let $\sqrt{n}\mathbf{e}\to_d\sqrt{n}\left(\widehat{\mathbf{\theta}}-\mathbf{\theta}\right)$. For each $J \in \mathbb{N}$ there is some $N\in \mathbb{N}$ such that if $n>N$, then:
    $$ \mathbb{P}\left[ d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e}) > x+\epsilon|\widehat{\theta}\right] \leq \mathbb{P}\left[d_J(\widetilde{U},\widehat{\theta} )> x\right]\qquad \forall x\in \mathbb{R}  $$

    Proof: Section \ref{proof:thm:Ru_thetahat}.
\end{theorem}

Theorem \ref{thm:Ru_thetahat} says that if we know the distribution of $ d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e}) $ then we can construct critical values for $ d_J(\widetilde{U},\widehat{\theta} )$. We can learn the distribution of $d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e})$ by the bootstrap.  We simply resample article-by-article to draw $\mathbf{e}^* = \widehat{\theta}^*-\widehat{\theta}$. The bootstrap is valid here because $\widehat{\theta}$ is a sample mean of uniformly bounded random variables and articles are assumed to have a bounded number of t-scores each. Resample many times to estimate the distribution of $d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e})$. By Theorem \ref{thm:Ru_thetahat}, this distribution is sufficient to construct critical values. If $F$ is the CDF of the bootstrapped $d_J(\widetilde{U},\mathbf{P}_{\widetilde{U}}\widehat{\theta}+\mathbf{e}^*)$, then the critical value is $\text{cv}(\alpha) = F^{-1}\left(1-\alpha\right)+\epsilon$. Our testing procedure is:
$$\text{Reject } \mathbf{H}_0 \text{ if } d_J(\widetilde{U},\widehat{\theta}) > \text{cv}(\alpha) $$

\subsection{Validity and Consistency}

Size control is Corollary \ref{cor:size_control} which is an immediate consequence of Theorem \ref{thm:Ru_thetahat}.

\begin{corollary}\label{cor:size_control}
 Let Assumption \ref{assum:articles} hold. If $\mathbf{H}_0$ is true and $T|h\sim N(h,1)$:

    $$\limsup_{n\to \infty}\mathbb{P}\left[d_J(\widetilde{U},\widehat{\theta} ) > cv(\alpha)\right]\leq \alpha$$
\end{corollary}

Finally we show consistency. Theorem \ref{thm:power} says that comparing $d_J(\widetilde{U},\widehat{\theta})$ to $cv(\alpha)$ is a consistent test given a sufficiently large $J$ and a sufficiently wide and dense grid for $\widetilde{U}$.

\begin{theorem}\label{thm:power}
 Let Assumption \ref{assum:articles} hold. For every $\Pi_0$, if $\mathbf{H}_0$ is false,  $T|h\sim N(h,1)$, and $J,L,\delta^{-1}$ are sufficiently large:

    $$\liminf_{n\to \infty}\mathbb{P}\left[d_J(\widetilde{U},\widehat{\theta} ) > cv(\alpha)\right]=1$$
\end{theorem}
\begin{proof}
   By Theorem \ref{thm:consistency_dj_large}, $d_J(U,\theta)>0$. Set $L$ large enough and $\delta$ small enough that $d_J(\widetilde{U},\theta) >\epsilon $. Then $d_J(\widetilde{U},\widehat{\theta}) \to_p d_J(\widetilde{U}_n,\theta) > \epsilon $. Since $\limsup_{n\to \infty}cv(\alpha)\leq \epsilon$, the test must reject with probability approaching 1.
\end{proof}



 A key consequence of Theorem \ref{thm:power} is that by increasing $J,L,\delta^{-1}$  with $n$, the test can be guaranteed to be consistent against any selective reporting process that is identified. In particular, we can always increase $J,L,\delta^{-1}$ slowly enough that all the results that assume these values are fixed still hold. Then $\liminf_{n\to \infty}\mathbb{P}\left[d_{J_n}(\widetilde{U}_n,\widehat{\theta} ) > cv(\alpha)\right]=1$  no matter how slowly $J,L,\delta^{-1}$  increase.



\begin{remark}\label{rem:symmeterizatoin}\normalfont
    In practice it is usually necessary to symmetrize the distribution of $t$-scores. $t$-scores of two-sided t-tests are usually reported as an absolute value. To account for this, take the sample of t-scores $\{t_i\}_{i=1}^n$ and replace it with $\{|t_i|,-|t_i|\}_{i=1}^n$. Under the null hypothesis this is equivalent to replacing  $\Pi$ with its symmetrized version. When resampling with the bootstrap, first resample from the original dataset $\{t_i\}_{i=1}^n$ and then symmetrize. Symmetrization is smooth.
\end{remark}

\begin{remark}\label{rem:shifting}\normalfont
   In practice it is helpful for power to shift all the $t$-scores by 1.96 or another relevant critical value. This is because the norm $||\cdot||_Y$ weights down distortions in $f_{T|R}$ far from zero. But we expect distortions to be near $1.96$. Therefore, power may increase if the meta-analyst simply subtracts $1.96$ from each $t$-score in the dataset (after symmetrization).
\end{remark}

\section{Avoiding Spurious Rejections}\label{sec:misspecification}

So far we have assumed that $T|h \sim N(h,1)$ exactly in the absence of selective reporting. In practice, this normality assumption can be violated. We wish to reject ${\mathbf{H}_0}$ only when the distortions in the $t$-curve were likely due to selective reporting and to avoid spurious rejections originating from non-normality of $T$. We consider two main causes of violations: (1) the $t$-score is actually distributed student-$t$ instead of exactly normal and (2) the $t$-score was reported after rounding and then de-rounded by the meta-analyst. By augmenting the critical value, the meta-analyst can avoid spurious rejections from either of these two sources. The ease of this augmentation is one of the main advantages of the projection test.

\subsection{The Approximately Normal Model}

From the outset, we have set up this test so that it is simple to account for misspecification of the exact distribution of $T|h$. We now generalize the setup from Section \ref{sec:setup} by allowing the noise term in the $t$-score to have a distribution $g$ that is close—--but not equal—--to the standard normal. Suppose that the meta-analyst knows that the total difference between the two densities is upper bounded by some constant $\Delta \geq 0$. Assumption \ref{assum:not_normal} makes this situation explicit.

\begin{assumption}\label{assum:not_normal}
     $ T= h+U $ where $U \protect\mathpalette{\protect\independenT}{\perp} h$ and $U$ is continuous with unknown density $g(t)$ that satisfies:
     $$ \sup_{h\in\mathbb{R}}\left|\left|g(t-h) - \varphi(t-h)\right|\right|_Y \leq \Delta,\qquad ||g||_\infty <\infty $$
     Where $\Delta \geq 0$ is a constant.
\end{assumption}

Under misspecification, the null hypothesis no longer guarantees the restriction in (\ref{eq:testable_restriction}), but instead implies the potentially untestable condition in (\ref{eq:true_restriction_misspec}).
\begin{equation}\label{eq:true_restriction_misspec}
    \inf_{\pi \in \mathcal{F}_\infty } \left|\left|f_{T|R}- \int_{-\infty}^\infty g(t-h)\pi(h)dh \right|\right|_Y =0
\end{equation}

Since the meta-analyst does not know $g$, they cannot test  (\ref{eq:true_restriction_misspec}). Instead, they run the misspecified projection onto the Gaussian convolutions and upper bound the error using the triangle inequality. The inequality below shows that even though the projection is misspecified, the resulting residual can still be bounded.
\begin{align*}
    \left|\left|f_T- \int_{-\infty}^\infty \varphi(t-h)\pi(h)dh \right|\right|_Y
    &\leq \sup_{h\in \mathbb{R}} \left|\left| (g(t-h)-\varphi(t-h))\right|\right|_Y\leq \Delta
\end{align*}

The null hypothesis $\mathbf{H}_0^{\text{miss}}$ for the non-normal model is that the residual of the misspecified projection is small enough that it can be explained by a deviation from $N(0,1)$ of $\Delta$.
\begin{equation}
    \mathbf{H}_0^{\text{\normalfont miss}} \: : \:  \inf_{\pi \in \mathcal{F}_\infty } \left|\left|f_{T|R}- \int_{-\infty}^\infty \varphi(t-h)\pi(h)dh \right|\right|_Y \leq \Delta
\end{equation}
 We test $\mathbf{H}_0^{\text{\normalfont miss}}$ using the same test statistic as before: $d_J(\widetilde{U},\widehat{\theta})$. The difference is that we now add $\Delta$ to the critical value.

 The main limitation of this strategy is that it requires a known upper bound on $\Delta$. When the researcher is unwilling to commit to a specific value of $\Delta$, they can instead report the {\it breakdown statistic}, or the smallest value $\widehat{B}$ of $\Delta$ that would overturn their detection of selective reporting.
\begin{equation}
    \widehat{B} \equiv d_J(\widetilde{U},\widehat{\theta})- cv(\alpha)
\end{equation}

\noindent $\widehat{B}$
  is the smallest deviation from normality that would be sufficient to explain away the observed lack of smoothness in the $t$-curve. In other words, it is the minimal value of $\Delta$ such that the test would no longer reject $\mathbf{H}_0$. The next two sections discuss how large $\widehat{B}$ needs to be for the meta-analyst to be confident that the rejection was probably not the spurious result of common distortions in the $t$-curve that do not come from selective reporting.


\subsection{The student-$t$ distribution}\label{sec:student_t}

In practice, the researcher does not know the true standard error of the point estimate with certainty. This means that when they construct the $t$-score, they divide the point estimate by an estimate of the standard error and the result is that $T-h$ is distributed student-$t$ with $\nu$ degrees of freedom, where $\nu$ is the effective sample size. We can calculate value of $\Delta$ associated with approximating the student-$t$ with the normal using numerical integration. For instance, for a study with effective sample size of $\nu=120$:
\begin{align*}
     \sup_{h\in\mathbb{R}}\left|\left|t_{120}(t-h) - \varphi(t-h)\right|\right|_Y = \left|\left|t_{120}(t) - \varphi(t)\right|\right|_Y < .0009
\end{align*}

Therefore, if we had a set of studies with effective samples sizes predominantly larger than 120 and $\widehat{B}>.0009$, we could conclude that the detection of $p$-hacking was not just a spurious outcome of approximating the normal distribution with the student-$t$ distribution.

\subsection{Derounding}\label{sec:derounding}

In practice,  $t$-scores  are always reported after some rounding in practice. To avoid spurious detections of $p$-hacking resulting from rounding discontinuities, the meta-analyst  must de-round the meta-data by adding random uniformly distributed noise commensurate to the number of significant digits \citep{elliott2024powertestsdetectingphacking}. This process does not leave the $t$-curve completely undistorted. It is therefore prudent to check whether $\widehat{B}$ is small enough to be explained by derounding distortions to the $t$-curve.

We can model the derounded $t$-score as $T=h+Z+V$ where $V$ is mean zero noise independent of $h,Z$ that comes from derounding. Since $V$ is small and of expectation zero, we can bound $\Delta$ by bounding the sup-norm of the difference between the density $f_{Z+V}$ of $f_Z = \varphi$. Proposition \ref{prop:derounding} bounds the distortion from normality $\Delta$ that can be incurred by rounding and derounding.

\begin{proposition}\label{prop:derounding}
    Suppose that $Z\sim N(0,1)$, but instead of observing  $T=h+Z$ we observe $T=h+Z+V$ where $V$ is continuous, $\mathbb{E}[V]=0$, $V\protect\mathpalette{\protect\independenT}{\perp} h+Z$, and $|V|<1$ almost surely. Then, $$\Delta < \sqrt{\frac{1}{2\pi}} \frac{\mathbb{E}[V^2]}{2}$$

    Proof: Section \ref{proof:prop:derounding}.
\end{proposition}


If the $t$-score is observed up to two decimal places (i.e. we can distinguish $T=1.96$ from $T=1.95$), then $|V|< 0.005$ with probability 1. Thus, when $\Delta \leq \sqrt{\frac{1}{2\pi}} \frac{.05^2}{2} \leq 0.00049$. This means that whenever $\widehat{B}> .00049$, we can conclude that the detection of $p$-hacking was not a spurious consequence of de-rounding. This is true of all of the breakdown statistics reported in our empirical application later on in Section \ref{sec:application}. We conclude that the projection test can be made robust to non-normality induced by derounding and that this robustness does not cost us enough power to matter in application.


\section{Simulations}\label{sec:sims}


This section presents two sets of simulations that compare the power and size control of this paper's projection test with the six tests studied by \cite{Elliott}. The takeaway message of these Monte Carlo exercises is that the projection test is as sensitive as the best of these six and controls size much better when $T|h$ is not exactly normal.

Figure \ref{fig:sims_powercurve} displays power curves showing that the projection test is as sensitive as the most sensitive of the six \cite{Elliott} tests. Here the DGP is as follows. The true effect $h$ equals 2 almost surely and $T$ is exactly normal centered on $h$. The researcher $p$-hacks with probability $q$ on the $x$-axis. So on the left part of the graph there is no $p$-hacking and on the right part of the graph everyone $p$-hacks. The $p$-hacking process matches Example \ref{ex:threshold} where the researcher draws $T_1,T_2$ which share an $h$ and are otherwise iid. The researcher reports $T_1$ if it is greater than 1.96 in absolute value and otherwise reports $\max\{|T_1|,|T_2|\}$. Since the power curves for the projection test and the maximum of the six \cite{Elliott} tests track each other almost exactly, we conclude that for this DGP no power is lost by using the projection test.



\begin{figure}[h!]
    \centering
    \includegraphics[width=0.9\linewidth]{figures/sims_powercurve_comparison.pdf}
    \caption{\footnotesize Simulated Power Curves. The $y$-axis is power and the $x$-axis is the fraction of researchers who $p$-hack. Here $p$-hacking means drawing two independent $t$-scores with the same $h$, reporting the first one if it is significant, and reporting the maximum of the two otherwise---see Example \ref{ex:threshold}. There are $n=5000$ studies and $h=2$ almost surely. The solid blue line is the power of the projection test and the dashed red line is the power of the upper envelope of all six tests from \cite{Elliott} (LCM, Binomial, Discontinuity, Fisher, CS1, CS2B) which mostly equals CS2B.    For the tests in \cite{Elliott} we use their values of the tuning parameters. For the projection test we set $J=30$, $\sigma_Y^2=1$,  $L=6.5$ and used a grid of $3000$ points which makes $\epsilon=.00037$.  We symmetrize about zero and shift by $1.96$ as described in Remarks \ref{rem:symmeterizatoin} and \ref{rem:shifting}. We run 1000 bootstrap repetitions and we run 500 simulations. To reduce computation time, we re-use the bootstrapped critical value from simulation 1 for the other 499 simulations. Projection used osqp package by \cite{osqp}.  }
    \label{fig:sims_powercurve}
\end{figure}


In a second simulation we show that the projection test in this paper can detect $p$-h

\begin{table}[h!]
\centering
\begin{tabular}{lcccccccc}

&Fraction hacking & Projection & CS2B & CS1 & LCM & Fisher & Disc. & Binomial \\
  \hline
Size &   0.00 & 0.02 & 0.02 & 0.00 & 0.00 & 0.00 & 0.05 & 0.00 \\
Power &1.00 & 1.00 & 0.03 & 0.01 & 0.00 & 0.00 & 0.05 & 0.00 \\
   \hline
\end{tabular}
\caption{\footnotesize Example of $p$-hacking detectable only by the projection test. Column 2 is the fraction of researchers who $p$-hack. Therefore row $1$ is the type-1 error rate of each test and row $2$ is power when all researchers $p$-hack. Column 3 is the rejection rate of the projection test, and columns 4-9 are the rejection rates of the six tests from \cite{Elliott}.  True effects were generated as: $h\sim N\left(1.96,0.7^2\right)$. For each $h$, the experimenter draws eight independent $t$-scores with $T\sim N(h,1)$ and reports only the largest one. Each meta-sample contains 100,000 iid reported $t$-scores and the simulation was run for 500 repetitions. This large $n$ was chosen to demonstrate that low rejection rates are not simply inadequate sample size, but does cost computation time. To speed up computation time, bootstraps were run for 100 iterations, $L=0.7$, and the critical value is not adjusted for small $L$ (we already know that $L$ is large enough). $J$ is set to 30 and $\sigma_Y^2=1$.  }
\label{tab:hard_detect}
\end{table}


\begin{comment}
    Figure \ref{fig:sims_size} displays size control as the number of studies in the meta-sample $n$ increases. Here there is no selective reporting at all and $h=0$. In addition $T-h$ is not exactly normal---instead it is the normalized sample mean of fifty iid uniforms. Since the support of $T$ is compact, the $p$-curve must slope upward near zero and the Cox-Shi tests of \cite{Elliott} must reject with probability approaching one as $n$ grows even though there is no $p$-hacking.  The graph in Figure \ref{fig:sims_size}  shows that the size of the CS2B test and the LCM test grow as $n$ increases while the size of the projection test remains controlled--whether we adjust for $\Delta=.0014$ or not.\footnote{Adjusting for $\Delta$ makes our test conservative (undersized). Fortunately, we know that this adjustment does not make the test too conservative because is not enough to overturn the rejections in the empirical application later on.} Since the CS2B test is almost always equal to the upper envelope of power in Figure \ref{fig:sims_powercurve}, we conclude that in order for the tests of \cite{Elliott} to match the sensitivity of our projection test, they must also accept some sensitivity to small deviations from normality of $T$.


\begin{figure}[h!]
    \centering
    \includegraphics[width=0.9\linewidth]{figures/sims_size_comparison.pdf}
    \caption{\footnotesize Simulated Size Curves. The $y$-axis is rejection rate (size) and the $x$-axis is the number of studies in the meta-sample $n$. We set $h=0$ almost surely and there is no $p$-hacking. Here $T$ is the normalized sample mean of 50 iid uniforms. The solid lines are the size of the projection test adjusting for $\Delta$ (blue) and unadjusted (orange). The adjustment of the critical values of $.0014$ is calculated in Example \ref{ex:uniform}. The dashed red line is the CS2B test from \cite{Elliott} and the dashed purple line is their LCM test. Other tests from their study (not shown) were properly sized. Otherwise the simulation parameterization matches Figure \ref{fig:sims_powercurve}}
    \label{fig:sims_size}
\end{figure}
\end{comment}




\section{Empirical Application}\label{sec:application}


This section applies the projection test to the data from \cite{bb}. This data consists of 21,740 two-sided t-statistics drawn from 684 articles. These articles were the universe of articles published in 25 top economics journals between 2015 and 2018 that used one of four causal methods. Each $t$-score corresponds to a two-sided $t$-test of a main hypothesis of interest. The coefficients in the numerators of the $t$-scores were estimated using Randomized Controlled Trials (RCT), Difference in Differences (DID), Regression Discontinuity Designs (RDD), or Instrumental Variables (IV). All  ``regression controls, constant terms, balance and robustness checks, heterogeneity of effects, and placebo tests" were excluded. Some tests were reported as $p$-values and \cite{bb} converted these to  $t$-scores using the standard normal distribution.

We run our tests on each of the four causal inference methods separately. The results are in Table \ref{tab:brodeur_results}. We report results for $J=30$ and $J=20$ to demonstrate that our conclusions are not very sensitive to the choice of smoothing parameter $J$. The $p$-values in columns 4 and 7 mostly confirm the findings from \cite{kudrinjmp24}. We detect selective reporting for RCTs, IVs, and DIDs.

Are these rejections spurious outcomes coming from (de)rounding or the limited degrees of freedom of the underlying studies? To determine this, we compute the breakdown statistic $\widehat{B}$ which is the smallest distortion in the $t$-curve from non-$p$-hacking sources that it would take to explain the observed $t$-curve and overturn our rejection. The smallest $\widehat{B}$ in Table \ref{tab:brodeur_results} is 0.0010. Section \ref{sec:student_t} shows that when the effective sample size of the underlying studies are greater than $\nu=120$, the student-$t$ distribution cannot explain a breakdown statistic of 0.0010 or larger. Are the effective sample sizes in this data small enough to explain this? To find out we took a random sample of 100 RCT papers, 100 IV papers, and 100 DID papers from \cite{bb}. We calculated their effective sample sizes as the number of clusters where the cluster was the level at which the standard error was clustered. The median effective sample sizes are reported in Table \ref{tab:brodeur_results}. We calculate $\Delta_\nu$ as the average over all the effective sample sizes $\nu_j$ of $\frac{1}{100}\sum_{j=1}^{100}||\varphi - t_{\nu_j}||_Y$, which can be computed directly by numerical integration.\footnote{We were unable to determine even a lower bound on effective sample size for 20/100 DID papers and 15/100 IV papers.} For all three methods, $\Delta_\nu < \widehat{B}$. This means that the effective sample sizes were overall too large  to explain the breakdown statistic and thus the rejection of $\mathbf{H}_0$ stands.

In Section \ref{sec:derounding}, we showed that if the $t$-score is observed up to two decimal places, (de)rounding can add at most $\sqrt{\frac{1}{2\pi}} \frac{.05^2}{2} \leq 0.00049$ to the test statistic. Since $\widehat{B} > 0.00049$ everywhere in Table \ref{tab:brodeur_results} we conclude that (de)rounding cannot explain the non-smoothness of the $t$-curves in the \cite{bb} data.

These empirical results add a novel insight to existing knowledge. The $t$-curves for RCTs, IVs, and DIDs are significantly distorted. These distortions are larger than can be explained by random chance, (de)rounding, or the student-$t$ distribution. We conclude that the most likely culprit is selective reporting of $t$-scores. The projection test proposed in this paper is the first (to our knowledge) that is simultaneously sensitive enough to detect these distortions in \cite{bb}'s sample and also robust to basic concerns related to (de)rounding, the unknown distribution of true effects, and finite degrees of freedom.


\begin{table}[h!]
    \centering
\begin{tabular}{lcccccccccc}
  & & & & &\multicolumn{2}{c}{J=30} && \multicolumn{2}{c}{J=20} \\
     \cline{6-7} \cline{9-10}
   & n & Articles & Med. $\nu$ & $\Delta_\nu$ & $p$ & $\widehat{B}$ & & $p$ & $\widehat{B}$ \\
  \hline
 RCT & 7569 & 145 & 185 & .0014 & .00 &.0030& & .00 &  .0028  \\
    IV & 5170 & 281 &287 &.0014& .00 & .0021&& .01 & .0022\\
   DID & 5853 & 241 &261 &.0012 & .02 & .0017 && .01 &  .0010\\
   RDD & 3148 & 85 &$\cdot$ &$\cdot$ & .10 & $\cdot$ && .06 &  $\cdot$  \\
   \hline
\end{tabular}

    \caption{Empirical Application Results. Columns 4 and 7 report the $p$-value of the meta-analyst's test of $\mathbf{H}_0$. Med $\nu$ is the median effective sample size as measured by the median number of clusters in a random sample of 100 articles for each causal inference method. In ambiguous cases, the smallest cluster count was always used. $\frac{1}{100}\sum_{j=1}^{100}||\varphi - t_{\nu_j}||_Y$ or the average difference between the PDFs of the standard normal and student $t$ for the over the sample of effective sample sizes collected. Results for $J=30$ and $J=20$ both included to show low sensitivity to the smoothing parameter. We form $\widetilde{U}$ using a grid of 3000 elements evenly spaced between $\pm 6.5$. Discretizations incurs error $\epsilon=0.00037$ which is already accounted for in $\widehat{B}$ and the $p$-value.  Computation used osqp package by \cite{osqp}. Data: \cite{bb}}
    \label{tab:brodeur_results}
\end{table}



\clearpage
\bibliographystyle{chicago}
\bibliography{biblio}