EconBase
← Back to paper

Does p-Hacking Mitigate or Exacerbate the Effects of Publication Bias?

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

69,522 characters

Does p-Hacking Mitigate or Exacerbate the Effects of Publication Bias?


\maketitle

\begin{abstract} \setstretch{1}
\noindent This paper studies the effects of p-hacking on the bias of published estimates when papers with statistically significant results are selectively published. We show that fast p-hacking---actions that lead to large changes in p-values---always exacerbates the bias from selective publication. On the other hand, slow p-hacking---actions that lead to small changes in p-values---exacerbates bias when selection is weak, but mitigates it when selection is strong.
In a model featuring both types of p-hacking, we show that a normality assumption identifies the true distribution of effects as well as the counterfactual mean that would obtain under selective publication without p-hacking. Applying the model to meta-analyses on the effects of behavioral nudges and development aid, we find suggestive evidence that both mitigation and exacerbation can arise in practice.
\looseness=-1
\end{abstract}

\newpage
\section{Introduction}

p-Hacking is widely considered to be a threat to the credibility of published research. Coined by \cite{simonsohn2014p}, the term refers to the collection of actions that empirical researchers take to increase the probability that their papers are published. A substantial body of work documents selective publication or publication bias \citep{sterling1959publication, rosenthal1979file}, in which papers that are more statistically significant are preferentially published (in economics, see, e.g., \citealt{brodeur2016star, christensen2018transparency, brodeur2020methods}). As such, the practice of p-hacking often entails increasing the statistical significance of empirical results, distorting reported results from the truth.

\cite{simonsohn2020fast} further distinguishes between \emph{fast} and \emph{slow} p-hacking. Fast p-hacking refers to processes that change p-values by a substantial amount from analysis to analysis. It is often associated with actions that lead to new draws of essentially independent data. A leading example is file-drawering, in which a researcher with a pre-registered analysis plan discards insignificant results and repeats an experiment wholesale. In contrast, slow p-hacking refers to processes that change p-values by a small amount from analysis to analysis. It typically arises because to analyze messy observational data, researchers have to make many small decisions, most of which have little impact on the final results. For example, researchers may have to choose between different ways of transforming a variable, discretization points of continuous variables, or thresholds for defining outliers. Changing these decisions, therefore, offers a way for researchers to make small adjustments to their p-values. Less innocuously, a researcher may also falsify a small fraction of their data.

p-Hacking, in either manifestation, has been linked to exaggerated effect sizes and excess false positives in published research. It is therefore considered an important contributor to the recent ``replication crisis", which began in psychology \citep{open2015estimating} and biomedical research \citep{prinz2011believe,begley2012raise} but has since been documented across fields, including in economics (see, e.g., \citealt{ioannidis2013s,baker20161, camerer2016evaluating, camerer2018evaluating, christensen2018transparency}). Given the potentially severe repercussions, it is important to understand the effects of p-hacking on the accuracy of published research. While a large and growing body of work tackles this issue, existing papers primarily take either an empirical or simulation-based approach to quantifying the distortions arising from p-hacking (see, e.g., \citealt{leamer1983let, ioannidis2005most,simmons2011false, masicampo2012peculiar,head2015extent,brodeur2016star,brodeur2020methods,friese2020p,moss2023modelling}).

To better understand the settings to which prevailing results apply, this paper formally studies the effects of p-hacking in a simple model with selective publication. We take the form of selective publication as given and assume that it is known to researchers. Specifically, publication probability is assumed to be a step function that jumps at some thresholds of significance. We say that selection is \emph{two-sided} when positive and negative significant results are preferentially published with the same probability. We say that selection is \emph{one-sided} when positive significant results are selectively published, but insignificant and negative significant results are published with the same probability. We focus on two- and one-sided selection for our theoretical analyses.

Selective publication results in biased estimates. It is well known, for example, that two-sided selective publication leads to an amplification bias: the published estimate has, in expectation, the same sign as the study-specific effect but a larger magnitude \citep{lane1978estimating,hedges1984estimation,iyengar1988selection}. For one-sided selective publication, the bias has the same sign as the direction of selection.

Building on these results, we evaluate the effects of p-hacking relative to a benchmark in which selective publication operates alone. This is motivated by the view that systematic p-hacking does not occur in the absence of selective publication. In our stylized model, researchers are endowed with an estimate and standard error. They can then increase the probability that their papers are published by engaging in p-hacking.
Researchers first decide whether or not to conduct fast p-hacking. This is modeled as the option to pay a fixed cost to redraw estimates from the same distribution as the original estimate.  Regardless of whether or not they draw new estimates, researchers then decide whether or not to conduct slow p-hacking. This is the option to pay a cost to report a manipulated estimate and/or standard error. The cost of slow p-hacking is convex and increasing in the amount of manipulation that researchers choose.

Under two-sided selection, Theorem \ref{thm:slow} shows that slow p-hacking exacerbates the amplification bias when selection is sufficiently weak. On the other hand, slow p-hacking mitigates the amplification bias when selection is sufficiently strong, such that the expected published estimate is sandwiched between the true effect and the selection-only benchmark. Meanwhile, Theorem~\ref{thm:fast} shows that fast p-hacking always exacerbates amplification bias, leading to more distorted estimates. Theorem~\ref{thm:combined} studies a generalized model with both types of p-hacking, finding---as in the slow p-hacking only case---that both mitigation and exacerbation can arise, depending on whether selection is strong or weak.
Similar results hold in the model with one-sided selection.

The effect of p-hacking depends on the type of p-hacking that is undertaken, as well as the selection environment that produced it. Although our results are proven in a stylized setting, we argue that both mitigation and exacerbation can arise in environments that are relevant for empiricists.
To that end, we show that under a normality assumption, the mean study effect, as well as the counterfactual mean of published estimates under selective publication only, are both identified. This allows us to assess the extent to which mitigation or exacerbation may have occurred in a given literature. Under additional parametric assumptions, we estimate a censored version of our one-sided model using data from two meta-analyses.
The first is \cite{mertens2022effectiveness}, which studies the effect of behavioral nudges on promoting desired choices. Here, we find that p-hacking may have mitigated the bias from selective publication by as much as 54\%. The second is \cite{doucouliagos2011ineffectiveness}, which concerns the effect of development aid on economic growth. Here, we find instead that p-hacking has exacerbated bias by 27\%.

Discussions about p-hacking and selective publication often treat the two phenomenon separately. Our results show that it is important to understand p-hacking in the context of the selection environment that produced it. That p-hacking could mitigate the effects of publication bias under extreme selection also hints at a self-correcting tendency in the scientific process. This echoes the idea that literatures as a whole may be more reliable than their constituent papers, a view in philosophy of science that goes at least as far back as \cite{kuhn1962structure}. However, this does not imply that p-hacking is desirable, since our paper is silent on the other negative effects of p-hacking, such as inflating the number of false positives,
the erosion of public trust in science, or distortions in the allocation of research funds, among others (see, e.g., \citealt{ioannidis2005most,button2013power,Wasserstein02042016}).

The remainder of this paper proceeds as follows. Section \ref{section--setting} describes the formal setting. Sections \ref{section:twosided} and \ref{section:onesided} contain our main theoretical results under two- and one-sided publication bias, respectively. Section \ref{sec:emp} discusses the identification of our proposed model and presents empirical results based on \cite{mertens2022effectiveness} and \cite{doucouliagos2011ineffectiveness}. Section \ref{section--conclusion} concludes.

\subsection{Related Literature}

This paper connects to the vast and growing literature on p-hacking. First, it speaks specifically to the line of work concerned with the effects of p-hacking on published estimates \citep{leamer1983let,ioannidis2005most, simmons2011false, gelman2013garden, brodeur2016star,ioannidis2017power,friese2020p,moss2023modelling,stefan2023big}. These papers typically provide quantitative evidence either empirically or through simulations. Complementing their approach, we develop a \textit{theoretical} framework to formally study how fast and slow p-hacking affects the bias of published estimates.\footnote{A contemporaneous paper, \cite{keane2026instrument}, also theoretically examines the effect of p-hacking on the bias of published estimates, focusing on the case where researchers p-hack instrumental-variable regressions.} We show that p-hacking can mitigate the effects of selective publication. This result contrasts with the prevailing views of p-hacking, with one notable exception: \cite{friese2020p}, which documents a similar mitigation effect in simulations of a statistical model in which researchers p-hack by drawing significant p-values from a triangular distribution.

Relatedly, a separate strand of work studies the optimal design of the publication rule in the presence of p-hacking.\footnote{Some other papers study optimal publication rules without allowing for p-hacking. See, e.g., \cite{frankel2022findings, kitagawa2023optimal}.} \citet{jagadeesan2024publication} designs the publication rule when slow p-hacking can distort researchers' incentives at the research-design stage. \citet{spiess2025optimal} designs estimation procedures when researchers' preferences over reported results are misaligned with society's. In contrast to the normative focus of these two papers, we take a \textit{positive} approach, as we take the publication rule as given and investigate the effect of p-hacking. In addition, our framework distinguishes between fast and slow p-hacking, which is not considered by \citet{jagadeesan2024publication} and \citet{spiess2025optimal}.

Second, our identification result contributes to an active strand of research on the detection of and correction for p-hacking and selective publication. Recent work has focused on the use of p-curves \citep{simonsohn2014p, ulrich2015p, simonsohn2015better, bruns2016p, elliott2022detecting, elliott2025power} caliper tests \citep{gerber2008statistical,gerber2008publication, kudrin2024robust} among other methods \citep{brodeur2016star,moss2023modelling,faridani2025p,kudrin2025testing} for detecting p-hacking. Closely related are methods for detecting and correcting for selective publication. Main approaches include selection modeling \citep{hedges1984estimation,iyengar1988selection, vevea1995general, mccrary2016conservative, andrews2019identification}, funnel asymmetry and related meta-regression methods \citep{egger1997bias, duval2000trim, stanley2008meta,stanley2014meta}, as well as Bayesian model averageing \cite{bartovs2023robust}.
Methods geared towards the detection of either p-hacking or selective publication often do not distinguish between the two sources of bias. One exception is \cite{brodeur2016star}, which proposes bounds for the proportion of p-hacked studies under the assumption of monotone selection on significance. Our identification result focuses on the selection-only counterfactual mean, which we obtain under parametric restrictions.

\section{Setting}\label{section--setting}

We adopt the setting of \cite{andrews2019identification} but allow for p-hacking of the estimate and standard error. A researcher faces a latent study $(\Theta, \Sigma) \sim \mu$, where $\Theta$ is the effect size and $\Sigma$ is the standard error. He goes through a potential publication process as follows.
\begin{enumerate}
\item \textbf{(Initial draw)} The researcher draws an estimate $X_1$ from the distribution $\mathcal{N}(\Theta, \Sigma^{2})$.\footnote{The assumption of normal distribution is consistent with empirical practice, which typically performs inference under the premise of asymptotic normality.} We let $\Phi(x)$ denote its cumulative distribution function (CDF) and $\phi(x)$ the corresponding probability density function (PDF).

\item \textbf{(Fast p-hacking)} The researcher can redraw estimates independently from $\mathcal{N}(\Theta, \Sigma^2)$. The cost of the $n$-th draw is $\kappa_f\cdot c_n$, where $\kappa_f > 0$ is a scaling parameter that reflects the general difficulty of fast p-hacking. For notational consistency, we let $c_1 = 0$, i.e., the initial draw is free. We assume that $c_n$ strictly increases in $n$, with $\lim\limits_{n\rightarrow\infty}c_n = \infty$. For simplicity, we assume that the researcher is naive and believes $\Theta=0$ throughout the process.\footnote{A similar assumption also appears in \citet{mccloskey2024critical}. Its purpose is to rule out the possibility that the researcher learns about $\Theta$ during fast p-hacking, thereby simplifying the analysis. Alternatively, we could assume that the researcher knows the value of $\Theta$ from the outset or naively believes that $\Theta=\mathbbm{E}_{\mu}(\Theta)$. All of our theoretical results continue to hold under these alternative assumptions.} We let $\boldsymbol{X}=\{X_1, X_2, ..., X_N\}$ denote the set of estimates if the researcher ends up making $N$ draws.

\item \textbf{(Slow p-hacking)} Among all the estimates in $\boldsymbol{X}$, the researcher can pick one estimate $X_n$, perform slow p-hacking, and report $(\hat{X}, \hat{\Sigma})$. The cost of slow p-hacking is $\kappa_s \cdot c(|\hat{X} - X_n|, |\hat{\Sigma} - \Sigma|)$, satisfying $\kappa_s>0$, $c(0,0)=c_1(0,0) = c_2(0,0) = 0$, $c_{11}>0$, $c_{22}>0$, and $c_{12}= 0$.\footnote{One example of slow p-hacking cost function is $c(|\hat{X} - X_n|, |\hat{\Sigma} - \Sigma|) = c_x \cdot (\hat{X} - X_n)^2 + c_\Sigma \cdot (\hat{\Sigma} - \Sigma)^2$.} Notice that the researcher has the option of not undertaking slow p-hacking by letting $(\hat{X}, \hat{\Sigma}) = (X_n, \Sigma)$.

\item \textbf{(Publication)} The reported result $(\hat{X}, \hat{\Sigma})$ is published (denoted by $D=1$) with probability $p(\hat{X}/\hat{\Sigma})$, which only depends on $\hat{X}/\hat{\Sigma}$ (the test statistic for the null hypothesis $\Theta = 0$ when the reported result is not manipulated), as we specify below. This publication rule is known to the researcher from the beginning of the entire process. The researcher receives a reward of $V > 0$ upon publication and zero otherwise.
\end{enumerate}

The researcher's expected payoff from making $N$ draws, using the $n$-th estimate, and reporting $(\hat{X}, \hat{\Sigma})$ is
\[V\cdot p(\hat{X}/\hat{\Sigma}) - \kappa_s \cdot c(|\hat{X} - X_n|, |\hat{\Sigma} - \Sigma|) - \kappa_f \cdot \sum_{n=1}^{N} c_n .\]

\paragraph{\underline{Model of p-hacking.}} \ Our model of p-hacking follows \cite{simonsohn2020fast}, which distinguishes between \emph{fast} and \emph{slow} p-hacking.

Fast p-hacking refers to practices that change p-values substantially from analysis to analysis. A leading example of which is file-drawering, whereby researchers discard insignificant data and repeat an experiment unchanged. We therefore model fast p-hacking as independent redrawing of estimates. The fixed $\Sigma$ reflects the idea that researchers cannot change the design of their study, either because they have already pre-registered their experiment, or have already committed to some data collection strategy. We assume that the cost of drawing new estimates is increasing, reflecting factors such as the scarcity of experiment sites or objects, as well as the increasing reputation cost of continued p-hacking. The cost of drawing new estimates diverges to infinity, capturing the fact that researchers do not have unlimited resources for conducting studies. Our model of fast p-hacking resembles \citet{mccloskey2024critical}, in which the researcher can draw up to $L$ estimates. It is also consistent with \citet[Supplement 3]{simonsohn2014p}, which models file-drawering as researchers obtaining a sequence of independent p-values.

Slow p-hacking refers to practices that change the p-value by a small amount from analysis to analysis. It arises because researchers often have to make many decisions when analyzing messy data. For example, they choose the thresholds for discretizing continuous variables, which outliers to drop, or the transformation of variables. Individual decisions probably do not affect the p-values by much. Collectively, they offer a way for researchers to approach the significance threshold through a sequence of tiny steps. As a result, the outcome of slow p-hacking rarely exceeds the target threshold by much. This produces the bunching behavior most associated with p-hacking, and is what caliper tests are designed to detect. As such, we model slow p-hacking as the direct and deterministic manipulation of estimates and standard errors. The manipulation cost increases with the distance between the reported result and the original one,\footnote{Our model is in line with the literature on communication with lying costs (e.g., \citet{kartik2009strategic}).} reflecting factors such as the increasing amount of time spent on specification search.
Our model of slow p-hacking is similar to \citet{jagadeesan2024publication}, in which researchers pay a linear cost to manipulate the estimate only.

As a side note, our model assumes that researchers engage in fast p-hacking before resorting to slow p-hacking. This ordering is without loss of optimality, as the outcome of slow p-hacking is deterministic.

\paragraph{\underline{Publication rule.}} \ In our model, the publication probability of a reported result $(\hat{X}, \hat{\Sigma})$ only depends on $\hat{X}/\hat{\Sigma}$, the test statistic for the null hypothesis $\Theta = 0$ when the reported result is not manipulated. Throughout the paper, we focus on the following publication rule:
\begin{equation}
    p(z) = \begin{cases}
        1 & \text{ if } z \geq t, \\
        1/\delta_0 & \text{ if } -t < z < t, \\
        1/\delta_- & \text{ if } z \leq -t~,
    \end{cases}
\label{eq:general}
\end{equation}
where $\delta_0 \geq \delta_{-} \geq 1$. Under this rule, a positive significant result at the significance level $t$ is published with probability normalized to one, (weakly) higher than a negative significant result's publication probability, which in turn is (weakly) higher than that of an insignificant result.\footnote{The assumption that a positive significant result is more likely to be published than a negative significant one is without loss of generality. In the empirical analysis, we normalize signs so that this assumption holds; that is, if a literature favors negative results, we flip the signs of the estimates.} This rule reflects two common phenomena in scientific publishing. First, empirical results significant at a certain level (such as 5\%) are more likely to be published compared to those that are marginally insignificant at the same level.\footnote{For instance, such a preference for significant results is seen in the fields of economics \citep{brodeur2016star, chopra2024null}, medicine \citep{ioannidis2007exploratory}, political science \citep{gerber2008statistical}, psychology \citep{masicampo2012peculiar}, sociology \citep{gerber2008publication}, and social sciences as a whole \citep{franco2014publication}.} Second, in certain literature, editors may preferentially publish results with a particular sign. Specifically, they may prefer either ``reasonable results,'' which support their prior beliefs, or ``surprising results,'' which contradict their prior beliefs. For example, in the minimum wage literature, \cite{card1995time} and \cite{andrews2019identification} find evidence that editors favor results showing negative effects of minimum wage on employment.

Note that this publication rule encompasses two special cases. It corresponds to \textit{two-sided selective publication} if $\delta_{-} = 1$ and \textit{one-sided selective publication} if $\delta_{-} = \delta_0$. These two publication rules stand as the extreme cases of the general rule we consider in the paper.

\paragraph{\underline{Assumption on p-hacking cost.}} \ Finally, we introduce a regularity assumption on the cost of slow p-hacking to make the researcher's slow p-hacking behavior consistent with reality.

\begin{remark}
\label{remark:uniquemin}
For $X\geq 0$, the problem $\min\limits_{ \tilde{X} \geq t\tilde{\Sigma} > 0} c(|\tilde{X} - X|, |\tilde{\Sigma} - \Sigma|)$ has a unique solution.
\end{remark}
\begin{proof}
See Appendix~\ref{pf:remark:uniquemin}.
\end{proof}

Remark~\ref{remark:uniquemin} implies the following. Under the publication rule we consider in this paper, if the researcher decides to slow p-hack a positive insignificant estimate to positive significance, there is a unique result that minimizes the p-hacking cost among all candidates. Hereafter, when $0\leq X<t\Sigma$, we let $(\tilde{X},\tilde{\Sigma})=(\xi(X), \xi(X)/t)$ denote this unique minimizer, where we fix $\Sigma$ and express it as a function of $X$. Based on that, we also let $C(X):=c(|\xi(X) - X|, |\xi(X)/t - \Sigma|)$ denote the minimal cost of slow p-hacking this positive insignificant estimate towards positive significance.

We can extend these two definitions to $X<0$, with $(\xi(X), -\xi(X)/t)$ being the optimal result chosen by the researcher if he slow p-hacks a negative insignificant estimate to negative significance, and $C(X) > 0$ being the corresponding slow p-hacking cost. Due to the symmetric structure of the slow p-hacking cost function $c(\cdot,\cdot)$ and the fact that positive and negative significances require the same level, we have that $\xi(X)$ is anti-symmetric around zero, and $C(X)$ is symmetric around zero.

We are now ready to state the regularity assumption.

\begin{assumption}
\label{assump:largekappas}
We have $\kappa_s \cdot C(0) > V$.
\end{assumption}

Assumption~\ref{assump:largekappas} states that $\kappa_s$ is sufficiently large such that it is never worthwhile to undertake slow p-hacking when $X=0$. This further implies that the researcher's optimal slow p-hacking strategy does not change the sign of the estimate; that is, it is not worthwhile to slow p-hack a negative estimate to positive significance or vice versa. Although this assumption is not crucial for the qualitative results of the paper, it is consistent with the empirical observation that insignificant results are also published. If, instead, we allow for $\kappa_s \cdot C(0) \leq V$, then it becomes possible that every published result is significant under two-sided selective publication.

\section{Two-Sided Selective Publication}\label{section:twosided}

This section focuses on \textit{two-sided selective publication}, i.e., $\delta_{-} = 1$. To save notation, we let $\delta_0 = \delta$ in this section, so the publication rule becomes
\begin{equation*}
        p(z) = \begin{cases}
1 & \text{ if } |z| \geq t, \\
1/\delta & \text{ if } |z| < t,
        \end{cases}
    \end{equation*}
with $\delta > 1$ being referred to as the \textbf{extent of selective publication}.

Since we are interested in the interplay of p-hacking and selective publication, we fix $(\Theta, \Sigma, \kappa_s, \kappa_f)$ and consider $\delta$, the extent of selective publication, as the primary parameter in the analysis. We let $\gamma_h(\delta):=  \mathbbm{E}(\hat{X}|D=1,\Theta, \Sigma)$,  in which the subscript ``h'' stands for hacking, denote the expected published result in our model, in which selective publication and p-hacking co-exist. Meanwhile, we let $\gamma(\delta)$ denote the expected published result when there is selective publication but \textit{no} p-hacking; equivalently, we have $\gamma(\delta) = \lim\limits_{ \kappa_s, \kappa_f\rightarrow\infty}\gamma_h(\delta)$, as the researcher does not undertake any p-hacking when it is too costly. Similarly, we let $\gamma_f(\delta)= \lim\limits_{ \kappa_s\rightarrow\infty}\gamma_h(\delta)$ and $\gamma_s(\delta)= \lim\limits_{ \kappa_f\rightarrow\infty}\gamma_h(\delta)$ denote the expected published result under the scenario of fast-only and slow-only p-hacking, respectively.

\begin{assumption}
\label{assump:tiebreak}
The researcher adopts the following tie-breaking rule in Section~\ref{section:twosided}. \\
(a) If he is indifferent about undertaking a certain p-hacking action (i.e., his continuation value remains the same with or without a certain p-hacking action), he does \textit{not} p-hack. \\
(b) If he is indifferent among multiple results at the time of reporting, he reports the one with the lowest p-value.
\end{assumption}

Assumption~\ref{assump:tiebreak} specifies the researcher's behavior in certain scenarios where he is indifferent. The paper's results remain intact under alternative reasonable tie-breaking rules (see, e.g., Footnotes~\ref{footnote:tiebreak1} and \ref{footnote:tiebreak2}).

Note that we defined our outcomes of interest conditional on $(\Theta,\Sigma)$, so all our theoretical results hold pointwise. Without loss of generality, we only consider the case where the effect size is positive, i.e., $\Theta > 0$; the case with $\Theta < 0$ follows by symmetry.\footnote{When $\Theta=0$, due to the symmetry of our model, neither selective publication nor p-hacking induces any publication bias.}

\subsection{Benchmark: Selective Publication Without p-Hacking}

It is well known that published results under two-sided selective publication tend to overestimate the magnitude of the treatment effect. This is shown by \cite{hedges1984estimation} for $F$-tests of mean differences and \cite{iyengar1988selection} for the $t$-test. For completeness, Proposition~\ref{prop:nohack} confirms that the same result holds in our setting.

\begin{proposition}
\label{prop:nohack}
When $\Theta > 0$, we have $\gamma(\delta)>\Theta$, $\forall \delta>1$.
\end{proposition}
\begin{proof}
See Appendix~\ref{pf:prop:nohack}.
\end{proof}

Taking this publication bias as the benchmark, we now explore whether p-hacking mitigates or exacerbates the bias.

\begin{definition}
\label{def:mora}
Let $\Theta>0$. We say that p-hacking \textbf{mitigates} publication bias if $\Theta < \gamma_h(\delta) < \gamma(\delta)$ and \textbf{exacerbates} publication bias if $\Theta < \gamma(\delta) < \gamma_h(\delta)$. Similar definitions apply to slow-only p-hacking using the term $\gamma_s(\delta)$ and fast-only p-hacking using the term $\gamma_f(\delta)$, respectively.
\end{definition}

\subsection{Slow-Only p-Hacking}
\label{sec:slow}

In this subsection, we study the implications of slow p-hacking by letting $\kappa_f \rightarrow \infty$. In this case, the researcher only draws one estimate, $X_1$, and decides whether to undertake slow p-hacking. The following proposition summarizes the optimal slow p-hacking strategy.

\begin{proposition}\label{prop:slow_strat}
The researcher's optimal slow p-hacking strategy can be represented by a threshold $\overline{X}_s\in (0, t\Sigma ]$ such that: \begin{enumerate}[(i)]
\item He undertakes slow p-hacking if and only if $|X_1| \in (\overline{X}_s, t\Sigma)$.
\item His reported result is
\begin{align*}
(\hat{X},\hat{\Sigma}) = \begin{cases}
            (\xi(X_1), |\xi(X_1)|/t) & \quad\text{if }|X_1|\in (\overline{X}_s, t\Sigma),  \\
            (X_1,\Sigma) & \quad\text{otherwise}.
        \end{cases}
\end{align*}
\end{enumerate}
Furthermore, the slow p-hacking threshold $\overline{X}_s$ (weakly) decreases in $\delta$ and converges to $t\Sigma$ as $\delta$ goes to one.
\end{proposition}
\begin{proof}
See Appendix~\ref{pf:prop:slow_strat}.
\end{proof}

Proposition~\ref{prop:slow_strat} shows that the researcher undertakes slow p-hacking only when the insignificant estimate $X_1$ is sufficiently close to being significant ($|X_1| \in (\overline{X}_s, t\Sigma)$). If publication is more selective (higher $\delta$), then the incentive to p-hack is stronger, leading to a lower threshold for slow p-hacking (lower $\overline{X}_s$). Importantly, the optimal p-hacking strategy is geared to the publication rule, and, in particular, is symmetric around zero.

Having characterized the researcher's optimal slow p-hacking strategy, we can further explore its implications for publication bias.

\begin{theorem}
\label{thm:slow}
There exist two thresholds $1<\underline{\delta}_s< \overline{\delta}_s<\infty$ such that slow p-hacking mitigates publication bias if $\delta > \overline{\delta}_s$ and exacerbates publication bias if $\delta < \underline{\delta}_s$.
\end{theorem}
\begin{proof}
See Appendix~\ref{pf:thm:slow}.
\end{proof}

For intuition, suppose researchers cannot p-hack $\Sigma$ and consider all published results with $X \geq 0$. We can think of the overall distribution as a mixture of insignificant and significant estimates, with the latter having a higher mean. Slow p-hacking introduces two effects. On the one hand, researcher manipulation reduces the mass of insignificant estimates, thereby increasing the bias. On the other hand, manipulated results land exactly on the threshold. They are published with higher probability than before and have the smallest possible value among significant estimates. This lowers the mean of estimates conditional on significance. The relative strength of these two effects determines the overall outcome. In the extreme case with most severe publication selection ($\delta = \infty$), no insignificant results are published, thereby shutting down the first effect. This explains why slow p-hacking can mitigate publication bias when selective publication is sufficiently severe.

\paragraph{Quadratic Costs.} Theorem \ref{thm:slow} does not stipulate what happens when the extent of selective publication is intermediate $\delta\in [\underline{\delta}_s,\overline{\delta}_s]$. In that region, there are countervailing forces at play, and our general model does not yield a clear prediction for the effect of slow p-hacking. However, we can sharpen Theorem \ref{thm:slow} by imposing additional structure on the p-hacking technology. For the following result only, suppose that the slow p-hacking cost is quadratic:
\[c(|\hat{X}-X|,|\hat{\Sigma}-\Sigma|)=c_X\cdot (\hat{X}-X)^2+c_\Sigma\cdot (\hat{\Sigma}-\Sigma)^2\,,\quad c_X,c_\Sigma>0\,.\]

\begin{proposition}
\label{thm:slow_quad}
Suppose the cost of slow p-hacking is quadratic and that the function $\phi(x)+\phi(-x)$ is log-concave on $[t\Sigma-\sqrt{V(c_Xt^2+c_\Sigma)/(\kappa_sc_Xc_\Sigma)}, t\Sigma]$. Then there exists a single threshold $\delta^*\in(1,\infty)$ such that slow p-hacking mitigates publication bias if $\delta > \delta^*$ and exacerbates publication bias if $\delta < \delta^*$.
\end{proposition}
\begin{proof}
See Online Appendix~\ref{pf:thm:slow_quad}.
\end{proof}

The log-concavity assumption on $\phi$ is always satisfied when $\Theta\leq \Sigma$. It also holds if the slow p-hacking threshold under the most selective publication rule is not too low: $\lim_{\delta\rightarrow\infty}\overline{X}_s\geq 0.663 \Sigma$. When $t=1.96$, this is the case when a researcher never moves an estimate by more than about 1.3 standard errors. That is, when researchers only slow p-hack results with p-values below 0.51.

\subsection{Fast-Only p-Hacking}
\label{sec:fast}

In this subsection, we investigate the implications of fast p-hacking by letting $\kappa_s\rightarrow \infty$. Essentially, the researcher chooses when to stop drawing new estimates.

Let $\Delta_f(\delta):= \mathbb{P}_{X\sim\mathcal{N}(0,\Sigma^2)}(|X|\geq t\Sigma) \cdot (1-\frac{1}{\delta}) V$ be the researcher's expected payoff gain from exactly one more draw if all his current results are insignificant. Let $n^*_f:= \max\{n \mid \kappa_f c_n < \Delta_f(\delta) \}$. The researcher's optimal strategy when he can only fast p-hack is as follows.

\begin{proposition}
\label{prop:fast_strat}
The researcher draws up to $n^*_f$ estimates until obtaining a significant result. Furthermore, $n^*_f$ is weakly increasing in the extent of selective publication $\delta$.
\end{proposition}
\begin{proof}
See Appendix~\ref{pf:prop:fast_strat}.
\end{proof}

Based on the researcher's optimal strategy, we have the following result.

\begin{theorem}\label{thm:fast}
There exists a threshold $\overline{\delta}_f \in (1, \infty)$ such that fast p-hacking exacerbates publication bias if $\delta > \overline{\delta}_f$ and has no effect otherwise. Moreover, the effect of fast p-hacking converges to zero when $\delta\rightarrow \infty$; that is, $\lim\limits_{\delta\rightarrow\infty}[\gamma_f(\delta) - \gamma(\delta)]=0$.
\end{theorem}
\begin{proof}
See Appendix~\ref{pf:thm:fast}.
\end{proof}

The intuition behind Theorem~\ref{thm:fast} is as follows. Fast p-hacking enables the researcher to hide insignificant estimates, making significant estimates more likely to be reported than insignificant ones. In terms of overestimating the magnitude of the treatment effect, this force works in the same direction as selective publication, thereby exacerbating the publication bias. Moreover, as $\delta\rightarrow\infty$, the effect of fast p-hacking on publication bias vanishes since all insignificant estimates are unpublished in that limit, regardless of the extent to which the researcher hides them. When $\delta \leq \overline{\delta}_f$, fast p-hacking is never undertaken and has no effect on the distribution of published results.

\subsection{General Model}
\label{sec:joint}

In this section, we consider situations where the researcher can undertake both types of p-hacking. Beyond generalizing our previous results, we emphasize how the two types of p-hacking interact.

As before, it is strictly optimal for the researcher to stop engaging in any form of p-hacking once he obtains a significant result. If all his current results are insignificant, the decision of whether to draw another estimate depends on the expected benefits. Let $X_k^*:= \text{arg}\max\limits_{X_\kappa: \kappa\leq k}|X_\kappa|$ denote the researcher's ``best current estimate'' if he has drawn $k$ estimates (i.e., the one with the lowest p-value).\footnote{Throughout the paper, we restrict attention to the generic situations where no two estimates have identical absolute values.} Additionally, we let $\Delta_h:=\mathbb{P}_{X\sim \mathcal{N}(0, \Sigma^{2})}\left[|X|>\overline{X}_s\right] \times \mathbb{E}_{X\sim \mathcal{N}(0, \Sigma^{2})}\left[V(1-\frac{1}{\delta}) - \kappa_s C(X)\mid |X|>\overline{X}_s \right]$ and $n^*:= \max\{n\mid k_f\cdot c_n < \Delta_h\}$.

\begin{proposition}\label{prop:combined_strat}
The researcher draws at most $n^*$ estimates. His optimal p-hacking strategy can be represented by a slow threshold $\overline{X}_s$ and a decreasing series of fast thresholds $\{\overline{X}_f^n\}_{n=1}^{n^*-1}$ satisfying
\[0< \overline{X}_s <\overline{X}_f^{n^*-1} < \cdots < \overline{X}_f^{2} < \overline{X}_f^{1}< t\Sigma. \]
Suppose he has already drawn $k<n^*$ estimates. He draws a $(k+1)$-th estimate if and only if his best current estimate $X_k^*$ satisfies $|X_k^*|<\overline{X}_f^k$. After he stops drawing new estimates, he undertakes slow p-hacking if and only if $|X_k^*|\in (\overline{X}_s,t\Sigma)$ and reports the following result:
\begin{align*}
(\hat{X}, \hat{\Sigma}) = \begin{cases}
(\xi(X_k^*), |\xi(X_k^*)|/t) & \quad\text{if }|X_k^*|\in (\overline{X}_s, t\Sigma),  \\
            (X_k^*,\Sigma) & \quad\text{otherwise}.
        \end{cases}
\end{align*}
\end{proposition}
\begin{proof}
See Appendix~\ref{pf:prop:combined_strat}.
\end{proof}

Proposition~\ref{prop:combined_strat} shows that, at any moment, the researcher's continuation p-hacking strategy depends on his best current estimate and how many draws have been made. Two features of his optimal p-hacking strategy are worth highlighting. First, the fast p-hacking threshold, $\overline{X}_f^n$, decreases in $n$, indicating that the researcher is less likely to fast p-hack if he has already drawn more estimates. This is driven by the fact that the fast p-hacking cost $c_n$ increases in $n$. Second, the slow p-hacking threshold is independent of $n$. The intuition is that once the researcher undertakes slow p-hacking, he must have decided to stop drawing new estimates, and therefore, only the best current estimate matters. Indeed, this threshold is the same as the p-hacking threshold in the special case with only slow p-hacking (see Section~\ref{sec:slow}).

\begin{theorem}
\label{thm:combined}
There exist two thresholds $1<\underline{\delta}_h< \overline{\delta}_h<\infty$ such that p-hacking mitigates publication bias if $\delta > \overline{\delta}_h$ and exacerbates publication bias if $\delta < \underline{\delta}_h$.
\end{theorem}
\begin{proof}
See Appendix~\ref{pf:thm:combined}.
\end{proof}

Theorem~\ref{thm:combined} can be understood as follows. When $\delta$ is sufficiently small, the researcher does not undertake fast p-hacking, so the overall effect of p-hacking follows its slow component. When $\delta$ is sufficiently large, the researcher engages in both fast and slow p-hacking. As we point out in Theorem~\ref{thm:fast}, the effect of fast p-hacking converges to zero as $\delta$ approaches infinity. Therefore, in the limit, the overall effect of p-hacking is also captured by the slow component.

While discussions about p-hacking generally suppose that it is undesirable, our result suggests the importance of considering p-hacking within the context of selective publication and that there exist simple, plausible models in which p-hacking may mitigate publication bias. Nevertheless, it is worth emphasizing that our result does not speak to other potential damaging effects of p-hacking, such as reducing trust in science and distorting the incentives of career researchers.

Theorems \ref{thm:slow} to \ref{thm:combined} compare $\gamma(\delta)$ with $\gamma_s(\delta)$, $\gamma_f(\delta)$, and $\gamma_h(\delta)$, respectively, aiming to investigate the effect of p-hacking on publication bias. The next theorem compares $\gamma_h(\delta)$ with $\gamma_s(\delta)$ and $\gamma_f(\delta)$, which enables us to understand the effect of banning one type of p-hacking, taking the existence of both types of p-hacking as the status quo.

\begin{theorem}
\label{thm:fastandslow}
(a) Banning fast p-hacking always reduces publication bias. That is, for any $\kappa_s$, we have $\gamma_h(\delta) \geq \gamma_s(\delta) > \Theta$, $\forall \delta$. The first inequality is strict if and only if $n^* > 1$ (i.e., the researcher sometimes draws more than one estimate when fast p-hacking is allowed). \\
(b) Banning slow p-hacking increases publication bias when publication is selective enough. That is,  for any $\kappa_f$, $\gamma_f(\delta) > \gamma_h(\delta)  > \Theta$ when $\delta$ is sufficiently large.
\end{theorem}
\begin{proof}
See Appendix~\ref{pf:thm:fastandslow}.
\end{proof}

As established in Theorem~\ref{thm:fastandslow}, holding the slow p-hacking technology fixed, banning fast p-hacking is unambiguously desirable in reducing publication bias. By contrast, for any given fast p-hacking technology, prohibiting slow p-hacking backfires when selective publication is sufficiently stringent.

\section{One-Sided Selective Publication}
\label{section:onesided}

This section analyzes the other extremal publication rule, \textit{one-sided selective publication}. To save notation, we let $\delta_{-} = \delta_0 =\delta$ in this section, so the publication rule becomes
\begin{equation*}
        p(z) = \begin{cases}
1 & \text{ if } z \geq t, \\
1/\delta & \text{ if } z < t,
        \end{cases}
    \end{equation*}
where the parameter $\delta > 1$ is the \textbf{extent of selective publication}. Unlike Section~\ref{section:twosided}, we do \textit{not} restrict attention to $\Theta > 0$ in this section, since the publication rule is not symmetric around zero.

The purpose of this section is to show that the main results of Section~\ref{section:twosided} continue to hold under one-sided selective publication: allowing p-hacking mitigates publication bias when publication selection is sufficiently severe, but exacerbates it when publication selection is sufficiently mild. To keep the analysis brief, we directly consider our general model where both fast and slow p-hacking are allowed. Also, since the proofs in this section follow similar reasoning to those in Section~\ref{section:twosided}, we relegate them to Online Appendix~\ref{onlineapx_secondaryproof}.

Similarly to Section~\ref{section:twosided}, we introduce the following tie-breaking assumption. As before, its purpose is to simplify the analysis, and all results remain intact if we relax it.

\begin{assumption}
\label{assump:tiebreak_onesided}
The researcher adopts the following tie-breaking rule in Section~\ref{section:onesided}. \\
(a) If he is indifferent about undertaking a certain p-hacking action (i.e., his continuation value remains the same with or without a certain p-hacking action), he does \textit{not} p-hack. \\
(b) If he is indifferent among multiple results at the time of reporting, he reports the one with the highest t-statistics.
\end{assumption}

The following proposition states the researcher's optimal p-hacking strategy under one-sided selective publication. Let $\breve{X}_k^*:=\max\limits_{l\leq k}X_l$ denote the researcher's best current estimate if he has drawn $k$ estimates.

\begin{proposition}
\label{prop:onesided_strat}
Under one-sided selective publication, the researcher draws at most $n^*$ estimates. His optimal p-hacking strategy can be represented by a slow threshold $\overline{X}_s$ and a decreasing series of fast thresholds $\{\overline{X}_f^n\}_{n=1}^{n^*-1}$ satisfying\footnote{For notational simplicity, we reuse $n^*$, $\overline{X}_s$, and $\{\overline{X}_f^n\}_{n=1}^{n^*-1}$ for both the two-sided and one-sided selective publication settings. Note that while the optimal slow p-hacking threshold $\overline{X}_s$ remains identical across both models, the values for $n^*$ and $\{\overline{X}_f^n\}_{n=1}^{n^*-1}$ are specific to the publication environment and differ between the two sections. \label{footnote1}}
\[0< \overline{X}_s <\overline{X}_f^{n^*-1} < \cdots < \overline{X}_f^{2} < \overline{X}_f^{1}< t\Sigma. \]
Suppose the researcher has already drawn $k<n^*$ estimates. He draws a $(k+1)$-th estimate if and only if $\breve{X}_k^* < \overline{X}_f^k$. After he stops drawing a new estimate, he undertakes slow p-hacking if and only if $\breve{X}_k^*\in (\overline{X}_s, t\Sigma)$ and reports the following result:
\begin{align*}
(\hat{X}, \hat{\Sigma}) = \begin{cases}
                        (\xi(\breve{X}_k^*), \xi(\breve{X}_k^*)/t) & \quad\text{if }\breve{X}_k^*\in (\overline{X}_s, t\Sigma),  \\
            (\breve{X}_k^*,\Sigma) & \quad\text{otherwise}.
        \end{cases}
\end{align*}
\end{proposition}
\begin{proof}
See Online Appendix~\ref{pf:prop:onesided_strat}.
\end{proof}

We let $\breve{\gamma}(\delta)$ denote the expected published estimate with one-sided selective publication and no p-hacking, and $\breve{\gamma}_h(\delta)$ denote the expected published estimate with both selective publication and p-hacking. It is not difficult to see that one-sided selective publication gives rise to publication bias, i.e., $\breve{\gamma}(\delta) > \Theta$ for $\delta>1$. We say that p-hacking mitigates publication bias if $\breve{\gamma}(\delta) > \breve{\gamma}_h(\delta) > \Theta$ and exacerbates publication bias if $\breve{\gamma}_h(\delta) > \breve{\gamma}(\delta) > \Theta$.

\begin{theorem}
\label{thm:onesided}
Under one-sided selective publication, there exist two thresholds $1<\underline{\delta}_1< \overline{\delta}_1<\infty$ such that p-hacking mitigates publication bias if $\delta > \overline{\delta}_1$ and exacerbates publication bias if $\delta < \underline{\delta}_1$.
\end{theorem}
\begin{proof}
See Online Appendix~\ref{pf:thm:onesided}.
\end{proof}

Theorem~\ref{thm:onesided} confirms that the main insights of Section~\ref{section:twosided} continue to hold under one-sided selective publication.

\section{Empirical Demonstration}\label{sec:emp}

Our analysis shows p-hacking can mitigate or exacerbate the effects of publication bias, depending on the selection environment. These results were obtained under simplifying assumptions that may be restrictive. We show in this section that both exacerbation and mitigation can arise under data-generating processes that are empirically plausible.

To better match data, we generalize our model to allow slow p-hacking to be potentially stochastic in Section \ref{section:emp_model}. Section \ref{section:identification} shows that it is possible to identify the selective-publication-only counterfactual at the literature level under the assumption that $\Theta \mid \Sigma = s \sim N(\theta(s),\sigma^2(s))$. Nonetheless, identification and estimation appear tricky in practice. Section \ref{sec:emp_mle} introduces parametric assumptions so that the restricted model can be estimated by maximum likelihood. In Section \ref{section:emp_results}, we estimate our model using two meta-analyses, finding suggestive evidence that both mitigation and exacerbation can arise in practice.

\subsection{Empirical Model with Stochastic Slow p-Hacking}\label{section:emp_model}

This section describes our empirical model, which generalizes the theoretical model by allowing slow p-hacking to be potentially stochastic. As will become clear in the next two sections, our empirical applications are more suited to models of one-sided selection. Therefore, we develop the empirical model based on one-sided selection.\footnote{Online Appendix \ref{app:twosided_emp} treats a more general empirical model in which $\delta_- \in [1, \delta_0]$.}

We modify the slow p-hacking setting as follows. Suppose a researcher slow p-hacks an estimate $X_n>0$ to positive significance. The associated cost is still $\kappa_s C(X_n)$. However, unlike the baseline model, we now let the generated output $(\hat{X}, \hat{\Sigma})$ be stochastic: it is drawn from an arbitrary distribution $H(\cdot, \cdot \mid X_n, \Sigma)$ and must satisfy $\hat{X}/\hat{\Sigma} \in[t, t_w]$ for some known $t_w \geq t$.
Note that if we let $H$ be the Dirac-delta function that assigns probability one to $(\xi(X_n), \xi(X_n)/t)$, this modified slow p-hacking setting is identical to that in the theoretical model. As discussed in \cite{simonsohn2020fast}, slow and fast p-hacking differ in how much control researchers have over their p-values. Allowing slow p-hacking to be stochastic captures the fact that this control is imperfect. Researchers nonetheless retain substantially more control under slow than fast p-hacking, as encoded by the support restriction on the former.

Notice that the researcher's optimal p-hacking strategy in this modified setting still follows Proposition~\ref{prop:onesided_strat}. Intuitively, allowing the outcome of slow p-hacking to be stochastic does not affect the researcher's incentives at any stage of the p-hacking process. Although the outcome of slow p-hacking is now stochastic, slow p-hacking yields the researcher the same payoff and incurs the same cost as in the theoretical model. Therefore, the researcher's slow p-hacking strategy remains unchanged. As a consequence, his incentive to engage in fast p-hacking also remains unchanged, since the continuation payoff from stopping drawing new estimates and potentially engaging in slow p-hacking is the same as in the theoretical model. Therefore, we can continue to use Proposition~\ref{prop:onesided_strat} for characterizing the researcher's p-hacking behavior in the empirical model.

\subsection{Identification under Normality}\label{section:identification}

Up to this point, our analysis has operated at the \textit{study level}, as our theoretical results characterized $\gamma_h$ and $\gamma$ conditional on the study $(\Theta, \Sigma)$. Going forward, we will turn to identifying parameters at the \textit{literature level}. We focus on the joint distribution of $(\Theta, \Sigma)$ and the mean effect size $\mathbb{E}(\Theta)$, as well as the means under p-hacking and selective publication and selective publication only. With some abuse of notation, we define the latter quantities as:
\begin{equation*}
    \mathbb{E}(\gamma_h) := \frac{\mathbb{E}_{\Theta, \Sigma}(q_h(\Theta, \Sigma; \delta)\gamma_h(\delta))}{\mathbb{E}_{\Theta, \Sigma}(q_h(\Theta, \Sigma; \delta))} \quad \mbox{ and } \quad \mathbb{E}(\gamma) := \frac{\mathbb{E}_{\Theta, \Sigma}(q(\Theta, \Sigma; \delta)\gamma(\delta))}{\mathbb{E}_{\Theta, \Sigma}(q(\Theta, \Sigma; \delta))} ~,
\end{equation*}
where $q_h(\Theta, \Sigma; \delta)$ and $q(\Theta, \Sigma; \delta)$ are the publication probabilities $\mathbb{P}(D = 1 \mid \Theta, \Sigma; \delta)$ under under the respective scenarios. We suppress the dependence of the above parameters on $\delta$ to lighten notation.

To simplify the analysis, we will assume that $\Theta \mid \Sigma = s \sim N(\theta(s), \sigma^2(s))$. This assumption nests the classic random effects model, in which $\Theta$ has a normal marginal distribution and is independent of $\Sigma$. This allows us to identify the parameters of interest:

\begin{theorem}\label{thm:identification_onesided}
   Suppose $t$ and $t_w$ are known and that $\Theta \mid \Sigma = s \sim N(\theta(s), \sigma^2(s))$. Moreover, suppose $\mathbb{P}(\hat{X}_i/\hat{\Sigma}_i > t_w), \mathbb{P}(|\hat{X}_i/\hat{\Sigma}_i| < t) > 0$ and that $\sigma^2(s) > 0$ for all $s \in \text{Supp}(\Sigma)$. Then, the joint distribution of $(\Theta, \Sigma)$, as well as $\delta$ are identified. In particular, $\mathbb{E}(\Theta)$, $\mathbb{E}(\gamma)$ and $\mathbb{E}(\gamma_h)$ are identified.
\end{theorem}
\begin{proof}
See Appendix~\ref{pf:thm:identification_onesided}.
\end{proof}

Our proof makes use of identification-at-infinity as well as deconvolution arguments. As such, identification may be weak and estimation may be challenging in practice. This motivates the additional restrictions that we describe in the next section.

Theorem \ref{thm:identification_onesided} requires two parametric assumptions. The first is the normality of $\Theta \mid \Sigma$, which allows us to identify the full distribution of $(\Theta, \Sigma)$ using only observations for which $\hat{X}_i/\hat{\Sigma}_i > t_w$. The second is the assumption that selection takes the form of a step function. This allows us to separate the effects of p-hacking from that of selection, to back out $\delta$ and hence $\mathbb{E}(\gamma)$. $\mathbb{E}(\gamma_h)$ is identified since it is the mean of published papers. We can thus infer the amount of bias arising from selective publication and assess whether p-hacking exacerbates or mitigates it.

Additionally, $t$ and $t_w$ need to be specified. Since a large empirical literature demonstrates the bunching of p-values at $5\%$ and $1\%$ across different fields of study \citep{masicampo2012peculiar, benjamin2018redefine, andrews2019identification}, it is natural to set $t = 1.96$. $t_w$ does not need to be exactly known. It only needs to be sufficiently large so that slow p-hacking is unlikely to exceed it. To accommodate p-hacking up to the 1\% level of significance, it is natural to set $t_w = 2.58 + \varepsilon$.

Note that the model is not fully identified. For example, the thresholds for slow p-hacking as well as the cost functions for fast and slow p-hacking are not identified.

\subsection{Maximum Likelihood Estimation}\label{sec:emp_mle}

The previous section shows that the counterfactuals of interest are semi-parametrically identified. In practice, however, identification and estimation may be challenging.

To make headway, we impose that researchers can draw at most one additional estimate, so that $n^*=2$. We set $t_w = 4$ to accommodate slow p-hacking up to the 1\% level of significance (critical value = 2.58) while ensuring sufficient mass in the tail for identification (see Theorem \ref{thm:identification_onesided}). The buffer allows researchers to overshoot their slow p-hacking target by a reasonable amount. Appendix \ref{app:emp_robust} shows that the results in the next section are not sensitive to the choice of $n^*$ and $t_w$.

Moreover, we will assume that $\Theta$ and $\Sigma$ are independent and let $(\theta, \sigma^2):=(\mathbb{E}(\Theta),\text{Var}(\Theta))$. \cite{andrews2019identification} require this assumption to identify the probability of selection using meta-analyses. In our case, independence reduces the nonparametric problem of estimating $(\theta(s), \sigma(s))$ to a parametric problem. Furthermore, as in the empirical sections of \cite{andrews2019identification}, we will assume that $\Sigma \sim \Gamma(\kappa, \lambda)$, where $\kappa$ and $\lambda$ are the shape and the rate parameters of the Gamma distribution to be estimated.

Finally, while we previously treated $\Sigma$ as given and formulated the slow p-hacking cost essentially as $C(X;\Sigma)$, our estimation adopts a more structured specification. We assume that
\[C(X; \Sigma)=\tilde{C}(X/\Sigma),\]
so that the cost of slow p-hacking depends only on the unmanipulated p-value associated with the estimate being manipulated. This specification reflects the idea that what matters for the difficulty of obtaining statistical significance is not the estimate itself, but its statistical significance relative to its standard error. Intuitively, it is easier to manipulate $X$ in an environment that is noisier. Combined with the assumption that researchers believe $\Theta=0$, this specification implies that $\overline{X}_f^n(s) = \overline{X}_f^n \cdot s$ for constant $\overline{X}_f^n$.

Under the above assumptions, we are able to specify a likelihood function for a censored version of the data with only a finite number of identified parameters. The likelihood function can be found in Appendix \ref{app:likelihood}. It describes censored data where observations in the bins $(-t,t)$ and $(t, t_w]$ are treated as indistinguishable from each other. As such, we do not have to specify the extent to which researchers slow p-hack $\hat{X}_i$ relative to $\hat{\Sigma}_i$, or any randomness that may be part of the slow p-hacking process, as long as the resulting t-statistic is upper bounded by $t_w$ in absolute value.

Additionally, we impose the constraints that $-t \leq \overline{X}_f^{n*-1} \leq \cdots \leq \overline{X}_f^1 \leq t$ as well as other model-implied constraints that are described in Appendix \ref{app:likelihood}. To avoid constrained optimization, as well as issues with parameters at the boundary, we implement them as strict inequalities via appropriate transformations of the variables.

Because of weak identification and error from numerical integration, the Hessian can be poorly conditioned. To compute standard errors for $\mathbb{E}(\gamma)$, we implement the Moore-Penrose inverse for the Hessian, dropping eigenvalues that are smaller than $10^{-9} \times$ the largest eigenvalue. In all cases, we verify that the dropped directions are orthogonal to those needed to compute the variance of $\mathbb{E}(\gamma)$, so that omitting them have minimal effects.

\subsection{Results}\label{section:emp_results}

We apply the above maximum likelihood procedure to data from two meta-analyses and find that both exacerbation and mitigation can be relevant in practice. Section \ref{section:emp_nudge} considers the effectiveness of nudges in promoting desired choices. We find that p-hacking may have mitigated the effects of publication bias in this literature. Section \ref{section:emp_aid} presents our results on the effectiveness of development aid on promoting growth, suggesting that p-hacking has exacerbated the effects of publication bias in this literature.

\subsubsection{Effect of Nudges on Choices}\label{section:emp_nudge}

The meta-analysis of \cite{mertens2022effectiveness} studies the effectiveness of nudges---interventions in the choice architecture---for promoting personally or socially desirable behavior. The analysis is based on 455 estimates from 334 studies\footnote{Studies refer to independent experiments, possibly belonging to the same paper. \cite{mertens2022effectiveness} treat studies as independent units of analysis, which is standard in meta-analysis in Psychology.}, obtained from 213 published and unpublished papers. The authors find that nudges are effective in altering behavior, with the headline result that $\mathbb{E}(\Theta) \approx 0.45$, measured in Cohen's $d$.\footnote{Cohen's $d$ is a standardized measure of effect size commonly used in meta-analysis.} \cite{maier2022no} argues instead that there is substantial selection in publication. Correcting for it using Bayesian model averaging, they estimate $\mathbb{E}(\Theta)$ to be between 0.04 and 0.11, with Bayes factor below $1$, implying weak evidence that the parameter is non-zero.

We apply our method to the data of \cite{maier2022no}.\footnote{This is the corrected version of \cite{mertens2022effectiveness}, though it assigns 2 papers to \texttt{publication\_id} 95. Fixing this discrepancy leads to a total of 213 papers.
}  Since it does not distinguish between published and unpublished papers, we use ChatGPT 5.6 Sol High to obtain the publication status of each paper, dropping the 6 unpublished papers from our sample\footnote{The unpublished papers have \texttt{publication\_id} 17,18,19,20,22 and 87.}. Following \cite{maier2022no}, we treat studies as independent units of analysis. Some studies present multiple estimates based on the same data. We choose the most significant estimates as our preferred sample. This corresponds to the notion that a paper needs at least one significant result to be ``publishable". Our estimate for $\mathbb{E}(\gamma_h)$ is then the simple mean of these preferred estimates.

Our final sample comprises 328 studies (from 207 published papers), each contributing one estimate. Histograms of the t-statistics for the full sample and our preferred sample are presented in Figure \ref{fig:emp_nudge}. The apparent bunching of the t-statistics at the threshold of 1.96 suggests appreciable slow p-hacking in this literature. We might therefore expect p-hacking to have a mitigative effect here. Comparing the two plots, we also see that retaining only the most significant estimate does not qualitatively alter the distribution of t-statistics.

\begin{figure}
\centering
\includegraphics[width=0.48\linewidth]{application/hist_nudge_full.png}
\includegraphics[width=0.48\linewidth]{application/hist_nudge_p1.png}
\caption{Histograms of the t-statistics from the full published sample with multiple estimates per study, and from our preferred sample, which keeps only the most significant estimates. Red dashed lines indicate $\pm 1.96$.}\label{fig:emp_nudge}
\end{figure}

As discussed in Section \ref{sec:emp_mle}, we set $n^* = 2$ and $t_w = 4$. Because the left-tail is non-existent, the two-sided selection model is not identified and we impose the one-sided selection model. Appendix \ref{appendix:emp_robust_nudge} shows that the results below are not sensitive to the choices of $n^*, t_w$ or our choice of preferred sample.

Results are presented in Table \ref{tab:emp_nudge}. We find that negative significant and insignificant results are published at 8.3\% the rate at which positive and significant results are published. Correcting for selective publication and p-hacking, we estimate $\mathbb{E}(\Theta)$ to be 0.171. It is insignificant at the 5\% level and comparable to the upper end of the values obtained by \cite{maier2022no}. We estimate $\mathbb{E}(\gamma)$ to be $0.959$ and $\mathbb{E}(\gamma_h)$ to be $0.531$, suggesting that p-hacking has mitigated the bias from selective publication by 54\%. That the mitigation from slow p-hacking dominates the exacerbation effects in this literature is in line with our intuition from the bunching in Figure \ref{fig:emp_nudge}.

\begin{table}[htbp]
  \centering
    \begin{tabular}{ccccc}\hline\hline
    $1/\delta$ & $\mathbb{E}(\Theta)$ & $\mathbb{E}(\gamma_h)$ & $\mathbb{E}(\gamma)$ & $\mathbb{E}(\gamma) - \mathbb{E}(\gamma_h)$ \\
    \hline
    0.083 & 0.171 & 0.531 & 0.959 & 0.428 \\
    (0.031) & (0.130) & (0.028) & (0.180) & (0.178) \\
    \hline\hline
    \end{tabular}
    \caption{MLE on the dataset of \cite{mertens2022effectiveness} allowing for selective publication and p-hacking. $1/\delta$ is the relative publication probabilities for negative significant and insignificant results compared to positive and significant results. $\mathbb{E}(\Theta)$ is the mean effect size under the true distribution of latent studies. $\mathbb{E}(\gamma_h)$ is the observed mean under both selective publication and p-hacking. $\mathbb{E}(\gamma)$ is the counterfactual mean under selective publication only. Standard errors computed using Moore-Penrose inverse are presented in round brackets.} \label{tab:emp_nudge}
\end{table}

\subsubsection{Effect of Development Aid on Economic Growth}\label{section:emp_aid}

\cite{doucouliagos2011ineffectiveness} conducts meta-analysis on the effect of development aid on economic growth. Using a dataset of 105 papers and 1217 estimates, the authors find that positive and significant results are selectively published, and that correcting for these results using meta-regressions leads to estimates of partial correlations\footnote{Partial correlations are obtained by standardizing regression estimates to make them more comparable across studies. They are commonly used in meta-analysis.} that are small in magnitude---between 0.02 to 0.04---and often statistically insignificant.

We apply our method to the data of \cite{doucouliagos2011ineffectiveness}, which was updated in \cite{doucouliagos2013robust} to include a total of 113 papers and 1347 estimates. As before, we select the most significant estimate in each study as the preferred estimate. We plot the t-statistics from the full sample as well as our preferred sample in Figure \ref{fig:emp_aid}. In contrast to the literature on nudges, there is no pronounced excess mass at the significance thresholds, so that we might expect less mitigative effects from slow p-hacking. Comparing the two plots in Figure \ref{fig:emp_aid}, we see that the distribution of our preferred estimates puts substantially less mass in the insignificant region than the full sample. This follows mechanically from our selection criterion. If we believe that a paper is assessed based on its strongest results, then our preferred estimates more closely reflect the statistics that are marginal for an editor's decision, and which are therefore candidates for p-hacking. Our theoretical analysis also suggests that reducing the mass of insignificant results will make exacerbation less likely. As such, if the most significant sample is not the marginal sample, we loosely interpret the results below as a lower bound on the exacerbation effect of p-hacking in this literature.

\begin{figure}
\centering
\includegraphics[width=0.48\linewidth]{application/hist_aid_full.png}
\includegraphics[width=0.48\linewidth]{application/hist_aid_p1.png}
\caption{Histograms of the t-statistics from the full published sample with multiple estimates per paper, and from our preferred sample, which keeps only the most significant estimates. Red dashed lines indicate $\pm 1.96$.}\label{fig:emp_aid}
\end{figure}

As before, we fix $n^* = 2$, $t_w = 4$. We also impose one-sided selection, since \cite{doucouliagos2011ineffectiveness} argues that the two leading sources of publication bias in this literature are (1) researcher idealism and (2) incentives for increasing development funding, both of which favor positive estimates. Appendix \ref{appendix:emp_robust_aid} shows that the results below are not sensitive to the aforementioned choices, although fitting a general model that estimates $\delta_-$ leads to estimates with substantially larger standard errors.

Results are presented in Table \ref{tab:emp_aid}. We find that negative significant and insignificant results are published at 15\% the rate at which positive and significant results are published. Correcting for selection and p-hacking, estimate for $\mathbb{E}(\Theta)$ to be $-0.113$,  which is negative but insignificant at the 5\% level. We estimate $\mathbb{E}(\gamma)$ to be $0.121$ and $\mathbb{E}(\gamma_h)$ to be $0.185$, suggesting that p-hacking has exacerbated the bias from selective publication by 27\%. \cite{doucouliagos2013robust} argues that researchers in the aid literature cherry-pick control variables in order to achieve positive and significant estimates. To the extent that this behavior is captured by our stylized model of fast p-hacking, it would explain why exacerbation dominates mitigation in this literature.

\begin{table}[htbp]
  \centering
    \begin{tabular}{ccccc}\hline\hline
     $1/\delta$ & $\mathbb{E}(\Theta)$ & $\mathbb{E}(\gamma_h)$ & $\mathbb{E}(\gamma)$ & $\mathbb{E}(\gamma) - \mathbb{E}(\gamma_h)$ \\
    \hline
     0.149 & -0.113 & 0.185 & 0.121 & -0.064 \\
    (0.057) & (0.061) & (0.032) & (0.038) & (0.017) \\
    \hline\hline
    \end{tabular}
    \caption{MLE on the dataset of \cite{doucouliagos2011ineffectiveness} allowing for selective publication and p-hacking. $1/\delta$ is the relative publication probabilities for negative significant and insignificant results compared to positive and significant results. $\mathbb{E}(\Theta)$ is the mean effect size under the true distribution of latent studies. $\mathbb{E}(\gamma_h)$ is the observed mean under both selective publication and p-hacking. $\mathbb{E}(\gamma)$ is the counterfactual mean under selective publication only. Standard errors computed using Moore-Penrose inverse are presented in round brackets.} \label{tab:emp_aid}
\end{table}

\section{Conclusion}\label{section--conclusion}

This paper studies the effects of p-hacking on the bias of published estimates in environments with selective publication. We show that under both one- and two-sided selection, fast p-hacking unambiguously exacerbates the effects of publication bias, whereas slow p-hacking can either exacerbate or mitigate it, depending on the extent of selection in the publication process. The possibility of mitigation echoes the work of \cite{kuhn1962structure}: communities of biased researchers can nonetheless produce useful work due to self-correcting dynamics in the scientific process.

Although these results are developed in a stylized setting, we show that the selection-only counterfactual mean is identified under a normality assumption. Taking the model to the data, we find suggestive evidence that p-hacking mitigates the effects of selective publication in the literature on the effects of nudges, but exacerbates it in the literature on the effects of aid on growth.

Discussions about p-hacking and selective publication often treat the two phenomena separately. Our results show that it is important to understand p-hacking in the context of the selection environment that produced it, echoing \cite{friese2020p}, which observes their interacting effects in simulations. We stress, however, that our results concern only the bias of published estimates. They say nothing about the other pernicious effects of p-hacking, such as inflating the number of false positives, eroding trust in science, or distorting the incentives of career researchers. What is clear is that p-hacking and selective publication cannot be understood in isolation. Only by studying them together can we hope to model the scientific process faithfully and design publication policies that preserve its credibility.

\newpage
\begin{singlespace}
\bibliographystyle{chicago}
\bibliography{phacking.bib}
\end{singlespace}

\clearpage
\newpage