EconBase
← Back to paper

Publication Design with Incentives in Mind

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

88,239 characters · 15 sections · 74 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Publication Design with Incentives in Mind

abstractThe publication process both determines which research receives the most attention, and influences the supply of research through its impact on researchers' private incentives. We introduce a framework to study optimal publication decisions when researchers can choose (i) whether or how to conduct a study and (ii) whether or how to manipulate the research findings (e.g., via selective reporting or data manipulation). When manipulation is not possible, but research entails substantial private costs for the researchers, it may be optimal to incentivize cheaper research designs even if they are less accurate. When manipulation is possible, it is optimal to publish some manipulated results, as well as results that would have not received attention in the absence of manipulability. Even if it is possible to deter manipulation, such as by requiring pre-registered experiments instead of (potentially manipulable) observational studies, it is suboptimal to do so when experiments entail high research costs. We illustrate the implications of our model in an application to medical studies.

\onehalfspacing

Introduction

Publication decisions shape the process of scientific communication. By selecting what to publish, journals affect which findings receive the most attention and can inform the public about the state of the world. The design of publication rules has therefore motivated recent debates on how statistical significance should affect publication when the goal is to direct attention to the most informative results abadie2020statistical,frankel2022findings.

However, the publication process also affects the supply of research by influencing researchers' incentives about how to conduct research. Researchers have many degrees of freedom about how to conduct their research, such as how and where to run an experiment allcott2015site,gechter2022combining; the size, cost, and effort associated with the study thompson2000should, grabowski2002returns; and which findings to report from a given study brodeur2020methods,elliott2022detecting. We refer broadly to their choices about each of these aspects of a study broadly as a research design.

Researcher's private incentives may influence how they choose their design. Yet, “while economists assiduously apply incentive theory to the outside world, we use research methods that rely on the assumption that social scientists are saintly automatons” glaeser2006researcher. This raises the questions of when and how researchers' incentives should impact the design of publication processes, and, more broadly, the optimal allocation of attention to research.

This paper studies optimal publication decisions when researchers choose research designs based on private costs and benefits. We frame this question as a mechanism design problem: a social planner (principal) optimizes a publication rule, taking into account the incentives of a researcher (agent). The planner aims to use the publication process to efficiently allocate the attention of the audience to research findings.

More specifically, as in frankel2022findings (building in turn on wald1950statistical), we suppose that research results impact the actions of an audience who has limits on how much attention they can devote to research. The social planner seeks to publish results that are most important for the audience, net of the cost of (or taking into account constraints on) publication or attention. Due to attention costs, not all results will be published, which leaves the planner with a non-trivial trade-off about which results and designs to publish. We then introduce a model in which publication decisions affect the supply of research in the first place (Section (ref)): given the publication rule, the researcher chooses the design that maximizes her value from publication (or other attention) net of research costs.

As a concrete example, consider a medical journal that is seeking to decide whether to publish results from a clinical study. The journal wants to convey accurate information on drug efficacy and direct the audience's attention to the most effective drugs. \footnote{ For example, the stated mission of the New England Journal of Medicine is “to publish the best research and information at the intersection of biomedical science and clinical practice and to present this information in understandable, clinically useful formats that inform health care practice and improve patient outcomes.” See also ana2004role and demaria2022role for further discussion of the role of a medical journal. } However, researchers respond to the design of the publication system through the size, length, cost of research studies, the composition of control groups thorlund2020synthetic, and in some cases, which specific findings to report riveros2013timing,shinohara2015protocol.

We draw a dichotomy between researchers' incentives about (i) whether or how to conduct a study and (ii) which findings to report (e.g., via data manipulation or selective reporting).

We first focus on (i) and abstract from data manipulation and selective reporting (which we defer to Section (ref)); that is, we first suppose that the research design is observable and verifiable as part of the publication process (Section (ref)). For example, a researcher could choose between experiments with different mean-squared errors and costs, and, constrained by a pre-analysis plan, truthfully report an unbiased estimate of a treatment effect. The planner can direct attention to the results of any executed study by publishing them; publication can depend on both a study's design and its results.

We suppose that the publication process affects the supply of research through an individual rationality constraint: for a researcher to be willing to conduct a study, they must be compensated with a large enough ex ante publication probability. Without these individual rationality constraints, the first-best publication rule would ignore researchers' private costs and publish only results from studies conducted with the lowest possible mean squared error; due to attention costs, not all such results would be published.

However, when the research cost of accurate designs is high enough, the first-best publication rule does not compensate the researcher enough to incentivize them to execute such designs. As a result, the planner faces a trade-off between directing the audience's attention to results that would not deserve it (e.g., treatments with negligible effects) and rewarding the researcher enough to make them willing to use a costly design. The planner may prefer publishing results from less accurate studies---if they are sufficiently less expensive---to avoid having to publish too many results from accurate and costly experiments.

Returning to our medical example, after providing a new drug to a treatment group, scientists can evaluate its efficacy by using experiments with a different number of participants. A larger experiment is more precise, but more costly for the researcher. Allowing smaller experiments can impact on the supply of medical research and new drugs by decreasing research costs. Our analysis shows that due to the interaction between attention constraints and supply effects, publishing results from smaller experiments can be desirable when experiments are sufficiently costly for researchers to execute.

We next turn in Section (ref) to settings in which researchers can engage in potentially arbitrary data manipulation or selective reporting. We suppose that the researcher can chose their research design after learning the (likely) results of their study. For instance, in the absence of pre-specification, researchers may engage in forms of $p$-hacking by, e.g., selecting regression specifications or control groups based on the observed outcomes. We may also be concerned that researchers select experimental environments where they expect treatment effects to be non-representatively large---termed $g$-hacking by niederle2025experiments.\footnote{This form of manipulation introduces a form of bias analogous to site selection bias.}

More concretely, we suppose that after observing the study's results, the researcher can report a biased statistic at a (reputational) cost increasing in the bias. The bias is unobserved by the planner. The audience is unaware of the possibility of manipulation in publishing findings, so takes published results at face value. The planner seeks to choose a publication rule to minimize the audience's loss, accounting for the possibility of manipulation.

Each publication rule generates a different degree of manipulation. For example, suppose that the planner used the publication rule that would be optimal without manipulation---i.e., publishing if the reported statistic is above a cutoff. The researchers would then manipulate their results to reach the cutoff, at a substantial loss for the audience. A different approach that has been proposed is to completely deter manipulation by making publication dependent only on the design, and not on results.\footnote{In practice, this approach can be implemented by committing to publication based on pre-analysis plans, as the Journal of Development Economics and the Journal of Clinical Epidemiology do (among others).} However, this approach can incur substantial costs by directing the audience's attention to results that would not substantially affect their actions.

We show that the optimal rule has three key features. First, it increases the cutoff for which findings always get published compared to settings without manipulation. Second, just below this cutoff, it randomizes publication decisions, making the researcher indifferent about whether to (or how much to) manipulate their results. In particular, the planner publishes some findings that would not be published without the possibility of manipulation. Third, some manipulation does occur in equilibrium for results that would merit attention absent manipulation. Unlike manipulation under standard cutoff rules, this form of manipulation benefits the planner by directing attention to results that impact the audience's action enough to merit attention.

To gain some intuition for these features of the optimum, consider first the optimal publication rule without manipulation. The planner can eliminate researchers' incentives to manipulate by publishing some results that are below the cutoff, but that would not merit attention. To publish fewer such studies, the planner should increase the cutoff for a result to be guaranteed publication. Increasing the cutoff however also reduces the number of studies that would be published in the absence of manipulation. Therefore, the planner then encourages researchers with results that would be published without manipulation, but are below the new cutoff, to engage in some small manipulation to increase their chances of publication. The loss from publishing some slightly manipulated studies is second-order relative to the gain from publishing more results that substantially impact the audience's action.

To formally characterize the optimal publication rule under manipulability, we formulate the social planner's problem as a mechanism design problem with “false moral hazard” due to the researcher choosing a manipulation after learning the true results. The absence of direct transfers and the inability to reward the researcher with a publication probability above one make the mechanism design problem effectively one with limited transfers. As a result, standard methods as in mirrlees1971exploration and myerson1981optimal do not apply. We solve the mechanism design problem by identifying the precise pattern of binding incentive constraints.

In Section (ref), we combine the models from Sections (ref) and (ref) to ask whether/when the planner should incentivize researchers to run costly experiments that adhere to pre-analysis plans, rather than allow for (cheaper) observational studies. \footnote{Our analysis also applies to whether the planner should mandate researchers to send a costly signal (e.g., a detailed pre-analysis plan) that deters them from engaging in manipulation.} Without accounting for the supply effects of making research more costly, the planner would always require a nonmanipulable experiment spiess2018optimal,kasy2023optimal. However, due to supply effects, the planner may prefer some observational studies, even when manipulation may occur.

Finally, in Section (ref), we bring our model to the data and study the optimal publication rules for medical studies, taking into account researchers' best response. We first focus on optimal publication rules in contexts with manipulation. We calibrate the model using about 800,000 $p$-values from studies in hundreds of medical and pharmaceutical journals collected by head2015extent. Setting attention costs to capture a standard $5\%$-level $t$-test, we find that the threshold at which results should always be published increases from 1.96 to 2.64. However, many results strictly smaller than 1.96 are published in equilibrium. Using the optimal rule makes the average bias of published findings drops by more than a factor of 2.

We then ask when an observational study with possible manipulation may be preferable to an (equally precise) clinical trial without manipulation, building on the framework of Section (ref). Using our calibration, we determine how costly the full clinical trial must be for an observational study to be preferable for the planner. We compare these costs with estimates of the expected cost-benefit ratio from clinical trials in 13 therapeutic areas. Our calculations illustrate two broad takeaways. First, although experiments dominate observational studies on average in several therapeutic areas, some areas with high research costs may merit the use of observational studies. However, achieving gains from allowing observational studies requires using a different publication standard for them that accounts for their manipulability.

Our results speak to a policy debate regarding the design of controls in clinical trials. After providing a new drug to a treatment group, scientists can evaluate its efficacy by using either an experimental placebo group or a synthetic control group obtained from historical medical records popat2022addressing,yin2022exploring. Using a synthetic control group can have large impact on the supply of medical research and new drugs by decreasing research costs jahanshahi2021use, FDA2023evaluation, wong2014examination. However, using a synthetic control group may increase the estimate's mean-squared error due to lack of randomization, as well as open the door to the manipulation of how the control is specified. Our analysis highlights that allowing synthetic controls may be desirable if the publication or approval rules is designed to account for their manipulability---a consideration that has not been raised in the current policy debate on the use of synthetic control groups.

\paragraph*{Related literature.} This paper connects to a growing literature that develops economic models to analyze statistical protocols. In the context of scientific communication, andrews2019identification, abadie2020statistical, andrews2021model, kitagawa2023optimal, and (most closely related to our paper) frankel2022findings have analyzed how research findings are or should be reported to inform the public. Our analysis builds on this literature by introducing a model that incorporates researchers' incentives. This allows us to study how researchers' incentives shape the optimal design of the optimal publication process.

We connect to a broad literature on statistical decision theory wald1950statistical,savage1951theory, manski2004, hirano2009asymptotics, tetenov2012statistical, KitagawaTetenov_EMCA2018 focusing in particular on settings with private researcher incentives. \footnote{In addition to the references discussed in detail below, related work in this line includes chassang2012selective, manski2016sufficient, banerjee2017decision, banerjee2020theory, williams2021preregistration, bates2022principal,bates2023incentive, frankel2022improving, libgober2022false, and yoder2022designing. } We develop a unified framework that allows us to analyze both settings in which researchers may choose the research design absent private information, and settings in which researchers can choose the design and manipulate reported findings with private information. This allows us to formally study ideas such as when/whether unsurprising results should be published, and whether manipulation should occur in equilibrium, as asked by glaeser2006researcher.

In particular, an important distinction from some of these models studying approval decisions, such as tetenov2016economic, bates2022principal, bates2023incentive, and viviano2021should, is that we consider data manipulation and selective reporting. We also incorporate costs for the researcher of manipulating (and of executing the design), unlike analyses with manipulation by spiess2018optimal and kasy2023optimal---which we show leads to qualitatively different optimal publication rules. mccloskey2022incentive and andrews2019identification propose statistical adjustments for $p$-hacking or publication bias holding researchers' behavior fixed. We instead study optimal publication rules in equilibrium, taking into account researchers' best responses. Last, henry2019research and di2017persuasion,di2021strategic study decisions with sequential access to the data and with selective sampling, which are different from the question of selective design (and reporting) choice studied here.

A large empirical and econometric literature has documented several aspects of the research process, including selective reporting, data manipulation, specification search, as well as site selection bias and observational studies' bias allcott2015site,olken2015promises, banerjee2020praise,brodeur2020methods, rosenzweig2020external, miguel2021evidence, elliott2022detecting, gechter2022combining, rhys2024much. Our contribution here is to provide a formal model that studies how incentives interact with several of these choices, shedding light on optimal publication rules and how they compare to ones used in practice.

Setup

Consider three agents: a researcher, a (representative) audience, and a social planner. The audience and social planner are interested in learning a parameter $\theta \in \mathbb{R}$. All agents share a common prior $\theta \sim \mathcal{N}(0, \eta^2)$, whose mean is normalized to 0 without loss of generality.\footnote{Our results continue to hold for $\theta \sim \mathcal{N}(\mu,\eta^2)$, with the audience adjusting their action accordingly.}

A researcher conducts a study to inform the audience about $\theta$, and seeks to publish their findings. A study is summarized by $(X, \Delta),$ where $\Delta$ denotes the design and $X$ the results observed in the study. If a study is conducted, it will be evaluated according to a publication rule $p(X,\Delta)$ with values in $[0,1]$. Here, $p(X,\Delta)$ represents the probability of publishing the study, which is assumed to be a Borel measurable function of $(X,\Delta)$. We assume that

equation[equation omitted — 115 chars of source]

Here, $S_\Delta^2$ is the variance of design $\Delta$. The quantity $\beta_\Delta$ represents the component of the mean of design $\Delta$ that the audience does not understand; we refer to $\beta_\Delta$ as the bias of design $\Delta$.

Specifically, conditional on publication, the audience forms posterior beliefs about $\theta$ using Bayes' Rule assuming $\beta_\Delta = 0$---i.e., $\theta \sim \mathcal{N}\big(\frac{X \eta^2}{S_\Delta^2 + \eta^2},\frac{S_\Delta^2 \eta^2}{S_\Delta^2 + \eta^2}\big)$. (We can recenter results to incorporate the component of $\mathbb{E}[X(\Delta)- \theta|\theta]$ that the audience understands.) Conditional on non-publication, the audience's posterior mean equals its prior mean (zero), as would arise from Bayes' Rule with $p(\cdot)$ symmetric in $X$ and $\beta_\Delta$ symmetric around 0. Thus, the audience's action $a_p^\star(X,\Delta)$ is $$ a_p^\star(X, \Delta) =

cases\frac{X \eta^2}{S_\Delta^2 + \eta^2} & if the study is published \\ 0 & otherwise

. $$

Given results $X=X(\Delta)$ for a design $\Delta$, and a parameter $\theta$, the planner incurs a loss

equation[equation omitted — 181 chars of source]

conditional on $X$ and $\theta$, where $\mathbb{E}_p$ denotes expectation with respect to any stochasticity in the publication decision rule. This is the (expected) loss of the audience, net of attention costs (or shadow costs of attention constraints) proportional to $c_a$.

We normalize the value of publication for the researcher to 1. Given a design $\Delta$, a publication rule $p$, and results $X$, the researcher's expected payoff (conditional on $X$) is

equation[equation omitted — 87 chars of source]

where $C_\Delta \le 1$ is the researcher's cost of executing design $\Delta$. \footnote{We assume $C_\Delta \le 1$ simply to rule out trivial cases in which design $\Delta$ is never chosen by the researcher.} As is standard, whenever the researcher is indifferent between two designs, we implicitly assume she chooses the design that minimizes the planner's expected loss.

Publication rules under verifiable designs

This section studies optimal publication when the planner can observe the research design, and condition publication on it. We focus on designs $\Delta$ that are unbiased in that $\beta_\Delta = 0$, and defer the analysis of designs that are biased due to manipulation to the following section.

figure[figure omitted — 3,360 chars of source]

Our analysis proceeds in three steps. We first characterize the optimal publication rule subject to the constraint of incentivizing the researcher to implement a particular design $\Delta$. We then characterize which designs are worth incentivizing relative to an outside option. Last, we characterize the optimal publication rule that chooses between multiple designs.

Preliminary analysis for implementing a particular design

As a first step, we characterize the constrained optimal publication when the planner must make implementing a particular design $\Delta$ individually rational for the researcher---i.e., when the planner must guarantee that the researcher is willing to conduct the study.

defnA constrained optimal publication rule for a design $\Delta$ is a publication rule $p_\Delta^\star$ that minimizes $\mathbb{E}\left[\mathcal{L}_p(X(\Delta), \Delta, \theta) \right]$ subject to $\mathbb{E}[v_p(X(\Delta), \Delta)] \ge 0.$

Our first result shows that the constrained optimal publication rule then takes a threshold form, where the threshold $t^\star_\Delta$ for publication depends on the prior, the mean squared error and research cost of the design $\Delta$, and the publication cost.

prop[Constrained optimal publication rule] If $\Delta$ is an unbiased design, then a constrained optimal publication rule for $\Delta$ is the threshold rule $p_\Delta^\star(X) = 1\left\{|X| \ge t^\star_\Delta\right\}$, where \footnote{The proof shows that this rule is uniquely optimal (up to sets of measure 0).} \[t^\star_\Delta = \min\left\{\frac{S^2_\Delta + \eta^2}{\eta^2} \sqrt{c_a},\left|\Phi^{-1}(C_\Delta/2)\right| \sqrt{S_\Delta^2 + \eta^2}\right\}.\]

Here, we write $\Phi$ for the cumulative distribution function of a standard normal.

The proof is in Appendix (ref). To understand the intuition behind Proposition (ref), first suppose that the research cost is $C_\Delta = 0$, so there is no individual rationality constraint for the researcher. Then, as in frankel2022findings, as the planner's publication cost $c_a$ is nonzero, the planner will publish results that move the audience's optimal action enough to justify incurring the attention costs $c_a$: i.e., results $|X| \ge \gamma^\star_\Delta,$ where \[\gamma^\star_\Delta = \frac{S^2_\Delta + \eta^2}{\eta^2} \sqrt{c_a}.\] When this cutoff rule guarantees an ex ante publication probability of at least $C_\Delta$, the individual rationality constraint does not bind. In this case, we say the design is cheap.

defnAn unbiased design $\Delta$ is cheap if $C_\Delta < \mathbb{P}\left(|X(\Delta)| \ge \gamma_\Delta^\star\right)$ and expensive otherwise.

Note that higher publication costs $c_a$, and lower prior variances $\eta^2$, both raise the threshold $\gamma^\star_\Delta$ and hence make designs more likely to be expensive. Whether a design is cheap depends on how informative it is (relative to attention costs).

For expensive designs, the cutoff rule from frankel2022findings does not provide a large enough ex ante publication probability to entice the researcher to conduct the study in the first place. Hence, the planner needs to commit to publishing more results in order to satisfy the researcher's individual rationality constraint. It is optimal for the planner to publish results that move the audience's action the most, even if these results do not move the audience's action enough to justify the attention cost $c_a$. Hence, the planner sets a cutoff that ensures an ex ante publication chance of $C_\Delta$---i.e., a cutoff of \[\left|\Phi^{-1}(C_\Delta/2)\right| \sqrt{S_\Delta^2 + \eta^2}.\] This second cutoff is below $\gamma^\star_\Delta$ for (and only for) expensive designs, and is the optimal cutoff for such designs. Thus, if research costs are large enough that the implemented design is expensive, the researcher's incentives play a central role in determining the optimal publication rule, unlike in frankel2022findings.

The following corollary summarizes and formalizes the preceding discussion.

cor\begin{enumerate}[label=(\alph*)] • If $\Delta$ is a cheap design, then the cutoff for a constrained optimal publication rule is $t^\star_\Delta = \gamma^\star_\Delta$. • If $\Delta$ is an expensive design, then the cutoff for a constrained optimal publication rule is $t^\star_\Delta = \left|\Phi^{-1}(C_\Delta/2)\right| \sqrt{S_\Delta^2 + \eta^2}$. \end{enumerate}

Which designs are ever worth incentivizing

As a second step, we study when a design is worth incentivizing relative to the outside option of no study. If it is not worth doing so, then it is not worth making it individually rational for the researcher to conduct research based on design $\Delta$ in optimum.

Let $\mathcal{L}^\star_\Delta = \mathbb{E}\left[\mathcal{L}_{p_\Delta^\star}(X(\Delta), \Delta, \theta) \right]$ denote the optimal expected loss for the planner once implementing design $\Delta$. (Here $p_\Delta^\star$ is a constrained optimal publication rule for $\Delta.$) The expected loss if no research is published is the prior variance $\eta^2$. Comparing these two quantities determines whether a design is worth incentivizing in the first place.

defnA design $\Delta$ is worthwhile if $\mathcal{L}^\star_\Delta \le \eta^2$.

We next characterize which designs are worthwhile. If a design is cheap, then the planner can selectively publish only results that move the audience's beliefs enough to justify incurring the attention cost. Thus, the ex post loss under the constrained optimal publication rule is always lower than $\eta^2$, and so the (ex ante) expected loss is less than $\eta^2$.

prop[When are cheap designs worthwhile?] Every cheap design $\Delta$ is worthwhile.

The proof is in Appendix (ref). Whereas cheap design are always worthwhile, for expensive designs, the situation is more delicate. Incentivizing the researcher to implement a design requires committing to publish results that the planner would ex post prefer not to publish. When attention costs $c_a$ are large enough, the cost of publishing these marginal results outweigh the benefit of publishing results that substantially affect the audience's action. How large $c_a$ needs to be for this to occur depends on the design's cost and variance.

To formalize this intuition, it will be convenient to express our results in terms of the difference between the posterior and prior variances conditional on publication of the results of a design $\Delta$, which we denote by \[\operatorname{PostVarRed}(\Delta) := \eta^2 - \frac{S_\Delta^2\eta^2}{S_\Delta^2 + \eta^2} = \frac{\eta^4}{S_\Delta^2 + \eta^2}.\] This quantity is a measure of the informativeness of a design: it represents how much learning the results of the design improves the expected utility of a Bayesian audience with $L^2$ loss. Note that $\operatorname{PostVarRed}(\Delta)$ is increasing in $\eta^2$ and decreasing in $S_\Delta^2$. Whether a design is worthwhile then depends on how $\operatorname{PostVarRed}(\Delta)$ compares to the product of the attention and research costs, up to a small remainder that vanishes with large research costs ($C_\Delta \approx 1$).\footnote{Proposition (ref) is sharp up to the term $\eta^2 (1 - C_\Delta)^3$; Lemma (ref) in Appendix (ref) provides a more involved condition that is both necessary and sufficient for worthwhileness.}

prop[When are expensive designs worthwhile?] Let $\Delta$ be an expensive design. \begin{enumerate}[label=(\alph*)] • If $\operatorname{PostVarRed}(\Delta) \ge C_\Delta c_a + \eta^2 (1 - C_\Delta)^3$, then $\Delta$ is worthwhile. • If $\operatorname{PostVarRed}(\Delta) < C_\Delta c_a$, then $\Delta$ is not worthwhile. \end{enumerate}

The proof is in Appendix (ref). In particular, Proposition (ref) shows that as $C_\Delta$ or $c_a$ increases, the posterior variance reduction must increase proportionally for a design to remain worthwhile (up to a small remainder that vanishes in the case of large research costs).

Choosing which design to incentivize

We next study the optimal publication rule when there is more than one possible design. Without loss of generality, we suppose that both designs are worthwhile, and that the one with a lower mean squared error has a higher research cost, so the planner faces a non-trivial problem about which design to incentivize.

settingResearchers can choose between two designs $E,O$ that are unbiased and worthwhile. The designs have mean squared errors $S_E^2 < S_O^2$ and costs $C_E > C_O$.

We think of $\Delta = E$ as a possibly expensive experiment and $\Delta = O$ as a lower-cost experiment or non-manipulable observational study (manipulation is studied in Section (ref)).

exmp[Low-cost and costly experiment] Suppose that $O$ corresponds to an experiment with fewer participants than $E$. In this case, we have $C_O < C_E$ and $S_O^2 > S_E^2$. \qed
exmp[Experiment versus nonmanipulable observational study] Suppose that $O$ corresponds to an observational study with no strategic manipulation of the results. Let $X(E) = \theta + \varepsilon_E$ and $X(O) = \theta + b_O + \varepsilon_O$, where $\varepsilon_E \sim \mathcal{N}(0, S_E^2)$ denote the estimation noise from the experiment, $\varepsilon_O \sim \mathcal{N}(0,\sigma_O^2)$ denotes the idiosyncratic noise from the experiment or observational study in $O$, and $b_O | \varepsilon_O \sim \mathcal{N}(0,\sigma_B^2)$ denotes a random effect, which captures unobserved bias drawn from a fixed (Gaussian) distribution. \footnote{For example, rhys2024much investigates the distribution of $b_O$ through a meta-analysis. } We then have that $X(O) \sim \mathcal{N}(\theta,S_O^2)$, where the mean-squared error $S_O^2 = \sigma_O^2 + \sigma_B^2$ includes both sampling uncertainty $\sigma_O^2$ and irreducible error $\sigma_B^2$ arising from the variance of the bias. \qed

We next use Proposition (ref) to study the optimal choice between the two designs. Because the design $\Delta$ is observable by the planner and verifiable as part of the publication process, the planner can incentivize their preferred design by setting

equation[equation omitted — 228 chars of source]

For instance, the planner may only accept experiments with a minimum level of precision. It is immediate that $p^\star(X, \Delta)$ minimizes the planner's expected loss. We therefore study the optimal design choice by comparing the minimized loss of the social planner when implementing the experiment versus implementing the observational study. More generally, we can use similar logic to compare the effectiveness of any two designs.

defnDesign $\Delta$ is planner-preferred to design $\Delta'$ if $\mathcal{L}^\star_{\Delta} < \mathcal{L}^\star_{\Delta'}$.

It is immediate that if a design is planner-preferred to a worthwhile design $\Delta'$, then $\Delta$ is worthwhile. In particular, Proposition (ref) implies that if a design $\Delta$ is planner-preferred to a cheap, unbiased design $\Delta'$, then $\Delta$ is worthwhile.

If the more precise experiment $E$ is cheap, then its higher research cost is irrelevant to the planner. Therefore, the experiment is planner-preferred to $O$.

propIn Setting (ref), if $E$ is cheap, then $E$ is planner-preferred to $O$.

The proof is in Appendix (ref). This result implies that it suffices to compare the mean-squared error of two cheap studies to identify which one is planner-preferred.

When the experiment $E$ is expensive, the situation is more delicate. Implementing $E$ requires committing to publish more results, which may be costly for a planner. When the publication or attention costs $c_a$ are large enough, the costs of publishing more results outweighs the benefits of a more precise design.

propIn Setting (ref), if $E$ and $O$ are expensive, then there exists a threshold $c_a^\star(E,O,\eta) > 0 $ such that $E$ is planner-preferred to $O$ if and only if $c_a < c_a^\star(E,O,\eta)$, where \[c_a^\star(E,O,\eta) = \frac{\operatorname{PostVarRed}(E) - \operatorname{PostVarRed}(O)}{C_E - C_O} - \eta^2 \frac{O\big((1-C_E)^3\big) - O\big((1-C_O)^3\big)}{C_E - C_O}.\]

The proof is in Appendix (ref). Proposition (ref) shows that to choose between $E$ and $O$, assuming $E$ and $O$ have high research costs ($(1 - C_O)^3 \approx 0$), it suffices to compare

equation[equation omitted — 134 chars of source]

up to a small remainder. That is, we must compare the difference in the posterior variance reductions to the difference in research costs, adjusted by the attention cost $c_a$. A larger $c_a$ favors less costly designs. The following theorem formalizes these intuitions and sharpens them to apply even for smaller research costs.

theorem[Comparing two designs] In Setting (ref), suppose that $E$ and $O$ are expensive. \begin{enumerate}[label=(\alph*)] • If $\operatorname{PostVarRed}(E) - \operatorname{PostVarRed}(O) \ge \big(1 - \frac{C_O}{C_E}\big)c_a,$ then $E$ is planner-preferred to $O$. • If $\operatorname{PostVarRed}(E) - \operatorname{PostVarRed}(O) \le \big(C_E - \frac{1 + 2C_O}{3} \big) c_a$, then $O$ is planner-preferred to $E$. \end{enumerate} If instead $O$ is cheap, then (a) and (b) hold with $C_O$ replaced by $P(|X(O)| \ge \gamma_O^\star)$.

The proof is in Appendix (ref). Intuitively the comparison between two studies must depend on the posterior variance reduction of each study (which itself depends their mean-squared error) and the costs of each study.

figure[figure omitted — 952 chars of source]

As the cost of attention $c_a$ increases, the planner's preference shifts from a more accurate design to a less accurate design with a smaller cost. This is because, for costly studies, the planner must internalize not only the effect of the mean-squared error on the audience's loss function, but also the research cost associated with the study. With high attention costs (large $c_a$), more costly experiments impose more stringent constraints on the publication rules, making those undesirable for the planner. For example, in this case, medical studies with smaller experiment size may be preferred over more precise experiments when variable costs are sufficiently large.

Theorem (ref) provides bounds that are not tight to enhance interpretability. However, more involved necessary and sufficient conditions for each design to be planner-preferred can be obtained directly from a result in the Appendix (Lemma (ref) in Appendix (ref)). Using these conditions, Figure (ref) reports the indifference curves between two experiments with different mean-squared errors and costs. Figure (ref) shows how the planner's preference shifts towards cheaper (and noisier) designs even when the experiment has zero variance.

rem[Calibrating parameters] As we illustrate in Section (ref), we can calibrate $\eta^2$ as the prior variance of parameters obtained from meta-studies rhys2024much,bartovs2023empirical. The cost $c_a$ can be calibrated to make the cutoff $\gamma^\star_E$ for cheap experiments match the critical value of a $t$-test for a particular level---for example, $\sqrt{c_a} \frac{\eta^2 + 1}{\eta^2} = 1.96$. \qed
rem[Internalizing research costs in the planner's objective] A variant of the model considered here is to let the planner internalize the research costs. For instance, one could consider a planner's objective $\mathcal{L}_\Delta^\star + C_\Delta$. Our main analysis continues to hold in this setting, by taking into account that the comparisons in Equation (ref) and Theorem (ref) must account for the additional cost component in the planner's loss function. \qed

Publication rules under non-verifiable designs

In this section, we turn to settings where researchers may engage in data manipulation and selective reporting. To this end, we investigate optimal publication rules when researchers choose the research design $\Delta$ using knowledge of the statistics drawn in the experiment. Here, the design $\Delta$, and its corresponding bias $\beta_\Delta$, are known to the researcher but not verifiable by the social planner. Specifically, we consider the following setting.

settingThe class of designs is $\Delta \in \mathbb{R}$, with $S_\Delta^2 = S^2$ common knowledge, and $\beta_\Delta = \Delta$ known to the researcher. Writing $X(\Delta) = \theta + \beta_\Delta + \varepsilon$ (where $\varepsilon | \theta \sim \mathcal{N}(0, S^2)$), the researcher observes $\theta + \varepsilon$ and chooses $\Delta$ to maximizes her realized payoff $v_p(X(\Delta), \Delta)$. Research costs are given by $C_\Delta = c_m |\beta_\Delta| + C_0$, where $0 < c_m < \infty$ and $C_0 < 1$. The social planner chooses a (Borel measurable) publication rule $p(X,\Delta) = p(X)$ as a function of $X$ only.

Figure (ref) illustrates the model: the researcher deterministically chooses the bias of the reported statistic. They, however, pay a cost $C_\Delta$ increasing in the bias. The component $c_m|\beta_\Delta|$ of the cost $C_\Delta$ captures reputational or computational costs associated with the manipulation, assumed to be increasing and linear in the magnitude $|\beta_\Delta|$ of the bias. The component $C_0$ captures a fixed cost.\footnote{It is possible to also incorporate fixed costs of manipulation (e.g., $C(\Delta) = c_m|\beta_\Delta| + c_f 1\{\beta_{\Delta} \not= 0\} + C_0$) to capture costs of introducing any manipulation, or to include nonlinear costs of manipulation, though the specific characterization of the optimal publication rule would then be different.} The researcher observes $\theta + \varepsilon$, and hence maximizes realized utility conditional on the observed statistics when choosing $\Delta$.

We think of the researcher's action of deterministically choosing the bias as a stylized description of data manipulation or selective reporting, whereby researchers can choose their research design after learning the results of potential studies $X$. In practice, in the absence of a precise pre-specification, researchers can change the covariates in a regression, winsorize the data in particular ways, or make other design choices functions of the statistics. These manipulations are all forms of $p$-hacking, and bias results in a way that is difficult or impossible for the planner to verify. \footnote{ For precise experiments (i.e., when $X \approx \theta)$, our model also speaks to a related form of manipulation, which niederle2025experiments terms $g$-hacking. In particular, researchers may choose the specific experimental environment to be one where they expect treatment effects to be non-representatively large using private information about $\theta$---which they may obtain from piloting, theory, or intuition. For example, researchers can choose the experimental site for a field experiment, or an experimental setting in a lab experiment, in which they expect treatment effects to be large. These manipulations introduce bias in results analogous to site selection bias, and are also difficult or impossible for the planner to verify. }

figure[figure omitted — 3,340 chars of source]

As we discuss in Section (ref), the audience updates their beliefs assuming that $\beta_\Delta = 0$. Thus, we assume that the audience is unaware of the possibility of data manipulation---i.e., that they take published findings at face value. For example, in our medical example, the audience may represent doctors or policymakers who are not familiar with (or would need to incur high costs to understand) experimental details. By contrast, the planner is aware of the possibility of manipulation, and that the audience is unaware of it, and minimizes the audience's loss taking both of these points into account.

We assume that the variance of the residual noise $\varepsilon$ equals $S^2$ for all designs $\Delta$. We interpret this assumption as stating that standard errors are verifiable as part of the publication process; hence, we focus on manipulation that introduces unverifiable bias in reported results.

Optimal publication rule under manipulation

The planner knows $S^2$, cannot observe or verify $\beta_\Delta$, and minimizes expected loss over $(\theta,\varepsilon)$ taking into account the researcher's (endogeneous) incentives to manipulate their results. Formally, writing $\mathcal{P}$ for the set of all Borel measurable functions $p(X,\Delta)$ that are constant in $\Delta$ (i.e., do not depend on the design), an optimal publication rule is defined by

equation[equation omitted — 275 chars of source]

Here, $\Delta_p^{\star}$ denotes an optimal response of the researcher to the publication rule given $\theta + \varepsilon$.

The main result of this section shows that the optimal publication rule is in a class of smoothed cutoff rules. Intuitively, a linearly smoothed cutoff rule is a deterministic publication rule below and above thresholds $X^\star - \frac{1}{k}$ and $X^\star$, respectively; it randomizes the publication chances between these two thresholds, with publication probability increasing linearly the value of the reported statistic $|X|$ with slope $k$.

defnA linearly smoothed cutoff rule with cutoff $X^\star$ and slope $k$ is defined by \[p_{X^\star,k}(X) = \begin{cases} 0 & \text{if } |X| \le X^\star - \frac{1}{k}\\ 1-k(X^\star - |X|) & \text{if } X^\star - \frac{1}{k} < |X| < X^\star\\ 1 & \text{if } |X| \ge X^\star \end{cases}.\]

The special case of slope $k = \infty$ and threshold $X^\star = \gamma^\star = \frac{S^2 + \eta^2}{\eta^2} \sqrt{c_a}$ corresponds to a publication rule for cheap experiments without manipulation (Corollary (ref)).

We next characterize the optimal publication rule in settings with manipulation.

theorem[Optimal publication rule under unverifiable designs] In Setting (ref): \begin{enumerate}[label=(\alph*)] • There exists a cutoff $X^\star \in \left(\gamma^\star,\gamma^\star + \frac{1 - C_0}{c_m}\right)$ such that the linearly smoothed cutoff rule $p_{X^\star,c_m}$ is optimal. • For each optimal publication rule $p$, there exists $X^\star \in \left(\gamma^\star,\gamma^\star + \frac{1 - C_0}{c_m}\right)$ such that $p(X) = p_{X^\star,c_m}(X)$ (resp. $p(X) \le C_0$) for almost all $X \ge 0$ with $p_{X^\star,c_m}(X) > C_0$ (resp. $p_{X^\star,c_m}(X) \le C_0$). \end{enumerate}

The proof is in Appendix (ref). As publication probabilities are between 0 and 1, the mechanism design problem is effectively one with limited transfers. As a result, standard techniques to eliminate transfers from the planning problem mirrlees1971exploration and myerson1981optimal do not apply, and we need to deal directly with both publication probabilities and equilibrium manipulation. To solve the mechanism design problem, we identify the precise pattern of binding incentive constraints, which are upward incentive constraints to and from type $\gamma^\star$. Constraining the equilibrium utility level for that type to be $u$, we show that the optimal publication rule is a linearly smooth cutoff rule. We then optimize over $u$ to bound the optimal cutoff $X^\star$. Although the optimal cutoff $X^\star$ does not admit a simple closed-form expression, it can be computed numerically---as we show in Figure (ref).

Interpretation and implications for published findings

To provide intuition for the structure of the optimal publication rule in Theorem (ref), we illustrate how the possibility of manipulation affects the optimal publication rule. For ease of exposition, we abstract from fixed research costs in our discussion (i.e., take $C_0 = 0$). Our formal results all apply to the case of general $C_0$, and all proofs are in Appendix (ref).

Suppose first that the social planner ignored the possibility of manipulation, and set a cutoff rule for a cheap experiment as in Corollary (ref). Then we would observe bunching around the publication cutoff $\gamma^\star$, as researchers with $|\theta + \varepsilon| \in \big(\gamma^\star - \frac{1}{c_m}, \gamma^\star\big)$ would introduce a bias to publish. Researchers with $|\theta + \varepsilon| < \gamma^\star - \frac{1}{c_m}$ would find it unprofitable to introduce any bias (as the cost would not compensate the benefits) and therefore would not publish. The first line of Table (ref) and the first two panels of Figure (ref) summarize this discussion.

table[table omitted — 1,194 chars of source]
figure[figure omitted — 1,085 chars of source]

Next, suppose that the planner introduces randomization in the publication rule whenever $|X| \in \big(\gamma^\star - \frac{1}{c_m}, \gamma^\star\big)$ as in a linearly smoothed cutoff rule with cutoff $\gamma^\star$. This randomization makes the researcher indifferent between manipulating and not manipulating the data, at the cost of publishing some results below $\gamma^\star$, which do not move the audience's action enough to justify incurring the attention cost. The second line of Table (ref) summarizes this discussion.

More generally, the optimal publication rule randomizes publication for some unmanipulated results below $\gamma^\star$. This is in stark contrast with the case without manipulation.

prop[Some results that do not merit attention are published despite not being manipulated] In Setting (ref), consider any optimal publication rule $p^\star$. For some types $|\theta + \varepsilon| < \gamma^\star$, we have $p^\star(X(\Delta^\star_{p^\star})) > C_0$ but $\beta_{\Delta^\star_{p^\star}} = 0$.

However, simply randomizing for results below $\gamma^\star$ is still suboptimal as too many unsurprising results are published in the randomization regime. The last step is to increase the threshold $X^\star$ to lower the loss from publishing results that do not move the audience's action enough. A consequence is that some results that merit attention are not published.

prop[Some results that merit attention are not published] In Setting (ref), consider any optimal publication rule $p^\star$. For some types $\theta + \varepsilon > \gamma^\star$, we have $p^\star(X(\Delta^\star_{p^\star})) < 1$.

Given that results are only guaranteed publication if they cross a higher threshold than $\gamma^\star$, some manipulation can be beneficial to the planner to increase the publication rate of surprising findings. Therefore, in the planner's preferred equilibrium, researcher types below the cutoff $X^\star$ and above $\gamma^\star$ engage in some manipulation. This form of manipulation is distinct from the manipulation that researchers engage in under (non-smoothed) cutoff rules, which involves results that should not be published and therefore hurts the planner.

prop[Manipulation in equilibrium] In Setting (ref), consider any optimal publication rule, and let $X^\star$ be as in Theorem (ref)(b). For almost all $\theta + \varepsilon \in (\gamma^\star, X^\star)$, we have $\beta_{\Delta_{p^\star}^\star} > 0$.

There is no point in manipulating beyond the cutoff $X^\star$, as results $X^\star$ are published with probability 1. As a result, there is bunching at the cutoff $X^\star$ in optimum.

prop[Bunching at $X^\star$ in equilibrium] In Setting (ref), consider any optimal publication rule, and let $X^\star$ be as in Theorem (ref)(b). There exists $\zeta > 0$ such that $\theta + \varepsilon + \beta_{\Delta_{p^\star}^\star} = X^\star$ for almost all $\theta + \varepsilon \in (X^\star - \zeta,X^\star)$.

The third line of Table (ref) and the third panel of Figure (ref) summarize Propositions (ref) and (ref).

remAn unrelated application of our analysis is to the structure of the lottery for top picks in the National Basketball Association (NBA) draft. To help weak teams become more competitive, the NBA seeks to offer a higher chance of top picks to the lowest-ranked teams. This creates incentives for the teams near the bottom to manipulate their quality ($\theta$) by losing matches so to be classified as the worst-team (i.e., lowering $X$). To disincentive this form of manipulation, in 2019, the NBA began to offer the same chances at top picks to the three worst-ranked teams. This form of randomization has a similar structure to the randomization features in our optimal publication rule under manipulation.

Taking stock: Research costs versus manipulability

In this section, we combine the analyses of the previous two sections to study when the planner may want to incentivize a manipulable observational study instead of a more costly, but pre-registered, experiment. More precisely, we consider the following setting.

setting[Pre-specification versus possible manipulation] Consider the following two possible families of designs, all of which have variance $S^2$. \begin{itemize} • Experiment with pre-analysis plan: As in Section (ref), there is an unbiased design $E$. Researchers cannot manipulate their findings, truthfully report $X = \theta + \varepsilon$, and pay a research cost $C_E$. The social planner chooses a publication rule $p_{E}^\star$ as in Definition (ref) and incurs expected loss $\mathcal{L}_E^\star$. • Possible manipulation: We are in Setting (ref). Thus, there is a family of designs $\Delta \in \mathbb{R}$. The manipulation cost is $c_m < \infty$, and researchers can manipulate their findings after observing $\theta + \varepsilon$. The fixed research cost is $C_0 = 0$. The social planner chooses an optimal publication rule $p^\star$, and incurs a corresponding expected loss $\mathcal{L}_M^\star = \mathbb{E}_{\theta,\varepsilon}\left[\mathcal{L}_{p^\star}(X(\Delta_{p^\star}^{\star}), \Delta_{p^\star}^{\star}, \theta)\right]$ under the planner's preferred equilibrium. \end{itemize}

In Scenario (A), researchers cannot manipulate their findings (as for, e.g., a pre-registered experiment), but pay a fixed cost $C_E$ of conducting an experiment. In Scenario (B), researchers can manipulate their findings---as for an observational study. \footnote{While we assume zero research costs in Scenario (B), similar results hold for an manipulable studies that have sufficiently low research costs.}

To compare the planner's loss across the two scenarios, we introduce a quantitative measure of how much loss the expensiveness of the design $E$ in Scenario (A) entails for the planner. We call this measure the incentive cost of $E$, as it captures the cost of incentizing the researcher to choose design $E$. Namely, we consider the difference in the planner's loss (under the planner's optimal publication rules) between the design $E$ and a hypothetical design $E'$ with the same variance but no research costs.

defn[Incentive costs] Given an unbiased design $E$, let $\mathrm{IC}(E) = \mathcal{L}_E^\star - \mathcal{L}_{E'}^\star$ where $E'$ is a cheap design with $C_{E'} = 0$ and variance $S_{E'}^2 = S_E^2$.

Whenever $E$ is a cheap design, we have $\mathrm{IC}(E) =0$, while when $E$ is an expensive design, we have $\mathrm{IC}(E) >0$. Also, note that $\mathrm{IC}(E)$ is non-decreasing in $C_E$ by Lemma (ref) in Appendix (ref), and can be readily computed by using the same lemma.

The next proposition provides a simple characterization of when the experiment in Scenario (A) is preferred by the planner to the observational study in Scenario (B) in terms of the incentive costs of the experiment in Scenario (A). High incentive costs for the experiment favor the observational study, and we provide a quantitative bound on incentive costs for observation studies to be preferred by the planner.

prop[Experiment versus manipulable observational study] In Setting (ref): \begin{itemize} • If $E$ is cheap (i.e., $\mathrm{IC}(E) = 0$), then $\mathcal{L}^\star_E < \mathcal{L}^\star_M$. • If $\mathrm{IC}(E) > \frac{1 + 2 S c_m}{c_m^2}$, then $\mathcal{L}^\star_M < \mathcal{L}^\star_E$. \end{itemize}

The proof is in Appendix (ref). When the cost of the experiment is high (and therefore $\mathrm{IC}(E)$ is large), whether the experiment or the observational study is preferred depends on on (i) the cost $C_E$ of the experiment and (ii) the cost of data manipulation $c_m$ in the observational study. Clearly, if $c_m$ is small, then an experiment with a pre-analysis plan may be preferred. Intuitively, with a low cost of manipulation, the planner must end up publishing a larger set of results that do not merit attention from the audience---thereby making the audience incur a possibly large attention cost.

Now suppose that $c_m$ is large. Then, an observational study with possible manipulation is preferred by the planner for sufficiently high research costs $C_E$ for the experiment---despite the cost of the experiment being private and paid only by the researcher. This preference arises because a sufficiently large experimental research cost $C_E$ increases the loss of the planner, who must publish results that do not merit the audience's attention. This conclusion contrasts with analyses that abstract from the costs of experiments with pre-analysis plans, where (pre-registered) experiments always dominate (manipulable) observational studies kasy2023optimal, spiess2018optimal.

Application to medical studies

In this section, we study the implications of our model for medical studies. We first calibrate our model from Section (ref) to meta-analyses of medical studies. We then illustrate several properties of the optimal publication decision rules, and investigate when studies with potential manipulation may dominate experiments with no manipulation.

Calibration

To map the model in Section (ref) to the data, we consider a family of studies $i$. We normalize the estimates and estimands so that $S^2 = 1$. Thus, we consider test statistics $X_i = \theta_i + \beta_i + \varepsilon_i$ where $\theta_i$ is the parameter of interest of the study $i$ defined as the size of the effect in standard deviation units, $\beta_i$ is a bias due to manipulation (potentially correlated with $\theta_i + \varepsilon_i$) and $\varepsilon_i \sim \mathcal{N}(0,1)$.\footnote{The normality assumption here captures a large-sample approximation.} As in previous sections, we consider a prior $\theta_i \sim \mathcal{N}(0, \eta^2)$; here, the variance captures any additional prior heterogeneity arising from different studies, and the assumption of mean zero is consistent with the empirical distribution of studies in the Cochrane Database of Systematic Reviews, the leading database for systematic reviews in health care.\footnote{\url{https://www.cochranelibrary.com/cdsr/about-cdsr}, see bartovs2023empirical.}

We use data from head2015extent to calibrate our key parameters $(\eta^2, c_a, c_m)$ for manipulable designs. As in Section (ref), we impose $C_0 = 0$ in our model---i.e., that fixed costs are sunk at the stage at which researchers decide about whether/how to manipulate.

head2015extent collected $p$-values via text-mining the PubMed database across several disciplines. We focus here on $827,115$ $p$-values published in medical and pharmaceutical journals. As head2015extent's head2015extent data set does not contain $t$-statistics, we construct $t$-statistics by inverting a two-sided $t$-test. That is, we construct $t$-statistics as $X_i:=\Phi^{-1}(1 - p_i/2)$, where $\Phi$ is the Gaussian CDF and $p_i$ is the $p$-value of study $i$.\footnote{Because the data does not contain information about $t$-tests directly, this assumes that the majority of $p$-values are constructed using two-sided $t$-tests. Although this may not always be the case, we note that $t$-tests are the standard practice in medical sciences fda. It is possible to conduct our analysis assuming one-sided instead of two-sided tests.} To adjust for publication bias in the data from head2015extent, we use estimates from vorland2024publication that $36\%$ of the medical studies are never published. As a simplifying assumption, we assume that all such studies have non-significant results ($p$-value above $5\%$).

\paragraph{Calibration of $c_m$.} For competitive journals, standard $1$ or $5\%$ significance levels correspond to publication rules of the form $p^\star(X) = 1\{|X| \ge q\}$ for $q \in \{1.96,2.56\}$ where $q$ is the critical Gaussian quantile at a given $5\%$ or $1\%$ significance level.\footnote{Although some journals publish non-significant results, researchers may have incentives to manipulate their results to publish in top journals, where most published results are significant laviolle2025trends.} We consider two calibrations that correspond to each of these two levels. In each case, we would observe bunching at $q$, with researchers manipulating the findings in the interval $(q - \frac{1}{c_m}, q)$. Taking a standard bunching approach, we can therefore use the data to estimate $c_m$ by solving $\Phi(q) - \Phi(q - \frac{1}{c_m}) = b$, where $b$ is the share of findings in a small neighborhood around $q$. We find that about $18\%$ of findings have a $p$-value close to $0.05$ and $10\%$ of $p$-values close to $1\%$ in the set of studies reported by head2015extent.\footnote{We choose $b$ as the share of $|X|$ between $1.95$ and $2$ for $q = 1.96$ and between $2.55$ and $2.6$ for $q = 2.56$. } About $27\%$ of trials published in PubMed between 2002 and 2015 have some form of pre-registration lamberink2022clinical. Assuming no manipulation of pre-registered findings, we estimate $b$ by multiplying these shares by $(1 - 0.36)/(1 - 0.27)$ to account for the $36\%$ of results that are not published, and $27\%$ of studies with pre-registration in PubMed. We find $c_m = 0.98$ for our 5%-level calibration (which considers manipulation of a $t$-test of level $5\%$), and $c_m = 0.83$ for our 1%-level calibration (which considers manipulation of a $t$-test of level $1\%$).

\paragraph{Calibration of $\eta^2$.} Manipulation may change the distribution of $X$, so we cannot directly estimate $\eta^2$ from second moments. However, for a publication rule of the form $1\{|X| \ge q\}$, under our model (from Setting (ref) in Section (ref)), manipulation only occurs below the significance threshold. To construct a measure robust to manipulation, we take the $95$th percentile of the distribution of $X$, denoted as $\bar{q}_{95}$. Taking into account that an additional $36\%$ of findings is unobserved (unpublished) but non-significant by assumption, we can see that $\bar{q}_{95} = \hat{q}_{95 - \gamma}$ where $\gamma = 0.05 \times 0.36/(1 - 0.36)$ and $\hat{q}$ is the empirical quantile of $t$-statistics from head2015extent. If $\hat{q}_{95 - \gamma} > 2.56$, our estimate will be robust to manipulation around the significance threshold under our model. We find that $\hat{q}_{95 - \gamma} = 3.43$---well above the critical values around which we expect manipulation. Using properties of the Gaussian distribution, $\eta^2 \approx (\bar{q}_{95}/2)^2 - 1$, which leads to an estimate of $\eta^2 = 1.94$.

\paragraph{Calibration of $c_a$.} We choose $c_a$ to reflect the critical values of tests that are typically used in practice. As (absent manipulation) the critical threshold for publication is $\sqrt{c_a} \frac{1 + \eta^2}{\eta^2}$ in our model (Proposition (ref)), we consider take $\sqrt{c_a} \frac{1 + \eta^2}{\eta^2} = 1.96$ for our 5%-level calibration (which captures a 5%-level $t$-test), and $\sqrt{c_a} \frac{1 + \eta^2}{\eta^2} = 2.56$ for our 1%-level calibration.

Optimal publication rules under \texorpdfstring{$p$}{p}-hacking

figure[figure omitted — 777 chars of source]

We next illustrate the impact of manipulation on the distribution of results, both in the data, and simulated in our model under the optimal publication rule. In Figure (ref), we plot the distribution that we would see in the absence of manipulation in our calibration (left panel), the empirical distribution (center panel), and the distribution under the optimal publication rule under our 5%-level calibration (right panel). The empirical distribution exhibits bunching around $1.96$ and $2.56$, consistent with a much higher chance of publication above these significance thresholds.\footnote{In Figure (ref) in Appendix (ref), we report the raw distribution of $p$-values. Given rounding to two or three digit levels for $p$-values in the data, some of the bunching should be interpreted around (not exactly equal) to the critical thresholds elliott2022detecting.} The simulated distribution under the optimal publication rule features much less bunching.

table[table omitted — 2,067 chars of source]

In Table (ref), we compare the simulated properties of the standard cutoff rule and the optimal rule in our calibrated model in more detail. We calculate the cutoff $X^\star$ under the optimal publication rule, and compare the share of published findings and share of manipulated published findings between the optimal rule and the standard rules $t$-test rules $1\{|X| \ge t^*\}$ for $t^* \in \{1.96, 2.56\}$, both with and without manipulation.

The first main takeaway is that the optimal publication rule substantially increases the threshold from $1.96$ to $2.64$ under the 5%-level calibration (and from $2.56$ to $3.39$ under the 1%-level calibration) due to the relatively small cost of manipulation. The new threshold $X^\star$ characterizes the point after which studies are published with probability 1.

The second key observation is that under the optimal publication rule, the amount of manipulation within such studies is substantially lower---but non-zero. As Table (ref) shows, under the standard 5%-level $t$-test rule $1\{|X| \ge 1.96\}$, about $54\%$ of the studies are published, and of these published findings the average bias is 0.31. In contrast, if there were no manipulation, only $25\%$ of the findings would be published under the same threshold rule (this is because $X_i \sim \mathcal{N}(0, 2.94)$ unconditional on $\theta_i$). The optimal publication rule publishes a similar number of findings as in the absence of manipulation ($25\%$), and the average bias across all such published findings is $0.11$---less than half than under a standard $t$-test rule.

Taken together, these results show that publication rules in leading medical journals should publish a similar number of findings if there was no manipulation under the standard 5%-level $t$-test rule $1\{|X| \ge 1.96\}$. However, the critical threshold for which studies are published with probability 1 increases from 1.96 to $2.64$. However, some results between 1.64 and 2.64 should still be published, albeit with probability less than 1. In particular, some results between 1.62 and 1.96 should be published to reduce incentives to substantially manipulate. Thus, publication decisions must depend on results through the “publication score” $$

alignedp^\star(X) = \begin{cases} 0 & if X < X^\star - \frac{1}{c_m} \\ 1 - c_m (X^\star - |X|) & if X^\star - \frac{1}{c_m} \le |X| \le X^\star \\ 1 & if X \ge X^\star \end{cases}

$$ where $(c_m, X^\star) = (0.98, 2.64)$, which indicates the relevance of findings for publication. For the 1\%-level case, one should instead take $(c_m, X^\star) = (0.83, 3.39)$ in this scoring rule.

Pre-registered experiment versus manipulable observational study

Returning to the exercise in Section (ref), we ask when the planner should incentivize an experiment with costs $C_E > 0$ and no manipulation (e.g., adherence to a pre-analysis plan) over a manipulable observational study with lower research cost $C_0 = 0$. \footnote{ The assumption that the observational study has no fixed research cost can be relaxed as long as its fixed costs do not affect the supply of research (individual rationality constraint is not binding).} We suppose that the two studies share the same estimand, and have equal precision absent bias from manipulation.

In Figure (ref) (left panel), we report the ratio between the planner's optimal loss for the experiment and the loss for the observational study, as a function of the cost of the experiment $C_E$. We consider two publication rules for the observational study: the standard $t$-test rule that is optimal without manipulation frankel2022findings, and the optimal publication rule accounting for manipulation (from Theorem (ref)).

figure[figure omitted — 705 chars of source]

Under a standard $t$-test rule for the observational study, due to manipulability the observational study is dominated by the experiment, no matter the experimental cost. When we consider instead the optimal publication rule for the same observational study, the observational study performs much better. Experiments with costs $C_E > 0.3$ (i.e., experiments whose research costs are above 30% of the value of a publication) become dominated by observational studies. This comparison reflects Proposition (ref), as high-cost experiments have high incentive costs. Quantitatively, the underperformance of the observational study relative to the experiment is close to zero even for lower experimental research costs ($C_E < 0.3$): the expected loss from the experiment is more than $99\%$ of that of the observational study.

To understand the magnitude of $C_E$ in different pharmaceutical industries, the right panel of Figure (ref) reports the ratio of total development cost of the drug to the expected costs (adjusting for cost of failures) using data from sertkaya2024costs. We interpret this as a proxy for $C_E$ for a researcher deciding whether to investigate the efficacy of a drug in a context with perfect competition (zero expected profits), which may (conservatively) capture private interests from drug companies when publication entails marketing and increased credibility modi202310.

We find substantial heterogeneity in whether an experiment or an observational study should be preferred under an optimal publication rule for the observational study. For most industries, $C_E$ is around $0.3$ or only slightly larger, suggesting that an experiment may be preferred. Three industries are however significantly above, with $C_E = 0.4$. These results suggests that there may be welfare gains from allowing observational studies in industries in which experiments are particularly costly to implement--- provided that publishers use different “significance rules” for observational studies that account for their manipulability.

Conclusion

This paper studies how researcher's incentives shape the optimal design of the scientific process. Ignoring the researcher's incentives, it is optimal to publish the most surprising results frankel2022findings. When researcher's incentives matter, we show that optimal publication rules depend on private costs of research and incentives for research manipulation.

In the absence of manipulation, we show that the planner prefers studies with larger mean-squared errors over sufficiently costly experiments. With manipulation, we show that it is optimal to (i) publish some unsurprising results and (ii) knowingly allow for manipulation at the margin. Observationally, the optimal policy would reduce the bunching of the findings around the publication cutoff. However, the optimal policy does not completely remove bunching. Even when the planner can eliminate manipulation by publishing only pre-specified experiments, this may not be the preferred policy when experiments entails large research costs. In our application to medical studies, we highlight the importance of setting different publication rules for experimental and observational studies.

Our results have implications for the policy debate regarding the use of synthetical control groups in clinical trials. Allowing synthetic control groups opens the door to manipulation of the specification of the control group, similar to the manipulability of observational studies. However, studies synthetic control groups are cheaper to implement as they can achieve similar power in a smaller experiment. Our analysis suggests that when experiments are costly, allowing for synthetic controls may also be desirable--- provided that publishers and approvers use different significance rules for them that account for their manipulability.

Future research should study more complex decisions by the planner and the researcher. For example, in contexts with pre-analysis plans, the planner may allow for the publication of non-prespecified findings. Finally, as our contribution lies at the intersection of econometrics and mechanism design, future research should study how other aspects of researchers' incentives impact the design of scientific communication.

\singlespacing