Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,037 characters · 14 sections · 36 citation commands
Optimal Pre-Analysis Plans: Statistical Decisions Subject to Implementability
\onehalfspacing
Keywords: Pre-analysis plans, Statistical decisions, Implementability\\ JEL codes: C18, D8, I23
\bgroup \renewcommand\fnsymbol{footnote}{\fnsymbol{footnote}} \renewcommand\fnsymbol{mpfootnote}{\fnsymbol{mpfootnote}} \footnotetext[0]{ We thank Stefano DellaVigna, Ted Miguel, Marco Ottaviani, and Davide Viviano, as well as Alex Frankel, Carlos Gonzalez Perez, Rohit Lamba, Ludvig Sinander, Alex Teytelboym, and participants at the BITSS 2022 meeting, the 2022 AEA meetings, and the 2022 conference in Honor of Jim Powell for helpful discussions and suggestions. \\ Maximilian Kasy was supported by the Alfred P. Sloan Foundation, under the grant “Social foundations for statistics and machine learning.” } \egroup
When writing up their studies, empirical researchers might cherry-pick the findings that they report. Cherry-picking distorts the inferences that we can draw from published findings. As a potential solution, pre-analysis plans (PAPs) have become a precondition for the publication of experimental research in economics, for both field experiments and lab experiments. \footnote{Just as in the case of randomized experiments, the adoption of PAPs in economics follows their prior adoption in clinical research; see for instance the guidelines of the \citetalias{fda1998} on PAPs, fda1998.} PAPs can enable valid inference by pre-specifying a mapping from the data to testing decisions or estimates, cf. christensen2018transparency,miguel2021evidence. This can prevent the cherry-picking of results, and thus provide a remedy for the distortions introduced by unacknowledged multiple hypothesis testing. The widespread adoption of PAPs has not gone uncontested, however, \footnote{See for instance CoffmanNiederle2015, Olken2015, and duflo2020praise, who discuss the costs and benefits of PAPs in experimental economics from a practitioners' perspective.} and has been criticized for constraining our ability to learn from experiments.
In this article, we clarify the benefits and optimal design of pre-analysis plans by modeling statistical inference as a mechanism-design problem myerson1986multistage,kamenica2019bayesian. To motivate this approach, note that, in single-agent statistical decision theory, rational decision-makers with preferences that are consistent over time do not need the commitment device that is provided by a PAP. This holds in particular when a single decision-maker aims to construct tests that control size, or estimators that are unbiased. Single decision-makers have no reason to “cheat themselves.” The situation is different, however, when there are multiple agents with conflicting interests. When there are multiple agents, not all statistical decision rules might be implementable. Furthermore, allowing for messages (PAPs) before the data are seen can increase the set of implementable rules, and thus improve welfare.\footnote{A separate argument for pre-analysis plans, which we do not pursue in this paper, might be based on dynamic inconsistencies in agent preferences, for instance because of present-bias.}
Our framework provides a theoretical justification of PAPs. In addition to our theoretical results, which are based on this framework, we also derive guidance for practitioners, including both decision-makers (e.g., readers, editors) and data analysts (e.g., study authors). From the decision-makers' perspective, we describe how tests, estimators, or other decision rules can be implemented by requiring pre-analysis plans. We then focus on hypothesis tests, and describe how to derive optimal pre-analysis plans from the analysts' perspective. These pre-analysis plans maximize power while controlling size and maintaining implementability. We furthermore provide software (an interactive web app) to facilitate the design of optimal pre-analysis plans.
\paragraph{Examples}
In our model, we consider the interaction between a decision-maker and an analyst. The analyst has private information and interests which differ from those of the decision-maker. One example of such a conflict of interest is between a researcher (analyst) who wants to reject a hypothesis, and a reader of their research (decision-maker) who wants a valid statistical test of that same hypothesis; the relevant decision here is whether to reject the null hypothesis. Another example is the conflict of interest between a researcher (analyst) who wants to get published, and a journal editor (decision-maker) who only wants to publish studies on effects that are large enough to be interesting; the relevant decision here is whether to publish a study. A third example is the conflict of interest between a pharmaceutical company (analyst) who wants to sell drugs, and a medical regulatory agency (decision-maker) who wants to protect patient health; the relevant decision here is whether to approve a drug.
\paragraph{Model and timeline}
The timeline of our model is as follows. Before observing the data, the analyst can send a message to the decision-maker. This message might for instance be in the form of a pre-analysis plan. Then the analyst observes the data. The data are given in the form of a set of statistics, such as the outcomes of different hypothesis tests, or estimates for different model specifications. The analyst chooses a subset of these statistics to report to the decision-maker.
The decision-maker observes the pre-analysis message and the statistics which the analyst reported, and makes a decision based on this information. We assume that this decision is real-valued, and that the analyst always prefers a higher value for this decision. We consider different objectives for the decision-maker, including statistical testing subject to size control.
In our model, the analyst can hide information from the decision-maker, by not reporting some statistics, but they cannot lie about the data that they report. The potential value of a pre-analysis message in this model comes from the fact that it allows the analyst to share private information (i.e., expertise) with the decision-maker. Sharing such information truthfully would not be incentive-compatible if the message could only be sent after seeing the data. The analyst might have private information regarding the availability of statistics, and regarding the state of the world.
To make it possible for the analyst to hide information, they need to have plausible deniability: The decision-maker does not know what statistics the analyst got to see. Experiments might not have been run, or data might not have been collected, for instance. The analyst might also have prior uncertainty over the availability of statistics, but this is not necessary for our conclusions.
The mechanism-design approach which motivates our model takes the perspective of a decision-maker who wants to implement a statistical decision rule. Not all rules are implementable, however, when the analyst has divergent interests and private information. This mechanism-design perspective allows us to stay close to standard statistical theory, while taking into account the implementability constraints that are a consequence of the social nature of research.
\paragraph{Implementable decision rules}
For this model, we first characterize the set of implementable statistical decision rules. This set is independent of decision-maker preferences. We show that implementable decision rules are such that reporting more results can never make the analyst worse off, given the pre-analysis message, and given the realization of the data. Formally, implementable decision rules need to be monotonic in the reported set of statistics, in terms of set inclusion.
Implementable decision rules furthermore need to be compatible with truthful revelation of analyst private information prior to observing any data myerson1986multistage. This condition is equivalent to the conditions satisfied by proper scoring rules savage1971elicitation, gneiting2007strictly.
Pre-analysis messages allow the decision-maker to implement a larger set of decision rules than would be available without such messages. Implementable rules can be implemented using different mechanisms, based on such pre-analysis messages. One possible implementation allows the analyst to choose from a restricted set of decision rules before seeing the data. Each of these rules needs to be monotonic in the set of reported statistics. This implementation corresponds to the actual practice of pre-analysis plans, where the analyst chooses a decision rule before the data becomes available.
The set of implementable rules can be characterized as a convex polytope. If the decision-maker's objective is convex, and in particular if it is linear, then the optimal implementable rule is necessarily an extremal point of this polytope vanderbei2020linear.
\paragraph{Optimal implementable hypothesis tests}
We next turn to the specific problem of finding optimal implementable hypothesis tests. Such tests are required to satisfy size control conditional on the state of the world and conditional on analyst private information that is available before observing the data. We show that the optimal implementable test, for the decision-maker, can be implemented by (i) requiring the analyst to choose an arbitrary full-data test, which is a function of all statistics that the analyst might observe, where this test controls size, and then (ii) implementing this test, making worst-case assumptions about any unreported statistics.
The analyst's problem of finding a full-data test that maximizes expected power for this mechanism can be cast as a linear programming problem. If the analyst knows the set of available statistics at the time of writing their pre-analysis plan, this problem reduces to the classic problem of finding a test (based on the full set of available statistics) with high expected power, subject to size control. The solution to this problem takes the form of a likelihood ratio test. More generally, the set of available statistics might not be known for sure at the time of writing the PAP. We provide an interactive app that allows the analyst to solve the linear programming problem for this case, based on their prior beliefs. The output of our app can serve as a basis for their pre-analysis plan.
\paragraph{Roadmap}
The rest of this article is structured as follows. We conclude this introduction with a review of some related literature. In (ref), we present a motivating example concerning statistical testing and p-hacking. In (ref), we introduce the general model. In (ref), we characterize implementable decision rules. In (ref), we characterize optimal implementable hypothesis tests. In (ref), we illustrate our results by applying them to the setting of dellavigna2018motivates, using expert forecasts to construct a prior distribution. In (ref), we summarize and discuss some limitations of our model. (ref) contains all proofs.
Our article speaks to the current debates around pre-registration -- and other possible reforms -- in empirical economics and other social- and life-sciences; cf. christensen2018transparency, miguel2021evidence, which are motivated by the distortions to statistical inference that might be induced by selective reporting, cf. publicationbias2019, andrews2019inference. In doing so, our article applies some of the insights from mechanism design and information design myerson1986multistage,kamenica2019bayesian,mechanismdesignnotes2023 to the settings of statistical decision theory and statistical testing, wald1950statistical,savage1951theory,lehmann2006testing.
More broadly, our article contributes to a literature that spans statistics, econometrics and economic theory, and which models statistical inference in multi-agent settings. We differ from other contributions to this literature, in that we focus on the role of implementability as a constraint on statistical decision rules, which rationalizes pre-analysis plans, and on the derivation of optimal decision rules subject to the constraint of implementability.
Drawing on classic references Tullock1959-be, Sterling1959-nu, Leamer1974-zl, Glaeser2006-yp considers the role of incentives in empirical research. A number of recent contributions model estimation and testing within multiple-agent settings, including glazer2004optimal, mathis2008full, Chassang2012-is,Tetenov2016-pw,Ottaviani2017-qq, Di_Tillio2017-ur,Spiess2018-nx,Henry2019-pt,McCloskey2020-pu,Libgober2020-ts,Yoder2020-fe,Williams2021-hg,Abrams2021-hz,Viviano2021-wt. In this literature, Banerjee2020-ql,Frankel_undated-xn, Andrews2020-cp, gao2022inference consider the communication of scientific results to an audience with priors, information, or objectives that might differ from the sender's.
The literature on Bayesian persuasion kamenica2011bayesian,kamenica2019bayesian, curello2022comparative, like the present article, considers a sender with information unavailable to a receiver, where sender and receiver have divergent objectives. One important way in which our model differs from that of Bayesian persuasion is that in our model the signal space of the analyst is restricted to the truthful but selective reporting of data. This restriction implies that the concavification argument central to Bayesian persuasion does not apply.
Before we introduce our general model, consider the following hypothesis-testing problem, as a motivating example and special case. The full data consists of two normally distributed statistics, $X = (X_1,X_2)$, with $X_i \sim \mathcal{N}(\theta,1)$, independently across components of the vector $X$. The $X_i$ might for instance correspond to experimental estimates of an average treatment effect, for two different experimental sites. There is a decision-maker and an analyst. The decision-maker wants to test the null hypothesis $H_0: \theta \leq 0$. The analyst, however, aims to simply maximize the probability of rejection.
The analyst might not always observe both statistics $X_1,X_2$. They instead observe the subvector $X_J$ for a random index set $J$. The possible values of the index set $J$ are $\emptyset$, $\{1\}$, $\{2\}$, and $\{1,2\}$. The statistic $X_i$, for $i \in \{1,2\}$, is observed with probability $P(i \in J)$. Observability is independent across statistics. $P(i \in J)$ is the decision-maker's a-priori probability that the analyst successfully implemented an experiment at site $i$.
The decision-maker does not know which statistics are actually available, that is, they do not know $J$. The analyst knows which statistics are available. This allows the analyst to selectively report (“p-hack”), with plausible deniability, since they might not have observed some unreported statistic. Upon learning the data $X_J$, the analyst chooses a subset $I\subseteq J$, and reports $(X_I,I)$ to the decision-maker. The decision-maker then rejects the null with probability $\mathbf{a}(X_I,I) \in [0,1]$. How should the decision-maker choose the testing rule $\mathbf{a}$ that maps the reported data to a rejection probability?
\paragraph{Five testing rules} We compare five different testing rules, $\mathbf{a}_1$ through $\mathbf{a}_5$. For each of these testing rules, (ref) shows the rejection probability as a function of $(X_1,X_2)$, assuming that $P(1 \in J) = 0.9$ and $P(2 \in J) = 0.5$. The rejection probability in (ref) conditions on $X$, but averages over the distribution of $J$, and takes into account the analyst's endogenous response to a given testing rule. The left panel of (ref) shows the corresponding power curves, i.e., the rejection probability as a function of $\theta$, averaging over the distribution of both $X$ and $J$.
Our benchmark is the optimal test using all the data. This test is not, in general, feasible, since not all statistics are always available. We have that $Z = \tfrac{1}{\sqrt{2}}(X_1 + X_2) \sim \mathcal{N}({\sqrt{2}} \cdot \theta,1)$ is a sufficient statistic for $\theta$. Since this statistic satisfies the monotone likelihood ratio property, the Neyman--Pearson Lemma implies that the uniformly most powerful test of level $\alpha$ is given by $\mathbf{a}_1(X) = \boldsymbol 1 (Z > z),$ where $z = \Phi^{-1}(1-\alpha)$; cf. Theorem 3.4.1 in lehmann2006testing.
Consider next the naive test which ignores potentially selective reporting by the analyst. This test acts as if the reported statistics $I$ are the full data available to the analyst, and implements the corresponding uniformly most powerful test, $$ \mathbf{a}_2(X_I, I) = \boldsymbol 1\left(\frac{1}{\sqrt{|I|}}\sum_{i \in I} X_i > z \right). $$ The best response of the analyst to this naive testing rule involves selective reporting (“p-hacking”), where $I^* \in \operatorname*{argmax\;}_{I\subseteq J} \mathbf{a}(X_I,I)$. The problem with the naive test is that it does not control size. Selective reporting by the analyst implies that the probability of rejection under the null is not bounded by $\alpha$.
We might correct for such selective reporting by making worst-case assumptions about all unreported statistics. This results in the conservative test, $$ \mathbf{a}_3(X_I, I) = \boldsymbol 1\left(\frac{1}{\sqrt{2}}(X_1 + X_2) > z \text{ and } I = \{1,2\}\right). $$ If there are statistics that are not reported, then the null is not rejected. This conservative test implies a probability of rejection given $X$ of $ P(J = \{1,2\}) \cdot \boldsymbol 1\left(\frac{1}{\sqrt{2}}(X_1 + X_2) > z\right).$ The conservative test controls size, but does not have good power properties.
As we show more generally in (ref) and (ref) below, the optimal test without a pre-analysis plan can be implemented by selecting a full-data test of level $\alpha$. When not all data are reported, the decision-maker needs to assume the worst about the unreported statistics, and then implements the corresponding full-data test. The decision-maker can choose the full-data test to maximize (ex-ante) expected power, averaging over their prior for $\theta$.
One possible full-data test ignores $X_2$, which is less likely to be observed in our numerical example, and rejects based on $X_1$ alone. This results in the test $$ \mathbf{a}_4(X_I, I) = \boldsymbol 1\left(X_1 > z \text{ and } 1 \in I\right). $$ This test implies a probability of rejection given $X$ of $ P(1 \in J) \cdot \boldsymbol 1\left( X_1 > z\right).$ This test is optimal for some parameter values, while in general, the optimal test depends on the decision-maker's prior.\footnote{ For the given prior over $J$, this test is for instance optimal when expected power is calculated using the degenerate prior $P(\theta = .3)=1$. More generally, whether this rule is optimal depends on the prior for both $\theta$ and $J$. } We lastly get to the optimal test with a PAP. The optimal test with a PAP is of the same form as the optimal test without a PAP, except that the analyst gets to choose the full data test, prior to seeing any data. Recall that in our example in this section the analyst knows the statistics $J$ that are available before possibly reporting a PAP, but we assume that they have no private information regarding $\theta$ or $X$. (We relax these assumptions in our general setup below.) The optimal implementable solution can be implemented as follows: The analyst communicates which statistics are available by sending the pre-analysis message $M = J$, and the test is given by $$ \mathbf{a}_5(M, X_I, I) = \boldsymbol 1\left(\frac{1}{\sqrt{|M|}} \cdot \sum_{i \in M} X_i > z \text{ and } M\subseteq I\right). $$ That is, the analyst commits to reporting all statistics in $J$, and for that set of statistics, the most powerful test is implemented.
\paragraph{Comparing size and power}
The left panel of (ref) plots the power curves for the five testing rules, for $n= \dim(X) = 2$, which is the case that we have considered thus far. The right panel shows analogous plots for $n = 10$, where the probability $P(i \in J)$ of observing each of the statistics $X_i$ is evenly distributed over a grid from $.5$ to $.9$. The latter case illustrates the differences between testing rules more starkly.
A number of observations are worth emphasizing here. First, the naive test does not control size. For $n=10$, the probability of rejection for $\theta = 0$ is close to $.5$, instead of the nominal size of $.05$. This is due to selective reporting (“p-hacking”). Second, the conservative test can be very conservative. Since it only rejects when all statistics of $X$ are reported, the probability of rejection under the alternative can be arbitrarily small, and remains below the nominal size of $.05$ for our example with $n=10$. Third, the optimal test without a PAP does considerably better than either of these rules. It controls size, and is in fact strictly conservative under the null. At the same time, it has non-trivial power, which greatly exceeds that of the conservative test. This test without a PAP remains itself far from optimal, however. The optimal test with a PAP, lastly, controls size exactly, under the null. Furthermore, its power under the alternative considerably exceeds that of the optimal test without a PAP.
\paragraph{From our example to the general model} Our motivating example is a special case of the general model that we lay out in (ref). The general model allows for cases where the researcher also has private information about $\theta$, and where the researcher only has partial information about availability $J$ of the data. The general model also covers decision problems other than testing, including estimation and treatment choice.
We next describe our general setup, which will be discussed for the rest of this paper. Our setup consists of a game between a decision-maker and an analyst. This game is summarized in (ref).\footnote{Our notation does not distinguish explicitly between random variables and their realizations. This should not cause any ambiguity. Where the distinction is important, we point this out explicitly.} The corresponding timeline is shown in (ref). Throughout, $X$ is a collection of statistics $X_i$, where $i \in \{1,\ldots,n\}$. $I$ and $J$ are (random) index sets, $I,J \subset \{1,\ldots,n\}$, and $X_I = (X_i)_{i\in I}$ denotes the subset of statistics corresponding to the index set $I$.
\paragraph{Discussion} This is a game of partial verifiability. The report $X_I$ is always truthful given $I$, but the non-availability of the statistics corresponding to $\{1,\ldots,k\} \setminus J$ cannot be verified by the decision-maker. Selective reporting, where not all available statistics are reported ($I \varsubsetneq J$), corresponds to p-hacking, or specification searching. Mis-reporting of $X_I$, which corresponds to scientific fraud, is not allowed in our setting.
The private signal $\pi$ corresponds to analyst expertise. The signal $\pi$ might be informative about $\theta$, corresponding to knowledge about which hypotheses are likely to be correct, about the likely magnitude of effect sizes, etc. The signal $\pi$ might also be informative about $J$, corresponding to knowledge about the viability of different identification approaches, the availability of experimental sites, etc.
There is prior uncertainty of the decision-maker regarding the availability $J$ of statistics $X_i$. Without such uncertainty, the mechanism design problem would be trivial, and the decision-maker could simply require the analyst to report everything, by threatening to take action $\min \mathcal A$ otherwise. Prior uncertainty allows for “plausible deniability,” because the decision-maker does not know the full set of results from which the reported results were selected.
In (ref), we have left the message space $\mathcal{M}$ for the pre-analysis message $M$ unrestricted. We will later encounter different, equivalent choices for $\mathcal{M}$: The message $M$ might directly communicate the analyst signal $\pi$, or their corresponding posterior, in the spirit of the revelation principle in mechanism design. Alternatively, and more realistically, the message $M$ might choose a decision function $\mathbf{a}$ from a restricted set, in the spirit of “aligned delegation” frankel2014aligned. This latter formulation corresponds more directly to the practice of pre-analysis plans.
\paragraph{Objectives} We have not yet described the objectives of either the decision-maker or the analyst; (ref) remains silent on these. We allow for conflicting objectives, which render the mechanism-design problem non-trivial. By contrast, we have already imposed common priors, so that there are no agency issues driven by divergent beliefs.
We leave the decision-maker's objective unspecified at this point. This allows us to first study implementability, as a general constraint on the set of decision-functions available to the decision-maker. This constraint does not depend on the decision-maker's objective. We also do not impose that the decision-maker is an expected utility maximizer. This allows us to study frequentist statistical decision problems subject to the constraint of implementability, including hypothesis testing and unbiased estimation, in addition to Bayesian decision problems.
By contrast, we do assume that the analyst is an expected utility maximizer. We furthermore impose the following restriction on their utility function for most of our discussion.
The analyst always prefers a higher outcome $A \in \mathcal{A}$. In the context of testing, the analyst always prefers to reject the null hypothesis. In the context of publication decisions, the analyst always would like their paper to be published. In the context of drug approval, the pharmaceutical company always would like their drug to be approved.
Conventional statistical decision theory considers decision functions that map the available information into statistical decisions wald1950statistical,savage1951theory. In our context, such decision functions $\bar{\mathbf{a}}(\pi,X_J,J)$ map the signal $\pi$, the available data $X_J$, and the set $J$ of available statistics into decisions $A$. We will call such functions $\bar{\mathbf{a}}$ reduced-form decision functions.
In our setting, not all such decision functions are available to the decision-maker, because of analyst private information and conflicting objectives. In this section, we will characterize the set of implementable reduced form decision functions $\bar{\mathbf{a}}$ which are consistent with analyst utility maximization. This leads to constrained versions of conventional statistical decision problems, including hypothesis testing and point estimation. We will show that implementation, in general, requires the use of pre-analysis messages.
The analyst's optimal message $M^*$ and reported set $I^*$ maximize analyst expected utility $\operatorname{E}[v(\mathbf{a}(M,X_I,I))]$, given the decision rule $\mathbf{a}$. Here $M^*$ and $I^*$ are random elements, where $M^*$ is measurable with respect to $\pi$, and $I^*$ is measurable with respect to $\pi,X_J,J$. Analyst expected utility maximization and strict monotonicity of $v$ imply
Consider now reduced-form decision functions $\bar{\mathbf{a}}(\pi,X_J,J)$ that map the information available to the analyst to a decision-maker action. We say that a function $\bar{\mathbf{a}}$ is implementable if it is consistent with analyst utility maximization.
The following theorem provides a complete characterization of implementable reduced-form decision rules in our setting. The proof of this theorem, and all subsequent proofs, can be found in (ref).\footnote{The function $\tilde \mathbf{a}$ is introduced in the following theorem as a technical device to deal with points $(\pi,X_J,J)$ outside the prior support.\\ It is worth noting that the revelation principle myerson1986multistage does not directly apply to our setting, since misreporting of analyst “types” is constrained by the verifiability of their reports $(X_I,I)$, and by $I \subseteq J$. See kephart2016revelation for a discussion of the revelation principle under partial verifiability and, more generally, for settings where misreporting is potentially costly.}
(ref) characterizes which reduced-form decision functions $\bar{\mathbf{a}}(\pi,X_J,J)$ can be implemented, but it does not tell us how to implement them. The following (ref) shows two different, canonical ways of implementing any such function. The first implementation uses truthful revelation of analyst signals. The second implementation uses delegation, where the analyst is allowed to choose the decision function from a pre-specified, restricted set $\mathcal{B}$. This second implementation corresponds closely to the actual practice of pre-analysis plans. In this implementation, the analyst pre-specifies a mapping $b$ from the reported data $(X_J,J)$ to the decision $A = b(X_J,J)$. (ref) shows that restricting attention to implementation by such pre-analysis plans is without loss of generality.
Having characterized implementable decision functions in general, we next discuss implementability for the special case of linear analyst utility $v$ and convex action space $\mathcal A$. We then discuss the connection of truthful revelation to proper scoring. We also consider variants of the model where decision-functions are constrained to be in some class of suitably simple functions.
\paragraph{The set of implementable rules as a convex polytope} In addition to Assumptions (ref) and (ref), assume for a moment that the action space $\mathcal{A} \subseteq \mathbb{R}$ is convex, and that analyst utility is linear -- without additional loss of generality, $v(A) = A$. The leading examples involve binary decisions, where we interpret $A$ as the probability of a positive decision. Binary decisions occur for statistical testing, as discussed in (ref) below, as well as for publication decisions, drug approval, etc. Linearity is without loss of generality for the case of binary decisions; in this case, it follows from expected utility maximization. Suppose finally that $\pi$ has finite support.
Under these additional assumptions, we get that every implementable decision functions $\bar{\mathbf{a}}$ is almost surely identical to a function $\tilde{\mathbf{a}}$ in the convex polytope characterized by the following constraints: \bals \tilde{\mathbf{a}}(\pi,X_J,J) &\in \mathcal{A},&&&(Support)\\ \tilde{\mathbf{a}}(\pi,X_{I},{I})- \tilde{\mathbf{a}}(\pi,X_{J},{J}) &\leq 0 &\forall\; \pi, X_J, J, I\subseteq J, &&(Monotonicity)\\ \sum_{X_J,J} \left(\tilde{\mathbf{a}}({\pi'},X_J,J){-}\tilde{\mathbf{a}}({\pi},X_J,J)\right)\: \operatorname{P}_{\pi}(X_J,J) &\leq 0 &\forall\; \pi', \pi.&&(Truthful message) \eals In the last inequality, $\operatorname{P}_\pi$ is a shorthand for the analyst's posterior distribution conditional on $\pi$. This characterization of the implementable set follows immediately from (ref).
If, furthermore, the decision-maker objective is linear in $\bar{\mathbf{a}}$, as is the case for a Bayesian decision-maker and binary actions, or if it is linear with an additional linear constraint, as is the case for expected power maximization subject to size control, then the problem of finding the optimal implementable reduced form decision function becomes a linear programming problem. Efficient algorithms exist for numerically solving such problems, cf. vanderbei2020linear. We will return to this point in (ref) below. We leverage such linear programming algorithms in our interactive app for finding optimal PAPs.
\paragraph{Truthful revelation of beliefs and proper scoring} Condition (ref) in (ref) ensures that the analyst reveals their relevant prior information truthfully. Condition (ref) is equivalent to the definition of a proper scoring rule, as introduced by savage1971elicitation. The theory of proper scoring rules has regained importance in the more recent statistics and machine learning literature, cf. gneiting2007strictly.
Let us elaborate on this equivalence. Given a reduced form decision rule $\bar{\mathbf{a}}$, define \be S(\pi', \pi) = \operatorname{E}_\pi[v(\bar{\mathbf{a}}(\pi',X_J,J))]. \ee The expectation $\operatorname{E}_\pi$ is taken over the conditional distribution $\operatorname{P}_\pi$ of $X_J,J$ given $\pi$. Here we assume for simplicity that $X$ has finite support, though the argument generalizes. Denote the Euclidean inner product for functions of $X_J,J$ by $\langle f( \cdot ), g( \cdot ) \rangle = \sum_{X_J,J} f(X_J,J) \cdot g(X_J,J)$, where the running indices $X_J,J$ are understood here as values, rather than random variables. $P_\pi$, the distribution of $(X_J,J)$ given $\pi$, is a vector in the space on which this inner product is defined. We obtain the following characterization, which was first stated by savage1971elicitation and is restated as Theorem 2 in gneiting2007strictly.
\paragraph{Simple pre-analysis plans} Item 2 of (ref) shows that reduced form decision rules can be implemented by delegation: The decision-maker offers a set $\mathcal{B} = \{b:\;(X_I,I) \mapsto \mathcal{A} \}$ of permissible pre-analysis plans (decision functions). The analyst then chooses and communicates one of the decision functions $b \in \mathcal{B}$ before gaining access to the data.
In practice, some pre-analysis plans may be unrealistically complicated, and we may wish to restrict attention to a smaller set $\mathcal{B}_0$ of simpler mappings. The decision-maker could be restricted to choosing $\mathcal{B}$ as a subset of this set of simple mappings, $\mathcal{B} \subseteq \mathcal{B}_0$.
One example of such a restricted set $\mathcal{B}_0$ are the index rules implemented in our app, which is described below. These index rules are of the form $$b(X_I,I) = \boldsymbol 1\left(I_b\subseteq I \text{ and } \sum_{i\in I_b} X_i \geq z_b\right),$$ where $I_b$ is the set of statistics included in the index, and $z_b$ is a critical value.
\paragraph{Aligned objectives} Why does implementability in our setting require a pre-analysis message, if that is not the case in conventional statistical decision theory? Assume for a moment that analyst and decision-maker share the same objective function. In this case, is there any need for a pre-analysis message? The answer is no.
To see this, consider the following variant of our setup. Suppose everything is as in (ref) ((ref)), except that the analyst gets to choose the message $M$ after they observe the data $X_J,J$. Put differently, the analyst cannot provide a verifiable time-stamp for their message $M$ to the decision-maker. The following observation states that in this modified setting, where there is no pre-analysis message, the decision-maker can still implement the first-best reduced-form decision rule, provided that preferences are aligned.
As (ref) shows, pre-analysis messages only become potentially useful in the presence of both private information and misaligned preferences.
\paragraph{Implementability without pre-analysis message}
We next characterize the set of decision functions $\bar{\mathbf{a}}$ that are implementable without a pre-analysis message, when objectives can be misaligned. In this case, the implementable functions are exactly the functions $\bar{\mathbf{a}}(\pi,X_J,J)$ that satisfy monotonicity, with respect to set inclusion for the index set $J$, and that do not depend on $\pi$. Analyst expertise can thus not be used to improve decisions at all, in the absence of a pre-analysis message. The proof of the following proposition parallels the proof of (ref).
We next specialize our general framework to the setting of frequentist hypothesis testing. In this setting, the decision-maker decides whether to reject a null hypothesis. We assume that the decision-maker wants to maximize expected power subject to size control. The analyst, however, always prefers a rejection of the null hypothesis.
Building on our previous results, we characterize the set of implementable testing rules that satisfy size control, in (ref). We furthermore provide a simple mechanism that allows the decision-maker to implement the optimal testing rule. This mechanism requires a pre-analysis plan, where the analyst may choose any full-data test that satisfies size control, and the decision-maker makes worst-case assumptions about any unreported data. This mechanism solves the decision-maker's problem.
In (ref) we then consider the analyst's problem of finding an optimal response to this mechanism, and show that they have to solve a linear programming problem to find the optimal pre-analysis plan. We provide software to solve this problem of the analyst. We also characterize the set of possible solutions to the analyst's problem, by describing the set of extremal points of their feasible set.
Throughout, we focus on the problem of testing a single (joint) hypothesis, and leave an extension to deciding which of multiple hypotheses to reject for future work.
Assume that the decision $A \in [0,1]$ represents the probability, given $(M,X_J,J)$, of rejecting the null hypothesis $\theta \in \Theta_0$. Suppose that the analyst is an expected utility maximizer, who (ex-post) only cares about the binary testing decision. Ex-ante, the analyst thus wants to maximize expected power. It follows that their utility is linear in $A$. We can then make the following normalizing assumption, without loss of generality.
The decision-maker also wants to maximize expected power, but subject to the constraint of size control under the null hypothesis.
Recall that we imposed, in (ref), that the conditional distribution of $X$ only depends on $\theta$, that is, $X|\theta,J,\pi \stackrel{d}{=} X|\theta.$ Under this assumption, the conditional expectation $\operatorname{E}[\bar{\mathbf{a}}(\pi,X_J,J) | \theta,\pi,J]$ is well-defined even outside the joint support of $\pi,\theta,J$, as long as $\theta$ is within its marginal support.
The implementability results of (ref) allow us to characterize optimal pre-analysis plans for hypothesis testing as follows.
This result builds on the general characterizations of (ref) and (ref). To get further intuition for (ref) note, first, that it is sufficient to verify size control for the full-data test $t$. The reason is that implementable reduced-form decision rules must fulfill the monotonicity constraint (ref). Subject to monotonicity in $I$, size control of $\bar{\mathbf{a}}$ in the sense of (ref) is equivalent to size control for the full-data test $\bar{\mathbf{a}}(\pi, X, \{1,\ldots,k\})$.
Note, second, that for optimal reduced-form testing rules the monotonicity constraint is in general binding, since both decision-maker and analyst aim to maximize expected power, subject to the constraints. For optimal rules it is therefore without loss of generality to assume $\bar{\mathbf{a}}(\pi,X_J, J) = \inf_{X';\;X'_J=X_J} t(X')$, which can be implemented by $b$ as in the statement of the theorem.
(ref) solves the optimal testing problem from the decision-maker's perspective: Let the analyst pre-specify a valid full-data test, and then make worst-case assumptions about unreported data. We next turn to the analyst's problem: What full-data test should they specify? This problem can be cast as a linear programming problem. The optimal value for any linear programming problem can be achieved on the set of extremal points of the feasible set. \footnote{The same holds more generally, for the maximum of a convex function on a convex set.} This insight, which is of central importance to mechanism design mechanismdesignnotes2023, allows us to characterize the set of potential solutions to the optimal testing problem subject to implementability.
\paragraph{Linear objective and linear feasible set}
For ease of exposition, we focus on point null hypotheses \(\Theta_0 = \{\theta_0\}\) in the following. Our results extend to compound hypotheses. Denote $K=\{1,\ldots,k\}$ the index set of all potentially available statistics. Let $\mathcal B$ be the set of measurable functions $b(X_J,J)$ defined by the following constraints. \bal \int b(X,K) d\operatorname{P}_{\theta_0}(X) &\leq \alpha, &&&(Size control)\nonumber \\ b(X_J,J)&\in [0,1] & \forall\; J,X,&&(Support)\nonumber \\ b(X_J,J) &\leq b(X,K) & \forall\; J,X.&&(Monotonicity) \eal This is the set of testing rules from which the analyst is effectively allowed to choose, after observing their private signal $\pi$. This characterization applies to both discrete and continuously distributed $X$. The set $\mathcal B$ is a convex polytope.
The (interim) analyst objective function is given by expected power, conditional on their private signal $\pi$, \bals \operatorname{E}_\pi[b(X_J,J)] &= \int b(X_J,J) d\operatorname{P}_\pi(X,J).&&(Interim expected power) \eals We provide code, in the form of an interactive app, which allows the analyst to easily solve the problem of maximizing expected power, subject to $b\in \mathcal{B}$.\footnote{This app is available at \href{https://maxkasy.github.io/home/pap_app/}{https://maxkasy.github.io/home/pap_app}. }
\paragraph{The case of known $J$} The analyst's problem simplifies to the standard problem of finding a test of maximal expected power subject to size control, if we assume that the analyst knows the value of $J$, at the time of specifying their PAP. Let $J'$ be this known non-random value of $J$. Under this assumption, the optimal implementable test is a function of $X_{J'}$ only, and can be written as a likelihood ratio test.
(ref) implies that the null should be rejected based on the value of the likelihood ratio test statistic $\frac{d\operatorname{P}_{\pi}(X_{J'}, J')}{d\operatorname{P}_{\theta_0}(X_{J'}, J')}$ (assuming this statistic is well defined). Note that the likelihood in the numerator $d\operatorname{P}_{\pi}(X_{J'}, J')$ is in fact the marginal likelihood under the interim prior given $\pi$, averaging over both the interim prior for $\theta$, $J'$, and over the sampling distribution of $X$ given $\theta$. See lehmann2006testing (Section 3.8) for a discussion of statistical tests that maximize weighted average power.
\paragraph{Potentially optimal tests: Extremal points of $\mathcal B$} Let us now return to the more general case, where the analyst does not necessarily know the value of $J$ after observing $\pi$. Suppose we maintain (ref), (ref), and (ref), but impose no further assumptions on the (interim) prior $\operatorname{P}_\pi$ of the analyst. What can we say about the set of potential solutions $b$ to the analyst's problem, in this case? The following proposition provides a characterization, based on the set of extremal points of the set $\mathcal B$, intersected with the set of rules $b$ for which monotonicity is binding.
In other words, we can restrict our attention to testing rules that partition values of the data $X$ into at most three regions: one where the test always rejects; one where the test never rejects; and one where it rejects with a single, intermediate probability. Furthermore, if there is more than one value for which the test takes this intermediate rejection probability, then the monotonicity constraint in the construction of the tests $b$ is binding for at least some subset $J$.
The result in (ref) characterizes the set of extremal points of $\mathcal B$ for which monotonicity is binding. The optimal analyst response is necessarily in this set. Can all of these points be rationalized as optimal for some analyst interim prior? The following proposition provides a partial answer.
This result shows that all testing rules that control size without an intermediate probability of rejection can be rationalized.
We next discuss a numerical example, to illustrate our results on optimal pre-analysis plans for hypothesis testing. Our example is calibrated to the data and the priors reported in dellavigna2018motivates, who experimentally evaluate 15 different treatments to induce costly effort, in addition to 3 control treatments. The outcome $X_i$ is the effect of treatment $i$ on the average number of button presses, in an Amazon Mechanical Turk task. We consider the effect relative to the control treatment where participants are paid 1 cent per 100 button presses. dellavigna2018motivates also report prior predicted treatment effects, as elicited from 208 academic experts.
We use these expert predictions to calibrate our prior for $\theta = (\theta_i)_{1\leq i \leq 15}$, where $\theta_i$ is the true effect of treatment $i$. We assume that $\theta \sim N(\mu, \Sigma)$ is jointly normal, with prior mean $\mu$ equal to the average of expert forecasts, and prior variance $\Sigma$ equal to the variance across forecasts. We furthermore assume that the estimated treatment effects have a sampling distribution of $X_i \sim \mathcal{N}(\theta_i, \sigma^2_i / n + \sigma_0^2 / n)$, where the sample size is $n = 100$, and $\sigma_0^2 / n$ is the variance of the mean outcome for the control treatment.\footnote{The variation across experts is different from the variance of the prior of any individual expert. We furthermore deviate the original sample size of around 550 per arm in the paper. For both these reasons, or numerical example should only be thought of as a calibration for the purpose of illustrating our theory.} The standard deviations $\sigma_i$ are assumed to be known and correspond to the standard errors reported in dellavigna2018motivates.
We lastly assume, for the purpose of illustration, that the analyst only intends to run experiments for two of the 15 experimental treatments, corresponding to arm 1 (4 cents per 100 presses), and arm 2 (a lottery with a chance of winning 1 dollar per 100 presses with 1% probability). We assume, for now, that the analyst knows ex-ante that $J=\{1,2\}$. Consider the null (joint) null hypothesis that there are no treatment effects for any of the incentive schemes, $\theta_i = 0$ for all $i$.
\paragraph{The optimal PAP} What is the optimal PAP for this null? The answer is given by (ref). The optimal PAP, for known $J$, pre-specifies a test which rejects whenever both components of $J$ are reported, and $\log\left(\tfrac{d\operatorname{P}_{\pi}(X_{J}, J)}{d\operatorname{P}_{\theta_0}(X_{J}, J)}\right)$ exceeds some critical value. Under our assumptions, \bals \log\left(\tfrac{d\operatorname{P}_{\pi}(X_{J}, J)}{d\operatorname{P}_{\theta_0}(X_{J}, J)}\right) &= const. +
' S^{-1}
-
' S_0^{-1}
\eals where $S_0$ is the sampling variance of $X_J$, and $S$ is the prior variance of $X_J$, which equals the sum of the prior variance of $\theta_J$ plus the sampling variance, \bals S_0 &= \tfrac1n
, & S =
+ S_0. \eals We visualize this test in (ref). The axes of the graph represent estimated treatment effects, normalized by their sampling standard error, $X_i / \sqrt{(\sigma_0^2 + \sigma_i^2)/n}$, for $i=1,2$. The blue ellipse (dashed line) represents the null distribution of $X_J$, with 95% of draws falling within the circle; and the purple ellipse (solid line) represents the prior marginal distribution of $X_J$. The optimal rejection region at a 5% size is shaded in yellow. The likelihood-ratio test of (ref) yields an ellipsoidal rejection region.
\paragraph{An optimal simple PAP} In practice, fully optimal tests may be hard to describe in a PAP. What is the optimal PAP subject to an additional simplicity constraint? Let us restrict attention to tests that reject if the test statistic $X_J' \cdot S_{0}^{-1} \cdot X_J$ exceeds some critical value $c_J$ and if all components in $J$ are reported, where both $J$ and the critical value are pre-specified. This is the standard Wald ($\chi^2$) test for the subset $J$. Subject to this restriction, it is optimal for the analyst to pre-register their true $J = \{1,2\}$, and a critical value of $6$ (for a test of size .05). We visualize this test in (ref). Restricting tests to be simple leads to a loss in average power, but this loss is small in our numerical example. Average power of the optimal test is approximately $.52$. The restriction to a Wald test reduces power to $.50$.
\paragraph{Analyst uncertainty about $J$} Assume now that the analyst is uncertain about which components $J$ will be available. Maybe some experiments are not always feasible, or data collected differ from those in the original plan. Assume that arm 1 is available with ex-ante probability .5, and arm 2 with probability .7, independently across arms. In this case, the rejection regions of the optimal PAP are more complex, and are given by the solution to the linear programming problem discussed in (ref). (ref) plots the optimal test, which solves this linear program. If one of the arms is not available in the end, then the decision-maker makes worst-case assumptions about this arm, and implements the corresponding testing decision. Because the components $i$ are not always available, overall expected power only equals $.32$ in this example.
We can also, again, consider simple PAPs, which specify Wald tests for some pre-selected set of components $J'$. The optimal simple PAP ignores arm 1 and specifies a two-sided t-test that rejects for $\frac{\sqrt{n} \left| X_{2}\right|}{\sqrt{\sigma^2_{2} + \sigma^2_0}} > 1.96$ ((ref)). That is, despite arm 1 being available some of the time, it is better to only consider arm 2 in this case. This result is driven by the different priors over the effect of these treatments, as well as by different availabilities, where arm 2 is more likely to lead to a rejection and is more likely to be available. Restricting attention to such a simple test reduces expected power from .32 (for the optimal test) to only .26.
Concluding our discussion of this numerical example, we plot in (ref) how the optimal set of pre-registered components $J'$, for a simple test, depends on the probability that data for either treatment is available. (ref) elaborates further.
We conclude by summarizing our main contributions, before discussing some limitations of our model and avenues for future research. We have proposed a principal-agent model of pre-specification in empirical research. In our model, a decision-maker relies on the examination and reporting of data by an analyst. The analyst can selectively report statistics that they observe, but they cannot lie about the observed statistics. The decision-maker does not know which data are available to the analyst. This allows for plausible deniability.
Our model provides a theoretical justification for pre-analysis plans (or, more generally, pre-analysis messages), which cannot be rationalized in traditional single-agent statistical decision theory. There is no need for sending messages prior to seeing data in the single agent framework -- in fact, there would not even be a recipient for such a message in this framework.
The constraint of implementability in our model leads to a constrained version of statistical decision theory. Constrained optimal decision functions generally require a PAP. PAPs allow the decision-maker to draw on analyst expertise. Such analyst expertise cannot be used under the alternative mechanism of unilateral specification of decision functions by the decision-maker.
Our model also allows us to derive practical guidance for the design of optimal PAPs. Optimal PAPs lead to constrained-optimal decision functions. We show that the decision-maker's optimal decision function can be implemented by allowing the analyst to choose from a restricted set of decision-functions, and communicating their choice in a PAP. For hypothesis testing, the analyst gets to choose any test that satisfies size control when all data are observed. If a statistic required by the pre-specified test is not reported, then the decision-maker later makes worst-case assumptions about this statistic. The analyst problem, for this mechanism, reduces to a linear programming problem. They have to maximize expected power subject to size control, and subject to the constraints implied by implementability. When the set of available statistics is known to the analyst in advance, then the solution to the analyst problem takes the form of a likelihood ratio test. More generally, we provide an app that allows the analyst to easily solve their optimization problem.
Our model is fairly general in describing the problem of selective reporting by an analyst with conflicting objectives and private expertise. There are some important considerations, however, which are not reflected in this model. First, we do not model the potential cost to researchers of documenting complex estimation and testing procedures in the PAP. This is a cost that has been emphasized by critics of the widespread adoption of PAPs CoffmanNiederle2015,Olken2015,duflo2020praise. Relatedly, we do not model the cost of communicating complex findings. Such costs likely play an important role in explaining why not all findings are published Frankel_undated-xn, Andrews2020-cp.
Second, there are a number of alternative mechanisms that might complement PAPs as tools to limit the adverse effects of conflicting interests and private information. One such mechanism is adversarial review, where reviewers might request additional statistics to be reported by researchers. Our model does not include a review stage. Another such mechanism is researcher reputation, and more generally the dynamics of repeated interactions. Our model is a one-shot game, which does not allow for such dynamics. We hope that future research will elaborate on mechanisms such as these, and the extent to which they might act as a substitute for PAPs.