Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
134,700 characters · 29 sections · 98 citation commands
“Post” Pre-Analysis Plans: Valid Inference for Non-Preregistered Specifications
Pre-analysis plans (PAPs) are time-stamped, publicly accessible documents that preregister hypotheses, experimental design, and planned statistical analyses before data are observed. Originating in clinical trials, PAPs have been championed as a key tool for improving credibility and transparency in experimental economics clinicalPAPs2004,olken2015promises. Following the 2013 launch of the American Economic Association’s RCT Registry—a central repository tracking ongoing, completed, and withdrawn trials—the number of preregistrations with PAPs has grown remarkably ofosu2023pre. Despite this growth, uncertainty remains about their requisite detail and comprehensiveness, as well as the extent to which researchers should be allowed to deviate from preregistered choices olken2015promises,banerjee2020praise. In a 2023 survey of experimental economists, imai2025 report that while 83% support deviations provided transparent disclosure, views diverge sharply on acceptable scope—ranging from unrestricted latitude to narrowly circumscribed departures. Overall, 74% expressed a desire for clearer norms and guidance for how PAPs should be drafted and reviewed.
At root, the core tension underlying arguments for and against PAPs turns on the permitted scope of deviation: On the one hand, strict adherence limits researchers' degrees of freedom for data mining, specification searching, and ex post rationalization Simmons2011FalsePositive,Wicherts2016dof. Formally, strict adherence sets the analysis rule ex ante, ensuring nominal frequentist guarantees (e.g., size, coverage) hold under identical replications of the experiment olken2015promises,kasy2023optimal. On the other hand, strict adherence constrains researchers' flexibility to handle unforeseen issues that arise during implementation, as well as ability to explore novel insights that emerge over the course of the study. Considering the high overhead costs to running a new experiment, not pursuing auxiliary analyses when the data suggest them also comes at a serious social cost Miguel2021TransparencyJEP, CoffmanNiederle2015, olken2015promises, banerjee2020praise. In practice, writing a PAP that is sufficiently exhaustive to anticipate all possible contingencies is extremely time and labor intensive, and deviations are extremely common OfosuPosner2020PAPsPAndP, ofosu2023pre, Brodeur2024reduce. The pervasiveness of such deviations presents an urgent need to understand how the reporting of preregistered estimates (henceforth, PAP estimates) formally differs from that of non-preregistered estimates (henceforth, post estimates), and by extension, how the latter ought to be presented in papers and interpreted by readers.
In this paper, we develop a framework to formalize the statistical properties of deviations from the PAP, thereby allowing us to provide certain inferential guarantees for the resulting post estimates. In doing so, we offer a means to resolve the core tension above: Researchers can still enjoy the benefits of PAP estimates, while leaving open the option to validly report and interpret post estimates. Our approach frames deviations as instances of selective reporting, which differ from prespecified analyses specifically in that the decision to report can depend on realized data. For example, a non-preregistered regression coefficient might be of interest only when it is found to be significant and large; a new specification pooling treatment arms might be reported only if the preregistered specification fails to reject the null; alternately, a non-preregistered economic model might be tested only if the data violate a core assumption of the preregistered model. In modeling deviations in this way, we formalize an existing norm where researchers are expected to provide a rationale for deviating. Here, this rationale is expressed through what we term the deviation set, which corresponds to the subset of data realizations satisfying the conditions for reporting a given post estimate. Using the deviation set, we show how one may obtain confidence intervals and point estimators that provide valid inference conditional on the event that the researcher deviates. We argue that researchers should leverage this deviation set to report corrected (i.e., conditional) inferences alongside conventional (i.e., unconditional) inferences. We therefore provide general inference procedures for obtaining confidence intervals that have correct conditional coverage and point estimators that are conditionally unbiased.
The first and primary contribution of this paper is the framing of PAP deviations as selective reporting, which provides a formal language for addressing the concerns faced by researchers when reporting post estimates. In viewing deviations through this lens, our framework is able to leverage results and algorithms from the conditional inference literature to correct for the statistical distortions arising from PAP deviations, ensuring valid inference for post estimates.
To develop our results, we consider a model where the estimates are normally distributed with a known covariance matrix. The inference procedures that we propose are based on the observation that, since a deviation only occurs in the data realizations from its deviation set, the corresponding post estimates are no longer normally distributed, owing to dependence between the post estimate and its deviation event. In the leading case where the deviation set is based on realizations of PAP estimates, this dependence is governed by the correlation between the PAP and post estimates. After accounting for this dependence, the distribution of post estimates corresponds to a class of truncated normal distributions, with truncation regions that depend on the deviation set. We show that for a broad class of polyhedral deviation sets, which encompass the deviations that occur in practice, this yields computationally tractable test statistics from which conditionally valid confidence intervals and point estimators can be derived---results from Pfanzagl1994 further imply the optimality of these procedures for our setting.
As a second contribution, we relax the assumptions of normality and known covariance to show that feasible analogues of the proposed procedures---where one plugs in asymptotically normal estimates and consistent covariance matrix estimators---are asymptotically valid in large samples, uniformly over a broad class of data generating processes. Thus, researchers can safely implement these plug-in procedures, which we detail explicitly. While such uniformity results have been established in markovic2017unifying; tian2017asymptotics; andrews2024inference; and mccloskey2024hybrid for the case of one polyhedra, our results establish uniformity for the general case of multiple polyhedra.\footnote{tibshirani2018uniform develop uniform asymptotics for a model where the events of interest take the form of unions of polyhedra, but (i) impose that underlying mean parameters shrink in accordance with the sample size and (ii) employ procedures that condition on individual polyhedra before aggregating. The latter can lead to efficiency loss in the normal model, relative to the optimal procedure that we employ fithian2014optimal.} In proving uniformity for unions of multiple polyhedra, we allow for the complexity and expressiveness of deviation sets that arise in practice—for instance, conditioning on two-sided significance requires conditioning on a union of two polyhedra. This uniformity result is of independent interest, beyond the PAP setting. For example, the selective inference literature is often interested in conditional inference for lasso model coefficients conditional on the lasso selection event, which can be represented as a union of polyhedra lee2016exact.
We apply our framework to bessone2021economic, which examines the economic effects of increased sleep among the urban poor. We construct three plausible deviation sets, based directly on text from the paper, and show that—depending on the deviation set and the realized data—accounting for conditional reporting can range from no change to economically meaningful departures from conventional practice. Intuitively, when the realized data lie near a reporting boundary implied by the deviation set, conditioning can have a large impact; when the data lie comfortably inside the region, the adjustments to conventional inference are minimal.
As an additional contribution, we examine the robustness of our procedure to certain forms of misspecification of the deviation set, encompassing cases of both strategic misreporting and earnest mistakes. First, we show that inference remains valid—albeit less precise—when the researcher can only describe a subset of their deviation set local to the realized data. Second, we present sensitivity analysis to assess robustness across alternative deviation sets, motivating potential norms that journals and researchers should adopt, such as agreed-upon significance cutoffs for reporting non-prespecified findings (e.g., a default significance level of 5%).
To ease the adoption of our approach, we give examples of how to specify deviation sets for the most common rationales, demonstrating that the costs to formal articulation are low. For example, if a post estimate is reported because it is significant at the $5\%$ level, then the deviation set is characterized by the significance cutoff. We show that an immediate implication of our framework is that conventional and conditional inferences are equivalent when deviations occur (i) for reasons independent of the data (e.g., arrival of new econometric methods, honest mistakes, or oversights in PAP specification) or (ii) in events that are independent of the post estimates (e.g., deviations based on the results from a pilot study on an independent sample). In these special cases, there is no need to adjust one's conventional inferences.
\paragraph{Related Literature.}
Our paper contributes to an ongoing dialogue among economists on the costs and benefits of PAPs; see, for instance, McKenzie2012, Humphreys2013, CoffmanNiederle2015, olken2015promises, Glennerster2017, ChristensenMiguel2018, ChristensenFreeseMiguel2019, banerjee2020praise. These discussions are accompanied by an active literature meta-analyzing the empirical success of PAPs in achieving their stated goals of curbing publication bias and $p$-hacking, such as BrodeurCookHeyes2020AERMethodsMatter, broduer2023unpack, ofosu2023pre, and Brodeur2024reduce. On the whole, existing work has primarily centered around questions of (i) whether to use a PAP in the first place and (ii) its appropriate length, detail, and contents. Less explored, however, are the specific consequences of deviating from a PAP after it is already preregistered (i.e., taken as fixed) and data observed. This question is the focus of our paper, isolated from considerations of publication bias and $p$-hacking.
There is also a literature discussing optimal PAP construction ludwig2019augmenting, banerjee2020theory, anderson2022highly, kasy2023optimal. Our conditional inference framework instead takes the PAP as given and asks how one can validly interpret deviations from the PAP. That said, while we do not make formal statements in this paper about optimal PAP specification, we do provide suggestions to researchers and journals for potential norms to adopt. We leave formal consideration of the implications of our framework for optimal PAP construction to future work.
Methodologically, this paper draws from the conditional inference literature, particularly from work on optimal conditionally quantile-unbiased estimation for polyhedral conditioning events pfanzagl1979optimal, fithian2014optimal, lee2016exact, andrews2024inference, mccloskey2024hybrid. The conditionally quantile-unbiased estimators form the building blocks for our proposed confidence intervals and point estimators. In terms of model setup, the asymptotic framework underlying our uniformity result most closely resembles the approaches of andrews2024inference and mccloskey2024hybrid, whose results can be used to prove the case of a single polyhedra. Here we prove the general case of multiple polyhedra.
\paragraph{Outline.} In Section (ref), we begin with a description of the research process surrounding PAP construction and deviations therefrom, using bessone2021economic as an illustrative application. In Section (ref), we introduce notation to formalizes this process; moreover, we interpret deviations through the lens of a simple Bayesian decision problem---though our inference results are valid without the Bayesian structure. In Section (ref), we introduce our conditional inference objectives in a model with normally distributed estimates, and present general procedures for constructing confidence intervals and point estimators that are conditionally valid. In Section (ref), we show how to implement these procedures in practice, allowing for non-normal estimates. In Section (ref), we apply our results to bessone2021economic. In Section (ref), we discuss robustness of our proposed procedures to misspecification of the deviation set. We provide all proofs and supplementary results in the Appendix.
We begin with a general description of how PAPs are used in the status quo, accompanied by examples from the experimental economics literature. This descriptive timeline will motivate the formal model presented in Section (ref). Throughout this paper, we focus on PAPs in the context of randomized control trials (RCTs), as (i) this is where PAPs are most used in economics, and (ii) it offers a well-defined division between preregistration and data collection. That said, these discussions in principle also apply to empirical research using preexisting datasets burlig2018improving,Miguel2021TransparencyJEP, though such cases generally raise the additional challenge of verifying that researchers are indeed preregistering studies before analyzing the data\footnote{There do exist non-experimental settings where researchers are required to prespecify analyses prior to data access. For instance, restricted-use microdata such as those accessed through the U.S. Census Bureau often require submission and approval of detailed research proposals before any data extracts are made, with each additional data request or modification of scope typically necessitating a new proposal.} ChristensenMiguel2018. To avoid distracting from the central focus of this paper, we take the delineation of “pre” and “post” data collection as given.
\paragraph{Priors.} Researchers embark on projects with priors informed by sources such as economic theory, existing literature, outcomes of a pilot study, discussions with experts, or first-hand observation. For example, bessone2021economic benchmark the anticipated magnitude of their results as follows:
At the highest level, priors can inform the researcher's choice of which hypotheses to test, estimands to consider, and outcomes to measure given theoretical salience or expected precision and effect size. Because empirical projects are constrained in scale and scope, priors can further discipline downstream design choices: how large of a sample to field; how many waves of follow-up to fund; which populations to emphasize; whether to invest in costly data enhancements; and how to allocate statistical power between estimating overall effects and probing heterogeneity or mechanism analysis.
\paragraph{Preregistration.} In advance of data collection, the researcher preregisters information about their study such as primary outcomes, experimental design, randomization method, randomization unit, clustering, and sample size (including total number of observations, number of clusters, and units per treatment arm). In addition to this information, the researcher may also register a PAP stipulating which specifications will be used to estimate quantities of interest. In practice, depending on the research topic and journal, there is variation in whether a PAP is required, and if so, to what level of specificity. At its simplest, a PAP might consist of a single list of linear regressions. On another extreme, a PAP might take the form of a complete contingent plan which maps every possible data realization to a particular set of specifications, mirroring the notion of a complete strategy in game theory.
To illustrate, consider a simple example, shown in Figure (ref). After estimating the average treatment effect (ATE) of a deworming intervention on school attendance, a complete contingent PAP might stipulate that if the attendance effect is statistically insignificant, the researcher will analyze a particular set of secondary outcomes such as health measures (e.g., incidence of anemia, weight-for-age) to assess whether the intervention improved child well-being through channels other than schooling. By contrast, if the effect is significant, the PAP specifies a set of persistence checks to perform, such as attendance in the following academic year.
\usetikzlibrary{arrows.meta,positioning,shapes.geometric}
In the sample PAP above, the researcher not only specifies their primary analysis, but also their intended secondary analysis for both possible contingencies based on their initial findings. Even the simplest of empirical exercises, however, can be far more complex than the contingencies shown in Figure (ref). Operationally, most PAPs in practice lie somewhere in the middle: The researcher has a core set of primary specifications and also stipulates how these specifications will be adjusted for some—but not all—contingencies.
An example of a partial contingent plan can be found in the (31-page) PAP for bessone2021economic. One component of their analysis tests whether sleep affects how attentive individuals are to earnings-related incentives. The authors approached this question by introducing a “salience variation” in a particular data-entry task. The plan was first to establish that participants respond more strongly when incentives are presented in a high-salience rather than a low-salience way, and then to assess whether the sleep intervention reduces this gap.\footnote{If individuals are fully attentive, their responses to incentives should be the same whether the incentives are displayed in a high- or low-salience way. A larger gap between the two conditions therefore signals inattention. As sleep is hypothesized to improve alertness, if the treatment reduces this gap, that provides evidence that sleep mitigates inattention to incentives.} However, it was unknown ex ante whether (i) the salience manipulation would succeed and (ii) what form the response would take. As directly stated in their PAP: \blockquote{ “We do not fully pre-specify this analysis since the appropriate analysis will depend upon first establishing that the salience variation works “as intended”—that is, the response to incentives is stronger under high-salience relative to low-salience. Moreover, the precise form of the response to salience—e.g. whether individuals notice and respond to the incentive change in 5 minutes versus 30 minutes—remains unknown at the time of this pre-registration, making it difficult to write down the full contingent analysis plan.” }
\paragraph{Data Collection and Analysis.} After completing the preregistered analyses, a researcher may wish to go beyond the PAP and compute additional results. Common motivations might include exploratory work prompted by unexpected findings, knowledge gained during data collection, new statistical tools, or unanticipated issues with the original specification. As an instance of the latter, bessone2021economic write: \blockquote{ “We preregistered daily net savings as our main variable of interest. However, this measure suffers from an unanticipated design issue: participants make large one-time withdrawals right before the study ends, which mechanically drives down net savings. We believe deposits more accurately reflect differences in savings behavior, and the accrued interest captures the benefit of savings.” } More generally, because PAPs reflect prior beliefs about where the most interesting or strongest effects are likely to lie, the desire for additional analysis is particularly likely when the realized data substantially shift the researcher's prior. The next section formalizes this idea.
\paragraph{Reporting of Results and Deviations from PAP.} The researcher publicly reports the PAP and the results of all preregistered analyses.\footnote{We assume that PAP results are always reported (i.e., the researcher does not hide any findings). For simplicity, we also do not model degrees of emphasis in reporting. In our framework, placing PAP results in an appendix is equivalent to including them in the main text.} At this time, there is broad uncertainty in the profession about whether and how to report non-preregistered results. Prescriptions range from strict “no reporting” to more permissive approaches that allow reporting conditional on clearly flagging non-PAP findings as exploratory or suggestive.
We refer to the reporting of a non-preregistered result as a deviation from the PAP. Uncertainty over their interpretation notwithstanding, deviations are extremely common in practice and arise for many reasons: understanding surprising results (exploration or mechanism probes), referee and editor requests (e.g., robustness or placebo checks), or the availability of improved statistical tools since PAP submission. We use bessone2021economic as a running example of such deviations.
This pooling decision is substantively motivated: When two night-sleep arms are practically indistinguishable, combining them yields a more precisely estimated quantity. However, at this stage there is uncertainty in the profession over how to proceed: A blanket ban on deviations would preclude learning from the pooled results, yet reporting conventional (unconditional) point estimators and confidence intervals would ignore that the pooling decision was made selectively after seeing the data, risking an overstatement of precision. While the researcher can flag such results as exploratory and meant to be taken with a grain of salt, it is still ambiguous how to interpret deviations from the PAP. In the coming sections, we express the stages outlined above with a formal model, then characterize the adjustments that should be made to account for the selective reporting of non-preregistered results.
We begin by describing the environment and introducing terms in Section (ref), formalizing the researcher's timeline discussed in the previous section. To ground motivation, in Section (ref) we contextualize the researcher's choice of specifications in terms of a Bayesian decision problem and show that even a fully Bayesian decision maker may wish to deviate if constrained to choose from a limited set of specifications. We emphasize, however, that our subsequent inferential framework in Section (ref) is valid without this Bayesian structure.
The researcher observes data $X \in \mathcal{X}$, for $\mathcal{X}$ a sample space. The researcher considers specifications $S: X \overset{}{\rightarrow}\mathbb{R}^{d}$ from some set $\mathcal{S}$; that is, each specification $S \in \mathcal{S}$ maps data $X$ to estimates $S(X)$. An example of $\mathcal{S}$ could be the set of all linear regressions.\footnote{This setup accommodates cases where the researcher considers regression specifications with differing numbers of regressors, since vectors in $\mathbb{R}^{\Tilde{d}}$ for $\Tilde{d} < d$ can also be represented as vectors in $ \mathbb{R}^{d}$.} In advance of observing the data, the researcher preregisters a set of specifications $\mathcal{S}_{pre} \subseteq \mathcal{S}$, where $\mathcal{S}_{pre}$ represents the pre-analysis plan (PAP). We denote specifications from the PAP by $S_{pre} \in \mathcal{S}_{pre}$. For example, $\mathcal{S}_{pre}$ might be the specific set of linear regressions the researcher intends to run, and $S_{pre}$ one such regression.\footnote{Observe that one could alternately define $\mathcal{S}$ to be the set of all possible PAPs, from which the researcher chooses a single element $S_{pre}\in\mathcal{S}$. Our choice of notation, while mathematically equivalent, is more practical for when we discuss inference in Section (ref), and seems better aligned with how researchers view PAPs.} After observing $X$, the researcher announces a set of non-preregistered specifications $\mathcal{S}_{post}^{X} \subseteq \mathcal{S} \setminus \mathcal{S}_{pre}$. We refer to $\mathcal{S}_{post}^{X}$ as deviations from the PAP, which may or may not depend on the observed data $X$. The researcher then reports (i) PAP estimates $\hat{\beta}_{pre} = S_{pre}(X)$ for all PAP specifications $S_{pre} \in \mathcal{S}_{pre}$ and (ii) post estimates $\hat{\beta}_{post} = S_{post}(X)$ for all deviations $S_{post} \in \mathcal{S}_{post}^{X}$. We depict this timeline in Figure (ref). To ground this notation, we revisit the example of bessone2021economic.
\noindentExample \ref*{d2:ex:sleep} (Continued). bessone2021economic cross-randomize workers to receive (i) one of two treatments designed to improve nighttime sleep (“Incentives” and “Encouragement”) and (ii) a treatment that provides nap breaks during workdays (“Nap”). They preregister a set of regression specifications $\mathcal{S}_{pre}$ regressing indices of work, well-being, and cognition outcomes on interactions of the two night sleep and nap treatments.
After conducting the study and observing data $X$, the authors compute the treatment effects of the five fully disaggregated treatment arms on preregistered outcomes, $\hat{\beta}_{pre} = S_{pre}(X)$. However, after observing that “those who received a night sleep treatment in addition to naps had very similar effects to those with naps only,” the authors reason that “naps have an overall positive effect on outcomes, whereas increases in night sleep do not.” To increase statistical power in discussion of effects on outcomes, the authors estimate an additional set of specifications $S_{post}\in\mathcal{S}_{post}^{X}$ which (i) pools the two night sleep treatments and (ii) removes the interaction term for night sleep and nap treatments, yielding $\hat{\beta}_{post} = S_{post}(X)$. The authors supplement their report of the PAP estimates $\hat{\beta}_{pre}$ (disaggregated) with post estimates $\hat{\beta}_{post}$ (pooled), found in Tables III and IV of their paper, respectively.
There are many reasons that deviations arise. In this section, we show that even a fully Bayesian researcher in this framework will sometimes want to deviate when they are limited to a constrained set of specifications for constructing their PAP.\footnote{This model is for motivation alone: The inference procedures in Section (ref) are valid without this Bayesian structure.} We offer discussion on where such constraints arise in practice, such as through labor and communication costs, tractability, cognitive limitations, and norms in the profession.
The distribution of $X$ is governed by the parameter $\theta \in \Theta$, $X\sim P_{\theta} \in \Delta(\mathcal{X})$, for $\Theta$ a parameter space and $\Delta(\mathcal{X})$ the set of probability distributions on $\mathcal{X}$. The researcher has a loss function $L(S(X), \theta) \geq 0$ that quantifies the consequences of reporting estimates $S(X) \in \mathbb{R}^{d}$ when the true parameter is $\theta \in \Theta$. For example, if the goal is to estimate a treatment effect $\tau(\theta) \in \mathbb{R}$, then a standard loss function is squared error:
The risk of specification $S$ is the expected loss from reporting $S(x)$ across possible realizations $x \in \mathcal{X}$ of the data:
Note that the risk depends on the unknown parameter $\theta \in \Theta$.
\paragraph{PAP Construction.} We consider the researcher's problem of choosing the set of specifications $\mathcal{S}_{pre}, \mathcal{S}_{post}^{X} \subseteq \mathcal{S}$ to estimate, focusing on the case where $|\mathcal{S}_{post}^{X}| \leq |\mathcal{S}_{pre}|=1$ for simplicity. We suppose the researcher has prior beliefs about the unknown parameter $\theta$, represented by a prior distribution $\pi \in \Delta(\Theta)$. The researcher determines which specification $S_{pre}$ to preregister in their PAP by minimizing the average risk of $S$ under their prior $\pi$. Stated formally, given a set of specifications $\mathcal{S}$, loss function $L$, and prior $\pi$, the researcher solves:
where $\Bar{R}(S, \pi)$ is the average risk of $S$ under $\pi$. After solving this optimization problem, the researcher preregisters $S_{pre}$ in their PAP.
\paragraph{Deviations from PAP.} The researcher observes $X$ and reports PAP estimates $\hat{\beta}_{pre} = S_{pre}(X)$, as planned. However, the data also lead the researcher to update their prior beliefs $\pi$ to form posterior beliefs, represented by the posterior distribution $\pi_{X} \in \Delta(\Theta)$ of $\theta$ conditional on $X$. Having gleaned new information about $\theta$, the researcher may now find their preregistered $S_{pre}$ to be suboptimal, leading the researcher to also report post estimates $\hat{\beta}_{post} = S_{post}(X)$, where $S_{post} \neq S_{pre}$. We formalize this decision to deviate below.
The analogous problem to minimizing prior average risk ((ref)) is to minimize posterior average loss:
where $\Bar{L}(S(X), \pi_{X})$ is the average loss of $S(X)$ under $\pi_{X}$. The researcher deviates from their PAP when $S_{post} \neq S_{pre}$. This decision to deviate depends on the coarseness of $\mathcal{S}$.
In a conventional Bayesian decision problem, $\mathcal{S}$ is taken to be the set of all functions from $\mathcal{X}$ to $\mathbb{R}^{d}$. In this case, a specification that minimizes prior average risk ((ref)) also minimizes posterior average loss ((ref)) almost surely, and vice versa lehmann2006theory. Intuitively, when $\mathcal{S}$ is unrestricted, a solution to ((ref)) perfectly prepares for every possible data realization $x \in \mathcal{X}$. In such cases, there is never a need to deviate.\footnote{ In Appendix (ref), we provide a variant of the decision problem where, instead of considering posterior average loss, the researcher solves (ref) with $\pi_{X}$ in the place of $\pi$. In this setup, $S_{post}$ is the counterfactual PAP that the researcher would have constructed had they entered the experiment with their posterior beliefs. If one takes $\mathcal{S}$ to be the relevant action space, this counterfactual PAP approach aligns with the conditional Bayes principle, which says to choose an action that minimizes average loss under one's current beliefs berger2013statistical. Interestingly, deviations arise in this setup even when $\mathcal{S}$ is unconstrained.}
We depart from the conventional Bayesian setup by allowing for a restricted $\mathcal{S}$. That $\mathcal{S}$ is constrained is not a modeling convenience, but a description of practice: At the preregistration stage, researchers face explicit and implicit constraints that narrow the permissible set of specifications. In general, researchers cannot write PAPs sophisticated and lengthy enough to prepare for every $x\in\mathcal{X}$. For one, researchers face time, cognitive, and communication constraints: For most experimental designs, it is not feasible to anticipate, formulate, and communicate optimal responses to every contingency olken2015promises, banerjee2020praise. As an example, the PAP for amy2012oregon—a 50 page paper—is 159 pages long. And, despite being exceptionally detailed, the initial PAP still fell short of exhausting all relevant analyses: From 2012 to 2020, ten additional PAPs (totaling an additional 368 pages) were pre-registered to explore new outcome measures in follow-up papers after collecting additional data amy_PAP_docs.
Norms in the profession also constrain $\mathcal{S}$, favoring parsimonious, interpretable specifications tied to the design and estimand. By default, this includes conventional functional forms and links (e.g., linear or low-order polynomials, logit/probit for binary outcomes, etc.), standard control sets motivated by identification, and prespecified clustering and variance estimators. Idiosyncratic specifications—such as cube-root transformations, high-order polynomials, post-hoc cutpoints, data-driven bins or knots, bespoke weighting schemes, stepwise selection, or open-ended machine learning screens—are typically viewed with skepticism unless accompanied by a compelling ex-ante justification grounded in theory, measurement scale, or experimental design (and appropriately powered).
The constraints on $\mathcal{S}$ lead to a classic dynamic inconsistency problem Strotz1955,KydlandPrescott1977, wherein a plan that is optimal ex ante under prior beliefs $\pi$ may no longer be optimal after updating to the posterior $\pi_X$.\footnote{In behavioral terms, present-bias and self-control generate the same logic Laibson1997,ODonoghueRabin1999,FrederickLoewensteinODonoghue2002. Related commitment perspectives appear in rules-vs-discretion and menu/temptation models BarroGordon1983,Schelling1960,Kreps1979,GulPesendorfer2001.} In a standard Bayesian decision problem with a rich $\mathcal{S}$, the researcher would be allowed to preregister a contingent rule (e.g., “after $X$ is observed, take whichever action meets criterion $C$”), so the plan chosen ex ante already pins down the ex-post action. When $\mathcal{S}$ excludes such contingent rules, however, this flexibility is lost. Concretely, if a researcher is restricted to naming one primary outcome in advance (employment or consumption), a contingent rule like “highlight whichever outcome best meets our welfare/precision criterion after we see $X$” is not in $\mathcal{S}$. Ex ante the researcher may pick employment under $\pi$, but ex post the data may favor consumption. Because $\mathcal{S}$ was insufficiently rich to accommodate switching, the ex ante choice remains binding even when the posterior $\pi_X$ would prefer the other outcome—hence the dynamic inconsistency.
This section proposes conditional inference procedures for post estimates. We begin in Section (ref) by introducing the problem of statistical inference in a model with normally distributed estimates. With the relevant notation and concepts established, we next motivate in Section (ref) the use of conditional inference for post estimates and provide examples in Section (ref). We then develop the conditional inference procedures in Sections (ref)--(ref) and provide an explicit example in Section (ref). The normality assumption is motivated by standard large sample approximations, such as the central limit theorem and the law of large numbers. In Section (ref), we describe how to implement analogues of the proposed conditional inference procedures using sample estimates and establish their uniform asymptotic validity.
Let $\hat{\beta}$ be a vector that contains (i) the PAP estimates $\hat{\beta}_{pre} = S_{pre}(X)$ for each PAP specification $S_{pre} \in \mathcal{S}_{pre}$ and (ii) the post estimates $\hat{\beta}_{post} = S_{post}(X)$ for each deviation $S_{post} \in \mathcal{S}_{post}^{X}$ made under $X$. We assume that $\hat{\beta}$ is normally distributed with unknown mean $\beta$ and known positive definite covariance matrix $\Sigma$:
Let $\mathbb{P}_{\beta}\left\{\cdot\right\}$ and $\mathbb{E}_{\beta}\left[\cdot\right]$ denote probabilities and expectations under this distribution.\footnote{We implicitly assume that estimators and confidence intervals of interest depend on $X$ through $\hat{\beta}$.} Each $\hat{\beta}_{pre}$ corresponds to a PAP estimand $\beta_{pre} = \mathbb{E}_{\beta}\left[\hat{\beta}_{pre}\right]$, while each $\hat{\beta}_{post}$ corresponds to a post estimand $\beta_{post} = \mathbb{E}_{\beta}\left[\hat{\beta}_{post}\right]$. By construction, $\beta$ contains all the PAP and post estimands. We consider the problem of conducting inference on parameters that can be expressed as linear combinations of these estimands: $v'\beta_{pre}$ and $l'\beta_{post}$, where $v, l \neq 0$.\footnote{This setup is quite general, since differentiable nonlinear functions of asymptotically normal estimates will also be asymptotically normal by the delta method. These nonlinear functions can be elements of $\hat{\beta}$.} For instance, in Example (ref), the effects of the various treatments correspond to different contrasts of the regression coefficients.
\paragraph{Inference for PAP Estimates.} Given PAP estimates $\hat{\beta}_{pre}$, let $\Sigma_{pre}$ denote the covariance matrix induced by $\Sigma$. By equation ((ref)), any linear combination of PAP estimates is normally distributed:
Given a desired significance level $\alpha \in (0,1)$, a confidence interval $CI_{\alpha}(X)$ has correct coverage for $v'\beta_{pre}$ across data realizations $X \in \mathcal{X}$ if it satisfies
In words, a confidence interval $CI_{\alpha}(X)$ that satisfies criteria (ref) contains the parameter of interest $v'\beta_{pre}$ with probability $1-\alpha$ for any possible value of the unknown $\beta$. This is a standard criteria for valid inference lehmann2024testing. One example is the conventional two-sided interval $[v'\hat{\beta}_{pre} \pm z_{1-\alpha/2}\sigma_{pre}]$, where $z_{\alpha}$ denotes the $\alpha$-quantile of the $N(0,1)$ distribution. This interval satisfies
and therefore yields valid inference for PAP estimates.
\paragraph{Inference for Post Estimates.} Unlike the preregistered specifications $\mathcal{S}_{pre}$, the researcher's deviations $\mathcal{S}_{post}^{X}$ may depend on the observed data $X$. We say that a data realization $X$ induces the deviation $S_{post}$ if $S_{post} \in \mathcal{S}_{post}^{X}$ in that data realization. For a given $S_{post}$, we may then collect the set of all data realizations that induce the deviation:
We refer to $\mathcal{X}_{post}$ as the deviation set for $S_{post}$. By definition, the researcher is only interested in post estimates $\hat{\beta}_{post}$ when $X\in\mathcal{X}_{post}$. In view of this, we propose confidence intervals that satisfy the following criteria: Given significance level $\alpha \in (0,1)$, a confidence interval $CI_{\alpha}(X)$ has correct conditional coverage for $l'\beta_{post}$ given $X \in \mathcal{X}_{post}$ if it satisfies
Intuitively, such intervals ensure correct coverage for $l'\beta_{post}$ across the set of data realizations $\mathcal{X}_{post} \subseteq \mathcal{X}$ where the researcher is interested in $\hat{\beta}_{post}$. This conditional coverage criteria generalizes the unconditional criteria in (ref). Indeed, since PAP estimates are always reported, the analogous “deviation set” for a PAP estimate yields $\{X \in \mathcal{X}: S_{pre} \in \mathcal{S}_{pre} \} = \mathcal{X}$, in which case conditional and unconditional coverage coincide.
\paragraph{Known Deviation Set.} To obtain $CI_{\alpha}(X)$ with correct conditional coverage, we assume the researcher can (and does) articulate the deviation sets $\mathcal{X}_{post}$ for each $S_{post} \in \mathcal{S}_{post}^{X}$ at hand. This is a natural first step for deriving conditional inference procedures. We relax this assumption in Section (ref), where we discuss robustness of conditional inference procedures to various forms of misspecification. To make the above concepts concrete, we return to the example deviation in bessone2021economic. In the following example, let $\text{se}(\cdot)$ denote the standard deviation of an estimator under normal distribution (ref).
\noindentExample \ref*{d2:ex:sleep} (Continued). bessone2021economic initially preregistered the fully interacted specification $\hat{\beta}_{pre} = (\hat{\tau}_{N}, \hat{\tau}_{NE}, \hat{\tau}_{NI}, \hat{\tau}_{E}, \hat{\tau}_{I})'$ , where
After seeing the data and estimating $\hat{\beta}_{pre}$, the authors report three additional post estimates, $\hat{\beta}_{post} = (\hat{\tau}_{E+I}, \hat{\tau}_{NE+NI}, \hat{\tau}_{N+NE+NI})'$, corresponding to the estimates which (i) pool incentives and encouragement into a single “night sleep” treatment ($\hat{\tau}_{E+I}$ and $\hat{\tau}_{NE+NI}$) and (ii) collapse interaction effects ($\hat{\tau}_{N+NE+NI}$).\footnote{The estimate $\hat{\tau}_{NE+NI}$ is not reported in the main paper, but rather in Online Appendix Table A.VIII (where the authors pool the two night sleep treatments but include a separate indicator for individuals who received a combination of either night sleep treatment along with the nap treatment).} The authors find that, “these results ($\hat{\beta}_{pre}$) provide evidence that naps have an overall positive effect on outcomes, while increases in night sleep do not. However... this analysis has limited statistical power.” The authors justify turning to a “simplified but higher-powered version of this analysis” with the following three reasons:
Taking these points at face value, a plausible characterization of the authors' deviation set for $\hat{\beta}_{post}=(\hat{\tau}_{E+I}, \hat{\tau}_{NE+NI}, \hat{\tau}_{N+NE+NI})'$ is the set of data realizations satisfying the intersection of these events, i.e., $$\mathcal{X}_{post} = \left\{X \in\mathcal{X}\,:\,\, \text{(A) and (B) and (C)} \right\}.$$
We argue that conditional coverage of $l'\beta_{post}$ should be the preferred inferential objective---not unconditional coverage, as is convention---since post estimates are only reported or acted upon in select states of the world. We see two main reasons for this. First, selective reporting can lead to distortions in the density of observed reports of $\hat{\beta}_{post}$, as we will demonstrate. Second, conditional validity is, in a certain sense, necessary and sufficient for confidence sets to be valid on average across studies, which is the very property underlying existing frequentist coverage objectives.
\paragraph{Conditional Coverage for Given Study.} Selective reporting can distort the density of observed reports $\hat{\beta}_{post}$ relative to the unconditional density assumed by conventional inference procedures. This is demonstrated in the following toy example.
\paragraph{Average Coverage Across Studies.} The conditional coverage criterion requires that $CI_{\alpha}(X)$ contain $l'\beta_{post}$ with high probability across data realizations $X \in \mathcal{X}_{post}$, where each data realization can be viewed as a particular draw of the same experiment in a given study. Thus, conditional coverage provides a natural inference criterion for a given study. However, one might ponder the relevance of conditional coverage when there are multiple studies, each with their own experiments and potential PAP deviations. In Appendix (ref), we establish a general sense in which conditional coverage for each study is necessary and sufficient for controlling average coverage across all studies.
Formally, we consider a sampling model where deviation sets are drawn from a rich class of distributions reflecting the unanticipated nature of PAP deviations. Under this structure, we show that to ensure average coverage across studies, it is both necessary and sufficient to ensure conditional coverage in each study. This equivalence result provides further motivation for the conditional inference approach to PAP deviations.
We now discuss leading examples of deviation sets. Section (ref) outlines deviations based on the significance of PAP estimates, which we discuss extensively in future sections. Section (ref) highlights classes of deviation sets that yield no inference distortions in our framework.
Researchers often consider post estimates $\hat{\beta}_{post}$ for which reasoning about $\mathcal{X}_{post}$ is predicated on values of PAP estimates $\hat{\beta}_{pre}$. In particular, $\hat{\beta}_{post}$ is often of interest when linear combinations of PAP estimates $v'\hat{\beta}_{pre}$ cross significance cutoffs $\kappa \geq 0$. Here we focus on $\kappa = z_{1-\eta/2}\sigma_{pre}$, where $\sigma_{pre} = \sqrt{v'\Sigma_{pre} v},$ which corresponds to a conventional two-sided statistical significance test with significance level $\eta \in (0,1)$. However, $\kappa$ can also be based on economic significance.
\paragraph{Statistical Significance.} The researcher deviates when $v'\hat{\beta}_{pre}$ is significantly large:
For example, consider dube2025cognitive, who study the effects of a cognitive-skills training program (Sit-D) for police officers on discretionary arrests. In their Section IV.B, they say
One can interpret the above quote as suggesting a deviation set based on statistical significance: $\hat{\beta}_{pre}$ is the vector of coefficients on the regression with the preregistered measure, $v$ is the unit vector that picks the coefficient on SitD, and $\hat{\beta}_{post}$ is the vector of coefficients for the regression with the non-preregistered measure, with $v'\hat{\beta}_{pre}$ significant at level $\eta = 0.05$ dube2025cognitive.
\paragraph{Statistical Insignificance.} The researcher deviates when $v'\hat{\beta}_{pre}$ is significantly small:
For example, in the quote from bessone2021economic presented in Example (ref), two of the reasons the authors give for pooling, (A) and (B), are based on statistical insignificance. We consider various representations of $\mathcal{X}_{post}$ for bessone2021economic in the empirical application in Section (ref).
An important and immediate implication of our framework is that when the deviation event $\{X \in \mathcal{X}_{post}\}$ is independent of the post estimates $\hat{\beta}_{post}$, one does not have to correct conventional inferences. In particular, since the conventional interval $[l'\hat{\beta}_{post} \pm z_{1-\alpha/2}\sigma_{post}]$ is a function of $\hat{\beta}_{post}$, independence yields
Such non-distortionary deviations can broadly be grouped into two cases, depicted in Figure (ref) and discussed below.
\paragraph{Unconditional Reporting.} The first case (left panel of Figure (ref)) is unconditional reporting, where even though $\hat{\beta}_{post}$ was not registered in the PAP, it would have been reported no matter the data observed. In other words, $\mathcal{X}_{post}=\mathcal{X}$. While seemingly a trivial result, there are nonetheless many examples of such deviations that make researchers uneasy in the status quo. Instances of $\mathcal{X}_{post}=\mathcal{X}$ might include: (i) honest mistakes, oversights in PAP specification, or adjustments for feasibility; (ii) unanticipated events (e.g., change in policy context, global pandemic); or (iii) newly acquired knowledge about the setting/context, econometric methods, or theoretical insights in the literature.\footnote{Examples of (i) include rafkin2021guidance; bhat2022long; alsan2024representation; jacobson2024price; agte2024investing; finkelstein2019take; kaur2024_financialconcerns; kelley2024monitoring. Examples of (ii) include kremer2009incentives; evsyukova2025linkedout. Examples of (iii) include bandiera2021allocation; field2021her; bessone2021economic; giacobinoschoolgirls2024.} For concreteness, we now consider an explicit example of (iii).
\paragraph{Independent Estimates.} The second case (center panel of Figure (ref)) is conditional reporting of $\hat{\beta}_{post}$ based on estimates $\hat{\beta}_{pre}$ that are independent of $\hat{\beta}_{post}$. There are mechanical ways the researcher can ensure this: A key example is pilot studies where researchers (i) register a PAP and conduct a pilot of their study on one population and (ii) make adjustments based on results from the pilot before conducting the main study on a larger population.\footnote{An example that broadly falls into this class is dean2024noise.} When the population of the pilot is sampled independently from that of the main study, any sort of deviation from the PAP based on findings in the pilot yields no inference distortions. Another example in this vein is sample splitting, where one computes estimates on one half of the data, uses those results to determine their deviations, and computes the corresponding $\hat{\beta}_{post}$ using the other half of the data---however, such procedures come at the cost of statistical precision.
Of course, independence need not be a mechanical feature of the experimental design. Consider an intervention disrupted by a lightning strike which changes the underlying sample, as in kremer2009incentives.\footnote{See also olken2015promises for a discussion.} This changes the estimand’s interpretation, as the conditioning event may index a different state/population $s$. However, conventional inference will be valid for this estimand. That is, we now have $\hat{\beta}|\{S=s\} \overset{d}{=} \hat{\beta}(s)|\{S=s\} \overset{d}{=} \hat{\beta}(s) \sim N(\beta(s), \Sigma(s))$, where the second equality follows from independence of selection variable $S$ (e.g., random lightning strike) and estimates $\hat{\beta}(s)$.
We now present general conditional inference procedures. In Section (ref), we state the formal inference objectives for confidence intervals and point estimators. In Sections (ref)-(ref), we detail the construction of optimal procedures that meet these objectives.
To perform valid conditional inference, it suffices to derive estimators $\hat{\mu}_{\alpha}$ that are $\alpha$-quantile conditionally unbiased in the sense that their overestimation probability for $l'\beta_{post}$ conditional on $X \in \mathcal{X}_{post}$ is equal to $\alpha$:
Given such $\hat{\mu}_{\alpha}$, the confidence interval $CI_{\alpha}(X) = [\hat{\mu}_{\alpha/2}, \hat{\mu}_{1-\alpha/2}]$ provides correct conditional coverage:
Alternatively, if we want a point estimator of $l'\beta_{post}$, we can use $\hat{\mu}_{1/2}$ as a median conditionally unbiased estimator:
In words, the conditional median of $\hat{\mu}_{1/2}$ is equal to $l'\beta_{post}$.
To derive $\hat{\mu}_{\alpha}$, we follow arguments from fithian2014optimal. Assume there exists a set $\mathcal{B}_{post}$ of positive measure such that
where the set $\mathcal{B}_{post}$ is implicitly allowed to depend on $\Sigma$. In other words, we focus on deviations that arise due to values of $\hat{\beta}$. This broadly accommodates the deviations that occur in practice, such as those based solely on PAP estimates $\hat{\beta}_{pre}$. For deviation sets satisfying (ref), we have
Thus, our goal is to obtain valid inference for $l'\beta_{post}$ conditional on $\hat{\beta} \in \mathcal{B}_{post}$.
\paragraph{Notation.} Let $l_{post}$ denote the vector induced by (i) matrix multiplication to select $\hat{\beta}_{post}$ from $\hat{\beta}$ and (ii) vector multiplication of $\hat{\beta}_{post}$ by $l$, so that $l'\hat{\beta}_{post} = l_{post}'\hat{\beta}$. Let $\Sigma_{post}$ denote the covariance matrix for $\hat{\beta}_{post}$, so that the variance of $l'\hat{\beta}_{post}$ is $\sigma_{post}^{2} = l_{post}'\Sigma l_{post} = l'\Sigma_{post} l$.
The main challenge is that $l'\hat{\beta}_{post}$ is no longer normally distributed after conditioning on $\hat{\beta} \in \mathcal{B}_{post}$, owing to correlation between $l'\hat{\beta}_{post}$ and $\hat{\beta}$:
To account for this correlation, consider the residual $\hat{r}$ from the regression of $\hat{\beta}$ on $l'\hat{\beta}_{post}$ under their joint unconditional distribution ((ref)):
The conditional distribution of $l'\hat{\beta}_{post}$ given $\{\hat{\beta} \in \mathcal{B}_{post}, \hat{r} = r\}$ is a $N(l'\beta_{post}, \sigma_{post}^{2})$ distribution truncated to the set $\mathcal{Z}_{post}(r) = \{ z \in \mathbb{R}: r + \gamma z \in \mathcal{B}_{post} \}$:
where the second equality follows from independence of $\hat{r}$ and $l'\hat{\beta}_{post}$. To proceed, let
denote the cumulative distribution function (CDF) of the $N(\mu, \sigma^{2})$ distribution truncated to a set $\mathcal{Z}$, where $\phi(z)$ denotes the probability density function (PDF) of the $N(0, 1)$ distribution. Letting $\widehat{\mathcal{Z}}_{post} = \mathcal{Z}_{post}(\hat{r})$, equation (ref) and the probability integral transform yields
The above holds for all $r$, so we can integrate out the residual $\hat{r}$ to obtain
Thus, conditional on $\hat{\beta} \in \mathcal{B}_{post}$, the probability of observing $F_{TN}(l'\hat{\beta}_{post}; l'\beta_{post}, \sigma_{post}^{2}, \widehat{\mathcal{Z}}_{post}) \geq 1-\alpha$ is equal to $\alpha$. The truncated normal distribution has strict monotone likelihood ratio in $\mu$, and hence its CDF is strictly decreasing in the potential values $\mu \in \mathbb{R}$ of $l'\beta_{post}$. Thus, given $\hat{\beta} \in \mathcal{B}_{post}$, there exists unique $\hat{\mu}_{\alpha}^{*}$ for which
The estimator $\hat{\mu}_{\alpha}^{*}$ is quantile conditionally unbiased in the sense of criteria ((ref)):
Based on estimator $\hat{\mu}_{\alpha}^{*}$, the proposed conditional inference procedures are as follows.
\paragraph{Confidence Intervals.} We propose $CI_{\alpha}^{*}(X) =[\hat{\mu}_{\alpha/2}^{*}, \hat{\mu}_{1-\alpha/2}^{*}]$ as confidence intervals for $l'\beta_{post}$, which yields
That is, $CI_{\alpha}^{*}(X)$ has correct conditional coverage for $l'\beta_{post}$.
\paragraph{Point Estimators.} We propose $\hat{\mu}_{1/2}^{*}$ as point estimators for $l'\beta_{post}$, which yields
That is, the conditional median of $\hat{\mu}_{1/2}^{*}$ is equal to $l'\beta_{post}$.
We now highlight two statistical properties of $\hat{\mu}_{\alpha}^{*}$. First, the estimator $\hat{\mu}_{\alpha}^{*}$ is optimal in the class of quantile conditionally unbiased estimators, in a broad sense formalized by Pfanzagl1994. A statement of this result in our setup and notation is as follows.
Thus, for a broad class of loss functions, the estimator $\hat{\mu}_{\alpha}^{*}$ yields lower expected loss than any other quantile conditionally unbiased estimator $\hat{\mu}_{\alpha}$. This means there is limited scope for improving upon $\hat{\mu}_{\alpha}^{*}$ for conditional inference: $\hat{\mu}_{1/2}^{*}$ is an optimal point estimator for $l'\beta_{post}$ and $CI_{\alpha}^{*}(X) =[\hat{\mu}_{\alpha/2}^{*}, \hat{\mu}_{1-\alpha/2}^{*}]$ is an optimal (equal-tailed) confidence interval for $l'\beta_{post}$.
The second property is that conditional inferences with $\hat{\mu}_{\alpha}^{*}$ will agree with conventional unconditional inferences when the conditioning event $\{\hat{\beta} \in \mathcal{B}_{post}\}$ occurs with high probability. That is, for deviations where the use of conventional point estimators $l'\hat{\beta}_{post}$ and confidence intervals $[l'\hat{\beta}_{post} \pm z_{1-\alpha/2}\sigma_{post}]$ would yield minimal distortions, the researcher does not pay a price when using the conditional analogues $(\hat{\mu}_{1/2}^{*}$, $CI_{\alpha}^{*}(X))$ instead.
To gain intuition for this property, note $\{\hat{\beta} \in \mathcal{B}_{post}\} = \{l'\hat{\beta}_{post} \in \widehat{\mathcal{Z}}_{post}\}$ and consider the case where $\mathcal{B}_{post} = \mathbb{R}^{\dim(\hat{\beta})}$, so that $\widehat{\mathcal{Z}}_{post} = \mathbb{R}$. In this case, formula (ref) yields
where $\Phi(z)$ is the CDF of the $N(0,1)$ distribution. This is solved by $\hat{\mu}_{\alpha}^{*} = l'\hat{\beta}_{post} + z_{\alpha}\sigma_{post}$. Thus, when $\mathbb{P}_{\beta}\left\{\hat{\beta} \in \mathcal{B}_{post}\right\} \approx 1$ so that $\mathcal{B}_{post} \approx \mathbb{R}^{\dim(\hat{\beta})}$, we expect $\hat{\mu}_{\alpha}^{*} \approx l'\hat{\beta}_{post} + z_{\alpha}\sigma_{post}$. We formalize this intuition in the following result, based on andrews2024inference.
To compute $\hat{\mu}_{\alpha}^{*}$ we focus on conditioning events $\hat{\beta} \in \mathcal{B}_{post}$ such that, for a finite set of matrices $A_{post, k}$ and cutoff vectors $c_{post, k}$ yielding polyhedra $\{\hat{\beta}: A_{post, k} \hat{\beta} \leq c_{post, k}\}$, $k=1,\ldots, K$, we have
That is, we restrict attention to deviations that arise from the estimates $\hat{\beta}$ falling into a union of polyhedra, where the cutoff vectors $c_{post, k}$ are implicitly allowed to depend on $\Sigma$. This setup yields computationally tractable $\hat{\mu}_{\alpha}^{*}$, and broadly accommodates the types of deviations that occur in practice, such as those based on statistical significance cutoffs.
\paragraph{Statistical Significance Cutoffs.} Suppose $\hat{\beta}_{post}$ is of interest when the linear combination $v'\hat{\beta}_{pre}$ of PAP estimates passes some significance threshold governed by $\eta$. Let $v'\hat{\beta}_{pre} = v_{pre}'\hat{\beta}$, where $v_{pre}$ denotes the vector induced by (i) matrix multiplication to select $\hat{\beta}_{pre}$ from $\hat{\beta}$ and (ii) vector multiplication of $\hat{\beta}_{pre}$ by $v$. A two-sided statistical significance cutoff yields
where $A_{post,1} = v_{pre}'$, $A_{post,2} = -v_{pre}'$, and $c_{post,1} = c_{post,2} = -z_{1-\eta/2}\sigma_{pre}$. On the other hand, a two-sided statistical insignificance cutoff yields
where $A_{post} = (-v_{pre}, v_{pre})'$ and $c_{post} = (z_{1-\eta/2}\sigma_{pre}, z_{1-\eta/2}\sigma_{pre})'$.
Under polyhedral conditioning, we obtain a convenient representation for the truncation set $\widehat{\mathcal{Z}}_{post}$, based on lee2016exact. In what follows, we define the maximum over the empty set as $-\infty$ and the minimum over the empty set as $+\infty$.
In addition to yielding computationally tractable estimators $\hat{\mu}_{\alpha}^{*}$, the polyhedral structure in (ref) allow us to better interpret the truncation set $\widehat{\mathcal{Z}}_{post}$, and hence the behavior of $\hat{\mu}_{\alpha}^{*}$ along relevant dimensions of the inference problem. The following section illustrates this behavior when the deviation event is characterized by a one-sided significance test.
Consider inference on post estimand $\beta_{post} \in \mathbb{R}$ conditional on PAP estimate $\hat{\beta}_{pre} \in \mathbb{R}$ crossing a one-sided $\eta = 0.05$ significance cutoff:
Here $\rho$ is the correlation between $\hat{\beta}_{pre}$ and $\hat{\beta}_{post}$, and we have assumed $\sigma_{pre} = \sigma_{post} = 1$. Let
denote the corrected (i.e., conditional) 95% interval. Figure (ref) depicts how this interval (in green) varies as the realized $\hat{\beta}_{pre}$ moves further from the reporting threshold of $1.96$. To give a sense of magnitudes, we (i) mark significance cutoffs $z_{1-\eta}$ on the horizontal axis for different values of $\eta$ and (ii) depict the conventional 95% interval (in orange). Each panel depicts this behavior over different values of the correlation $\rho$ between $\hat{\beta}_{pre}$ and $\hat{\beta}_{post}$.
There are two major takeaways from these plots of what drives distortions. (i) Distance: The closer $\hat{\beta}_{pre}$ is to the cutoff of the conditioning event (e.g., the distance between $\hat{\beta}_{pre}$ and 1.96), the greater the distortion from the conventional intervals. By contrast, when the realized $\hat{\beta}_{pre}$ is far above the cutoff, accounting for conditional reporting has a negligible effect. (ii) Strength of Correlation: When the dependence (governed by $\rho$) between $\hat{\beta}_{post}$ and the conditioning event $\{\hat{\beta}_{pre} \geq 1.96\}$ grows, the distortion becomes more pronounced. As seen in the first panel, the corrected and conventional intervals are equal when $\hat{\beta}_{post}$ and $\hat{\beta}_{pre}$ are independent ($\rho=0$).
We have thus far assumed that the vector $\hat{\beta}$ of PAP and post estimates is normally distributed with mean $\beta$ and a known covariance matrix $\Sigma$. This assumption is motivated by large-sample asymptotic results that yield asymptotic normality of estimates (e.g., under the central limit theorem) and consistency of covariance matrix estimators (e.g., under the law of large numbers). These asymptotic results underlie conventional procedures used for unconditional inference, which replace $(\hat{\beta}, \Sigma)$ with analogues $(\hat{\beta}_{n}, \widehat{\Sigma}_{n})$ constructed from a sample of size $n$. Following this same logic, Section (ref) shows how to implement plug-in versions of the conditional inference procedures proposed in Section (ref). In Section (ref), we establish the uniform asymptotic validity of these plug-in procedures over a broad class of probability distributions as $n \overset{}{\rightarrow}\infty$. We prove this asymptotic validity in Appendix (ref).
Given a sample of size $n$ from some unknown distribution $P_{n}$, the researcher constructs PAP and post estimates, which we collect into a vector $\hat{\beta}_{n}$. The researcher also constructs a corresponding covariance matrix estimator $\widehat{\Sigma}_{n}$. The researcher is interested in a linear combination $l'\hat{\beta}_{post,n}$ of some vector of post estimates $\hat{\beta}_{post,n}$. As before, we let $l_{post}$ denote the vector induced by (i) matrix multiplication to select $\hat{\beta}_{post,n}$ from $\hat{\beta}_{n}$ and (ii) vector multiplication of $\hat{\beta}_{post,n}$ by $l$, so that $l'\hat{\beta}_{post,n} = l_{post}'\hat{\beta}_{n}$. The corresponding variance estimator is $\hat{\sigma}_{post,n}^{2} = l_{post}'\widehat{\Sigma}_{n}l_{post} = l'\hat{\Sigma}_{post,n} l$. As we formalize below in Section (ref), we can think of $\hat{\beta}_{n}$ as corresponding to some underlying estimand $\beta(P_{n})$. Thus, the goal is to conduct inference on parameter $l'\beta_{post}(P_{n}) = l_{post}'\beta(P_{n})$.
The feasible conditional inference procedures are plug-in versions of the procedures from Section (ref). That is, the feasible plug-in procedures replace all expressions that depend on $(\hat{\beta}, \Sigma)$ in Section (ref) with analogous expressions that depend on $(\hat{\beta}_{n}, \widehat{\Sigma}_{n})$. We explicitly describe these plug-in procedures in Section (ref), and provide a concrete example in Section (ref).
\fbox{Step 1. Derive matrices and vectors $(A_{post,k}, \hat{c}_{post,k,n})$, $k = 1, \ldots, K$.}
Let $X_{1:n} \in \mathcal{X}_{n}$ represent the sample data. The researcher deviates to post estimates $\hat{\beta}_{post,n}$ when $X_{1:n} \in \mathcal{X}_{post,n}$, where
The polyhedra $\{\hat{\beta}_{n}: A_{post,k}\hat{\beta}_{n}\leq \hat{c}_{post,k,n}\}$, $k = 1, \ldots, K$, are based on nonrandom matrices $A_{post,k}$ and random cutoff vectors $\hat{c}_{post,k,n}$ that may depend on $\widehat{\Sigma}_{n}$.
\fbox{Step 2. Compute truncation set $\widehat{\mathcal{Z}}_{post,n}$.}
To account for conditioning, compute the residual
and compute the truncation quantities
As in Proposition (ref), the event $\{\hat{\beta}_{n}\in \hBpostn\}$ is equivalent to $\{l'\hat{\beta}_{post,n} \in \widehat{\mathcal{Z}}_{post,n}\}$, where
\fbox{Step 3. Solve for the plug-in quantile unbiased estimator $\hat{\mu}_{\alpha,n}^{*}$.}
Given $\hat{\beta}_{n}\in \hBpostn$, the plug-in quantile unbiased estimator is the unique $\hat{\mu}_{\alpha,n}^{*}$ that solves
Computation is fast, since this amounts to finding the root of strictly monotone function. The corresponding conditional inference procedures are
Section (ref) establishes the asymptotic validity of these feasible plug-in procedures for inference on $l'\beta_{post}(P_{n})$ as $n \overset{}{\rightarrow}\infty$.
Consider $\hat{\beta}_{n}= (\hat{\beta}_{pre,n}, \hat{\beta}_{post,n})' \in \mathbb{R}^{2}$ and $\beta(P_{n}) = (\beta_{pre}(P_{n}), \beta_{post}(P_{n}))' \in \mathbb{R}^{2}$. We want inference on post estimand $\beta_{post}(P_{n})$ conditional on the PAP estimate $\hat{\beta}_{pre,n}$ crossing a two-sided statistical significance cutoff governed by $\eta \in (0,1)$:
where $\hat{\sigma}_{cov,n}$ is the covariance estimator for $\hat{\beta}_{pre,n}$ and $\hat{\beta}_{post,n}$. In this case, $l = l_{post} = (0,1)'$.
\fbox{Step 1. Derive matrices and vectors $(A_{post,k}, \hat{c}_{post,k,n})$, $k = 1, \ldots, K$.}
In this case, the set $\hBpostn$ takes the form of (ref):
Thus, $K=2$ and
\fbox{Step 2. Compute truncation set $\widehat{\mathcal{Z}}_{post,n}$.}
The regression step yields
For $k=1$, we obtain $\widehat{Z}^{0}_{post,1,n} = 0$ and
For $k=2$ we obtain $\widehat{Z}^{0}_{post,2,n} = 0$ and
Thus, $\{|\hat{\beta}_{pre,n}| \geq z_{1-\eta/2}\hat{\sigma}_{pre,n}\} = \{\hat{\beta}_{post,n} \in \widehat{\mathcal{Z}}_{post,n}\}$, where
\fbox{Step 3. Solve for the plug-in quantile unbiased estimator $\hat{\mu}_{\alpha,n}^{*}$.}
Given $|\hat{\beta}_{pre,n}| \geq z_{1-\eta/2}\hat{\sigma}_{pre,n}$, formula (ref) yields
The estimator $\hat{\mu}_{\alpha,n}^{*}$ is the unique solution to $F_{TN}(\hat{\beta}_{post,n}; \hat{\mu}_{\alpha,n}^{*}, \hat{\sigma}_{post,n}^{2}, \widehat{\mathcal{Z}}_{post,n}) = 1-\alpha$.
We now show that inference based on $\hat{\mu}_{\alpha,n}^{*}$ is uniformly asymptotically valid, formally stated in equation (ref) below. Existing results by andrews2024inference and mccloskey2024hybrid can be used to prove uniformity the case of $K=1$. Here, we prove the general case for $K \geq 1$. In the proof we assume the rows of $A_{post,k}$ are nonzero, i.e., $(A_{post,k})_{j} \neq 0$ for all $(j,k)$, which rules out, for example, a deviation event of the form $\{0 \leq \hat{\sigma}_{pre,n} - \hat{\sigma}_{post,n}\}$.
\paragraph{Environment.} We suppose that a sample of size $n$ is drawn from some unknown distribution $P_{n} \in \mathcal{P}_{n}$, where $\mathcal{P}_{n}$ is a class of probability distributions corresponding to a sample of size $n$. For example, given a class of distributions $\mathcal{P}_{0}$ with bounded moments, if we have i.i.d. draws of size $n$ from some fixed $P_{0} \in \mathcal{P}_{0}$, then $P_{n} = (P_{0})^{n}$ is the product distribution and $\mathcal{P}_{n}$ is the set of all products of a fixed distribution with bounded moments. As we take $n \overset{}{\rightarrow}\infty$, there is a corresponding sequence of unknown distributions $\{P_{n}\} \in \times_{n=1}^{\infty}\mathcal{P}_{n}$, where $\times_{n=1}^{\infty}\mathcal{P}_{n}$ is the sequence of distribution classes, and the notation $\{P_{n}\} \in \times_{n=1}^{\infty}\mathcal{P}_{n}$ means that $P_{n} \in \mathcal{P}_{n}$ for all $n$. Below we let $P \in \cup_{n=1}^{\infty}\mathcal{P}_{n}$ index elements belonging to $\mathcal{P}_{n}$ for some $n$.
\paragraph{Asymptotic Normality.} We first assume that the scaled estimates $\tilde{\beta}_{n}= \sqrt{n}\hat{\beta}_{n}$ are uniformly asymptotically normal. Let $BL_{1}$ denote the set of real-valued functions that are bounded above in absolute value by one and have Lipschitz constant bounded above by one van1996weak. Furthermore, given a candidate covariance matrix $\Sigma$, let $\lambda_{min}(\Sigma)$ and $\lambda_{max}(\Sigma)$ denote its minimum and maximum eigenvalues.
Assumption (ref) requires that $\sqrt{n}(\hat{\beta}_{n}- \beta(P))$ converges to $N(0, \Sigma(P))$ in bounded Lipschitz metric, uniformly in $P$. This is a standard way to define uniform convergence in distribution.\footnote{Examples include andrews2024inference and mccloskey2024hybrid.} For example, if the components of $\hat{\beta}_{n}$ are sample averages of unit-level observations, then uniform convergence follows from bounds on the moments of the observations and bounds on dependence across observations. Assumption (ref) also requires the eigenvalues of $\Sigma(P)$ to be uniformly bounded above and away from zero, which ensures that $\sqrt{n}(\hat{\beta}_{n}- \beta(P))$ is stochastically bounded with nonzero asymptotic variance.
\paragraph{Consistent Covariance Matrix Estimation.} We next assume that the scaled covariance matrix estimator $\widetilde{\Sigma}_{n}= n\widehat{\Sigma}_{n}$ is consistent for $\Sigma(P)$, uniformly over $P$.
If $\widehat{\Sigma}_{n}$ is appropriate for the setting at hand (e.g., sample covariance for iid data, long-run covariance for time series data), then Assumption (ref) follows from the same kind of sufficient conditions that justify Assumption (ref).
\paragraph{Consistent Cutoff Vector Estimation.} Finally, we assume that the scaled cutoff vectors $\tilde{c}_{post,k,n}= \sqrt{n}\hat{c}_{post,k,n}$ are uniformly consistent, in similar fashion to $\widetilde{\Sigma}_{n}$.
For example, often the cutoff vectors take the form $\hat{c}_{post,k,n}= q_{post,k}\sqrt{C_{post,k}'\widehat{\Sigma}_{n}C_{post,k}}$, where $C_{post,k} \neq 0$ and $q_{post,k} \neq 0$ are nonrandom vectors. Assumption (ref) then implies
where the first inequality follows from $|\sqrt{x} - \sqrt{y}| \leq \sqrt{|x - y|}$ for $x, y \geq 0$, and the last equality follows from the matrix operator norm bound. Thus, under Assumption (ref), the above $\tilde{c}_{post,k,n}= q_{post,k}\sqrt{C_{post,k}'\widetilde{\Sigma}_{n}C_{post,k}}$ satisfies Assumption (ref) with $c_{post,k}(P) = q_{post,k}\sqrt{C_{post,k}'\Sigma(P) C_{post,k}}$.
\paragraph{Uniform Asymptotic Validity.} Let $\beta_{post}(P)$ denote the components of the centering vector $\beta(P)$ corresponding to $\hat{\beta}_{post,n}$. Under Assumptions (ref)-(ref), we show in Appendix (ref) that inference based on $\hat{\mu}_{\alpha,n}^{*}$ is uniformly asymptotically valid in the sense that
That is, $\hat{\mu}_{\alpha,n}^{*}$ satisfies an asymptotic analogue of the quantile conditional unbiasedness criteria (ref) for the normal model $\hat{\beta} \sim N(\beta, \Sigma)$.\footnote{The multiplication by $\mathbb{P}_{P}\left\{\hat{\beta}_{n}\in \hBpostn\right\}$ accounts for sequences where the probability of the conditioning event converges to zero. This a standard way to account for such cases andrews2024inference, mccloskey2024hybrid.} The required uniformity in the class of distributions $\mathcal{P}_{n}$ is the asymptotic analogue of requiring that procedures in the normal model be valid regardless of the unknown mean $\beta$.\footnote{The impossibility results of leeb2006can do not apply here, since the above approach does not attempt to consistently estimate the conditional distribution of $l'\hat{\beta}_{post,n}$ given deviation. Rather, it accounts for conditioning by using the plug-in analogue of pivotal quantity (ref) for $l'\beta_{post}$ from the normal model.}
\paragraph{Confidence Intervals.} The above convergence implies uniformly valid conditional coverage for our proposed confidence intervals:
which is the asymptotic analogue of criteria (ref) from the normal model.
\paragraph{Point Estimators.} The above convergence also implies uniformly valid conditional median-unbiasedness for our proposed point estimators:
which is the asymptotic analogue of criteria (ref) from the normal model.
We now return to Example (ref) and apply our approach to the empirical results of bessone2021economic, using the authors' stated reasons for deviating to consider different possible deviation sets. Recall that the authors cite three reasons for reporting the non-preregistered pooled coefficient $\hat{\tau}_{N+NE+NI}$:
With these three plausible deviation sets in mind, we can consider the empirical consequences of correcting for the conditional reporting of $\hat{\tau}_{N+NE+NI}$ by conditioning on sequential intersections of these events. Table (ref) shows the corrected 95% confidence intervals for the three discussed deviation sets alongside the original CI reported in the paper. The same significance cutoff of $\eta=0.05$ is used for all constraints in (A), (B), and (C). We show in subsequent tables how the choice of significance cutoff $z_{1-\eta/2}$ used for (A), (B), and (C) impacts these CIs.
\paragraph{Conditioning on (A) vs. No Conditioning.} We first observe that conditioning on (A) alone yields a $CI^*_{0.95}$ almost exactly equal to the conventional $CI_{0.95}$. Observe that (A) is made up of four inequalities:
Recall from the toy example of a single constraint from Figure (ref) that the corrected confidence intervals $CI^*_{0.95}$ converge to conventional confidence intervals at a sufficient distance from the cutoff for reporting. To contextualize this distance, consider the constraint closest to binding in the data, $$\hat{\tau}_{N} - \hat{\tau}_{NI} \leq 1.96\cdot \text{se}(\hat{\tau}_{N}-\hat{\tau}_{NI}).$$ Loosely speaking, the unchanged confidence intervals in this case reflect that the realized estimates in bessone2021economic are sufficiently “far” from the boundary of $\mathcal{X}_{post}^A$.\footnote{The corrected intervals $CI^*_{0.95}$ also do not change for different choices of significance cutoffs, $z_{1-\eta}$, for $\alpha \in\{0.01,0.05,0.1\}$ (not shown given redundancy).} To provide a sense of the scale of this distance, we can consider an exercise of holding estimates $\hat{\tau}_{N}, \hat{\tau}_{NI}$ fixed and varying $\text{se}(\hat{\tau}_{N}-\hat{\tau}_{NI})$ through $\hat{\sigma}^2_{NI}$ alone. To see any appreciable difference between $CI_{0.95}^*$ and the conventional $CI_{0.95}$, one would need $\hat{\sigma}^2_{NI}$ to be around 5 times smaller than its observed value.
\paragraph{Conditioning on (A) and (B).} Adding the condition that both $\hat{\tau}_E$ and $\hat{\tau}_I$ are near-zero imposes a constraint which is much closer to binding, and has an appreciable impact on the result of conditioning. For further intuition, Table (ref) shows how the confidence intervals change when one varies the the threshold for significance in (B), i.e., $z_{1-\eta/2}$.
In the data, this change is driven by the insignificance constraint on $\hat{\tau}_E$, which grows closer to binding as $\eta=0.1$ increases, i.e., the standard for significance becomes more lax.
\paragraph{Conditioning on (A), (B), and (C).} As with the addition of constraint (A) relative to no conditioning, the addition of constraint (C) to $\{$(A) and (B)$\}$ does not have a meaningful impact. This is because $\hat{\tau}_N$ is sufficiently far from the boundary of significance for conditioning to have a meaningful impact---the same holds true for conditioning either on $\{$(A) and (C)$\}$ or on (C) alone. Overall, the above empirical results demonstrate that, depending on the reason for deviating, the adjustments from our procedures can range from having no difference to an economically significant difference relative to conventional practice.
We have so far assumed that for any $\hat{\beta}_{post} = S_{post}(X)$, the reported deviation set $\mathcal{X}_{post} = \{X \in \mathcal{X}: S_{post} \in \mathcal{S}_{post}^{X}\}$ is correct. This assumption is justified when the researcher is honest and capable of articulating $\mathcal{X}_{post}$. In some settings, however, either of these assumptions may seem implausible. For example, an earnest researcher may face cognitive or communication costs when reporting $\mathcal{X}_{post}$ (much resembling the costs faced during initial PAP specification). Alternately, a nefarious researcher may know their true $\mathcal{X}_{post}$, but strategically report a different deviation set (e.g., to obtain shorter confidence intervals).
In this section, we consider the possibility that the researcher reports some deviation set different from the truth, $\Tilde{\mathcal{X}}_{post}\neq\mathcal{X}_{post}$, leading to potentially biased point estimators $\hat{\mu}_{1/2}^{*}(X,\Tilde{\mathcal{X}}_{post})$ and confidence intervals $CI_{\alpha}^{*}(X,\Tilde{\mathcal{X}}_{post})$ with incorrect coverage. We refer to $\Tilde{\mathcal{X}}_{post}$ as the reported deviation set. In cases where the reported deviation set varies with the realized data $x\in\mathcal{X}$, $\Tilde{\mathcal{X}}_{post}(x) \subseteq \mathcal{X}$ is defined more generally as a correspondence $\Tilde{\mathcal{X}}_{post} : \mathcal{X} \rightrightarrows \mathcal{X}$.\footnote{Note that defining results in terms of this correspondence abstracts from underlying assumptions about the researcher (e.g., incentives, utility, accuracy) generating the reporting behavior. Therefore, while we will point out possible assumptions about the researcher that generate certain reporting behavior (i.e., properties of correspondence $\Tilde{\mathcal{X}}_{post}(x)$), our results are agnostic to these assumptions.}
To frame upcoming discussion, we open with an impossibility result: Proposition (ref) states that no non-trivial conditional inference procedure can guarantee valid conditional coverage when the researcher is fully unrestricted in reporting $\Tilde{\mathcal{X}}_{post}$. For instance, no non-trivial procedure can insure against a nefarious researcher strategically choosing a particular $\Tilde{\mathcal{X}}_{post}$ to exclude some $\beta_{0}$.
Reframed, Proposition (ref) implies that some form of additional structure on reporting behavior must be imposed to guarantee valid coverage. Therefore, while the assumptions in previous sections of the researcher being honest and accurate in specifying their deviation sets may feel unpalatably strong, some type of assumption on $\Tilde{\mathcal{X}}_{post}(x)$ must be made. Observe also that a reinterpretation of this result can be understood as implying that an arbitrarily skeptical reviewer can always find an $\Tilde{\mathcal{X}}_{post}'$ which invalidates the researcher's reported results, should they wish to. In this sense, structure on allowable deviation sets also allows the researcher to defend their results against arbitrarily unfavorable counter-assertions of their “true” deviation set, which we demonstrate more concretely in Section (ref).
We now discuss two possible sources of such structure: Section (ref) considers partial reports $\Tilde{\mathcal{X}}_{post} \subseteq \mathcal{X}_{post}$, and Section (ref) presents sensitivity analysis for assessing the robustness of results to $\Tilde{\mathcal{X}}_{post}\neq\mathcal{X}_{post}$, more generally.
It may happen that an honest researcher can describe their deviation behavior local to the realized $x\in\mathcal{X}$, but has difficulty articulating deviation behavior across all of $\mathcal{X}_{post}$. For instance, when $\mathcal{X}_{post}$ consists of many disjoint components, the researcher may have a clear understanding of components “local” to the realized data draw $x\in\mathcal{X}$, but not for those “distant” from $x$. This disjoint structure is particularly likely when there are many potential motives for reporting $\hat{\beta}_{post}$, but only certain motives are relevant at any given draw $X\in\mathcal{X}$. In such a setting, articulating the full $\mathcal{X}_{post}$ requires the researcher to enumerate all hypothetical motives for reporting $\hat{\beta}_{post}$, otherwise reporting an incomplete description of $\Tilde{\mathcal{X}}_{post} \subseteq \mathcal{X}_{post}.$
We can show that under certain “local consistency” conditions in reporting behavior for some partial component $\Tilde{\mathcal{X}}_{post}\subseteq\mathcal{X}_{post}$ containing the data realization (i.e., $X\in\Tilde{\mathcal{X}}_{post}$), estimators $\hat{\mu}_{\alpha}^{*}(X,\Tilde{\mathcal{X}}_{post})$ and confidence intervals $CI_{\alpha}^{*}(X,\Tilde{\mathcal{X}}_{post})$ that condition on said partial component yield valid inferences. Note, however, that these inferences will be less precise than those based on procedures that condition on the full $\mathcal{X}_{post}$. For concreteness, we discuss this result in terms of a stylized example.
The following proposition gives us that under certain reporting conditions, it is valid to condition on just $\mathcal{X}_1$ when $X\in\mathcal{X}_1$ and $\mathcal{X}_2$ when $X\in\mathcal{X}_2$.
The crucial requirement here is that the researcher behaves “locally coherently” in reporting the incomplete deviation set $\mathcal{X}_m$ for all $X\in\mathcal{X}_m$. To be concrete, consider again the partial deviation sets from Example (ref). Define $\mathcal{X}_{1}\subseteq \mathcal{X}_{\mathrm{post}}$ by the following two sample-based conditions:
Together, these define \[ \mathcal{X}_{1} =\Big\{\,X\in\mathcal{X}:\; \widehat{\operatorname{Var}}_{j}\!\left(\overline{\text{PeerScore}}_{j}\right)>\delta \ \text{and}\ p(\hat\rho)>0.05 \Big\}. \]
This form of “coherence” can conceptually be broken into two conditions. The first condition is that for every data draw $X\in\mathcal{X}_{1}$, the researcher reports the same $\mathcal{X}_{1}$ as their deviation set. For this example, the first condition would be violated if there existed some counterfactual draw $X\in\mathcal{X}_{1}$ for which the researcher reported a different threshold $\delta'\neq\delta$ or different significance cutoff $\eta'\neq 0.05$ for $\hat{\rho}$.
The second condition requires that the partial reports $\{\mathcal{X}_m\}_m$ be disjoint subsets of the sample space $\mathcal{X}$, i.e., $\mathcal{X}_m \cap \mathcal{X}_{m'} = \varnothing$ for all $m \neq m'$. Disjointness fits settings where the researcher has multiple, distinct interpretive frames for reporting $\hat{\beta}_{post}$, each tied to a separate region of the data, so that at any realization $X$ only one interpretation applies. This requirement is naturally satisfied when interpretations are mutually exclusive, as in Example (ref).
In settings where the deviation set takes the form of a cutoff rule, the researcher may not be able to discern the exact value of the cutoff. In this sense, the true cutoff is fuzzy. For example, suppose $\hat{\beta}_{post}$ is of interest because a PAP estimate $\hat{\beta}_{pre}$ was observed to be small. That is, there exists some cutoff $\kappa_{0} > 0$ for which
However, the exact value of $\kappa_{0}$ may not be obvious to the researcher. After some introspection, the researcher reports $\tilde{\kappa}$ as an approximation to $\kappa_{0}$. This yields reported deviation set
In such settings, a natural robustness exercise is to plot conditionally valid intervals for a range of $\kappa$ above and below the reported $\tilde{\kappa}$. Formally, given $\varepsilon > 0$, let $\Tilde{\mathcal{K}}_{\varepsilon}$ be a set of $\kappa$ such that $\tilde{\kappa} - \varepsilon \leq \kappa \leq \tilde{\kappa} + \varepsilon$ for each $\kappa \in \Tilde{\mathcal{K}}_{\varepsilon}$. One can plot
This allows one to assess the sensitivity of conclusions to different potential values of $\kappa_{0}$.
This robustness exercise is formally justified under the condition that $|\tilde{\kappa} -\kappa_{0}| \leq \varepsilon$ for known $\varepsilon$. Under this condition, if the conclusions that the researcher reaches with $CI_{\alpha}^{*}(X, \Tilde{\mathcal{X}}_{post})$ can also reached with $CI_{\alpha}^{*}(X,\mathcal{X}_{\kappa})$ for each $\kappa \in \Tilde{\mathcal{K}}_{\varepsilon}$, then we know the researcher would reach those conclusions with the true $CI_{\alpha}^{*}(X,\mathcal{X}_{0})$. While knowledge of $\varepsilon$ may be a strong assumption, one can in practice vary $\varepsilon$ to determine the largest $\varepsilon$ for which the researcher's conclusions persist. Intuitively, if we believe that $\tilde{\kappa} \approx \kappa_{0}$, then we expect $\kappa_{0}$ to fall into $\Tilde{\mathcal{\kappa}}_{\varepsilon}$ for some $\varepsilon$ that is not too large. Plausible departures from the reported $\tilde{\kappa}$ ought to not lead to massive changes in results.
As intuition for this exercise, see Figure (ref), which considers $(\hat{\beta}_{pre},\hat{\beta}_{post})$ from Example (ref), where $\rho=0.25$. Recall the plots from Figure (ref). We consider the exercise of fixing a particular data realization $(\hat{\beta}_{pre},\hat{\beta}_{post})$ and plotting the estimates and confidence intervals as a function of cutoff $\tilde{\kappa}$. The green lines show the corrected $CI^*_{0.95}$ for $\hat{\beta}_{post}$, with the first panel corresponding to $\mathcal{X}_{post} = \{\hat{\beta}_{pre} \geq \tilde{\kappa}\}$. As one can see, the reported $\tilde{\kappa}$ can yield very different intervals: as one reports $\tilde{\kappa}$ further and further from $\hat{\beta}_{pre}$, the confidence intervals move from including to excluding zero.
\paragraph{Implications for Reporting Conventions.} For deviations based on statistical significance cutoffs $\kappa_{\eta} = z_{1-\eta/2}\sigma_{pre}$, a conventional choice of significance level is $\eta = 0.05$. By sticking to convention, there is less need for the scrutiny in (ref). But for cutoffs with potentially no obvious conventions, such as those based on economic significance, the robustness exercise in (ref) is useful. To avoid such scrutiny, one can preregister definitions of economic significance for their primary outcomes. Note that the above analysis is also valid for any deviation set $\mathcal{X}_{0}$ known up to a finite set of fuzzy cutoffs $\kappa_{0}$. Moreover, while we focused on confidence intervals $CI_{\alpha}^{*}(X,\mathcal{X}_{\kappa})$, the same sensitivity analysis can be applied to point estimators $\hat{\mu}_{1/2}^{*}(\mathcal{X}_{\kappa})$.
This paper considers the statistical consequences of deviating from prespecified analysis. We first develop a general model of preregistration in the research process, which we use to demonstrate that PAP deviations can be viewed as a form of conditional reporting that, if left unacknowledged, yields invalid inference. Our framework yields two recommendations. First, researchers should adopt conditional inference as the relevant criteria for non-prespecified analysis. Second, given this conditional criteria, researchers should articulate their corresponding reasons for deviating, then leverage $\mathcal{X}_{post}$ to report corrected (i.e., conditional) inferences alongside conventional (i.e., unconditional) inferences. To this end, we provide general and tractable inference procedures for obtaining confidence intervals with correct conditional coverage and point estimators that are conditionally unbiased. We formalize our conditional inference objectives in a model with normally distributed estimates, and present general procedures for constructing confidence intervals and point estimators that are conditionally valid. We then show how to implement these procedures in practice, providing uniformity guarantees for non-normal estimates. Using data from bessone2021economic, we demonstrate that, depending on the deviation set and data realization, accounting for conditional reporting can range from having no difference to an economically significant difference relative to conventional practice. In particular, when the data puts one close to the boundary of reporting, the impact of conditioning can be large, whereas when the data realization is sufficiently far from the boundary, there is no change from unconditional inference. This framework has direct implications for considering the validity of past non-preregistered results reported in papers.
We conclude with a discussion of the robustness of our procedures to misspecification of the reported deviation set. Our results suggest possible directions for future work. In particular, this framework may have implications for the optimal length and detail of PAP specification, yielding possible rules of thumb for researchers and journals. This framework may also suggest new paradigms for adaptive experiments with data-driven selection.