Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
84,989 characters · 34 sections · 81 citation commands
Dynamically Consistent Statistical Decisions
\onehalfspacing
Statistical decision theory à la Wald_1950_SDT provides a framework for studying the ex ante choice of data-contingent actions. In the context of treatment choice, for instance, the decision maker (DM) must commit to a decision rule that either approves or rejects the treatment as a function of the realized data Manski_2004_ECTA, Stoye_2009_JoE. This ex ante perspective is most natural when the DM can commit ex ante to a decision rule before observing the realized data.
In practice, however, DMs often revise their intended action after seeing the data. For instance, a regulatory agency with a publicly listed approval rule for new drugs may reject a drug after receiving its clinical trial data, even when the data satisfies all the conditions required for approval under the rule. In the context of empirical research, a researcher may commit to using a nonparametric estimator in a pre-analysis plan, but instead opt for a more parametric specification after observing that the resulting estimate is too noisy. If these DMs had fully specified their data-contingent actions at the ex ante stage, such behavior constitutes instances of dynamic inconsistency: after observing the data, the DM wishes to deviate from the decision rule chosen ex ante.
A prerequisite for discussing notions of dynamic consistency is an interim optimality criterion, which describes the DM's preferences after the realization of data.\footnote{Our terminology distinguishes between the interim stage, which occurs after the data realization but before the realization of the parameter space, and the ex post stage, which occurs after the realization of all uncertainty.} These interim preferences determine whether the DM wishes to deviate from a given decision rule once the data has realized. In Bayesian analysis, the interim problem is well-defined by the DM's posterior beliefs after the realization of data. By contrast, in frequentist decision theory, the interim problem is less clearly defined. There are no beliefs to update in the Bayesian sense, and standard ex ante criteria, such as minimax regret, do not admit a natural decomposition into interim preferences.
The first main contribution of this paper is to provide a formal analysis of the interim problem associated with frequentist ex ante optimality criteria. This fills an important gap in the literature, given the prominence of decision rules justified by such criteria. Through a series of examples, we illustrate that dynamic inconsistency is not merely a theoretical possibility. Across a variety of empirically relevant settings, many “reasonable" researchers and policymakers may find it optimal, after observing the data, to deviate from the ex ante optimal rule.\footnote{Related concerns appear in the statistics literature (e.g., Chapter 1.6 of Berger_1985_SDT). }
The second main contribution of this paper is to propose and axiomatize a class of ex ante optimality criteria that are {dynamically consistent}, and therefore do not suffer from the aforementioned problems. These correspond to the ex ante preferences of a sophisticated agent in the sense of Strotz_ReStud_1955, who correctly anticipates the interim preferences she will have after the data realization. Normatively, such criteria are appealing because they select decision rules that are optimal not only ex ante, but also after each possible data realization.\footnote{Computationally, this means that optimal rules can be computed via backwards induction.} Hence, such decision rules are credible: the DM has no incentive to deviate from the action prescribed by the ex ante rule after the data realization. We show that our dynamically consistent optimality criteria nest, as special cases, the as-if optimization of Manski_2021_ECTA and Gamma${}^*$-minimax criterion of Lim_2026_PIA.
At a high level, the dynamically consistent optimality criteria we characterize can be described as semi-Bayesian. They resemble Bayesian criteria in that they evaluate decision rules by aggregating continuation payoffs across possible data realizations. They differ from ordinary Bayesian criteria, however, in the role assigned to prior beliefs. Here, the DM is assumed to have enough prior information to form an ex ante marginal distribution over the data, but the interim criterion used after observing the data need not be obtained by Bayesian updating of a prior over the state space. Instead, the DM can use any statistical procedure (e.g., worst-case evaluation over a confidence region) after the data is realized.
This distinction is important because the underlying state space may have “no concrete reality in terms of physical quantities that are easily accessible to intuition, and yet the phenomenon under study may be quite familiar to the investigator" Berger_1985_SDT. That is, requiring the DM to have well-defined prior beliefs over the state space, in which case the marginal equals the prior predictive, is a more restrictive assumption than assuming that she has a subjective marginal distribution. For instance, a development economist evaluating a cash-transfer program may be unable to specify a prior over the collection of all joint distributions of potential outcomes, but she may have historically grounded beliefs about the distribution of study results likely to arise.
Finally, we remark that our framework can also describe a Bayesian DM with prior beliefs who nevertheless performs frequentist procedures using the data. This interpretation captures a familiar feature of empirical practice, where researchers typically adhere to the frequentist paradigm despite having substantive prior information about the phenomenon under study.
Since the early contributions of Manski_2004_ECTA and Dehejia_JoE_2005, a growing literature in econometrics applies statistical decision theory to economically meaningful problems, such as treatment choice and policy learning. In the frequentist setting, finite-sample minimax-regret optimal treatment rules have been derived across a variety of settings Stoye_2009_JoE, tetenov2012statistical, yata2021optimal, chen2025note, guggenberger2026minimax. Many decision rules based on asymptotic minimax-regret guarantees have also been proposed hirano2009asymptotics, Kitagawa_Tetenov_2018_ECTA, athey2021policy, mbakop2021model, Chernozhukov_et_al_2025, Sun_2026_EWM. Recent work by fernandez2024robust, Giacomini_Kitagawa_Read_2025, and Christensen_et_al_Restud_2026 studies related problems from the robust Bayesian perspective.
The aforementioned papers focus primarily on ex ante optimality. Our contribution is to ask whether rules justified by such ex ante criteria remain optimal after the DM observes the data.\footnote{This question also connects to recent work on pre-analysis plans, which studies the value of ex ante commitment and the consequences of post-data deviations from preregistered analyses olken2015promises, kasy2024optimalpreanalysisplansstatistical, sarfati2025post.} In the context of robust Bayesian analysis, the issue of dynamic (in)consistency has been studied extensively across both the statistical and axiomatic decision theory literatures vidakovic2000gamma, ES_recursive, Klibanoff_Hanany_2007_TE, Klibanoff_Hanany_2009_BEJTE, Stoye_2012_AR, fernandez2024robust, Lim_2026_PIA. By contrast, the dynamic inconsistency of frequentist optimality criteria has received far less attention, likely because the corresponding interim problem is not well-defined. We address this gap by formalizing the interim problem associated with such criteria, while proposing a diagnostic for determining whether ex ante optimal decision rules provide credible prescriptions in the interim period.
More broadly, we contribute to work at the intersection of statistical and axiomatic decision theory. As advocated for in Stoye_2012_AR, one principled approach to selecting an optimality criterion is axiomatic: the practitioner first considers the behavioral characterizations of the preferences induced by each criterion, then selects the criterion whose implications appear the most intuitive. In this regard, Hayashi_2008_JET, Stoye_2011_JET,Stoye_2011_TD, Stoye_2012_AR, and andrews2026misspecification provide axiomatic foundations for various statistical decision criteria. We contribute to this literature by characterizing the behavioral foundations of our dynamically consistent optimality criteria, which correspond to the choices of a sophisticated DM who anticipates her interim preferences.
The remainder of the paper is organized as follows. Section (ref) formally introduces the model. Section (ref) presents our framework for studying the interim problem associated with frequentist minimax criteria and proposes the diagnostic procedure for assessing their interim credibility. Sections (ref) and (ref) present examples of this diagnostic exercise. Section (ref) defines and axiomatizes our class of dynamically consistent optimality criteria. Section (ref) compares the interim prescriptions of various optimality criteria using both real and simulated data. Section (ref) concludes. All proofs are in the appendix, unless stated otherwise.
We first introduce the statistical decision problem in general form. Let $\Theta$ denote the state space, $ Z$ the data (sample) space, and $\mathcal A$ the action space. Unless otherwise stated, these spaces are endowed with the Borel $\sigma$-algebras generated by their respective topologies, which will be made explicit in specific applications. Let $\Delta(\cdot)$ denote the set of Borel measures over the relevant space, and let $P_\theta\in\Delta (Z)$ be the sampling distribution under $\theta\in\Theta$, with corresponding density $p_\theta$.
A decision rule is a measurable map $\delta: Z\to\Delta(\mathcal A)$, where $\mathcal A$ is an action space and $\Delta(\mathcal A)$ allows for randomized actions. Let $\mathcal D$ denote the set of feasible decision rules, which we assume is closed under measurable pasting.\footnote{That is, for any $\delta_1,\delta_2\in\mathcal D$ and measurable $B\subseteq Z$, the rule that agrees with $\delta_1$ on $B$ and with $\delta_2$ on $B^c$ also belongs to $\mathcal D$. This condition is satisfied, for instance, when $\mathcal{D}$ is the collection of all decision rules $\delta: Z \to \Delta(\mathcal{A})$.} Given the payoff (i.e., negative loss) function $u:\mathcal A\times\Theta\to\mathbb R$, we write $u(\delta( z),\theta)$ to denote the payoff averaged over the randomized action $\delta( z)$. The expected payoff (i.e., negative risk) from rule $\delta$ in state $\theta$ is $U(\delta,\theta)=\mathbb{E}_\theta[u(\delta( z),\theta)]$, where $\mathbb{E}_\theta$ denotes integration against $P_\theta$. Regret is defined relative to the optimal feasible rule given oracle knowledge of $\theta$, such that $R(\delta,\theta) \equiv \sup_{d\in\mathcal D}U(d,\theta)-U(\delta,\theta).$
In Bayesian decision theory, the DM is assumed to have a prior $\pi \in \Delta(\Theta)$. The corresponding ex ante expected payoff from $\delta$ is $\mathbb{E}_\pi[U(\delta,\theta)]$, where $\mathbb{E}_\pi$ denotes integration against the prior $\pi$. The Bayes regret of $\delta$ is $\mathbb{E}_\pi[R(\delta,\theta)] = \mathbb{E}_\pi[\sup_{d\in\mathcal D} U(d,\theta)]-\mathbb{E}_\pi[U(\delta,\theta)]$. Since the first term does not depend on $\delta$, minimizing Bayes regret is equivalent to maximizing expected payoff. We refer to any such rule as the Bayes rule under prior $\pi$, denoted $\delta_\pi\in\arg\max_{\delta\in\mathcal D}\mathbb{E}_\pi[U(\delta,\theta)]$.
The robust Bayesian generalization allows for multiple priors.\footnote{See Giacomini_Kitagawa_Read_2025 for a recent review of robust Bayesian analysis.} Let $\Pi \subseteq\Delta(\Theta)$ denote the set of priors the DM regards as feasible. The Gamma-minimax loss (written in payoff form) and regret rules solve, respectively,
The distinction between loss- and regret-based rules is salient with multiple priors because the oracle term $\mathbb{E}_\pi[\sup_{d\in\mathcal D}U(d,\theta)]$ varies with $\pi$. With singleton $\Pi = \{\pi\}$, both rules reduce to the Bayes rule under $\pi$.
Two canonical frequentist optimality criteria can be obtained as special cases of the Gamma-minimax criteria. Minimax loss (i.e., maximin payoff) evaluates rules by their worst-case payoff across states and solves
while minimax regret does the same using regret to solve
These coincide with the Gamma-minimax criteria when $\Pi=\Delta(\Theta)$, since worst-case expectations over $\Delta(\Theta)$ reduce to worst-case evaluations over states. In this sense, minimax loss and regret can be viewed as fully prior-free. They are also maximally conservative in the sense that no restrictions are posed on the collection of distributions over $\Theta$, and each rule is evaluated at the least favorable state.
Let $\mathcal D( z)\equiv \{\delta( z)\in\Delta(\mathcal A):\delta\in\mathcal D \}$ denote the set of feasible actions after the realization of $z \in Z$. At each $z$, the DM's interim problem concerns the choice of an optimal action $a\in\mathcal D( z)$. An optimality criterion is dynamically consistent if, after observing $z$, the DM still finds it optimal to implement the action prescribed by the ex ante optimal decision rule.
In the Bayesian setting, the DM's interim problem after observing the realization of $ z\in Z$ is well-defined. Given the prior $\pi$, Bayes rule induces a posterior $\pi(\cdot\mid z)$, and the DM chooses $a_{\pi(\cdot\mid z)}$ to maximize the conditional expected payoff, such that $a_{\pi(\cdot\mid z)} \in \mathcal{A}^*( z) = \operatorname*{arg\,max}_{a\in \mathcal{D}( z)} \mathbb{E}_{\pi(\cdot\mid z)}[u(a,\theta)]$. As in the ex ante case, conditional expected regret yields the same optimal action. It is well-known that the Bayes rule $\delta_\pi$ is dynamically consistent with interim problem (see, e.g., Chapter 4 of Berger_1985_SDT). That is, we have that $\delta_\pi( z) \in \mathcal{A}^*( z)$ $P_\pi$-almost surely, where $P_\pi =\int_\Theta P_\theta \, d\pi(\theta)$ is the prior predictive over $ Z$.
With multiple priors, each $\pi\in \Pi$ is updated prior-by-prior, generating the posterior set $\Pi_ z \equiv \{\pi(\cdot\mid z):\pi\in\Pi \}$. The conditional Gamma-minimax loss (written in payoff form) and regret criteria solve, respectively,
Unlike in the single-prior setting, Gamma-minimax rules are not generally dynamically consistent vidakovic2000gamma, ES_recursive.\footnote{Lim_2026_PIA clarifies the source of this dynamic inconsistency and proposes an alternative ex ante criterion that restores dynamic consistency. We return to this point in Section (ref).}
For the frequentist criteria, we face a more fundamental difficulty: there is no canonical interim problem to compare the ex ante rule against. Minimax loss and minimax regret can be represented as $\Gamma$-minimax criteria with $\Pi=\Delta(\Theta)$, but updating the full simplex prior-by-prior does not yield a meaningful Bayesian analogue. If the realized data has positive likelihood/density at every state, then the prior-by-prior update of $\Delta(\Theta)$ remains entire simplex, i.e., no learning from the data occurs.\footnote{This would be the case, for instance, if $Z \sim N(\theta, I_m)$, such that $p_\theta(z) > 0$ for all $z \in \mathbb{R}^m$.} Otherwise, prior-by-prior updating is not well-defined, since there exist Dirac priors under which the realized data have probability zero.
In the following section, we formalize the interim problem that is dynamically consistent with frequentist ex ante optimality criteria. We thereby provide a framework for evaluating the interim credibility of decision rules selected for their ex ante frequentist minimax guarantees.
When solving for minimax optimal rules, a common proof strategy interprets the decision problem as a zero-sum game against Nature and invokes tools from game theory to solve for equilibrium.\footnote{See, e.g., the discussion in Section 4.1 of Stoye_2012_AR and Chapter 5 of Berger_1985_SDT.} In the game, Nature chooses the distribution $ \pi \in \Delta (\Theta)$ over states, while the DM chooses $\delta\in\mathcal D$. Under minimax loss and regret, respectively, Nature's payoff is $-U(\delta,\theta)$ and $R(\delta, \theta) = \sup_{d\in\mathcal D}U(d,\theta)-U(\delta,\theta)$. In either case, we then see that Nature's payoff is of the form $c(\theta)-U(\delta,\theta),$ where $c(\theta)=0$ for minimax loss and $c(\theta)=\sup_{d\in\mathcal D}U(d,\theta)$ for minimax regret.
If the corresponding zero-sum game admits a Nash equilibrium $(\delta^*,\pi^*)$, then $\delta^*$ is minimax optimal for the relevant criterion, and $\pi^*$ is a least favorable prior in the sense that
Note that the latter equality holds because $c(\theta)$ does not depend on the choice of $\delta$. In other words, we see that once a least favorable prior $\pi^*$ has been identified, both the minimax loss and minimax regret rules are Bayes against $\pi^*$.
The least favorable prior representation suggests a natural candidate for defining the interim criterion: update $\pi^*$ and solve the corresponding conditional Bayes problem. Under the game-theoretic interpretation, this is equivalent to Nature's initial choice of $\pi^* \in \Delta(\Theta)$ being irrevocable, such that the remaining uncertainty after the data realization is reflected by its posterior update.\footnote{This idea resembles the dynamically consistent updating rules studied by Klibanoff_Hanany_2007_TE,Klibanoff_Hanany_2009_BEJTE, which preserve the optimality of ex ante decisions after the realization of information. The distinction is that a least favorable prior is a global saddle-point object for the ex ante minimax problem, so their posterior updates need not coincide with the full set of posteriors that support dynamic consistency.}
Let $P_{\pi^*} \equiv \int_\Theta P_\theta \, d\pi^*$ denote the prior predictive distribution over $ Z$ induced by $\pi^*$. The following proposition shows that any interim optimality criteria dynamically consistent with frequentist minimax rules agrees with the conditional Bayes problem induced by the least favorable prior.
The proof of Proposition (ref) is formally analogous to the standard argument that Bayes rules are dynamically consistent. The interpretation, however, is quite different. In the Bayesian case, the prior $\pi$ is an exogenous parameter that describes the DM's beliefs, and dynamic consistency holds for any decision problem. By contrast, the least favorable prior $\pi^*$ need not be unique, and it is not an exogenous object. Rather, $\pi^*$ is specific to the decision problem at hand, depending on the both the optimality criterion and the available class of decision rules $\mathcal D$.
Proposition (ref) should be read as a rationalization of interim criteria that are dynamically consistent with the minimax rule, rather than as an analogue of Bayesian dynamic consistency in the usual preference-based sense epstein1993dynamically, Ghirardato_BEU. Given a least favorable prior $\pi^*$ that supports the ex ante minimax rule, updating $\pi^*$ produces a conditional Bayes problem for which the prescribed continuation action $\delta^*( z)$ is optimal. However, $\pi^*(\cdot\mid z)$ need not represent the DM's actual interim beliefs or preferences.\footnote{The same observation applies to Gamma-minimax criteria: an ex ante Gamma-minimax rule is supported by a least favorable prior, whose posterior provides a Bayes rationalization of the continuation action. But in the robust Bayesian setting, the prior set $\Pi$ is specified exogenously, and the DM's interim beliefs are represented by the posterior set $\Pi_ z$, rather than the posterior induced by some least favorable prior.}
This perspective suggests a way to assess the interim credibility of decision rules justified by ex ante minimax guarantees. Given a least favorable prior $\pi^*$ that rationalizes the ex ante minimax rule $\delta^*$, one can ask whether $\delta^*( z)$ aligns with the action a researcher would find compelling after observing $ z$. When the two disagree, the question becomes whether it is “reasonable" to make decisions as if $\pi^*(\cdot\mid z)$ represented the relevant interim uncertainty.\footnote{Also related is the ML-II approach to prior selection (e.g., Chapter 3.5.4 of Berger_1985_SDT).} As we discuss in Section (ref), a related exercise is considered in the conclusion of Stoye_2009_JoE.
Our diagnostic exercise is naturally most informative when the interim choice set $\mathcal D( z)$ is sufficiently rich. Otherwise, even an implausible posterior may appear reasonable simply because the set of available alternatives leaves little scope for disagreement. The next section shows that the collection of least favorable posteriors can indeed be “unreasonable" across various examples.
Finally, we note that even though least favorable priors are central to game-theoretic proofs of minimax optimality, their implications beyond this role as a proof device have received little attention. To the best of our knowledge, this paper is the first to use least favorable priors to formalize the interim problem for frequentist decision making.
We now specialize the model to a problem of binary treatment choice, following Stoye_2009_JoE. Let $T\in \{0,1\}$ denote treatment, and suppose each individual $j \in J$ has binary potential outcomes $y^j(t)\in \{0,1\}$ for $t \in \{0,1\}$.\footnote{Everything in this section can be generalized to the setting with arbitrary bounded potential outcomes. We assume binary potential outcomes for simplicity of exposition.} Assigning treatment $t$ to an individual induces the random variable $Y_t$. Because the payoff and the sampling distribution depend on the law of $(Y_0,Y_1)$ only through its two marginal means, we take the state to be the pair of success probabilities $\theta=(\mu_0,\mu_1)\in\Theta\equiv[0,1]^2$, with $\mu_t \equiv \Pr(Y_t=1)=\mathbb{E}[Y_t].$
The action space is $\mathcal A=[0,1]$, where $a\in[0,1]$ is the probability of assigning treatment $t=1$.\footnote{Equivalently, take $\mathcal{A} = \{0,1\}$ and consider randomized decision rules $\delta : Z \to \Delta(\mathcal{A})$.} The DM observes $ z=(t_n,y_n)_{n=1}^N \in Z\equiv \{0,1\}^{2N}.$ A treatment rule is then $\delta: Z\to \mathcal{A}$, and the expected payoff of rule $\delta$ in state $\theta=(\mu_0,\mu_1)$ is
If $\theta$ were known, the DM would assign treatment $1$ whenever $\mu_1>\mu_0$, treatment $0$ whenever $\mu_0>\mu_1$, and would be indifferent when $\mu_0=\mu_1$.
We consider a random-assignment design where treatment assignments are independent fair coin flips, such that $T_j$ is independent of potential outcomes and satisfies $P(T_j=1)=P(T_j=0)= \frac{1}{2}$. Conditional on $T_j=t$, the observed outcome $y_j$ is an independent realization of $Y_t$.
Let $N_t$ denote the number of sample observations assigned to treatment $t$, and let $\bar y_t$ be the corresponding sample average. The following fact restates Proposition 1(ii) of Stoye_2009_JoE in the present notation.
The least favorable prior used in Stoye_2009_JoE's (Stoye_2009_JoE) proof of the minimax optimality of $\delta^*$ is $\pi^* = \frac{1}{2}\delta_{(a,1-a)} + \frac{1}{2}\delta_{(1-a,a)}$ for some $a > \frac{1}{2}$. Even though the least favorable prior is not unique, least favorability imposes a common restriction.
By Proposition (ref), the posterior update of any least favorable prior rationalizes the action chosen by any interim optimality criteria that is dynamically consistent with minimax regret. As discussed in the previous section, the diagnostic exercise regarding the credibility of $\delta^*$ centers on whether this posterior provides a “reasonable" interim rationale. In the context of our treatment choice setup, Proposition (ref) implies that any least favorable posterior is supported on states with $|\mu_1-\mu_0|=d_N$, regardless of the realized data.
For concreteness, consider the example in the conclusion of Stoye_2009_JoE. Suppose $N=1100$, with $N_0=1000$ observations assigned to $t=0$ and $N_1=100$ observations assigned to $t=1$. Assume further that $n_0=550$ and $n_1=99$, such that $\bar y_0=0.55$ and $\bar y_1=0.99$. The estimated treatment effect is $\bar y_1-\bar y_0=0.44,$ with standard error approximately $\sqrt{ \frac{0.99(0.01)}{100} + \frac{0.55(0.45)}{1000} } \approx 0.019.$
A conventional comparison of sample success rates therefore provides overwhelming evidence in favor of treatment $1$. However, the minimax regret rule $\delta^*$ assigns treatment $0$, since
We therefore see that the action $\delta^*( z) = 0$ prescribed by the ex ante optimal rule disagrees with the action that most DMs would take after observing $ z$.
The posterior rationale for the action $\delta^*(z) = 0$ is difficult to defend. Here, direct evaluation of the expression for $d_N$ in the proof of Proposition (ref) yields $d_N\approx 0.023$ (see Example (ref) in Appendix (ref)). Any least favorable posterior then assigns probability one to states in which $|\mu_1 - \mu_0 | \approx 0.023$. The realized data, however, estimate a treatment effect of $0.44$ with standard error less than $0.02$. Thus, every least favorable posterior is concentrated on effect sizes that are far from the empirical estimate.\footnote{In the conclusion of Stoye_2009_JoE, a similar discussion is presented using a specific least favorable prior. Proposition (ref) allows us to extend the argument to all least favorable priors.}
This example suggests that the ex ante minimax regret rule $\delta^*$ is unlikely to be dynamically consistent. After observing $ z$, a decision maker who evaluates the evidence using a conventional comparison of sample success rates would regard treatment $t=1$ as strongly favored. The posterior induced by any least favorable prior rationalizes the prescription $\delta^*(z) = 0$ only by concentrating mass on treatment-effect magnitudes that the realized data make highly implausible. Hence, the minimax rule $\delta^*$ lacks credibility as a decision rule that would be adhered to in the interim period, absent external commitment technologies.
Next, we consider the evidence aggregation problem, following ishihara2021evidence and yata2021optimal. A DM must decide whether to introduce a new policy in a target population. There are two study populations, indexed by $i \in \{1,2\}$, and the DM observes $ z=( z_1, z_2)\in Z\equiv\mathbb R^2$. The two studies are drawn independently, so under state $\theta=(\theta_1,\theta_2,\theta_3)$, we have $ z \sim N \left((\theta_1,\theta_2), \sigma^2 I_2\right).$
Here, $\theta_i$ is the welfare effect of the policy in study population $i \in \{1,2\}$, while $\theta_3$ is the welfare effect in the target population. Consider the state space $\Theta = \{ \theta\in\mathbb R^3: |\theta_1-\theta_3|\leq C_1,\ |\theta_2-\theta_3|\leq C_2 \}$, where $C_i$ are known constants that measure how different study population $i$ may be from the target population. The action space is $\mathcal A=[0,1]$, where $a\in[0,1]$ is the probability of introducing the policy, and a decision rule is the mapping $\delta: Z\to\mathcal A$. Normalizing the payoff of not introducing the policy to zero, the expected payoff of $\delta$ at $\theta$ is
If $\theta$ were known, the DM would introduce the policy whenever $\theta_3>0$, reject it whenever $\theta_3<0$, and be indifferent when $\theta_3=0$.
Suppose $C_1 > C_2$, such that study population $2$ is more externally valid than study population $1$, and define $\epsilon^* = \arg\max_{\epsilon\geq 0} (\epsilon+C_2)\Phi\left(-\frac{\epsilon}{\sigma}\right) $. The following fact restates the application of Theorem 1 to Example 1 in yata2021optimal.
That is, the minimax regret rule ignores the first study and bases the policy decision entirely on the more externally valid study. The following result characterizes the common support restriction that all least favorable priors must satisfy.
As in the previous section, we see that the class of least favorable priors share a support restriction. In particular, all least favorable priors are supported on states in which the more externally valid study is separated from the target population by exactly $C_2$, and the target welfare effect has magnitude $|\theta_3|=\epsilon^*+C_2$.
By Proposition (ref), the posterior update of any least favorable prior rationalizes the action chosen by any interim optimality criterion that is dynamically consistent with minimax regret. As before, the diagnostic question is whether the class of such posteriors provides a reasonable interim rationale. In the context of our evidence aggregation problem, Proposition (ref) implies that any least favorable posterior remains supported on $\Theta^+\cup\Theta^-$, regardless of the realized data.
For concreteness, suppose that $\sigma=1$, $C_1 = 2$, and $C_2 = 1$, in which case $\epsilon^* \approx 0.132$. Now consider the data realization $ z_1=5$, $ z_2=-0.05$, such that the minimax rule assigns $ \delta^*( z) = \mathds{1}\{ z_2\geq 0 \} = 0$, i.e., the policy is not implemented.
On the other hand, most conventional analyses of the data would regard the first study as strong evidence in favor of implementing the policy. Even allowing for the external validity bound $C_1=2$, the realization $z_1=5$ suggests that the effect in the target population is positive by a wide margin. For example, a conservative lower confidence calculation yields
while the second study is only slightly negative and statistically insignificant. In other words, the action $\delta^*(z)$ prescribed by the ex ante optimal rule disagrees with the action that most practitioners would take after observing $z$.
As in the previous section, we address this disagreement using the diagnostic exercise proposed in Section (ref). First, observe that every least favorable posterior is supported on states that satisfy $|\theta_3| = \epsilon^*+C_2 \approx 1.132$ and
However, the realized data from the first study (a random sample from $N(\theta_1, 1)$) is $ z_1 = 5$. The collection of least favorable posteriors then rationalizes rejecting the policy by treating the first study as a tail event nearly two standard deviations above the largest admissible value of $\theta_1$, making the posterior rationale for the minimax action difficult to defend.
This example illustrates that the minimax regret rule $\delta^*$ is unlikely to be dynamically consistent. A DM who regards $ z_1=5$ as compelling evidence about the target population would have strong reason to deviate from the ex ante minimax prescription and choose $a=1$ instead. As in the treatment choice example, we see that minimax rule is “too conservative" and lacks credibility as a decision rule that would be adhered to in the interim period.
The previous sections illustrated that decision rules justified by ex ante minimax guarantees may fail to provide credible interim prescriptions. In this section, we address this issue by proposing a class of ex ante optimality criteria that are dynamically consistent with their corresponding interim problems. We also axiomatize these criteria, thereby providing a behavioral characterization of the preferences and choice correspondences they induce.
Let $\mu\in\Delta( Z)$ denote the marginal distribution over the data space. If the DM has a prior $\pi\in\Delta(\Theta)$, then $\mu$ corresponds to the prior predictive $P_\pi\equiv\int_\Theta P_\theta d\pi(\theta)$ over $ Z$. For each realization of $ z\in Z$, let $\mathcal{Q}_z \subseteq\Delta(\Theta)$ denote the collection of “beliefs" the DM anticipates using at the interim stage, assuming without loss that each $\mathcal{Q}_z$ are closed and convex. We emphasize that the set $\mathcal{Q}_z$ need not have a Bayesian interpretation. It may, for instance, correspond to the confidence region or identified set for the parameter of interest.
We now have sufficient notation to formally introduce our two dynamically consistent optimality criteria. Their game-theoretic interpretation is that after the DM chooses $\delta$ and the data $z$ realizes, a malevolent Nature chooses the worst-case distribution over $\Theta$ from $\mathcal{Q}_z$.
Both criteria are dynamically consistent by construction. Conditional on observing $ z$, the corresponding interim minimax loss and regret problem is to solve
and the ex ante function aggregates these interim objectives using the marginal $\mu$.
A DM making decisions according to either criteria is sophisticated, in the sense of Strotz_ReStud_1955. She anticipates, at the ex ante stage, the procedure she will use in the interim stage after the data realization, and she chooses a decision rule whose prescriptions are optimal according to that anticipated interim procedure.
The marginal distribution $\mu$ can be interpreted in two ways. In the classical Bayesian setting, the DM has a prior $\pi\in\Delta(\Theta)$, and $\mu = P_\pi$ is the prior predictive over $Z$. On the other hand, the DM may have subjective beliefs directly over $Z$ instead, without having a prior over $\Theta$. As discussed in the introduction, the latter scenario is more realistic when $\Theta$ is high-dimensional or lacks a concrete interpretation, while the empirical phenomenon itself is familiar enough to the DM to support judgments about the marginal distribution of sample outcomes.\footnote{See Chapter 3.5.2 of Berger_1985_SDT for a more detailed discussion on the various sources of information about the marginal distribution $\mu$.} Our optimality criteria only require the existence of a subjective marginal distribution $\mu$, together with an anticipated interim procedure induced by the mapping $z \mapsto\mathcal \mathcal{Q}_z$. Bayesian decision making is then a special case with $\mu=P_\pi$ and $\mathcal Q_ z= \pi(\cdot\mid z)$.
In the following two examples, we show that the as-if optimization of Manski_2021_ECTA and Gamma${}^*$-minimax criteria of Lim_2026_PIA can be nested as special cases of our optimality criteria. For illustrative purposes, we use minimax regret in the former and minimax loss in the latter; analogous arguments also hold in reverse.
We now turn to providing a behavioral characterization of our two optimality criteria. Fixing some utility function $u: \mathcal{A} \times \Theta \to \mathbb{R}$, we abuse notation by associating each decision rule $\delta : Z \to \mathcal{A}$ with the utility act $\delta \in \mathbb{R}^{Z \times \Theta}$ defined by $\delta(z,\theta) = u(\delta(z), \theta)$.\footnote{The utility function can be recovered from preferences if we instead associate decision rules with Anscombe-Aumann acts that map from $Z \times \Theta$ to some mixture space.} If $\delta( \cdot, \theta) = \delta(\cdot, \theta')$ for all $\theta, \theta' \in \Theta$, we say that the utility act $\delta$ is Z-measurable. Likewise, if $\delta(z, \cdot) = \delta(z', \cdot)$ for all $z,z' \in Z$, then $\delta$ is $\Theta$-measurable. Throughout the remainder of this section, we assume for simplicity that both $Z$ and $\Theta$ are finite.
Most axiomatizations of statistical optimality criteria proceed by taking preferences or choice correspondences over risk acts as the primitive.\footnote{See, e.g., Stoye_2012_AR for a review. A recent exception is the axiomatic work of andrews2026misspecification, who consider the full space of utility (i.e., loss) acts for a different reason.} That is, they fix a likelihood function $\ell: \Theta \to \Delta(Z)$ and associate each decision rule with the (negative) risk act $f_\delta \in \mathbb{R}^\Theta$ defined by $f_\delta(\theta) = \mathbb{E}_{\ell(\cdot|\theta)} [u(\delta(Z), \theta)]$. This construction reduces decision rules to $\Theta$-measurable acts. As observed by Lim_2026_PIA, this reduction is with loss when considering dynamically consistent optimality criteria.
To axiomatize the dynamically consistent minimax loss criterion, we take the preference $\succsim$ over utility acts in $\mathcal{D} \cong \mathbb{R}^{Z \times \Theta}$ as the primitive relation.
Our first axiom requires that $\succsim$ satisfies the subjective expected utility (SEU) axioms over the collection of $Z$-measurable acts. Intuitively, this will allow us to recover a unique marginal $\mu \in \Delta(Z)$ that captures the DM's subjective beliefs about the marginal distribution over the data space. We also assume for simplicity that the recovered SEU measure $\mu$ has full support.
For $\delta, \beta \in \mathcal{D}$ and some $z' \in Z$, define $\delta_{ \{z' \} } \beta$ as the spliced act that satisfies $ (\delta_{ \{z' \} } \beta ) (z, \theta) = \delta(z,\theta)$ when $z = z'$ and $ (\delta_{ \{z' \} } \beta ) (z, \theta) = \beta(z, \theta)$ when $z \neq z'$. The next axiom imposes singleton separability with respect to this splicing operation over the data space $Z$.
When comparing two acts that map to different actions only when the data realization $\{Z=z\}$ occurs, the $Z$-separability axiom requires that the ranking must not depend on what happens outside of $\{Z=z\}$. It can thus be interpreted as a weakening of the Sure Thing Principle, imposed only over data realizations. Equivalently, if the DM anticipates evaluating actions after observing $z$, then prescriptions at unrealized data values are counterfactual and should not affect her comparison of the actions prescribed at $z$.
Let $\succsim_z$ be the induced preference over $\Theta$-measurable acts defined by $\delta \succsim_z \beta \Leftrightarrow \delta_{ \{z\} } \phi \succsim \beta_{ \{z\} } \phi$, where $\phi$ is any utility act. Note that $\succsim_z$ is well-defined under $Z$-separability, and it can naturally be interpreted as the conditional preference relation given the realization of $z$. The final axiom requires that every $\succsim_z$ satisfies the minimax expected utility (MEU) axioms of GS_MEU.
The $\Theta$-MEU axiom provides a behavioral interpretation of the interim belief correspondence $z \mapsto \mathcal \mathcal{Q}_z$. At each realized $z$, the induced preference $\succsim_z$ behaves as if the DM evaluates $\Theta$-measurable acts by their worst-case expected utility according to some set of beliefs over $\Theta$. The axiom permits ambiguity at the interim stage, while requiring that it only enters conditional on each data realization.
As shown by the following theorem, the above three axioms are equivalent to the DM having a dynamically consistent minimax loss representation.
It is well-known that optimality criteria based on minimax regret do not satisfy the Independence of Irrelevant Alternatives (IIA) axiom Hayashi_2008_JET, Stoye_2011_JET. For this reason, we must take choice correspondences $C(\cdot)$ that map finite menus $M \subsetneq \mathcal{D}$ to $C(M) \subseteq M$ as the primitive relation. We write $C_{(z,\theta)}(M)$ to denote the choice made after learning $(z,\theta)$ has realized.\footnote{Equivalently, $C_{(z,\theta)}(M) = C(M_{z,\theta})$, where $M_{z,\theta}$ contains constant acts $\delta(z,\theta)$ with $\delta \in M$.} Mixtures between menus are defined with respect to Minkowski addition, such that $\lambda M + (1-\lambda) N = \{ \lambda m + (1-\lambda) n : m \in M, \, n \in N\}$. We say that a menu has state-independent outcome distributions if $\{ p \in \mathbb{R} : \delta(z, \theta) = p \text{ for some } \delta \in M\}$ does not vary with $(z, \theta)$. That is, the collection of feasible utilities is constant across states.
Our axiomatization of the dynamically consistent minimax regret criterion proceeds by introducing additional structure on the endogenous prior minimax regret representation of Stoye_2011_JET. The axioms that we adopt without modification are as follows. We refer the reader to Section 2 of Stoye_2011_JET for their interpretation.
We need two additional axioms to characterize the dynamically consistent minimax regret criterion. The first axiom imposes IIA on $Z$-measurable acts, strengthening Stoye_2011_JET's (Stoye_2011_JET) IIA for constant acts. Intuitively, the axiom will require that $C(\cdot)$ induces a binary relation over the collection of $Z$-measurable acts, in which case monotonicity, independence, and mixture continuity axioms will allow us to recover the unique SEU measure $\mu \in \Delta(Z)$ as in the proof of Theorem (ref). Again, we assume for simplicity that the recovered SEU measure $\mu$ has full support.
Finally, we need to adapt the $Z$-Separability axiom from the case of minimax loss to the setting of choice correspondences. The following axiom does exactly this, requiring that the choice correspondence $C(\cdot)$ respect the Sure Thing Principle when splicing acts over data realizations. Its interpretation is analogous to that of $Z$-Separability in the previous section.
The following theorem illustrates that the above three axioms behaviorally characterize the dynamically consistent minimax regret representation.
In this section, we turn to evaluating the optimality criteria considered thus far. Our first application uses bursztyn2020misperceived's (bursztyn2020misperceived) study of information treatments on social norms to show that ex ante minimax regret can generate irrational interim prescriptions in treatment choice, while dynamically consistent minimax regret avoids these issues. The second application does the same using simulated data in the context of an evidence aggregation problem with limited external validity.
To evaluate the interim credibility of decision rules, we study the interim prescriptions generated by each rule. For the dynamically consistent minimax regret criterion, this means that only the correspondence $z \mapsto \mathcal{Q}_z$ is relevant. The DM's prior beliefs about the marginal distribution of data do not affect the interim action, provided that the ex ante choice domain is the set of all decision rules.
Specifically, we use the as-if minimax regret (as-if MMR) criterion of Manski_2021_ECTA as the interim objective, taking confidence sets $\hat \Theta(z)$ as set estimates of the parameter $\theta$ and setting $\mathcal{Q}_z = \Delta(\hat \Theta(z))$. This approach follows recent work that applies Manski_2021_ECTA's (Manski_2021_ECTA) as-if framework to decision problems with confidence sets Andrews_Chen_2025, Chernozhukov_et_al_2025, ben2025safe. Since dynamically consistent minimax regret provides the ex ante foundations of such as-if criteria, the results in this section can also be interpreted as illustrating the appeal of the as-if framework.
bursztyn2020misperceived study whether Saudi Arabian men's willingness to help their wives search for jobs outside the home is constrained by misperceived social norms. The authors find that most men privately support women working outside the home yet underestimate how many other men do. In their main experiment, the authors correct this misperception for a randomly selected half of the 500 participants.\footnote{The authors electronically randomized treatment values across all one thousand three-digit combinations before the experiment and assigned each man the value corresponding to the last three digits of his phone number (bursztyn2020misperceived, footnote 17).} We refer to this belief correction as the information treatment and to the no-information treatment as the status quo.
The outcome variable is an incentivized choice at the end of the session: whether to sign one's wife up for a job-matching service instead of receiving a cash bonus. At the aggregate level, the information treatment increases the sign-up rate from $0.235$ to $0.320$, a difference of $0.085$ ($z=2.1$). To study the minimax regret treatment choice problem with covariates, we form covariate subgroups from the twelve pre-treatment characteristics in the experiment.\footnote{These include the men's elicited second-order beliefs about how many other men support women working (in general, in semi-segregated environments, and on a minimum wage), the wedge between each belief and the truth, the confidence attached to it, and education, employment, and the wife's employment.} We form cells from covariates and their two- and three-way interactions, retaining only cells with at least five participants in each treatment arm (see Appendix (ref) for details).
For each covariate cell, the DM must decide whether to provide the information treatment $(a=1)$ or withhold it $(a=0)$ as a function of the data.\footnote{The DM may be, for example, a policymaker deciding whether to institute targeted information campaigns to correct Saudi Arabian men's misperceptions about other men's approval of women working outside the home.} Given the covariate partition $\mathcal{X}$, the DM's ex ante problem is then to choose a treatment rule $\delta: Z \to [0,1]^\mathcal{X}$. Let $z_x$ denote the data from individuals in covariate cell $x \in \mathcal{X}$, and let $\delta_x(z)$ denote the treatment probability assigned to them after the realization of $z \in Z$. The conditional average treatment effect $\tau_x$ denotes the effect of the information treatment on job-matching service sign-up.
Below, we consider the interim prescriptions of various decision rules after observing the realization of bursztyn2020misperceived's (bursztyn2020misperceived) study.
\paragraph{Minimax Regret} The minimax regret rule from Section (ref) extends naturally to the setting with covariates: Proposition 3 of Stoye_2009_JoE implies that applying the rule separately within each covariate cell is minimax-regret optimal. In other words, if $\delta$ denotes the minimax regret rule from Fact (ref), then the rule $\delta^*: Z \to \{0,1\}^\mathcal{X}$ defined by $\delta_x^*(z) = \delta(z_x)$ for every $x \in \mathcal{X}$ is minimax regret optimal.
\paragraph{Hypothesis Testing} A conventional one-sided size-$\alpha$ hypothesis testing rule with $H_0:\tau_x \le 0$ provides the information treatment only when there is statistically significant evidence of benefit within each covariate cell. We denote this rule by $T_\alpha$, where $(T_\alpha)_x(z) = 1$ if the test rejects in cell $x$, while $(T_\alpha)_x(z) = 0$ otherwise.
\paragraph{As-If MMR} The as-if MMR rule minimizes regret over a confidence set for the cell-specific treatment effect. Let this confidence set be $\hat \Theta_x(z)=[\hat\tau_x -c \cdot \widehat{\mathrm{se}},\hat\tau_x +c \cdot \widehat{\mathrm{se}}]$. Within each cell, the regret of providing treatment is $\max\{0,-\tau_x\}$, while the regret of withholding treatment is $\max\{0,\tau_x\}$. Over the confidence set $\hat \Theta_x(z)$, the worst-case regrets are $ \max\{0, -\hat\tau_x + c \cdot \widehat{\mathrm{se}}\}$ and $\max\{0,\hat\tau_x +c \cdot \widehat{\mathrm{se}}\}$, respectively. The former is strictly smaller if and only if $\hat\tau_x>0$, with indifference when $\hat\tau_x=0$. Thus, the as-if MMR rule coincides with the conditional empirical success rule of Manski_2004_ECTA: it assigns treatment to individuals in cell $x$ if and only if the estimated treatment effect in that cell is positive, i.e., $\hat\tau_x=\bar Y_{1x}-\bar Y_{0x}>0$.\footnote{Note that the as-if MMR rule is distinct from the conditional empirical success rule in general. The evidence aggregation application we consider in Section (ref) is one such example.}
\paragraph Since none of the above three rules pool data across covariate cells, we henceforth suppress the subscript $x$ and treat all decisions as cell-by-cell.
Observe that the minimax regret and as-if MMR rules are closely related. Writing $\Delta \equiv N_1-N_0$, the minimax regret index $I_N$ from Fact (ref) can be written as $I_N=N_0\hat\tau+\Delta(\bar Y_1-\tfrac12)$. The two rules then coincide whenever the treatment arms are balanced, i.e., $\Delta = 0$, and they differ only when the latter term overturns the sign of $\hat\tau$, i.e., $|\hat\tau|\lesssim |\Delta|\,|\bar Y_1-\tfrac12|/N_0$. Under fair-coin assignment, $|\Delta|$ is of order $\sqrt N$, so the two rules disagree only for estimated treatment effects in an $O_p(N^{-1/2})$ neighborhood of zero. Because this band is of the same stochastic order as the standard error itself, such disagreements can nevertheless coincide with statistically significant estimates.
Moreover, the following lemma illustrates the as-if MMR rule provides treatment whenever the hypothesis testing rule does. This eliminates the type of interim disagreement studied in Section (ref), where minimax regret recommends withholding treatment despite statistically significant evidence in favor of treatment.
We begin by focusing on disagreements between the minimax regret and hypothesis testing rules of the kind discussed in Section (ref), where the minimax regret rule $\delta^*$ is “too conservative” and withholds treatment. As shown in Table (ref), across 18 distinct covariate subgroups, $\delta^\ast$ withholds treatment even though the hypothesis testing rule $T_\alpha$ provides; 2 of these disagreements occur at the $\alpha=0.05$ level and 16 more at the $\alpha = 0.10$ level.\footnote{These subgroups are the distinct samples behind the cell counts reported in Appendix (ref); Appendix (ref) explains the overlapping structure of the cells and the scope of the exercise.} Many of these disagreements arise in the presence of treatment arm imbalance (i.e., $|\Delta| \gg 0$). Moreover, as predicted by Lemma (ref), the as-if MMR rule agrees with $T_\alpha$ for every specification reported in the table.
Consider, for instance, the first row, which corresponds to the covariate cell containing men whose wives are not employed, who know few other participants, and who place others' support for women's work outside the home in the lowest range. In this subgroup, the average sign-up rate for the job-matching service is $0.233$ $(N_1=30)$ and $0.062$ $(N_0=16)$ for those with and without the information treatment, respectively. The corresponding one-sided hypothesis test has $p$-value $0.041$, so $T_\alpha$ at level $\alpha=0.05$ provides the treatment. By Lemma (ref), the as-if MMR rule also provides, while the minimax regret rule $\delta^*$ withholds.
As in Section (ref), we interpret this disagreement as an empirically relevant instance of the interim incredibility of $\delta^*$. After seeing evidence that is strong enough to justify information provision under both hypothesis testing and as-if MMR, a reasonable DM may wish to deviate from the ex ante minimax-regret prescription.
Figure (ref) provides a broader view of the disagreement among the three rules. Each point is a covariate cell, plotted by cell-specific standardized arm imbalance $(N_1-N_0)/\sqrt{N}$ and Wald statistic $\hat\tau/\widehat{\mathrm{se}}$. The dashed and dotted horizontal lines mark the rejection thresholds for $T_\alpha$ at $\alpha=0.05$ and $0.10$, respectively.
The left panel illustrates that most disagreements between $T_\alpha$ and minimax regret arise because hypothesis testing is conservative: many blue points lie below the $T_\alpha$ threshold, so minimax regret provides treatment while $T_\alpha$ withholds. This is the phenomenon emphasized by Manski_2004_ECTA and discussed in Remark (ref), highlighting the conservatism of the hypothesis testing rule. However, we also see that minimax regret can provide treatment even when the estimated treatment effect is negative. This occurs in cells with arm imbalance in favor of the status quo, such that $(N_1-N_0)/\sqrt{N} \ll 0$. Such prescriptions are difficult to justify at the interim stage: after observing a negative estimated treatment effect, the DM is nevertheless instructed to provide treatment.
The kinds of disagreements reported in Table (ref) occur when arm imbalance favors the information treatment, such that $(N_1-N_0)/\sqrt{N} \gg 0$. The minimax-regret boundary where the DM is indifferent between providing and withholding treatment slopes upward with realized arm imbalance. As a result, treatment-heavy cells in the upper-right region of the left panel can be assigned to withhold even when the Wald statistic exceeds the $T_\alpha$ threshold.
The right panel shows that the as-if MMR rule eliminates this dependence on realized arm imbalance. As discussed previously, as-if MMR coincides with the conditional empirical success rule in this binary treatment-choice problem, and therefore provides treatment exactly when $\hat\tau>0$. In this sense, as-if MMR avoids both the excessive aggressiveness and the excessive conservatism induced by arm imbalance under the minimax-regret rule.\footnote{Of course, the same argument holds for the conditional empirical success rule, which we know from Manski_2004_ECTA to be asymptotically minimax-regret optimal. The evidence aggregation example we consider in the following section illustrates the value of the as-if MMR approach in settings where there is no obvious way to formalize an empirical success rule.}
We now return to the evidence aggregation problem introduced in Section (ref), where the DM must decide whether to adopt $(a=1)$ or reject $(a=0)$ the policy after observing the realization of $z=(z_1,z_2)$. Recall that $z_i\sim N(\theta_i,\sigma^2)$, where $\theta_i$ is the welfare effect of the policy in study population $i$, and $\theta_3$ is the welfare effect in the target population. The external validity restrictions imply $|\theta_i-\theta_3|\leq C_i$ for $i\in\{1,2\}$. We will continue to maintain the assumptions stated in Fact (ref), such that study 2 is more externally valid than study 1.
Next, consider the interim prescriptions of the three decision rules.
\paragraph{Minimax Regret} Recall from Fact (ref) that yata2021optimal's (yata2021optimal) ex ante minimax regret rule is $\delta^*(z)=\mathds{1}\{z_2\geq 0\}$, i.e., only the more externally valid study is used.
\paragraph{Hypothesis Testing} A conventional one-sided size-$\alpha$ hypothesis testing rule for $H_0: \theta_3\leq 0$ introduces the policy only when there is statistically significant evidence that the target-population welfare effect is positive. To construct the relevant lower confidence bound, first note that each study $i \in \{1,2\}$ induces the interval $[z_i-c\sigma-C_i,\, z_i+c\sigma+C_i]$, where $c$ is the appropriate critical value. Combining the two studies yields the confidence set
as long as the intersection is non-empty.\footnote{When the two study-implied intervals are disjoint, we expand the level to the smallest value of $c$ at which the intersection is nonempty and then apply the same rule to the resulting reconciled value of $\theta_3$. This keeps the rule defined for all data realizations and, in particular, keeps it using study $1$ precisely in the region where the two studies disagree and the ex ante minimax rule discards it. Such disjointness is rare: the two intervals fail to intersect only when $|z_1-z_2|>2c\sigma+C_1+C_2$. Even when the studies lie at opposite edges of their external validity bounds, this event occurs with probability $\Pr(Z>c\sqrt{2})$ ($\approx 0.3\%$ for a $95\%$ confidence set), and it is negligible otherwise.} The hypothesis testing rule, denoted $T_\alpha$, introduces the policy if and only if $L>0$.\footnote{ Because $T_\alpha$ rejects when either study's lower bound is positive, a single-study critical value would inflate the size of the combined test. Since the studies are independent, taking $c=\Phi^{-1}(\sqrt{1-\alpha})$ gives each study one-sided coverage $\sqrt{1-\alpha}$, so the lower bound $L=\max\{z_1-c\sigma-C_1,\,z_2-c\sigma-C_2\}$ covers $\theta_3$ with probability $1-\alpha$ and $T_\alpha$ is an exact level-$\alpha$ test of $H_0:\theta_3\le 0$. Figure (ref) shows the boundary $L=0$ at $\alpha=0.05$ ($c\approx1.95$) and $\alpha=0.10$ ($c\approx1.63$).} In other words, $T_\alpha$ introduces the policy only when the confidence set rules out non-positive target effects.
\paragraph{As-If MMR} The as-if MMR rule minimizes regret over the confidence set $\hat\Theta_3(z)=[L,U]$. For a given value of the target effect $\theta_3$, the regret of introducing the policy is $\max\{0,-\theta_3\}$, while the regret of rejecting the policy is $\max\{0,\theta_3\}$. Over the confidence set $[L,U]$, the worst-case regrets of adopting or rejecting the policy are then $\max\{0,-L\}$ and $\max\{0,U\}$, respectively. The as-if MMR rule introduces the policy whenever the former is smaller than the latter. Equivalently, for a nonempty interval $[L,U]$, it introduces the policy if and only if the midpoint is positive (i.e., $(L+U)/2>0$), with indifference when $L+U=0$.
\paragraph Unlike the minimax regret rule $\delta^*$, the hypothesis testing and as-if MMR rules use information from both studies. The hypothesis testing rule is more conservative: it introduces the policy only when the lower endpoint $L$ is positive. By contrast, the as-if MMR rule introduces the policy whenever the confidence set is centered above zero. Hence, whenever the hypothesis testing rule adopts the policy, the as-if MMR rule does so as well, proving the following lemma.
Figure (ref) visualizes the interim prescriptions of the three decision rules over the $(z_1,z_2)$ plane when $\sigma = 1$, $C_1 = 2$, and $C_2 = 1$. Because the minimax regret rule $\delta^*$ ignores the less externally valid study, its prescription depends only on $z_2$. It adopts the policy when $z_2 \geq 0$, i.e., in the region above the dashed black line. By contrast, the as-if MMR rule adopts the policy when $(L+U)/2 \geq 0$, corresponding to the region above the solid black line. The hypothesis testing rule $T_\alpha$ adopts treatment when the lower confidence bound $L$ is nonnegative. In the figure, this is the upper-right region relative to the green boundaries, where the solid line corresponds to $\alpha=0.05$ and the dashed line corresponds to $\alpha=0.10$.
The interim incredibility of the minimax regret rule discussed in Section (ref), where $\delta^*$ is “too conservative” relative to $T_\alpha$, corresponds to the region in which $T_\alpha$ adopts but $\delta^*$ rejects. In the figure, this is the pink region lying in the adoption region of $T_\alpha$ and below the dashed black line $z_2=0$. For example, $(z_1,z_2)=(5,-0.05)$ is the data realization considered in Section (ref). Here, the minimax regret rule rejects the policy because $z_2<0$, despite significant evidence in favor of the policy from the less externally valid study. As predicted by Lemma (ref), the as-if MMR rule adopts whenever $T_\alpha$ adopts, thus avoiding this form of excessive conservatism.
The figure also illustrates the opposite form of conservatism, in the spirit of Manski_2004_ECTA's (Manski_2004_ECTA) critique that hypothesis testing can be too conservative in treatment choice. This occurs in the beige and blue regions in the interior of the green boundaries, where $T_\alpha$ rejects even though the minimax regret rule adopts. The as-if MMR rule distinguishes between the two cases. In the blue region, it adopts despite rejection by $T_\alpha$, consistent with the prescription of $\delta^*$. In the beige region, it instead corrects the excessive aggressiveness of $\delta^*$ and rejects. Although $z_2 \geq 0$ leads $\delta^*$ to adopt, the evidence from the less externally valid study is sufficiently negative that the as-if MMR rule rejects.
Taken together, our results illustrate that the as-if MMR rule balances the competing concerns underlying minimax regret and hypothesis testing. It avoids the excessive conservatism of $\delta^*$ when strong evidence from the less externally valid study favors adoption, while also avoiding the excessive aggressiveness of $\delta^*$ when that evidence points against adoption. We interpret this as demonstrating the interim credibility of the as-if MMR prescription, thus supporting the rationale for the dynamically consistent minimax regret criterion.
This paper studies the dynamic consistency of statistical decision rules justified by ex ante optimality criteria. We show that, for frequentist minimax criteria, dynamically consistent interim preferences can be defined by updating a least favorable prior. This observation yields a diagnostic for assessing whether ex ante minimax rules remain credible after the data realization. In applications to treatment choice and evidence aggregation, this diagnostic reveals that minimax-regret rules can prescribe actions that are difficult to justify at the interim stage in empirically relevant settings. To address this issue, we propose and axiomatize a class of dynamically consistent minimax criteria that provide credible prescriptions in the interim period. Applications to both real and simulated data illustrate their potential value as a tool for guiding statistical decision making.