Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
57,176 characters · 10 sections · 103 citation commands
Robust Bayes Treatment Choice with Partial Identification
\sloppy
\onehalfspacing
A policy maker must decide between implementing a new policy or preserving the status quo. Her data provide information about the potential benefits of these two options. Unfortunately, these data only partially identify payoff-relevant parameters and may therefore not reveal, even in large samples, the correct course of action. Such treatment choice problems with partial identification have recently received growing interest; for example, see d2021policy, ishihara2021, yata2021, christensen2022optimal, kido2022distributionally or manski2022identification. Several interesting problems that arise in empirical research can be recast using this framework. See MQS and the end of this section for references.
This paper applies the robust Bayes approach---which interpolates between Bayesian and agnostic minimax analyses by evaluating minimax risk over a set of priors---to a class of treatment choice problems with partial identification. The use of the robust Bayes approach has drawn recent attention in problems that feature partial identification GiacominiKitagawa,GKR,christensen2022optimal. Indeed, due to partial identification, integrating Bayesian and minimax elements into decision making can be particularly attractive (see Poirier,MoonSchorfheide and references therein).
The robust Bayes approach can be applied ex-ante or ex-post, depending on whether the multiple priors are used to evaluate payoffs before or after seeing the data.\footnote{GKR discuss both notions; they refer to the ex-ante and ex-post problems as “Gamma-minimax” and “Conditional Gamma-minimax”, respectively. christensen2022optimal focus on the ex-post problem for treatment choice problems with partial identification in a restricted class of decision rules.} These concepts represent two different ways of resolving model ambiguity and sampling uncertainty, and both have been proposed to improve Bayesian robustness in decision problems.
By the well-known dynamic consistency of Bayesian decision making, ex-ante and ex-post robust Bayes coincide with each other---and with standard Bayes optimality---if the set of priors is a singleton. Indeed, this equivalence is used to calculate Bayes optimal decisions in practice because ex-post rules are usually computed but ex-ante Bayes optimality is claimed. As pointed out for the present context by GKR, they do not in general agree otherwise.\footnote{For an estimation problem with a quadratic loss, kitagawa2012estimation derives the ex-post $\Gamma$-minimax estimator and shows that it is not ex-ante $\Gamma$-minimax optimal.} However, to what extent this inequivalence affects treatment choice problem with partial identification is far from clear. The class of priors that we consider has a Cartesian product structure resembling rectangularity, a condition under which maximin welfare loss is known to be dynamically consistent.\footnote{See epstein2003recursive and also wakai2007note, amarante2019recursive, and references therein.} While this result does not apply here---the priors are not rectangular in the strict technical sense, and the theoretical results were not established for regret loss---one might wonder if the conclusion holds anyway or else, what qualitative and quantitative relationships exist between ex-ante and ex-post robust Bayes in the presence of partial identification. Using a convenient class of priors allows us to exhaustively answer these questions in a case that we believe holds some interest.
To this end, we use the same framework as yata2021 and MQS but impose a simple instance of GiacominiKitagawa's (GiacominiKitagawa) set of priors, namely a symmetric and uniform two-point prior for reduced-form parameters and no restriction at all on unidentified parameters given reduced-form parameters. Working with regret, we formally define ex-ante and ex-post robust Bayes and, following berger1985statistical, label them as “$\Gamma$-minimax regret” ($\Gamma$-MMR) and “$\Gamma$-posterior expected regret” ($\Gamma$-PER).\footnote{Among others, see Savage51, manski2004statistical, stoye2012new, and MQS for justifications of focusing on regret in treatment choice problems. In particular, while minimax loss can be an attractive alternative to minimax regret, it leads to trivial recommendations in treatment choice settings including our examples.} We then precisely characterize when these notions coincide and when they disagree. The main qualitative insights are as follows:
Our point is not to advocate for either notion of robust Bayes criterion. We also do not aim to solve for robust Bayes criteria for more general sets of priors as this would get much more involved but (we suspect) not much more instructive to illustrate the points discussed above. What we hope to illustrate is when, and how, $\Gamma$-MMR and $\Gamma$-PER criteria differ. We also relate these results to timing assumptions in a fictitious game between the policy maker and an adversarial Nature.
Several auxiliary findings might be of independent interest. First, for $\Gamma$-MMR, even if we restrict the set of decision rules to be a class of non-randomized threshold rules based on the “efficient” linear index, the optimal threshold is not always zero. Given the apparent symmetry of the problem, we find this feature rather surprising. Second, whenever the dimension of the signal is larger than $1$, there always exist (regardless of the parameter space and the variance of the signals) non-randomized linear-index threshold rules (with a threshold equal to zero) that are $\Gamma$-MMR optimal (among all decision rules). This is in stark contrast to MQS, in which no linear index rule is globally minimax regret optimal if the degree of partial identification is severe. The intuition is that the prior much reduces the state space; the signal space then becomes so rich relative to the state space that even linear threshold rules can effectively mimic randomization.
The literature on treatment choice with partially identified parameter has been growing since manski2004statistical and Dehejia2005. For partial identification with known distribution of data, Manski2000,manski2005social,manski2007identification and Stoye07 find minimax regret optimal treatment rules. For finite-sample minimax regret results with model ambiguity and sampling uncertainty, see stoye2012minimax,stoye2012new, yata2021, ishihara2021 and MQS; kido2023locally's (kido2023locally) analysis is asymptotic. Bayes and robust Bayes approaches are analyzed by chamberlain2012, GiacominiKitagawa, GKR, christensen2022optimal, among others. Earlier investigations of ex-ante and ex-post $\Gamma$-minimax estimators include dasgupta1989frequentist and betro1992conditional. See vidakovic2000gamma for a review. For different settings with point-identified welfare, finite- and large-sample results on optimal treatment choice rules were derived by canner, ChenGug, HiranoPorter2009,HiranoPorter2020, \citet*{kitagawa2022treatment}, schlag2006eleven, stoye2009minimax, and tetenov2012statistical. There is also a large literature on optimal policy learning with covariates containing results with point identified BhattacharyaDupas2012,kitagawa2018should,KT19,MT17,KW20,AW20,KST21,ida2022choosing as well as partially identified kallus2018confounding,ben2021safe,ben2022policy,d2021policy, christensen2022optimal,adjaho2022externally,kido2022distributionally,lei2023policy parameters. GugQuantile and ManskiTetenovJER analyze related problems but focus on quantile, as opposed to expected, loss; song considers partial identification but mean squared error regret loss.
The rest of this paper is organized as follows. Section (ref) sets up the problem, provides examples, and defines both versions of robust Bayes optimality. Section (ref) contains complete solutions for all aforementioned scenarios and relates them to timing assumptions in the “Games against Nature” interpretation of minimax theory. Section (ref) concludes. Proofs and auxiliary results are collected in the Appendix.
Our setup follows MQS, who in turn follow Ferguson67 and others. Consider a policy maker who needs to choose an action $a\in[0,1]$ interpreted as probability of assigning treatment in the target population.\footnote{Randomization could be i.i.d. across future potential treatment recipients, fractional in the sense of randomly assigning a certain fraction of the treatment population (in this sense, $a\in[0,1]$ can also be interpreted as the fraction of the population receiving the treatment), or an “all or nothing” randomization for the entire treatment population. While these might not be practically equivalent in all applications, they are in the current decision theoretic framework. See manski2007admissible for an exception in the related literature.} Her payoff when taking action $a\in[0,1]$ is captured by the welfare function
where $\theta\in\Theta$ is an unknown state of the world or parameter and the functions $W(1,\cdotp):\Theta\rightarrow\mathbb{R}$ and $W(0,\cdotp):\Theta\rightarrow\mathbb{R}$ are known. Here, we may interpret $W(1,\theta)$ and $W(0,\theta)$ as the welfare of actions $a=1$ (treating everyone in the population) and action $a=0$ (treating no one in the population). Therefore, (ref) implies that welfare is linear in actions, a standard assumption in the literature. Denote by $U(\theta):=W(1,\theta)-W(0,\theta)$ the welfare contrast at $\theta$. If $U(\theta)$ were known to the policy maker, her optimal action would simply be
The policy maker does not know $U(\theta)$ but can learn about $\theta$. Specifically, we assume that she observes a random vector $Y\in\mathbb{R}^{n}$ with multivariate normal distribution
where the function $m(\cdotp):\Theta\rightarrow\mathbb{R}^{n}$ and the positive definite matrix $\Sigma$ are known.
Our focus is on the case when the data is not entirely informative about the sign of $U(\theta)$: Even if the policy maker perfectly learned $m(\theta)$, she could not (necessarily) pin down the sign of $U(\theta)$. To formally model such treatment choice problems with (decision-relevant) partial identification, let
collect all the means of $Y$ that can be generated as $\theta$ ranges over $\Theta$. We refer to elements $\mu\in M$ as $\emph{reduced-form}$ parameters because they are identified in the statistical model (ref) without further assumptions. Define the identified set for the welfare contrast given $\mu$ as
and the corresponding upper and lower bounds as
Henceforth, when we refer to a treatment choice problem with partial identification, we mean there exists some nonempty open set in $M$ such that for all $\mu$ in that open set, $\underline{I}(\mu) <0<\overline{I}(\mu)$. For simplicity, we also assume that the infimum and supremum in (ref) are attained.
A decision rule $d:\mathbb{R}^{n}\rightarrow[0,1]$ is a (measurable) mapping from data $Y$ to the unit interval $[0,1]$. We call $d$ non-randomized if it (almost surely, a.s.) maps into $\{0,1\}$; otherwise, we call $d$ randomized, including if it randomizes for some but not all data realizations. We use $\mathcal{D}_n$ to denote the set of all decision rules, and we consider decision rules the same if they a.s. agree. As a result, a rule is $\emph{unique}$ only up to a.s. agreement. The oracle policy $\textbf{1}\{U(\theta) \geq 0\}$ is of special interest and for any given $\theta$ is contained in $\mathcal{D}_n$, but is not feasible in the statistical sense because $U(\theta)$ is not known. In general, there will not be an unambiguously best feasible decision rule, a problem that gave rise to a large literature on different optimality criteria and their implementation. Before introducing the robust Bayes approach, we give two examples that fit into our general framework.
Our setting up to here is as in MQS. We now connect it to the robust Bayes literature by imposing a set of priors $\Gamma$ on $\theta$. Following GiacominiKitagawa, we choose a particular single proper prior $\pi_{\mu}$ for $\mu\in M$ but leave the conditional prior of $\theta$ given $\mu$, denoted as $\pi_{\theta \mid \mu}$, unrestricted except for
Then, the class of priors $\Gamma$ consists of all priors on $\theta$ induced by the single prior $\pi_{\mu}$ and any conditional prior $\pi_{\theta \mid \mu}$ that meets (ref). Intuitively, we choose a single prior on the point-identified parameter and place no new restriction on the partially identified parameter $U(\theta)$. One can pick any proper prior $\pi_{\mu}$; for tractability, we let $\pi_{\mu}$ be supported on two symmetric points $\{\bar{\mu},-\bar{\mu}\}$ with equal probability, where $\bar{\mu}\in\mathbb{R}^n$ is chosen by the decision maker. Henceforth, $\Gamma$ is understood to refer to the implied set of priors:
While the set of priors $\Gamma$ broadly puts us into the “robust Bayes” territory, it still does not pin down a uniquely best decision rule because the sign of $U(\theta)$ can remain ambiguous. Let \[ L(a,\theta):=\sup_{a^{\prime}\in[0,1]} W({a^{\prime},\theta}) - W(a,\theta)=U(\theta)\left\{ \mathbf{1}\{U(\theta)\geq0\}-a\right\} \] be the regret of action $a\in[0,1]$. We evaluate decision rules $d$ by their expected regret, defined as
where for any $x\in\mathbb{R}^n$, $\mathbb{E}_{x}[\cdot]$ denotes expectation taken over $Y\sim N(x,\Sigma)$.
Even for given set of priors $\Gamma$ given and commitment to expected regret, the Robust Bayes literature contains multiple decision criteria that do not in general agree. The difference lies in when (and how) the expectations regarding the unknown parameter $\theta\in\Theta$ are taken. For decision rule $d\in \mathcal{D}_n$, let \[r(d,\pi):=\int_{\theta\in\Theta} R(d,\theta)d\pi(\theta) \] be its Bayes expected regret under a prior $\pi\in\Gamma$. Following berger1985statistical, we introduce the first robust Bayes optimality notion.
An alternative “posterior” robustness notion is also common in Bayesian analysis. For each action $a\in [0,1]$, define posterior expected regret under prior $\pi$ berger1985statistical \[ \rho(a,\pi_{\theta\mid Y}):=\int_{\tilde{\theta}\in\Theta}L(a,\tilde{\theta})d\pi_{\theta \mid Y}(\tilde{\theta}), \] where $\pi_{\theta\mid Y}$ is the posterior distribution of $\theta$ given prior $\pi$ and data $Y$.\footnote{In our setting, information from data $Y$ does not revise the conditional prior $\pi_{\theta\mid\mu}$ GiacominiKitagawa. For any event $A$ in the $\sigma$-algebra of $\Theta$, we therefore have $\pi_{\theta\mid Y}(A)=\int\pi_{\theta\mid\mu}(A)d\pi_{\mu\mid Y}$, where $\pi_{\mu\mid Y}$ is the posterior distribution of $\mu$ given $Y$.} Then we have the following, alternative optimality criterion berger1985statistical:
The labeling of $\Gamma$-MMR as “ex-ante” versus $\Gamma$-PER as “ex-post” can be related to the timing of a fictitious game against an adversarial Nature; see Section (ref) for additional discussion. If $\Gamma$ were a singleton, the criteria would coincide and would also agree with (single-prior) Bayes optimality. They do not in general agree otherwise. The term “Gamma minimax” usually (and even “robust Bayes” more often than not) refers to $\Gamma$-MMR; for example, see berger1985statistical.\footnote{\citet*{GKR} discuss both criteria for general loss functions and refer to Definitions (ref) and (ref) as the “unconditional $\Gamma$-minimax” and “conditional $\Gamma$-minimax” problems, respectively. For treatment choice problems with partial identification, christensen2022optimal optimize the $\Gamma$-PER criterion, restricting the action space to be $\{0,1\}$.}
While it is not our agenda to advocate for either criterion, some possible considerations are as follows. The ex-ante approach may be perceived as more suitable if the planner has commitment power and has also been justified axiomatically hayashi2008regret,stoye2011axioms. We also find some numerical and theoretical evidence that the $\Gamma$-MMR rule may have desirable frequentist properties, e.g. in states of the world off the prior's support; see discussions at the end of Section (ref). Regarding computational feasibility, the ex-post approach is often easier due to the applicability of backward induction and is routinely employed to quantify posterior robustness of statistical decisions. On the other hand, recent work on numerical discovery of minimax rules fernandez2024epsilon,guggenberger2025numerical may render the ex-ante approach more scalable. See additional discussions in Appendix (ref).
The two-point structure of $\pi_\mu$ in (ref) appears in several related treatment choice problems with minimax regret optimality criteria. For example, the least favorable prior takes such a form in a completely unconstrained minimax regret problem with point identified stoye2009minimax or partially identified stoye2012minimax,yata2021,MQS welfare contrast. In analogously constrained minimax regret problems in which we restrict $\mu\in[-\left|\bar{\mu}\right|,\left|\bar{\mu}\right|]$, extending the analyses in the preceding literature would also imply a similar two-point symmetric structure for the least favorable prior. Given these precedents, we think our choice of $\pi_\mu$ represents a general feature of this class of minimax regret decisions, in addition to offering computational tractability.
In this section, we solve for the two versions of robust Bayes optimality under the set of priors (ref). Following yata2021, we assume:
These conditions are restrictive but encompass many examples of empirical relevance; see this paper's introduction, MQS, and yata2021. Exploiting symmetry of the setting, we also set \[\overline{I}(\bar{\mu})+\underline{I}(\bar{\mu})>0.\] This is a normalization because, by Lemma (ref), Assumption (ref) implies $\overline{I}(-\bar{\mu})+\underline{I}(-\bar{\mu})=-(\underline{I}(\bar{\mu})+\overline{I}(\bar{\mu}))$ and we could always replace $m(\cdot)\mapsto-m(\cdot)$; furthermore, the case of $\overline{I}(\bar{\mu})+\underline{I}(\bar{\mu})=0$ gives rise to trivial solutions.\footnote{In this case, the identified set for $U(\theta)$ does not change with true value of $\mu$, and under all decision criteria considered here, the solution will be a no-data rule that simply flips a coin.}
We say a rule is a linear-index threshold rule if it has the form $\mathbf{1}\left\{ \beta^{\top}Y\geq c\right\} $ for some $\beta\in\mathbb{R}^{n}$ and $c\in\mathbb{R}$. Linear-index threshold rules are nonrandomized. They are of particular interest because they form a complete class when $U(\theta)$ is point-identified karlin1956theory and have received particular attention in the recent literature ishihara2021,MQS. For both $\Gamma$-MMR and $\Gamma$-PER, we will clarify when linear-index threshold rules are optimal and when they are not. For reasons that will become obvious, the following linear-index threshold rule is of particular interest:
For each vector $\beta\in\mathbb{R}^n$, let $\left\Vert \beta\right\Vert_{\Sigma}:=\sqrt{\beta^{\top}\Sigma\beta}$. Thus, $\lVert w \lVert_{\Sigma} =\sqrt{w^{\top}\Sigma w}=\sqrt{\bar{\mu}^{\top}\Sigma^{-1}\bar{\mu}}$. Denote by $\Phi(\cdot)$ the standard normal c.d.f. and by $\Phi^{-1}(\cdot)$ its inverse, i.e. the corresponding quantile function.
Theorem (ref) reveals that the $\Gamma$-MMR rules qualitatively change depending on whether condition ((ref)) is met or not. This condition admits an intuitive interpretation: Up to clamping to the unit interval (i.e., values outside this interval are mapped to its edges), the left-hand side of ((ref)) equals the unique minimax regret optimal rule for known $\mu$ Manski2007; therefore, it arguably measures the model's identification strength. The right-hand side of ((ref)) can be interpreted as the informativeness of the data about which of $\{-\bar{\mu},\bar{\mu}\}$ obtained. Therefore, Theorem (ref) says that if the model's identification power is sufficiently large compared to the informativeness of the data, then, the unique $\Gamma$-MMR optimal rule is $d_{w,0}^{*}$, a non-randomized linear index rule with threshold $0$ that effectively ignores the partial identification issue. In contrast, if the model's identification power is small compared to the informativeness of data, there are infinitely many $\Gamma$-MMR optimal rules, all of which satisfy ((ref)) and ((ref)). Examples include suitably smoothed versions of $d_{w,0}^{*}$ like $d_{\text{RT}}^{*}$ and $d_{\text{linear}}^{*}$ (the functional forms of which showed up in MQS) as well as $d_{\text{step}}^{*}$ (which is a new result).
To understand how these two “regimes” arise, it is helpful to think of a zero-sum game against an adversarial Nature in which a MMR decision rule and a distribution over parameter values (the least favorable prior) form a Nash equilibrium.\footnote{Analyzing this game is also how results are formally proved. See wald45 for what may be the first clear statement of this and lehmann2006theory for a formalization. Within the literature on decisions under partial identification, the proof technique was first explicitly used in Stoye07 and the equilibrium structure with a noninformative and an informative regime was first encountered in stoye2012minimax.} If identification is strong in the sense of (ref), this game has an informative equilibrium: the least favorable prior evenly randomizes over two points $(\mu,U(\theta)) \in \{(\bar{\mu},\overline{I}(\bar{\mu})),(-\bar{\mu},\underline{I}(-\bar{\mu}))\}$, to which the unique Bayes response (and, therefore, uniquely optimal MMR rule) is $d_{w,0}^*$. In all other cases, the equilibrium is uninformative: the least favorable prior is supported on four points $(\mu, U(\theta))\in\{(\bar{\mu},\overline{I}(\bar{\mu})),(\bar{\mu},\underline{I}(\bar{\mu})),(-\bar{\mu},\overline{I}(-\bar{\mu})),(-\bar{\mu},\underline{I}(-\bar{\mu}))\}$ with a probability profile such that the posterior expectation of $U(\theta)$ always equals $0$. While any decision rule best responds to this prior, not every decision rule is MMR because the least favorable prior must also best respond to the decision rule. In this very structured setting, the latter is guaranteed by conditions (ref) and (ref), establishing claim (ii) and nonuniqueness of optimal decision rules. Indeed, even the nonrandomized threshold rules from Theorem (ref)(iv) do not reflect any updating along the game's equilibrium path; they rather use uninformative features of the data as randomization device.
In the uninformative equilibrium, we encounter several additional findings. First, despite the problem's apparent symmetry, $d_{w,0}^*$ is not optimal even among linear-index threshold rules that use the index $w^{\top}Y$. Instead, Theorem (ref)(iii) characterizes exactly two optimal thresholds, one positive and one negative. Second, while the particular linear-index threshold rule $d_{w,0}^*$ is not $\Gamma$-MMR optimal, in higher dimensional problems ($n>1$), there do exist linear-index threshold rules that are. This finding is in stark contrast to MQS, who find that, in a large class of special cases, no linear index threshold rule is MMR optimal. The crucial difference in settings is that the set of priors much constrains the decision theoretic problem's state space; as a result, the signal space is much richer than the state space, and this can be exploited to mimic randomization without nominally randomizing. Compare manski2022identification's (manski2022identification) abstract observation that, if a policy maker is not allowed to explicitly randomize, sampling uncertainty can be beneficial by providing an implicit randomization device. We note that this phenomenon is reminiscient of classic “purification” results in game theory DWW,purify.\footnote{We thank Elliot Lipnowski for reminding us of this literature.} In contrast, it is not deeply related to the general intuition that “Bayesians don't randomize.”
We next apply Theorem (ref) to Example (ref) and immediately get the following results.
Note that there is no analog to Theorem (ref)'s case (iv); indeed, no linear threshold rule is optimal in case (ii). This is because the scalar nature of the signal $Y$ shuts down the aforementioned purification mechanism. See Figure (ref) for an illustration of different $\Gamma$-MMR optimal rules in Example (ref) for selected parameter values of $k,\sigma$ and $\bar{\mu}$.
The results of Theorem (ref) offer some important clarifications regarding $\Gamma$-PER in treatment choice problems with partial identification. First, even if regret is evaluated according to the posterior distribution, it is not necessarily true that optimal rules are non-randomized. In fact, whenever there is model ambiguity regarding the sign of $U(\theta)$ (i.e., $\underline{I}(\bar{\mu})<0<\overline{I}(\bar{\mu})$), the unique $\Gamma$-PER optimal rule is randomized. Therefore, restricting the action space to $\{0,1\}$ in such problem is not without loss of generality even under the $\Gamma$-PER criterion. Comparing results with Theorem (ref) also allows for instructive observations on when $\Gamma$-PER and $\Gamma$-MMR optimal rules agree or disagree; we elaborate these in Corollary (ref). Applying Theorem (ref) to Example (ref), we finally obtain:
In Figure (ref), we depict $\Gamma$-PER optimal rules for Example (ref) with the same parameter values considered in Figure (ref). We see clearly that $\Gamma$-MMR and -PER optimal rules coincide (if randomization is allowed) only in the special case when $k\leq\bar{\mu}$, an observation we generalize in Corollary (ref). More specifically, in the top left panel of Figure (ref), as $k\leq\bar{\mu}$, $\Gamma$-MMR and -PER coincide and are the non-randomized threshold rule $d^*_{0}$. In the top right panel, $k>\bar{\mu}$ and the $\Gamma$-PER optimal rule becomes $d^*_{\text{PER}}$. However, since (ref) still holds, $d^*_0$ is still $\Gamma$-MMR optimal. For the bottom two panels, as it still holds $k>\bar{\mu}$, $d^*_{\text{PER}}$ is still $\Gamma$-PER optimal. However, the associated parameter values imply (ref) fails. As a result, $d^*_{0}$ is no longer $\Gamma$-MMR optimal and many $\Gamma$-MMR rules exist. But even in theses cases, $\Gamma$-MMR and -PER rules differ, as among the class of step function rules (which contain $d^*_{\text{PER}}$), only $d^*_{\text{step}}$ is $\Gamma$-MMR optimal, still different from $d^*_{\text{PER}}$.
Applying MQS to the current setting, we can conclude that all rules, including both $\Gamma$-MMR and $\Gamma$-PER optimal rules, are at least admissible. Moreover, by definition, the $\Gamma$-MMR rule will have the lower value of the (ex ante) game. Therefore, it might be more useful to compare the (frequentist) profiled regrets MQS of ex-post $\Gamma$-PER and other rules as a function of the true but unknown mean $\mu$ of data $\hat{\mu}$. In the context of Example (ref), we can report $\bar{R}(d,\mu):=\sup_{\mu^*\in[\mu-k,\mu+k]}R(d,\mu,\mu^*)$ as $\mu$ varies for each rule $d$, where $R(d,\mu,\mu^*)$ is the expected regret of rule $d$ as a function of $\mu$ and $\mu^{*}$. We report several findings (we emphasize that these are not obvious: we took an ex-ante perspective but did not restrict ourselves to those values of $\mu$ that the prior allows). First, in both panels of Figure (ref), the profiled expected regret of the $\Gamma$-PER rule exceeds that of the $\Gamma$-MMR rule. Therefore, in this particular example, $\Gamma$-MMR arguably outperforms $\Gamma$-PER from a broader frequentist point of view. Second, we in fact prove that, whenever $\bar{\mu}$ is sufficiently small and $k$ is sufficiently large, the $\Gamma$-PER rule is profiled-regret dominated, i.e., there exists a rule $d\neq d^*_{\text{PER}}$ such that
for all $\mu\in\mathbb{R}$ with the inequality strict for some $\mu$; see Lemma (ref) in Appendix (ref) for an exact statement. Intuitively, the $\Gamma$-PER rule mixes between a coin flip rule and the naive threshold rule ($d_0^{*}$). When $\bar{\mu}$ is small, $\Gamma$-PER rule is more analogous to the coin flip rule, which MQS show to be dominated in terms of profiled regret; also, its profiled regret fails to vanish as $\mu\to\pm\infty$ even though the optimal treatment is known ex ante in this case. Furthermore, Figure (ref) reveals that the $\Gamma$-MMR rule $d^{*}_{\text{linear}}$, though not globally MMR optimal, in some cases has essentially the same profiled risk function as the least randomizing global MMR optimal rule MQS. This may lend a Robust Bayes interpretation to the latter.
The different approaches analyzed here can all be expressed as different timing assumptions in the statistical game. The distinction is quite obvious for ex-ante versus ex-post: In the former (and original Waldian) perspective, this game is simultaneous move; in particular, Nature moves before data $Y$ are realized. In the ex-post perspective, Nature sees $Y$ before choosing a prior. It is immediately clear that this may be easier to solve because it allows for backward induction. There is also an immediate sense that solutions might not agree, as we indeed found.
But whether the decision maker is allowed to (or at least wants to) randomize or not can equally be thought of as changing the game's timing, providing another way to think about the action space for the ex-ante and ex-post approaches. The more standard perspective is that the decision maker may randomize over decision rules and Nature must move before learning the outcome of this randomization. By the nature of zero-sum games, this setup will frequently yield randomized solutions. In consequence, it is essential to define the decision maker's action space as $[0,1]$. In contrast, if Nature is allowed to move after the decision maker's randomization is realized, then any incentive to randomize is gone and we may as well restrict the action space to $\{0,1\}$.
Theorems (ref) and (ref) clarify that these distinctions actually matter in an interesting example. We spell this out in Corollary (ref), which considers both cases when randomization is allowed and not allowed. The bottom line is that the assessment is quite sensitive to how the problem is set up. If underlying parameters lead to sufficiently small identification power, the criteria disagree.
Thus, if randomization is allowed, $\Gamma$-MMR and $\Gamma$-PER optimal rules coincide only in the somewhat trivial case in which we a priori know that $U(\theta)$ and $\mu$ have the same sign, so that optimal treatment choice reduces to Bayesian inference on the point identified $\mu$. They disagree in all other cases, including in settings where there are infinitely many $\Gamma$-MMR optimal rules.
One might expect more agreement once randomization is excluded; after all, this leads to a much simpler action space. Part (ii) shows that there is some truth to this: The condition for agreement changes from $\underline{I}(\bar{\mu})\geq0$ to the strictly weaker (ref). However, the criteria continue to disagree in many cases. It may be instructive to think of these cases in terms of “comparative statics.” For example, consider holding all parameters of the problem fixed but scaling the signal variance $\Sigma$ by a positive scalar, say to reflect a change in sample size. Then (ref) will hold if, and only if, $\Sigma$ is large enough; hence, as long as $\underline{I}(\bar{\mu})\geq0$, increasing sample size will eventually cause disagreement between $\Gamma$-MMR and -PER even if randomization is excluded. Similarly, for fixed $\Sigma$, $\Gamma$-MMR and -PER rules will always disagree if the model's identification power is small enough.
We studied treatment choice problems that display partial identification through the lens of the robust Bayes criteria. To do so, we take the general framework of yata2021 and others and embed in it a simple example of the set of priors advocated by GiacominiKitagawa. We describe and contrast (ex-ante) $\Gamma$-minimax regret and (ex-post) $\Gamma$-posterior expected regret and analytically derive optimal solutions with and without randomization.
Our results contain two key messages that we think are valuable to the literature. First, with partial identification and multiple priors, ex-ante and ex-post assessments do not agree in general, whether or not randomized rules are allowed. This may at first seem expected due to dynamic inconsistency of multiple prior Bayes criteria, but was not obvious in view of the set of prior's specific structure. Second, randomization can be optimal in both ex-ante and ex-post problems---it is with loss of generality to exclude them even when regret is evaluated ex-post. The contrast between the results also illustrates a need to better understand the comparative advantages---whether from a theoretical or practical perspective---of using one criterion over the other.
An obvious limitation lies in our use of a convenient but restrictive prior on $\mu$. As we discovered in Section (ref), the $\Gamma$-MMR rule also performs well across other values of $\mu$; however, this may be related to the fact that the globally least favorable prior in this setting has a two-point structure on $\mu$ as well and would obviously not generalize to arbitrary uses of restrictive priors. We provide some insight on more general priors in Appendix (ref). In short, the tractability advantage of $\Gamma$-PER may become pronounced in such settings, although we hope that recent computational developments fernandez2024epsilon,guggenberger2025numerical will attenuate this concern.