EconBase
← Back to paper

Decision Theory for Treatment Choice Problems with Partial Identification

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

68,293 characters · 11 sections · 106 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Decision Theory for Treatment Choice Problems with Partial Identification

\sloppy

abstractWe apply classical statistical decision theory to a large class of treatment choice problems with partial identification. We show that, in a general class of problems with Gaussian likelihood, all decision rules are admissible; it is maximin-welfare optimal to ignore all data; and, for severe enough partial identification, there are infinitely many minimax-regret optimal decision rules, all of which sometimes randomize the policy recommendation. We uniquely characterize the minimax-regret optimal rule that least frequently randomizes, and show that, in some cases, it can outperform other minimax-regret optimal rules in terms of what we term profiled regret. We analyze the implications of our results in the aggregation of experimental estimates for policy adoption, extrapolation of Local Average Treatment Effects, and policy making in the presence of omitted variable bias. Keywords: Statistical decision theory, treatment choice, partial identification.

\onehalfspacing

Introduction

A policy maker must decide between implementing a new policy or preserving the status quo. Her data provide information about the potential benefits of these two options. Unfortunately, these data only partially identify payoff-relevant parameters and may therefore not reveal, even in large samples, the correct course of action. Such treatment choice problems with partial identification have recently received growing interest; for example, see d2021policy, ishihara2021, yata2021, christensen2022optimal, kido2022distributionally or manski2022identification. Several interesting problems that arise in empirical research can be recast using this framework. A non-exhaustive list includes extrapolation of experimental estimates for policy adoption ishihara2021,menzel2023transfer, policy-making with quasi-experimental data in the presence of omitted variable bias \citep*{diegert2022assessing}, and extrapolation of Local Average Treatment Effects \citep*{mogstad2018using,mogstad2018identification}.

In this paper, we analyze such problems in terms of Wald50's (Wald50) Statistical Decision Theory. We do so in a finite-sample framework characterized by a Gaussian likelihood and partial identification.

{ Main Results:} Three optimality criteria are routinely used to endorse or discard decision rules: admissibility, maximin welfare, and minimax regret.

Admissibility. A decision rule is (welfare-)admissible if one cannot improve its expected welfare uniformly over the parameter space. This is usually considered a weak requirement for a decision rule to be considered “good”. Our first result shows that, whenever problems in our setting exhibit partial identification, every decision rule, no matter how exotic, is admissible (Theorem (ref)). As we discuss later, this result stands in stark contrast with applications of the admissibility criterion to point-identified treatment choice problems, in which admissibility meaningfully refines the class of decision rules (by, for example, discarding rules that randomize the policy recommendation). We also show that our result is not tied to the choice of Gaussian likelihood, but instead to the bounded completeness of the statistical model for the available data (Theorem (ref)).

Maximin Welfare. A decision rule is maximin(-welfare) optimal if it attains the highest worst-case expected welfare. Echoing critiques from Savage51 to manski2004statistical, our second result shows that maximin decision rules will preserve the status quo regardless of the data (Theorem (ref)).

Minimax Regret. A decision rule is minimax-regret (MMR) optimal if it attains the lowest worst-case expected regret, where an action's regret is its welfare loss relative to the action that would be optimal if payoff-relevant parameters were known. In some point-identified treatment choice problems, the MMR rule is both essentially unique and nonrandomized canner,stoye2009minimax,tetenov2012statistical. We show this may not be true under partial identification. To make this point, we specialize our framework to the class of treatment choice problems recently studied by yata2021. Our third result shows that for cases where the identified set for payoff relevant parameters is large enough, there are infinitely many MMR optimal decision rules, all of which randomize the policy recommendation for some or all data realizations (Theorem (ref)). This presents an important challenge for the application of MMR, at least if one hopes for the resulting recommendation to be unique. Moreover, as we explain later by means of an example, different MMR optimal rules can lead to meaningfully different policy choices for the same data.

Least Randomizing MMR rule: Finally, we refine the set of MMR optimal rules by identifying the least randomizing (in a sense we make precise) MMR optimal rule. Our main motivation is the following trade-off. Recall that, for a wide range of parameter values, any MMR optimal rule must randomize for some data realizations. At the same time, despite wide adoption of randomized treatment allocations in economics and the social sciences for the purpose of experimentation, it might be difficult in many policy applications to randomize one's policy. Thus, we look for a rule that recommends such randomization as infrequently as possible. We explicitly characterize (in Theorem (ref)) an essentially unique least randomizing rule for the problems considered by yata2021.

We analyze the regret of the least randomizing MMR rule after profiling out some parameters of the risk function. We specifically analyze profiled regret, by considering a parameter of interest and reporting worst-case expected regret at each of these parameter values (in a sense we make precise). We show, in the context of our running example, that our least-randomizing rule can profiled-regret dominate the MMR rule suggested by stoye2012minimax and recently extended by yata2021 for a general class of treatment choice problems with partial identification (Proposition (ref)). More generally, we show that the use of profiled regret renders the rule that uniformly randomizes the policy recommendation inadmissible (Proposition (ref)), where the profiled regret function we consider reports the worst-case regret for a fixed value of the problem's point-identified parameters.

We also discuss the extent to which our least-randomizing rule can be obtained by explicitly penalizing randomized policy recommendations in the policy maker's welfare function. We show that, under some conditions, the least-randomizing rule is minimax regret optimal (among a suitably defined class of decision rules) for a penalty function that penalizes all randomized assignments equally (Proposition (ref)).

{ Applications:} We illustrate the practical implications of our results for three problems that recently arose in applied work.

First, we analyze in detail a running example based on ishihara2021's (ishihara2021) “evidence aggregation” framework. Here, a policy maker is interested in implementing a new policy in country $i=0$. She has access to estimates of the effect of the same policy for other countries $i=1,...,n$ and attempts to extrapolate results using baseline covariates. We give explicit MMR treatment rules for this example. An interesting finding is that, when the identified set is large enough, the least randomizing rule can be related to the estimated bounds on the treatment effect of interest and randomizes only (though not always) if these bounds contain both positive and negative values. This illustrates how an estimator of the identified set can be used for optimal decision making.

Second, we study extrapolation of Local Average Treatment Effects mogstad2018using,mogstad2018identification with binary instrument and no covariates. Here, the payoff-relevant parameter is a “policy-relevant treatment effect” heckman2005structural that corresponds to expanding the complier subpopulation. We show that Theorem (ref) applies in this example, so that all decision rules are admissible. In particular, a decision rule that implements the policy for large values of the usual instrumental-variables estimator is not dominated.

Third, we consider a policy maker who uses quasi-experimental data to decide on a new policy. She is willing to assume a constant treatment effect model and unconfoundedness given a set of covariates $(X,W)$; however, only $X$ is observable and $W$ is not. In this setting, diegert2022assessing argue that researchers may be interested in how much selection on unobservables is required to overturn findings that are based on a feasible linear regression. The least randomizing MMR rule can inform a complementary, decision-theoretic breakdown point analysis: For a given estimated effect of the policy, what is the largest effect of unobserved confounding under which it is still optimal to adopt the seemingly better policy without any hedging? We show that this breakdown point tolerates more confounding than the one of diegert2022assessing.

Related literature. The econometric literature on treatment choice has grown rapidly since manski2004statistical and Dehejia2005. When welfare is partially identified, Manski2000,manski2005social,manski2007identification and Stoye07 provide optimal treatment rules assuming the true distribution of the data is known. stoye2012minimax,stoye2012new, yata2021, and ishihara2021 focus on finite sample MMR optimal rules, and \citet*{twopointprior} on multiple prior MMR rules, in such settings. For different settings with point-identified welfare, finite- and large-sample results on optimal treatment choice rules were derived by canner, ChenGug, HiranoPorter2009,HiranoPorter2020, \citet*{kitagawa2022treatment}, schlag2006eleven, stoye2009minimax, and tetenov2012statistical. christensen2022optimal extend HiranoPorter2009's (HiranoPorter2009) limit experiment framework to a class of partially identified settings; see on this also kido2023locally. Treatment choice is furthermore related to a large literature on optimal policy learning that contains many results for point identified BhattacharyaDupas2012,kitagawa2018should,KT19,MT17,KW20,AW20,KST21,ida2022choosing as well as partially identified kallus2018confounding,ben2021safe,ben2022policy,d2021policy, christensen2022optimal,adjaho2022externally,kido2022distributionally,lei2023policy treatment choice that may condition on covariates. Bayesian aspects of treatment choice are discussed in chamberlain2012.

The remainder of this paper is organized as follows: Section (ref) introduces the formal framework and the running example. Section (ref) is devoted to the aforementioned main results on admissibility, maximin wellfare, minimax regret and the least randomizing MMR rule. The applications are presented in Section (ref). Section (ref) concludes. Appendix (ref) contains proofs of our main results and Appendix (ref) discusses the notion of profiled regret. Additional proofs and results can be found in Online Appendix (ref).

Notation and Framework

Statistical decision theory calls for three ingredients: the menu of actions available, their consequences as a function of an unknown state of the world, and a statistical model of how the data distribution depends on that state. We now present these elements and lay out an example that will be used to illustrate objects, terms, and results throughout.

The policy maker can choose an action $a\in[0,1]$, which we interpret as the proportion of a population that will be randomly assigned to the new policy. Thus, $a=1$ means that everyone is exposed to the new policy and $a=0$ means that the status quo is preserved. Under this interpretation, $a=.5$ means that 50% of the population will be exposed at random to the new policy; however, the formal development equally applies to either individual or population-level randomization. Our interpretation abstracts from integer issues arising with small populations.

The payoff for the policy maker when taking action $a\in[0,1]$ is captured by a welfare function

equation[equation omitted — 77 chars of source]

where $\theta\in\Theta$ is an unknown state of the world or parameter (possibly of infinite dimension) and $W(1,\cdotp):\Theta\rightarrow\mathbb{R}$ and $W(0,\cdotp):\Theta\rightarrow\mathbb{R}$ are known functions. Thus, welfare is linear in the action, a common assumption in the literature.\footnote{In particular, this applies if $W(\cdot,\theta)$ is an expectation and, for the case where randomization is interpreted as fractional assignment, there are no spillover effects or externalities. These assumptions are the default in the literature. An exception is manski2007admissible, who consider the welfare of an action to be a concave transformation of $W(\cdot,\theta)$.} The form of ((ref)) also implies that if $\theta$ were known to the policy maker, her optimal choice of action would simply be

equation[equation omitted — 139 chars of source]

Following HiranoPorter2009, we refer to $U(\theta)$ as the welfare contrast at $\theta$. Thus, the policy maker's optimal action in (ref) is to expose the whole population to the new policy if the welfare contrast is nonnegative and to preserve the status quo otherwise.

The policy maker observes a realization of $Y \in \mathbb{R}^n$ distributed as

equation[equation omitted — 66 chars of source]

Here, the function $m(\cdotp):\Theta\rightarrow\mathbb{R}^n$ and the positive definite covariance matrix $\Sigma$ are known. However, $m(\cdot)$ need not be injective: $m(\theta) = m(\theta')$ does not imply $\theta=\theta'$. As a result, even perfectly identifying $m(\theta)$ (from infinite data) may not identify the optimal action.

In economics applications, the normality assumption in (ref) is unlikely to hold exactly; however, the data can often be summarized by statistics that are asymptotically normal and whose asymptotic variances can be estimated. Treating the limiting model as a finite-sample statistical model then eases exposition and allows us to focus on the core features of the policy problem. Working directly with such a limiting model is common in applications of statistical decision theory to econometrics; see muller2011efficient and the references therein for theoretical support and applications in the context of testing problems and ishihara2021, stoye2012minimax, or tetenov2012statistical for precedents in closely related work.

We finally define a decision rule, $d:\mathbb{R}^{n}\rightarrow[0,1]$, as (measurable) mapping from the data $Y$ to the unit interval. We let $\mathcal{D}_n$ denote the set of all decision rules. We call $d\in \mathcal{D}_n$ non-randomized if $d(y)\in\{0,1\}$ for (Lebesgue) almost every $y\in\mathbb{R}^n$ and randomized otherwise.

{ Running Example:} Our running example is a special case of ishihara2021's (ishihara2021; see also manski2020towards) “evidence aggregation” framework. A policy maker is interested in implementing a new policy in country $i=0$ and observes estimates of the policy's effect for countries $i=1,...,n$. Let $Y=(Y_1,...,Y_n)^\top \in \mathbb{R}^n$ denote such experimental estimates and let $(x_0,\ldots,x_n)$ be nonrandom, $d$-dimensional baseline covariates. The policy maker is willing to extrapolate from her data by assuming that $W(1,\theta)=\theta(x_0)$, $W(0,\theta)=0$, $U(\theta)=\theta(x_0)$ and that

equation*[equation* omitted — 250 chars of source]

where $\theta: \mathbb{R}^d \rightarrow \mathbb{R}$ is an unknown Lipschitz function with known constant $C$. For simplicity, in the example we write $\mu_i$ for $\theta(x_i)$ henceforth, thus $Y_i \sim N(\mu_i,\sigma_i^2)$. We also assume w.l.o.g. that countries are arranged in nondecreasing order of $\left\Vert x_i-x_0 \right\Vert$. While our analysis extends to $x_1=x_0$, we focus on the case of $x_1 \neq x_0$, so that the sign of $\mu_0$ is not necessarily identified. We assume that $(x_1,\ldots,x_n)$ are distinct. Even if this were not the case in raw data, one would presumably want to induce it (by adding fixed effects, whose size can be bounded) because the Lipschitz constraint would otherwise imply that countries with same $x$ exhibit no heterogeneity whatsoever. Finally, we assume $\Vert x_1-x_0\Vert<\Vert x_2-x_0\Vert$, i.e., the nearest neighbor of country $0$ is unique.\footnote{If two or more countries are nearest neighbors, their signals can be merged into one more precise signal.} $\square$

Main Results

In statistical decision theory, three criteria are commonly used to recommend decision rules: admissibility, maximin welfare, and minimax regret.\footnote{See stoye2012new and references therein for theoretical trade-offs between these criteria.} In this section, we show that the application of these criteria to our setting presents nontrivial challenges.

Everything is Admissible

Let $\mathbb{E}_{m(\theta)}[\cdot]$ denote expectation with respect to $Y\sim N(m(\theta),\Sigma)$. Recall the following definition:

defnA rule $d\in\mathcal{D}_{n}$ is (welfare-)admissible if there does not exist $d^{\prime}\in\mathcal{D}_{n}$ such that \[ \mathbb{E}_{m(\theta)}[W(d^{\prime}(Y),\theta)] \geq \mathbb{E}_{m(\theta)}[W(d(Y),\theta)], \quad \forall \theta \in \Theta, \] with strict inequality for some $\theta\in\Theta$.

Thus, a rule is admissible if it is not dominated (in the usual sense of weak dominance everywhere and strict dominance somewhere). This is generally considered a minimal but compelling requirement for a decision rule to be “good” and goes back at least to Wald50. Admissibility can be used to recommend classes of rules whose payoff cannot be uniformly improved and/or whose members improve uniformly on non-members, as in complete class theorems karlin1956theory,manski2007admissible; conversely, one may use it to caution against particular (classes of) decision rules, as was recently done by andrews2022gmm.

Our first result shows that, under mild assumptions, admissibility cannot serve either purpose in our problem. This is because any decision rule is admissible. To formalize this, let

equation[equation omitted — 97 chars of source]

collect all means that can be generated as $\theta$ ranges over $\Theta$. We refer to elements $\mu \in M$ as $\emph{reduced-form}$ parameters because they are identified in the statistical model (ref).\footnote{The notation is consistent with our running example, in which the observable moments are $(\mu_1,\ldots,\mu_n)$.} Define the identified set for the welfare contrast as function of $\mu$ as

equation[equation omitted — 117 chars of source]

and the corresponding upper and lower bounds as

equation[equation omitted — 105 chars of source]

When we refer to models as “partially identified,” we henceforth mean that partial identification obtains on an open set in parameter space (i.e., not almost nowhere).\footnote{For simplicity of exposition, Definition (ref) is stated in a way that forces $M$ to have a non-empty interior. This would, for example, exclude equality constraints. For the purpose of Theorem (ref) below, we can weaken the assumption to allow such cases as long as $\mathcal{S}$ is rich within $M$.}

defn[Nontrivial partial identification] The treatment choice problem with payoff function (ref) and statistical model (ref) exhibits nontrivial partial identification if there exists an open set $\mathcal{S}\subseteq M \subseteq \mathbb{R}^{n}$ such that \begin{align*} I(\mu) & <0<\overline{I}(\mu), for all \mu\in\mathcal{S}. \end{align*}

{ Running Example---Continued:} The identified set for the welfare contrast $\mu_0$ is

equation*[equation* omitted — 136 chars of source]

Its extrema can be written as intersection bounds: \[ \underline{I}(\mu) = \max_{i=1,\ldots,n} \{ \mu_i - C\left\Vert x_i-x_0 \right\Vert \}, \quad \overline{I}(\mu) = \min_{i=1,\ldots,n} \left \{ \mu_i + C \left\Vert x_i-x_0 \right\Vert \right \}. \] For $\mu$ sufficiently close to the zero vector, we therefore have nontrivial partial identification. $\square$

We are now ready to state the first main result.

thmIf a treatment choice problem with payoff function (ref) and statistical model (ref) exhibits nontrivial partial identification in the sense of Definition (ref), then every decision rule $d\in\mathcal{D}_{n}$ is (welfare-)admissible.
proofSee Appendix (ref).

For a proof sketch, suppose by contradiction that some rule $d$ is inadmissible. Then some other rule $d'$ dominates it. This $d'$ must perform weakly better at every $\theta \in m^{-1}(\mathcal{S})$, where $\mathcal{S}$ is the set that appears in Definition (ref). Because of nontrivial partial identification, all $\theta \in m^{-1}(\mathcal{S})$ are compatible with positive and negative welfare contrast $U(\theta)$. This implies that \[ \mathbb{E}_{m(\theta)}[d(Y)] = \mathbb{E}_{m(\theta)}[d'(Y)] \quad \textrm{for each } \theta \in m^{-1}(\mathcal{S}).\] By i) completeness of the Gaussian statistical model in ((ref)) and ii) mutual absolute continuity of the Gaussian and Lebesgue measures in $\mathbb{R}^{n}$, we then have $d(\cdot) = d'(\cdot)$ (Lebesgue) almost everywhere in $\mathbb{R}^{n}$, a contradiction.\footnote{A family $\mathcal{P}$ of distributions $P$ is complete if $\mathbb{E}_{P}[f(X)]=0$ for all $P\in\mathcal{P}$ implies $f(x)=0$ $P$-almost everywhere, for every $P \in \mathcal{P}$. See, for example, lehmann05testing. }

While the statement of Theorem (ref) makes reference to the Gaussian model in (ref), the above proof sketch only uses this model's bounded completeness (and the mutual absolute continuity of the Gaussian and Lebesgue measures in $\mathbb{R}^n$). Indeed, Theorem (ref) in Appendix (ref), establishes a stronger version of Theorem (ref) that applies to boundedly complete statistical models (in a sense we make precise).\footnote{We would like to thank Tim Christensen and the Editor for pointing this out.} It is known that many exponential family distributions lehmann2006theory as well as some location families mattner1993some are boundedly complete. However, our proof does not apply to distributions that are not boundedly complete even if they are close to Normal; for example, if $Y=Z+\epsilon$, where $Z\sim N(m(\theta),\Sigma)$ and $\epsilon$ is independent of $Z$ and contains i.i.d components with characteristic functions containing zeros such as the uniform distributions in $[-\eta,\eta]$ for any $\eta>0$ (c.f. mattner1993some and references therein). While we do not know whether bounded completeness is strictly necessary for this result, it cannot just be dropped. To see this, consider the preceding counterexample with a uniform error term, in which we further impose $n=1$, $\Sigma =0$ and $\eta=1$. Furthermore, suppose $\overline{I}(\mu)=\mu+1$ and $\underline{I}(\mu)=\mu-1$. Then the coin-flip rule $d_{\text{coin-flip}}(Y)=1/2$ would be dominated by the rule that chooses $0$ if $Y<-2$, chooses $1$ if $Y>2$, and $1/2$ otherwise.

remTheorem (ref) shows that the notion of admissibility does not have any refinement power on the class of decision rules considered by the policy maker. We think this result is quite surprising, given that in treatment choice problems with point identification, admissibility does meaningfully refine the class of decision rules. For example, let $n=1$, $\Theta = \mathbb{R}$, $m(\theta) = W(1,\theta) = \theta$, and $W(0,\theta) = 0$. In this case, the policy maker observes a noisy signal, $Y \sim N(\theta, \sigma^2)$, of the payoff-relevant, point-identified parameter $\theta \in \mathbb{R}$. By karlin1956theory's (karlin1956theory) classic result, any decision rule that is not a threshold rule (i.e., is not of form $\mathbf{1}\{ Y > c \}$ for some fixed $c \in \mathbb{R} \cup \{-\infty,\infty\}$) is dominated. In fact, the class of all threshold rules is complete (in the sense that any non-threshold rule is dominated by a threshold rule). Therefore, any decision rule that is not a threshold rule can be dismissed by appealing to the notion of admissibility alone. We think it is quite surprising that introducing partial identification renders all decision rules admissible. Indeed, Theorem 1 implies that, no rule, regardless how eccentric it may appear, is (welfare-)dominated. This makes it much more challenging to recommend a rule or a class of rules, at least without commitment to a more specific decision-theoretic optimality criterion (such as maximin welfare or minimax regret).

Theorem (ref) also admits a more optimistic interpretation. The positive interpretation is that the procedures suggested in the related literature will all perform well relative to one another in some parts of parameter space. This is an important observation because some of these suggestions contain novel or nonstandard components. For example, ishihara2021 place ex-ante restrictions on the class of decision rules, while christensen2022optimal transform the original loss function by profiling out partially identified parameters. By Theorem (ref), all these approaches are at least admissible. This was arguably obvious in the former case (since for any linear threshold rule, it is easy to find a prior that it uniquely best responds to) but certainly not the latter one.

Maximin Welfare is Ultra-Pessimistic

We next analyze the maximin welfare criterion. Our main result echoes earlier findings by Savage51 and manski2004statistical: Maximin typically leads to “no-data rules” that preserve the status quo.

defnA rule $d_{\text{maximin}}\in\mathcal{D}_{n}$ is maximin optimal if \[\inf_{\theta\in\Theta}\mathbb{E}_{m(\theta)}[W(d_{\text{maximin}}(Y),\theta)]=\sup_{d\in\mathcal{D}_{n}}\inf_{\theta\in\Theta}\mathbb{E}_{m(\theta)}[W(d(Y),\theta)]. \]
thmSuppose that there exists $\theta \in \Theta$ such that $U(\theta) \leq 0$. If \begin{equation} \inf_{\theta\in\Theta}W(0,\theta) = \inf_{\theta\in\Theta:U(\theta)\leq0}W(0,\theta), \end{equation} then the no-data rule $d_{\text{no-data}}(y):=0$ is maximin optimal. The maximin value is \[ \inf_{\theta\in\Theta}W(0,\theta).\]
proofSee Appendix (ref).

This result can be seen as follows. When $U(\theta)\leq 0$, it is optimal to preserve the status quo; substituting in for this response, we find that $\inf_{\theta\in\Theta}\mathbb{E}_{m(\theta)}[W(d(Y),\theta)]\leq \inf_{\theta\in\Theta,U(\theta)\leq0}W(0,\theta)$ for any rule $d\in\mathcal{D}_n$. Under condition ((ref)), this upper bound is attained by $d_{\text{no-data}}$.

{ Running Example---Continued:} Theorem (ref) applies to the running example. In particular, the example's maximin welfare equals $0$ and is achieved by never assigning the new policy. $\square$

A similar result was shown by manski2004statistical for testing an innovation with point-identified welfare contrast (a result that we generalize\footnote{In manski2004statistical's (manski2004statistical) example, $W(0,\theta)$ does not depend on $\theta$, so that condition ((ref)) is trivially satisfied.}), and the concern can be traced back at least to Savage51. There is a discussion of whether such “ultrapessimism” occurs, in a technical sense, more generically with maximin utility versus minimax regret ParmigianiTD,SadlerTD. However, a string of more optimistic results regarding MMR canner,stoye2009minimax,stoye2012new,tetenov2012statistical,yata2021 suggests that, with state spaces that describe real-world decision problems, the concern is more salient for maximin.\footnote{Given that we next elaborate on multiplicity of MMR rules, it is of interest to also discuss the potential multiplicity of maximin rules: if $W(0,\theta) $ is treated as known (e.g., equal to 0), the maximin rule may be unique. However, in settings where both $W(1,\theta)$ and $W(0,\theta)$ are unknown and can be arbitrarily “bad”, the typical result is that all decision rules are maximin. }

Minimax Regret Admits Many Solutions

In view of Theorem 1 and Theorem 2, it seems natural to consider the minimax regret (MMR) optimality criterion. It is known that some point-identified treatment choice problems admit nonrandomized, essentially unique MMR optimal rules. In contrast, our results below will show that uniqueness of MMR optimal rules should not be expected to hold generally with partial identification. To make this point, we consider a class of problems where finite-sample minimax optimal results are available. Within this class, we show that, if the identified set for the welfare contrast is sufficiently large relative to sampling error, then i) we can find infinitely many MMR optimal decision rules, all of which depend on data through an optimal linear index and ii) under weak additional conditions, any MMR rule that depends on the data through the optimal linear index must be randomized.

Our numerical and analytic findings below (see Figure (ref) and Proposition (ref)) demonstrate that, beyond having the same worst-case regret, two MMR optimal rules might be qualitatively and quantitatively very different. This presents a practical challenge for a policy maker as, given the same data realizations, two MMR optimal rules may recommend different policy actions.

The expected regret of a decision rule $d\in\mathcal{D}_n$ in state $\theta$ is its expected welfare loss compared to the oracle rule:

eqnarray[eqnarray omitted — 142 chars of source]
defnA rule $d^*\in \mathcal{D}_{n}$ is minimax regret (MMR) optimal if \begin{equation} \sup_{\theta \in \Theta}R(d^*,\theta)=\inf_{d\in\mathcal{D}_{n}}\sup_{\theta \in \Theta}R(d,\theta). \end{equation}

Solving minimax regret problems is often hard, both analytically and algorithmically. Algorithms exist for certain cases yu1995min, chamberlain2000econometric, fernandez2024epsilon, guggenberger2025numerical, but it is known that obtaining the minimax solution of a decision problem---and sometimes even deciding whether a minimax solution exists---is NP hard in general; see daskalakis2021complexity. In important, recent work, yata2021 characterizes MMR optimal rules for a large class of binary action problems. He imposes the following restrictions on the parameter space and welfare function.

assumption\begin{itemize} • $\Theta$ is convex, centrosymmetric (i.e., $\theta\in\Theta$ implies $-\theta\in\Theta$) and nonempty. • $m(\cdot)$ and $U(\cdot)$ are linear.\end{itemize}

{ Running Example---Continued:} Our example satisfies Assumption (ref). In particular, the space of $C$-Lipschitz functions, $\theta:\mathbb{R}^{d} \rightarrow \mathbb{R}$ is convex as well as centrosymmetric, and the functions $U(\cdot)$ and $m(\cdot)$ are linear in $\theta$ as they simply report the values of $\theta$ at points $(x_0, \ldots, x_n)$. $\square$

Under Assumption (ref), yata2021 shows existence of an MMR rule that depends on the data only through $(w^*)^{\top} Y$, where the unit vector $w^*$ can be approximated by solving a sequence of tractable optimization problems. When the identified set for the welfare contrast at $\mu= \mathbf{0} := 0_{n\times1}$ is large enough, yata2021's (yata2021) MMR rule can be expressed as

equation[equation omitted — 131 chars of source]

for some uniquely characterized $\tilde{\sigma} > 0$.\footnote{For readability, we slightly abuse notation and, when a decision rule $d$ depends on data $Y$ only through a simple feature like $w^{\top}Y$ or $Y_1$, we write $d(Y)$ as $d(w^{\top}Y)$ or $d(Y_1)$. } Moreover, it then has two important algebraic properties:

equation[equation omitted — 130 chars of source]

and

equation[equation omitted — 159 chars of source]

In words, these features are as follows: First, if the mean function $m(\cdot)$ equals zero, expected exposure to the new policy is $1/2$. Second, worst-case regret occurs precisely at this point.\footnote{yata2021 also provide sufficient conditions under which these conditions hold true. Our proof is constructive, and our results below apply to a large class of models with these properties and shall not be considered as a specific example. } The latter is due to careful calibration of the MMR decision rule, and one might conjecture that it renders this rule unique. However, the following result establishes the contrary.

thmConsider a treatment choice problem with payoff function (ref) and statistical model (ref) that exhibits nontrivial partial identification in the sense of Definition (ref). Suppose that Assumption (ref) holds and that there is a MMR optimal rule $d^*$ that depends on the data only through $(w^*)^\top Y$ and that satisfies (ref) and (ref). If there exists $\mu\in M$ such that $\overline{I}(\mu)>\overline{I}(\mathbf{0})$ and $\overline{I}(\mathbf{0})$ is large enough, then \begin{itemize} • There are infinitely many MMR optimal rules. • Any MMR rule that depends on the data only through $(w^*)^{\top}Y$ (and is weakly increasing in this argument) must randomize for some data realizations. • If $\overline{I}(\mu)$ is differentiable at $\mu=\mathbf{0}$, then no linear threshold rule, i.e., no rule of form $\mathbf{1}\{w^{\top}Y\geq c\}$ for some $w\in\mathbb{R}^{n}$ and $c\in\mathbb{R}\cup\{-\infty,\infty\}$, is MMR optimal. \end{itemize}
proofSee Appendix (ref).

To be clear, this finding applies if the problem is sufficiently far from point identification, with an exact condition given in the proof.\footnote{This condition will always be met if $\Sigma$ is sufficiently small, holding other parameter values fixed. Heuristically, this means that our multiplicity result is always relevant as sample size becomes large. However, to what extent our result is a useful characterization in a limit experiment is beyond the scope of the current paper and an interesting question that we leave for future research. } Close to point identification and under mild additional assumptions, yata2021 shows MMR optimality of a linear threshold rule.

Part (i) of Theorem (ref) is established constructively: We show that, whenever $d^*_{\text{RT}}$ is MMR optimal, then so is the piecewise linear rule

equation[equation omitted — 246 chars of source]

where $\rho^*>0$ is characterized in Appendix (ref), Equation (ref). This implies existence of infinitely many MMR rules because the set of such rules is closed under convex combination.

Next, if the identified set is large enough for given $\Sigma$ (or as $\Sigma$ vanishes for given identified set), all of the above rules randomize for some data realizations. A natural question to ask is whether this feature is shared by all MMR rules. Parts (ii) and (iii) give qualified affirmative answers: If we focus on decision rules that increase in $(w^*)^{\top}Y$ and if $\overline{I}(\mathbf{0})$ is large enough, then randomization is necessary for MMR optimality; if bounds are furthermore differentiable in reduced-form parameters at $\mathbf{0}$, randomization is necessary for any MMR rule that depends on a linear index of the data.\footnote{The class of nonrandomized but otherwise unrestricted rules is not interestingly different from the class of all rules due to the possibility of purifying randomized rules. See Remark (ref) for discussion and references. Similarly, the differentiability condition is needed to preclude that some component of $Y$ can effectively be used as randomization device.} In particular, the threshold rule

equation[equation omitted — 86 chars of source]

is not MMR optimal.

{ Running Example---Continued:} Theorem (ref) applies to our running example. We next improve on this observation by providing explicit MMR optimal rules for the example. Broadly speaking, these are characterized by a weighting of studentized signals that resembles a triangular kernel for small enough $C$ and turns into a nearest neighbor kernel as $C$ becomes large. $\square$

propIn the running example, the following statements hold true. \begin{itemize} • If \begin{equation} C \left\Vert x_1 - x_0 \right\Vert < \sqrt{\pi / 2} \cdot \sigma_1, \end{equation} then the following decision rule is uniquely (up to almost sure agreement) MMR optimal: \begin{eqnarray*} d_{m_0^*}& :=& \mathbf{1}\{w_{m_0^*}^\top Y \geq 0\}, \\ w_{m_0^*}^\top &:= & \left(1, \frac{\max \{m_0^* - C\left\Vert x_2-x_0 \right\Vert,0\}/\sigma_2^2}{(m_0^* - C\left\Vert x_1-x_0 \right\Vert)/\sigma_1^2},\ldots,\frac{\max \{m_0^* - C\left\Vert x_n-x_0 \right\Vert,0\}/\sigma_n^2}{(m_0^* - C\left\Vert x_1-x_0 \right\Vert)/\sigma_1^2}\right), \end{eqnarray*} where $m_0^* > C\left\Vert x_1-x_0 \right\Vert$ solves a simple fixed point problem ((ref) in Online Appendix (ref)). • If \begin{equation*} C \left\Vert x_1 - x_0 \right\Vert = \sqrt{\pi / 2} \cdot \sigma_1, \end{equation*} then $\mathbf{1}\{ Y_1 \geq 0\}$ is MMR optimal. • If \begin{equation} C \left\Vert x_1 - x_0 \right\Vert > \sqrt{\pi / 2} \cdot \sigma_1, \end{equation} then the rule \begin{equation} d^{*}_{linear}(Y_1):=\begin{cases} 0, & Y_1<-\rho^{*},\\ \frac{Y_1+\rho^{*}}{2\rho^{*}}, & -\rho^{*}\leq Y_1\leq\rho^{*},\\ 1, & Y_1>\rho^{*}, \end{cases} \end{equation} where $\rho^* \in (0,C \left\Vert x_1-x_0 \right\Vert)$ is uniquely defined by $\rho^*=C \left\Vert x_1-x_0 \right\Vert (1-2\Phi\left(\rho^{*}/\sigma_1\right))$ is MMR optimal. So is the rule $d_{\text{RT}}^*(Y_1):=\Phi\bigl(Y_1/\tilde{\sigma})$ with \begin{equation} \tilde{\sigma} = \sqrt{ 2C^2 \left\Vert x_1-x_0\right\Vert^2/\pi - \sigma_1^2 } \end{equation} as well as all convex combinations of these rules. • If Equation (ref) holds, no linear threshold rule is MMR optimal. \end{itemize}
proofSee Online Appendix (ref).
remAll decision rules above converge to the one from case (ii) as $C \Vert x_1 - x_0 \Vert \to \sqrt{\pi / 2} \cdot \sigma_1$.
remThis solution relates to the literature as follows. The problem is within the framework considered by yata2021 (who uses results from stoye2012minimax), and his analysis applies; in particular, our linear index differs from his $w^*$ only by a more explicit characterization. Although some of our proof steps use algebra from stoye2012minimax, the alternative solutions in (iii), the uniqueness statement, and part (iv) are entirely new. ishihara2021 numerically find a solution within the class of symmetric threshold rules (i.e., rules of form $\mathbf{1}\{w^{\top}Y\geq 0\}$). This in principle recovers the global solution if $C \left\Vert x_1 - x_0 \right\Vert \leq \sqrt{\pi / 2} \cdot \sigma_1$ but will exclude all globally MMR optimal decision rules otherwise. That said, ishihara2021's (ishihara2021) solution approach applies considerably more generally. This is because we view their approach primarily as a numerical strategy for finding a minimax solution over a constrained class. Indeed, given a statistical model and a risk function, one can always try to find a minimax decision rule numerically within a class of threshold rules, which may or may not pin down the true minimax rule.
remIf $\mu$ is exogenous and known, then the decision rule $d^*_{\text{known }\mu}(\mu):= \max\{\min\{\overline{I}(\mu)/(\overline{I}(\mu)-\underline{I}(\mu)),1\},0\}$ uniquely attains MMR Manski2007. Only our new rule $d^*_\text{linear}$ converges to $d^*_{\text{known }\mu}$ in certain special cases. Similarly, ishihara2021 discuss that an analogous convergence fails for any MMR rule that they propose. This may appear puzzling; however, any presumption that MMR rules “should” converge to such limits delicately depends on how one conceives the limit of the decision problem. Hence, it is not clear that we observe failure of any convergence that “should” have occurred.
figure[figure omitted — 720 chars of source]

We conclude that different MMR rules can lead to rather different policy actions for the same data. This difference is illustrated in Figure (ref) for parameter values calibrated to ishihara2021's (ishihara2021) empirical example. How serious a challenge it is depends on one's view. If one truly thinks of MMR as encoding a decision maker's complete preferences, and hence of competing optimal rules as mutually indifferent, it is not much of a concern. However, it may make it harder to communicate MMR-based decisions to policy makers. In addition, our next results below will show that one may plausibly have preferences among the different MMR rules.

{ Running Example---Continued:} Recall that, for some parameter values, there are infinitely many MMR rules depending on the data only through $Y_1$. Thus, to better compare and visualize different rules, consider the $w^*$-profiled regret for $w^*=(1,0,\ldots,0)^{\top}$ defined as follows (also see the general definition and more detailed discussions of profiled regret in Appendix (ref)):

eqnarray*[eqnarray* omitted — 215 chars of source]

and we can easily verify that $ \gamma \in \mathbb{R}$. $\square$

For the same parameters used in Figure (ref), Figure (ref) depicts $w^*$-profiled regret over the range $\gamma \in [-30,30]$ for four decision rules. The blue (solid) line is $d^*_{\text{linear}}$ (see Equation (ref)); the red (dotted) line is $d^*_{\text{RT}}$ (see Equation (ref)); the black (bimodal, solid) line is the symmetric threshold rule $d^*_0(Y)=\mathbf{1}\{Y_1 \geq 0\}$; and the green (dashed) line is $d_{\text{coin-flip}}=1/2$.\footnote{Appendix (ref) presents algebraic and computational details that underlie this figure.}

figure[figure omitted — 249 chars of source]

An immediate use of Figure (ref) is to compare decision rules in terms of their worst-case regret. For example, consistent with Theorem (ref)-(ii), $d^*_0$ is not MMR optimal (the maximum of the black curve is clearly above those of $d^*_{\text{RT}}$ and $d^*_{\text{linear}}$). For the specific parameter values used here, we can furthermore show that $d_0^*$ is minimax regret optimal among linear threshold rules; thus, the black line also illustrates the minimax regret efficiency loss from restricting attention to linear threshold rules. Figure (ref) furthermore reveals interesting differences among MMR optimal rules. In particular, $d^*_{\text{linear}}$ appears to have smaller $w^*$-profiled regret than $d^*_{\text{RT}}$ virtually everywhere. Indeed, Proposition (ref) in Appendix (ref) shows that whenever condition (iii) of Proposition (ref) holds, $d^*_{\text{linear}}$ dominates $d^*_{\text{RT}}$ in terms of $w^*$-profiled regret for a large range of values of $C \left\Vert x_1-x_0 \right\Vert$ and $\sigma_1$, including those used in Figure (ref) (though there exists parameter values under which the dominance does not hold for small nonzero $\gamma$). Moreover, the considerable difference between profiled regret functions in Figure (ref) continues beyond the figure: The ratio $\overline{R}_{w^*}(d^*_{\text{linear}},\gamma)/\overline{R}_{w^*}(d^*_{\text{RT}},\gamma)$ decays to zero at exponential rate as $\gamma \to \pm \infty$. Finally, Figure (ref) illustrates that $d_{\text{coin-flip}}$ is $w^*$-profiled regret inadmissible in this example. In fact, profiled-regret inadmissibility of $d_{\text{coin-flip}}$ holds in a more general framework; see Proposition (ref) in Appendix (ref) for additional results.

Least Randomizing MMR Optimal Rules

We next argue that further refining the MMR criterion presents an interesting research opportunity and may even lead to unique recommendations. To this purpose, we propose consideration of, and characterize, the least randomizing MMR rule. Our main motivation is that, despite the wide adoption of randomized treatment allocations in economics and the social sciences, policy makers might shy away from exposing only a fraction of a population to the new policy. Thus, we attempt to recommend actions $a \in (0,1)$ as infrequently as possible. We justify our approach by a decision maker with a lexicographic preference (See Remark (ref) below). Alternatively, one may also explicitly incorporate aversion to randomization in the loss function and try to solve the MMR criterion under a modified risk function. See Section (ref) for detailed discussions.

To formalize this, observe that both $d^*_\text{RT}$ and $d^*_{\text{linear}}$ can be considered smoothed versions of $d^*_0$ in a sense that we now make precise. Let $F:\mathbb{R}\rightarrow[0,1]$ be a c.d.f. and consider a decision rule of form $F \circ w^*:=F((w^*)^{\top}Y)\in \mathcal{D}_n$; that is, the step function $d^*_0$ was smoothed into a c.d.f. We will restrict attention to c.d.f.'s whose associated distributions are symmetric (i.e., $F(-x)=1-F(x)$) and unimodal (i.e., $F(\cdot)$ is convex for $x\leq 0$ and concave otherwise). Let $\mathcal{F}$ be the set of all such c.d.f's and let \[ \tilde{\mathcal{D}}_{n}:=\{F\circ w^*\in\mathcal{D}_{n}: F\in \mathcal{F}\}. \] Note that each rule $F\circ w^* \in \tilde{\mathcal{D}}_n$ depends on the data only via $(w^*)^\top Y$ and is nondecreasing in $(w^*)^\top Y$. Moreover, for each $F\circ w^*\in \tilde{\mathcal{D}}_n$, the interval on which treatment assignment is randomized equals (up to closure)

equation[equation omitted — 169 chars of source]

All MMR decision rules considered in this paper are in $\tilde{\mathcal{D}}_{n}$. We next show that $d^*_{\text{linear}}$ is least randomizing among them and among all other MMR decision rules that might exist in this class.

thmSuppose all conditions of Theorem (ref) hold. If $F\circ w^*\in\tilde{\mathcal{D}}_n$ is MMR optimal, then $V(d^*_{\text{linear}}\circ w^*)\subseteq V(F\circ w^*)$, with equality if and only if $F=d^*_{\text{linear}}$.
proofSee Appendix (ref).

In words, any symmetric, weakly increasing and unimodal MMR optimal rule that depends on data only via $(w^*)^{\top} Y$ must have a randomization area that is wider than that of $d^*_{\text{linear}}$, strictly so if it is a meaningfully distinct rule. Thus, the least randomizing criterion provides a pragmatic and unique refinement among the set of known MMR optimal rules.

To establish Theorem (ref), we first show that for any rule $F\circ w^* \in \tilde{\mathcal{D}}_n$, its expected regret at any $\theta$ for which $(w^*)^\top m(\theta)=0$ equals the MMR value of the problem. If $F\circ w^* \in \tilde{\mathcal{D}}_n$ is MMR optimal, its expected regret must therefore be maximized at $(w^*)^\top m(\theta)=0$. For any symmetric and unimodal c.d.f. $F$ with $V(d^*_{\text{linear}}\circ w^*)\nsubseteq V(F\circ w^*)$, we can show that a necessary condition for this maximization fails.

{ Running Example---Continued:} Recall that, applied to the running example and for $C$ large enough, $d^*_{\text{linear}}$ can be expressed as (ref). Of note, an identified set for $\mu_0$ given the true mean of $Y_1$ (i.e., $\mu_1$) is \[ \left [ \: \mu_1 - C \left\Vert x_1-x_0 \right\Vert\:,\: \mu_1 + C \left\Vert x_1-x_0 \right\Vert\: \right], \] which can be estimated naturally by

equation[equation omitted — 155 chars of source]

As $d^{*}_{\text{linear}}$ only randomizes when $\left\vert Y_1 \right\vert < \rho^*$, we see that interval (ref) contains $0$ whenever $d^*_{\text{linear}}$ randomizes. Equivalently, $d^*_{\text{linear}}$ always (never) implements the new policy when (ref) is to the right (left) of zero. In this sense, the estimated identified set is explicitly used for decision making.\footnote{One may consider a scenario in which, to aid interpretability of decision rules, the decision maker does not wish to randomize whenever the estimated identified set for $\mu_0$ (ref) does not contain zero, i.e. the estimate suggests that the sign of $\mu_0$ is identified. This restricts decision rules to the following set:

equation[equation omitted — 231 chars of source]

As (ref) is a subset of $\tilde{\mathcal{D}}_n$ containing $d^*_{\text{linear}}$, Theorem (ref) implies that $d^*_{\text{linear}}$ is also least-randomizing MMR in (ref). Moreover, $d^*_{\text{linear}}$ is so far the only known rule in the literature in set (ref) that is also globally MMR optimal. } In contrast, $d^*_\text{RT}$ always randomizes the policy recommendation, although for large $Y_1$ the fraction of population assigned to treatment will be large. $\square$

remConsider a decision maker who wishes to pick an optimal rule according to MMR but is hesitant to implement randomized rules due to additional inconvenience cost from implementation or concerns over ex-post fairness (i.e., units in the same population may receive different treatments). In this case, it is indeed possible to define the MMR problem among nonrandomized rules. However, if we do not place any restriction to the class of nonrandomized rules, game-theoretic purification arguments dww,purify suggest existence of nonrandomized solutions that emulate arbitrarily well the risk profile of the randomized MMR optimal rule. Moreover, these solutions would be unattractive and unnatural; for example, they cannot be monotone due to Theorem (ref)(iii). In light of these observations, we think there is little value in pursuing solutions over nonrandomized rules in our setup, unless one focuses on a specific class of benign or reasonable nonrandomized rules, e.g. as in ishihara2021. Furthermore, within the class of rules $\tilde{\mathcal{D}}_n$, our least randomizing MMR rule can be justified by a decision maker who possesses a lexicographic preference with a priority given to minimizing MMR criterion over minimizing the inconvenience cost of randomization (in our case, the cost is proxied by the Lebesgue measure of the realization of $(w^*)^\top Y$ taking a value in $(0,1)$). Such lexicographic ordering of statistical decisions has ample precedents in the literature, e.g., for statistical decision rules, we first remove dominated rules and then choose among undominated ones according to further optimality criteria Wald50,manski2021econometrics; In the Neyman-Pearson paradigm of selecting hypothesis tests, one first controls size and then maximizes power; In choosing estimators, it is customary to focus on unbiased estimators, among which the one that minimizes variances is regarded as optimal.

Further Applications

Extrapolating Local Average Treatment Effects

We next apply our analysis to extrapolation of Local Average Treatment Effects mogstad2018using,mogstad2018identification. Let $Z\in\{0,1\}$ be a binary instrument, $D\in\{0,1\}$ a binary treatment assignment, and $(Y(1),Y(0))$ potential outcomes under treatment and control. As usual, the observed outcome is $Y=DY(1)+(1-D)Y(0)$. To simplify exposition, we assume that there are no covariates and that $Y(1),Y(0)\in\{0,1\}$. Following heckman1999local,heckman2005structural,\footnote{See, for example, Assumption I and Equation (2) in mogstad2018identification. Also see imbens1994identification for additional references.} let $p(z):=P\{D=1 \mid Z=z\}$ be the propensity score and write $D = \mathbf{1}\{ V \leq p(Z)\}$, where $(V\mid Z=z) \sim \textrm{Unif}[0,1]$. The parameter space $\Theta$ contains all tuples $\theta:=(p(1),p(0),\text{MTE}(\cdot))$, where $p(1)\in[0,1]$, $p(0)\in[0,1]$, $p(1)\geq p(0)$, and $\text{MTE}(\cdot)$ is the marginal treatment effect function \[\text{MTE}(v):=\mathbb{E}[Y(1)-Y(0)\mid V=v].\] The policy maker observes

equation[equation omitted — 191 chars of source]

where

eqnarray*[eqnarray* omitted — 185 chars of source]

are the population reduced-form and first-stage coefficients and $\Sigma$ is positive definite. We assume that the policy of interest would expand the complier subpopulation through an additive shift of size $\alpha>0$ in the propensity score. mogstad2018using show that the payoff relevant parameter then is the “policy-relevant treatment effect” heckman2005structural \[\text{PRTE}(\alpha)=\mathbb{E}[Y(1)-Y(0)\mid V\in(p(0),p(1)+\alpha]]. \] Suppose the welfare contrast $U(\theta)$ equals $\text{PRTE}(\alpha)-\text{PRTE}(0)$, which can be written as

equation[equation omitted — 189 chars of source]

Hence, the decision maker wants to find an optimal treatment policy given partial identification of parameter (ref) in model (ref). In Online Appendix (ref), we verify that Theorem (ref) applies. Therefore, any decision rule is admissible in this example. For example, implementing a policy for large values of the IV estimator would be admissible, as would be the approach of christensen2022optimal, who discuss the same application.

Decision-theoretic Breakdown Analysis

Consider a policy maker who uses quasi-experimental data but is worried about confounding. More specifically, she assumes a constant treatment effect model and unconfoundedness given covariates $(X,W)$, motivating the linear regression model \[ Y=\gamma_{0}+\beta_{\text{long}}D+\gamma_{1}^{\top}X+\gamma_{2}^{\top}W+e, \] where $Y$ is observed outcome, $D$ is the binary treatment, and $e$ is a projection residual. The infeasible optimal treatment policy is $\mathbf{1}\{\beta_{\text{long}}\geq0\}$. However, $W$ is unobserved, so that the policy maker can only estimate the “medium” regression \[ Y=\pi_{0} +\beta_{\text{med}} D+\pi_{1}^{\top} X+u, \] where $u$ is a projection residual.\footnote{We express all regressions as projections for alignment with the literature and because only projection algebra is used. However, motivating $\mathbf{1}\{\beta_{\text{long}}\geq 0\}$ as optimal usually requires causal interpretation and therefore slightly stronger assumptions on $e$; in other words, readers may want to think of the long regression as causal and the medium one as best linear prediction. See Hansen, whose notation we also borrow, for a lucid discussion.} In general, if there exists selection on unobservables (that is, $D$ is correlated with $W$), then $\beta_{\text{long}}$ is only partially identified. Specifically, diegert2022assessing show that the identified set of $\beta_{\text{long}}$ given $\beta_{\text{med}}$ is \[ \beta_{\text{long}}\in[\beta_{\text{med}}-k,\beta_{\text{med}}+k], \] where \[ k:=

cases\sqrt{\frac{\operatorname{var}\left(Y^{\bot D,X}\right)}{\operatorname{var}\left(D^{\bot X}\right)}\frac{\overline{r}_{D}^{2}R_{D\sim X}^{2}}{1-R_{D\sim X}^{2}-\overline{r}_{D}^{2}}},& \quad if 0\leq\overline{r}_{D}< \sqrt{1-R^2_{D\sim X}},\\ \infty,&\quad if \overline{r}_{D}\geq \sqrt{1-R^2_{D\sim X}},

\] and where $\operatorname{var}\left(Y^{\bot D,X}\right)$ is the variance of the residual from projecting $Y$ onto $(1,D,X)$, $\operatorname{var}\left(D^{\bot X}\right)$ is the variance of the residual from projecting $D$ onto $(1,X)$, $R_{D\sim X}^{2}$ is the $R^{2}$ from projecting $D$ onto $(1,X)$, and $\overline{r}_{D}\geq0$ is a user-specified sensitivity parameter that measures the relative importance of selection on unobservables versus selection on observables.

diegert2022assessing use this result to ask: How strong does omitted variables bias have to be to potentially overturn findings based on $\beta_{\text{med}}$? At population level, the answer is that this can happen if $\left\vert k \right\vert > \left\vert \beta_{\text{med}}\right\vert$, a condition that can be related to primitive parameters through the above display and for which diegert2022assessing provide estimation and inference theory.

Suppose now that there is an estimator $\hat{\beta}_{\text{med}} \sim N(\beta_{\text{med}},\sigma^{2})$. The results of diegert2022assessing imply an estimated breakdown point $\tilde{k}(\hat{\beta}_\text{med}):=\hat{\beta}_{\text{med}}$ for positive $\hat{\beta}_{\text{med}}$. Our results apply upon letting $\theta =(\beta_{\text{long}},\beta_{\text{med}})^\top\in \mathbb{R}^2$, $U(\theta)=\beta_{\text{long}}$ and $m(\theta)=\beta_{\text{med}}$. In particular, when $k>\sqrt{\frac{\pi}{2}}\sigma$, there are infinitely many MMR optimal rules, with the least randomizing one among known ones being

equation[equation omitted — 308 chars of source]

where $\rho^*>0$ uniquely solves $\rho^*=k(1-2\Phi(-\rho^*/\sigma))$. When $k\leq\sqrt{\frac{\pi}{2}}\sigma$, we know from stoye2012minimax that $d^*_0(\hat{\beta}_{\text{med}}) = \mathbf{1}\{\hat{\beta}_{\text{med}}\geq0\}$ is essentially uniquely MMR optimal.

figure[figure omitted — 214 chars of source]

These results motivate a complementary breakdown analysis guided by statistical decision theory. For given $\hat{\beta}_{\text{med}}>0$, we can ask: How large could $k$ have to be so that the MMR optimality criterion still supports assigning the new policy without any hedging?\footnote{Informal exploration of this question goes back at least to stoye2009partial.} Due to its least randomizing property, $d_{\text{linear}}^*$ implies the tightest possible answer to this question. Specifically, MMR supports non-randomized policy assignment up to the “decision theoretic breakdown point”

eqnarray*[eqnarray* omitted — 352 chars of source]

In contrast, the implied breakdown point of $d^*_{\text{RT}}$ is a constant across all values of $\hat{\beta}_{\text{med}}$, as $d^*_{\text{RT}}$ always randomizes regardless of data realizations when $k>\sqrt{\frac{\pi}{2}}\sigma$. Figure (ref) displays both $\tilde{k}(\hat{\beta}_\text{med})$ and $\bar{k}(\hat{\beta}_\text{med})$ when $\sigma=1$. It turns out that the decision theoretic breakdown point tolerates more ambiguity; this difference is salient for smaller values of $\hat{\beta}_{\text{med}}$ and vanishes as $\hat{\beta}_{\text{med}}$ diverges.

Conclusion

In this paper, we used statistical decision theory to study treatment choice problems with partial identification. For a large and empirically relevant class of such problems, we show that every decision rule is admissible, that maximin welfare optimality criterion often select no-data decision rules, and that there are infinitely many minimax regret optimal rules, all of which randomize the policy action at least for some data realizations. These results stand in stark contrast with treatment choice problems with point-identified welfare.

We also provide a decision rule that is least randomizing in a large class of MMR optimal rules including all known ones. We show, in the context of our running example, that our least-randomizing rule can profiled-regret dominate other MMR rules.

We illustrate our results in three applications that arise in applied work: extrapolation of experimental estimates for policy adoption, policy-making with quasi-experimental data when omitted variable bias is a concern, and extrapolation of Local Average Treatment Effects.