EconBase
← Back to paper

Asymptotics for Treatment Choice with Partial Identification

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

99,937 characters

Local Asymptotics for Treatment Choice with Partial Identification


\title{Local Asymptotics for Treatment Choice \\ with Partial Identification\thanks{We thank Jack Porter for encouragement at an early stage of this project and Karun Adusumilli, Tim Armstrong, Xinyue Bei, Tim Christensen, Kei Hirano, Toru Kitagawa, Brendan Kline, Han Xu, and Kohei Yata as well as audiences at Bonn, TAMU, UT Austin, as well as participants at EcoSta 2026 and the Iowa Econometrics Conference for feedback. William Pan, Elliott Serna and Sara Yoo provided excellent research assistance. We gratefully acknowledge financial
support from the NSF under grant SES-2315600.}}

\author{Jos\'e Luis Montiel Olea\thanks{Email: [email removed]}\and Chen Qiu\thanks{Email: [email removed]}  \and J\"{o}rg Stoye\thanks{Email: [email removed]}}

\date{August 2026}

\maketitle


\vspace{-0.5cm}

\begin{abstract}
We provide a new  asymptotic framework to
derive approximately optimal treatment assignments when sampling noise from data is compounded by fundamental uncertainty due to partial identification. We recenter the reduced-form parameter around its \emph{least-favorable} configuration and consider drifting parameter sequences that yield both diminishing levels of sampling uncertainty and of partial identification. We characterize the limiting decision problem as a normal location shift model with a suitable limiting identified set.
We apply our results to treatment choice problems with contaminated outcomes, to robust welfare analyses with partially identified consumer surplus, and to the problem of aggregating experimental estimates for policy adoption.



\end{abstract}
\medskip
\textsc{Keywords}: statistical decision theory, treatment assignment, minimax regret, limit experiment, partial identification.

\newpage{}

\onehalfspacing

\section{Introduction}

A policy maker must decide between implementing a new policy or preserving the status quo. Her data provide information about the potential benefits of these two options. Unfortunately, these data only \emph{partially identify} payoff-relevant parameters and may therefore not reveal, even in large samples, the correct course of action. The question of how to best use the available data in such \emph{treatment choice problems with partial identification} has received recent attention in the literature; see, for example, \cite{ishihara2021}, \cite{yata2021}, \cite*{christensen2022optimal}, and \cite{kido2023locallyasymptoticallyminimaxstatistical}. Several interesting problems that arise in empirical research can be recast  using this framework; see \cite{MQS2023decision}.

In this paper, we use Le Cam's limits of experiments framework \citep{le2000asymptotics,le2012asymptotic} to propose a simple local asymptotic approximation for a general class of treatment choice problems with partial identification.  In the problems that we study, the decision maker observes a random sample of size $n$ of a vector-valued outcome variable $Y_i$. There is a point-identified parameter $\gamma \in \mathbb{R}^{d_m}$ that flexibly determines the distribution of the data. There is also a payoff-relevant parameter, $U^{*} \in \mathbb{R}$, that is partially identified. As in \cite{christensen2022optimal}, we assume that the decision maker can use $\gamma$ to deduce restrictions on $U^*$. The loss function for the decision problem is given by the usual regret loss $L(a,U^*,\gamma) := U^* ( \mathbf{1}\{U^* \geq 0\}-a)$, where we interpret the action $a \in [0,1]$ as the fraction of the population exposed to the new policy.

We show that the treatment choice problems that we have just described above can be conveniently approximated by a simpler problem. In the limiting decision problem, the decision maker observes a $d_m$-valued multivariate normal vector $\Delta$ with unknown mean, $h$, but known variance. There is a payoff-relevant parameter $\mu^* \in \mathbb{R}$ that, for each $h$, is known to belong to a nonempty interval $I_{\infty}(h):=[\underline{I}_{\infty}(h),\overline{I}_{\infty}(h)]$. We refer to $I_{\infty}(h)$ as the \emph{limiting (local) identified set}. The endpoints of the interval are possibly nonlinear transformations of $h$. The limiting loss function takes the form $L_{\infty}(a,\mu^*,h):= \mu^*(\mathbf{1}\{\mu^* \geq 0\}-a)$.

Our Theorem \ref{thm: general} then shows that if we find a minimax rule for the limiting decision problem, we can  construct a sequence of decision rules in the original problem that are \emph{locally asymptotically minimax regret} optimal (in a sense we make precise). We then provide specific solutions to the limiting problem when the endpoints of $I_{\infty}(h)$ are affine (Theorems \ref{thm:full-diff} and \ref{thm:full-diff-asymptotic}), which occurs when the bounds of the original identified set for $U^*$ are differentiable. Since in many partially identified models the bounds of the original identified set are only directionally differentiable, we also consider the limiting problem without imposing the affinity of the bounds of the set $I_{\infty}(h)$. In this case, we show that when the parameter space for the original problem is i) convex and centrosymmetric (in a sense we make precise) and ii) the original bounds for the identified set are directionally differentiable, we can also solve the limiting problem (Theorems \ref{thm:non-diff} and \ref{thm:non-diff-asymptotic}).\footnote{In contrast, it is well known that directional differentiability causes irregularity for inference  problems \citep{hirano2012impossibility}.}

We use two key ideas to justify our asymptotic approximations. The first key idea is to consider drifting parameter sequences over which the identified set for the payoff-relevant parameter shrinks to a point (at a root-$n$ rate) as the sample size grows large. We argue that local asymptotic approximations based on sequences that lead to \emph{near point-identification} ensure a nontrivial trade-off between identification and estimation in the limit as the sample size grows large. The idea of shrinking the identified set as the sample size grows large is based on the seminal work of \cite{armstrong2021sensitivity}, who use this device to conduct sensitivity analysis in point and partially-identified models defined by moment conditions which are only required to hold in an approximate sense. Drifting parameter sequences that lead to near point-identification are also discussed in \cite{song2014point} and \cite{Stoye2009ecma}.

The second key idea is that we make an explicit recommendation about which reduced-form parameter value to localize about.  In point-identified treatment choice problems, one often localizes at a point such that the payoff-relevant parameter equals zero  \citep{HiranoPorter2009,HiranoPorter2020}. This choice is usually motivated as being the hardest (or least-favorable) for the decision maker. In partially-identified treatment choice problems, choosing where to localize is a more delicate matter. Indeed, any reduced-form parameter value at which the identified set for the payoff-relevant parameter contains zero makes the treatment choice problem hard,  but different values of the reduced-form parameter imply different degrees of partial identification. We argue that in our partially-identified treatment choice problem, there also exists a natural notion of a least-favorable or hardest case, which we formalize and advocate to localize around.

We apply our results to three concrete problems. First, we consider treatment choice problems with binary outcomes that may be contaminated or corrupted \citep{horowitz1995identification}. Second, we present a stylized example in robust welfare analyses \citep{kang2025robustness}. Third, we revisit an example in the evidence aggregation framework in \cite{ishihara2021}.

{\scshape Related Literature.} \cite{Manski2000,manski2004statistical} and \cite{Dehejia2005} advocated using a decision-theoretic framework to study treatment choice problems. \cite{Manski2000,manski2005social,manski2007identification} and \citet{Stoye07} provide optimal treatment rules assuming the true distribution of the data is known.
\cite{stoye2012minimax,stoye2012new}, \cite{yata2021}, \cite{ishihara2021} and \cite{MQS2023decision} focus on finite-sample minimax regret optimal rules; \citet{aradillas2024robust} on multiple prior minimax regret rules; and \cite{kitagawa2023treatment} on mean-squared regret optimal rules.

In important recent work, \cite{christensen2020robust,christensen2022optimal} pioneered a local asymptotic approach for discrete choice problems when payoffs depend on a partially-identified parameter (binary treatment choice problems with partial identification being a special case). Their main innovation is to \emph{profile out} the partially-identified parameter in the loss function and recenter a decision rule's risk around that of an oracle who knows a reduced-form parameter that restricts the payoffs for the different choices. While they focus on a Bayes criterion based on profiled loss, \cite{xu2026asymptotic} extends their methodology to a minimax criterion and also allows for randomized decisions. Our paper shares a similar motivation to these two papers and continues the research agenda of \cite{christensen2022optimal}. An important difference is that we do not consider a loss function that profiles out the partially-identified parameter. Our asymptotic analysis with a shrinking identified set enables us to respect the problem's original loss function and retain a more traditional minimax regret perspective on the finite-sample problem. However, to achieve this, we must be specific about where to localize the reduced-form parameter, whereas \citet{christensen2022optimal} and \citet{xu2026asymptotic} can be very flexible in this choice. Another difference from \citet{xu2026asymptotic} is that we find that approximately minimax regret optimal decisions even when the bounds of the identified set are only directionally differentiable. Thus, we view our work as a natural complement to the work of \cite{christensen2020robust,christensen2022optimal,xu2026asymptotic}. \cite{song2014point} studied local asymptotic minimax decisions for interval-identified parameters by profiling out the mean-squared error risk function; see also \cite{kido2023locallyasymptoticallyminimaxstatistical} for analogous treatments in   treatment choice problems. We think all of these papers (and ours) contribute to creating a bridge to existing results concerning limit experiments in point-identified settings \citep{HiranoPorter2009,HiranoPorter2020}.

Drifting parameter sequences of parameter values analogous to the near point-identification asymptotics used in this paper have a long history in econometrics: for example, the local-to-zero asymptotics in the linear instrumental variables model of \cite{staiger1997instrumental}, local-to-unit-root asymptotics in the study of nearly integrated autoregressive processes in \cite{phillips1988regression} and \cite{dou2021generalized}, local-to-identification-failure analysis for Generalized Method of Moments models in \cite{andrews2022optimal}. Local asymptotics are also used in the work of \cite{hirano2017forecasting} to compare the performance of different forecasting procedures under model uncertainty. See also \cite{powell2017identification} for further examples.

An important motivation for our paper is the work of \citet{stoye2012minimax}, \citet{ishihara2021}, \citet{yata2021}, and \citet{MQS2023decision}. All of these papers consider treatment choice problems with partial identification where a point-identified (reduced-form) parameter and  a payoff-relevant but partially-identified parameter.  Crucially, the analysis of minimax-regret optimal decision rules is conducted under the assumption that the statistical information for the point-identified parameter is revealed only through the mean of a Gaussian location shift model. \cite{yata2021} makes significant progress characterizing exact finite-sample minimax-regret optimal treatment assignment rules by assuming that the parameter space is convex and centrosymmetric. As noted by \cite{hirano25}, \emph{``Given the shifted normal form for the reduced-form model, it is tempting to think of this statistical setup
as corresponding to the limit experiment in the locally asymptotically normal case, so that these results
can be interpreted as simultaneously providing asymptotic optimality results when the reduced-form
model is regular and parametric. However, making this connection is somewhat complicated.''} Our approach justifies the aforementioned Gaussian construction as a \emph{bona fide} limit experiment (subject to a suitable adjustment of the limit identified set). While our results assume that the statistical model is parametric (in the sense of being indexed by a finite-dimensional parameter), the results in \citet[][Appendix C and part of Section 4]{armstrong2021sensitivity} can be used to provide local asymptotic justifications  for treatment choice problems in semiparametric models where the identified set for the payoff-relevant parameter is characterized via  moment conditions. Beyond the use of a parametric statistical model, we note that a key difference is that our asymptotic analysis  simplifies the geometry of the identified set considerably in the limit. As a result, we are able to provide solutions to the limiting problem even beyond a centrosymmetric parameter space.

Finally, our main results are most helpful to a researcher who believes that identification and  estimation-induced uncertainty are roughly equally important for their finite-sample problem.  If the dataset at hand is so large that they believe the lack of identification is the sole issue, one is free to set a slower drifting rate for the identification-induced uncertainty. In this case, we simply recommend the  plug-in or ``as-if'' \citep{manski2021econometrics} approach that just solves population-level problems after plugging in consistent estimators for identifiable quantities. Alternatively, if a researcher thinks that the model-induced uncertainty vanishes faster than estimation-induced uncertainty (because data is so scarce and the degree of partial identification is tiny in the model), our results equally apply by simply setting the bounds of the identified set in the limit experiment as approaching zero.


{\scshape Outline.} The rest of the paper is organized as follows. In Section \ref{sec:illustrate}, we explain the key idea of our paper via a simple and  stylized example. Section \ref{sec:frame} introduces our limit treatment choice problem in a general setup and formal notion of asymptotic minimax regret optimality. Section \ref{sec:results}
presents our results on solving the limit problem and constructing asymptotically optimal rules.
We apply our results to three applications in Section \ref{sec:applications}. Section \ref{sec:conclude} concludes.  All proofs and additional technical results are reserved in the Appendix and Online Appendix.


\section{Illustrative Example}\label{sec:illustrate}

In this section, we illustrate the key idea of the paper via a stylized example that also demonstrates some general difficulties of asymptotics with partially identified payoff-relevant parameters. Suppose a decision maker (DM) must decide what fraction $a\in[0,1]$ of a target population to assign to a new policy. Denote by $\mu^{*}\in\mathbb{R}$ the true treatment effect on the target population. Applying action $a$ yields welfare $\mu^{*}a$ (status quo welfare is assumed to be known and normalized to zero). Thus, the infeasible optimal rule is $\mathbf{1}\left\{ \mu^{*}\geq0\right\} $. DM observes one realization of $Y\sim\mathcal{N}(\mu,\sigma^{2})$, where $\sigma>0$ is known and $\mu$ is an $\textit{identified}$ treatment effect; it is related to $\mu^{*}$ by
\begin{equation*}
\mu^{*}\in I(\mu):=\left[\mu-k,\mu+k\right]~~\text{ for some known }k>0.\label{eq:id.illustrative.eg}
\end{equation*}
Here, the parameter space is
\begin{equation}\label{eq:stylized}
 \Theta:=\{(\mu,\mu^{*})\mid\mu\in\mathbb{R},\mu^{*}\in[\mu-k,\mu+k]\}.
\end{equation}
A statistical decision rule $d$ maps $Y$ to the unit interval, and the goal is to find a decision that minimizes worst-case expected regret defined as
\begin{equation}
\sup_{\mu^{*}\in I(\mu),\mu\in\mathbb{R}}{\mu^{*}\mathbf{1}\{\mu^{*}\geq0\}-\mu^{*}\mathbb{E}_{\mu}d(Y)}.\label{eq:illustrative.regret}
\end{equation}
This minimax regret (MMR) problem is tractable due to (i) normal likelihood of $Y$ and (ii) convexity and centrosymmetry of $\Theta$ \citep{stoye2012minimax,yata2021}. The qualitative and quantitative features of the optimal decision depend on the tradeoff between $\sigma$ (indicating estimation uncertainty) and $k$ (indicating severity of partial identification): if $k\leq\sigma\sqrt{\pi/2}$, then $\mathbf{1}\{Y\geq0\}$ is MMR optimal, uniquely so if the inequality is strict; if $k>\sigma\sqrt{\pi/2}$, then infinitely many MMR optimal rules exist, one of which is $\Phi\left(Y/\sqrt{2k^{2}/\pi-\sigma^{2}}\right)$ and all of which in a large regularity class involve randomized treatment assignment.

These findings are instructive, but in practice, data are rarely exactly normal.  As noted in  \cite{hirano25}, rigorously formalizing the idea that the above Gaussian construction represents a limit experiment result is challenging.
\begin{figure}[http]
 \centering
\begin{tikzpicture}[
    axis/.style={->, thick, black},
    grayline/.style={very thick, gray},
    redline/.style={very thick, red},
    myredline/.style={very thick, myred},
    blueline/.style={very thick, blue},
    greenline/.style={very thick, mygreen},
    tangent/.style={blue, ultra thick}
]


\begin{scope}[xshift=0cm]
  \begin{scope}
    \clip (-2.8,-2.8) rectangle (2.8,2.8);
    \fill[orange!30]
      (-2.8, {-2.8-1.5})
      -- (-2.8, {-2.8+1.5})
      -- (2.8, {2.8+1.5})
      -- (2.8, {2.8-1.5})
      -- cycle;
    \draw[grayline]  (-2.8, {-2.8+1.5}) -- (2.8, {2.8+1.5});
    \draw[grayline] (-2.8, {-2.8-1.5}) -- (2.8, {2.8-1.5});
  \end{scope}

  \draw[axis] (-2.8,0) -- (2.8,0) node[right] {$\mu$};
  \draw[axis] (0,-2.2) -- (0,2.8) node[above] {$\mu^*$};

  \draw[redline] (1.0, {1.0-1.5}) -- (1.0, {1.0+1.5});
  \filldraw[red] (1.0, {1.0+1.5}) circle (1.5pt);
  \filldraw[red] (1.0, 0)           circle (1.5pt);
  \filldraw[red] (1.0, {1.0-1.5}) circle (1.5pt);
  \node[red, font=\large\bfseries] (mu0label) at (1.0+0.7, 0.7) {$\mu_0$};
    \draw[->, red, thick] (mu0label.south west) -- (1.0+0.07, 0.07);

  \filldraw[black] (0, 1.5) node[left=3pt] {$k$};
  \filldraw[black] (0,-1.5) node[right=3pt] {$-k$};

  \draw[greenline] (-0.7, {-0.7-1.5}) -- (-0.7, {-0.7+1.5});
  \filldraw[mygreen] (-0.7, {-0.7+1.5}) circle (1.5pt);
  \filldraw[mygreen] (-0.7, 0)             circle (1.5pt);
  \filldraw[mygreen] (-0.7, {-0.7-1.5}) circle (1.5pt);
  \node[mygreen, font=\large] (mulabel) at (-0.7-1, 1.4)
    {$\mu_0 + \dfrac{h_1}{\sqrt{n}}$};
    \draw[->, mygreen, thick] (mulabel.south) -- (-0.7-0.08, 0.07);

  \draw[->, very thick, black]
      (-0.7+0.15, -2.4) -- (1.0-0.15, -2.4)
      node[midway, below, font=\normalsize] {as $n \to \infty$};
\end{scope}


\begin{scope}[xshift=7.5 cm]
  \begin{scope}
    \clip (-2.8,-2.8) rectangle (2.8,2.8);

    \fill[orange!30]
      (-2.8, {-2.8-1.5})
      -- (-2.8, {-2.8+1.5})
      -- (2.8, {2.8+1.5})
      -- (2.8, {2.8-1.5})
      -- cycle;

    \fill[blue!30]
      ({-0.9}, {-0.9-1.2})
      -- ({0.9}, {0.9-1.2})
      -- ({0.9}, {0.9+1.2})
      -- ({-0.9}, {-0.9+1.2})
      -- cycle;

    \begin{scope}
      \clip
        ({-0.9}, {-0.9-1.2})
        -- ({0.9}, {0.9-1.2})
        -- ({0.9}, {0.9+1.2})
        -- ({-0.9}, {-0.9+1.2})
        -- cycle;
      \draw[tangent] (-2.8, {-2.8+1.2}) -- (2.8, {2.8+1.2});
      \draw[tangent] (-2.8, {-2.8-1.2}) -- (2.8, {2.8-1.2});
    \end{scope}

    \draw[very thick, gray] (-2.8, {-2.8+1.5}) -- (2.8, {2.8+1.5});
    \draw[very thick, gray] (-2.8, {-2.8-1.5}) -- (2.8, {2.8-1.5});
  \end{scope}

  \draw[axis] (-2.8,0) -- (2.8,0) node[right] {$\mu$};
  \draw[axis] (0,-2.8) -- (0,2.8) node[above] {$\mu^*$};

  \draw[redline] (0,-1.5) -- (0,1.5);
  \filldraw[red] (0, 1.5) circle (1.5pt) node[left=3pt, black] {$k$};
  \filldraw[red] (0,-1.5) circle (1.5pt) node[right=3pt, black] {$-k$};
  \filldraw[red] (0, 0)  circle (1.5pt);
  \node[red, font=\large\bfseries] at (0.45, 0.43) {$\mu_0$};

  \node[blue, font=\large] at (1.75, 1.25) {$\Theta_n$};
  \draw[blue, ->, thick] (1.4, 1.2) to[out=210, in=30] (1, 1);
\end{scope}

\end{tikzpicture}
\caption{Two different asymptotics in the stylized example. Left: we localize at any $\mu_0$. As a result, $I\left(\mu_0+h/\sqrt{n}\right)\rightarrow I(\mu_0)$ as $n\rightarrow\infty$. Right: we localize at  $\mu_0=0$ and consider  a local parameter space $\Theta_n$ implying a shrinking identified set.}
\label{fig:illustrate}
\end{figure}


\subsection{The Challenge}\label{sec:eg.challenge}

Suppose now the DM observes a random sample $\{Y_i\}_{i=1}^{n}$, where $Y_i\sim\mathcal{N}(\mu,\sigma^{2})$ with both $\mu$ and $\sigma$ unknown. The parameter space becomes
\[
\Theta=\{(\mu,\sigma^{2},\mu^{*})\mid\mu\in\mathbb{R},\sigma^{2}>0,\mu^{*}\in I(\mu)\}.
\]
As $n$ becomes large, the location of $\mu$ and $\sigma^{2}$ becomes easier to learn but that of $\mu^*$ may still remain ambiguous. Suppose that, in the spirit of \citet{HiranoPorter2009,HiranoPorter2020} and work that they build on, we recenter the reduced-form parameter around some $\mu=\mu_{0}$ and $\sigma^{2}=\sigma_{0}^{2}$:\footnote{The typical notion of localization, including that in  \cite{HiranoPorter2009,HiranoPorter2020},  focuses on point-identified problems and does not apply to partially identified problems.}
\begin{align*}
\mu & =\mu_{0}+\frac{h_{1}}{\sqrt{n}},~~\sigma^{2}=\sigma_{0}^{2}+\frac{h_{2}}{\sqrt{n}},
\end{align*}
where $h=(h_{1},h_{2})^{\top}\in\mathbb{R}^2$.
Then, the identified set for $\mu^{*}$ is $I\left(\mu_{0}+h_1/\sqrt{n}\right)$. As $n\rightarrow\infty$, the location  $h$ will be signaled by an \emph{asymptotically sufficient statistic} $\Delta\sim\mathcal{N}(h,\mathbf{I}_0^{-1})$ with Fisher Information $\mathbf{I}_0$, while it is clear that (see also the left panel of Figure \ref{fig:illustrate})
\begin{align*}
\mu_{0}+\frac{h_{1}}{\sqrt{n}} & \rightarrow\mu_{0},
~~I\left(\mu_{0}+\frac{h_{1}}{\sqrt{n}}\right)\rightarrow I\left(\mu_{0}\right).
\end{align*}
In other words, in the limit, the DM would observe $\Delta\sim\mathcal{N}(h,\mathbf{I}_{0}^{-1})$
and \emph{know} $\mu^{*}\in[\mu_{0}-k,\mu_{0}+k]$. In this limit, a rule  $d_\infty$  maps $\Delta$ to $[0,1]$ and the MMR problem becomes
\begin{equation}\label{eq:default.asymptotics.limit.game}
\min_{d_{\infty}}\sup_{\mu^*\in [\mu_{0}-k,\mu_{0}+k],h\in\mathbb{R}^{2}}\mu^{*}\left[\mathbf{1}\left\{ \mu^{*}\geq0\right\} -\mathbb{E}d_\infty(\Delta)\right].
\end{equation}
The signal $\Delta$ has become pure noise, revealing no information regarding $\mu_0$ or $\mu^*$. It follows that \eqref{eq:default.asymptotics.limit.game} is solved by the no-data rule  \citep{manski2007identification}
\[
\left[\frac{\mu_{0}+k}{2k}\right]_{0}^{1}:=\begin{cases}
0, & \mu_{0}<-k\\
\tfrac{\mu_{0}+k}{2k}, & -k\leq\mu_{0}\leq k\\
1, & \mu_{0}>k
\end{cases}.
\]
This finding reflects partial identification overwhelming sampling noise in the limit. Indeed, under these asymptotics, sampling uncertainty vanishes at the standard $1/\sqrt{n}$ rate whereas the degree of partial identification is constant even in the limit. As a result,  there is no meaningful trade-off between estimation uncertainty and the severity of partial identification, as opposed to what we saw earlier in the ``finite-sample'' version of the problem. Moreover, these asymptotics cannot (sufficiently) inform the selection of good rules because too many rules will be asymptotically MMR optimal in the sense of worst-case regret converging to the same limit. For example, this applies to $\left[(\hat{\mu}+k)/(2k)\right]_0^1$ for any consistent estimator $\hat{\mu}$ of $\mu$.




\subsection{Introducing Local Asymptotics}\label{sec:eg.key.idea}
Our asymptotic approach embodies two major innovations; see also the right panel of Figure \ref{fig:illustrate} for a visualization. First, we are more explicit about which parameter values to localize about. In point-identified settings, one often picks a point at which the payoff-relevant parameter equals zero, representing a \emph{hardest case} where it is difficult to determine the best
treatment even with large sample sizes \citep[][p.1687]{HiranoPorter2009}.  In our example,  that would correspond to $\mu^*=0$. However, as $\mu^*$ is partially identified, any point of $(\mu,\sigma^2)$ such that $I(\mu)$ contains $0$ is consistent with $\mu^*=0$, and it is not immediately clear which point in the space of   $(\mu,\sigma^2)$ we should pick. We argue that, even in the partially identified setting, there is an analogous notion of the  \emph{hardest case}  that  motivates our choice of localization for the reduced-form parameter. In this stylized example, that corresponds to $\mu=0$ (and the point of localization for $\sigma^2$ can be arbitrary). Intuitively, if the sample size is sufficiently large to inform the location of $\mu$, then $\mu=0$ is the most challenging case as it involves the greatest   ambiguity regarding  $\mu^*$. See Section \ref{sec:local.point} for the definition of our ``most difficult case'' in a general setting.

Second, to ensure a nontrivial trade-off between the severity of partial identification and estimation difficulty, we let the former vanish at an appropriate rate \citep{armstrong2021sensitivity}. That is, we look at drifting sequences that imply ``near point identification.''  Specifically,  we consider a local parameter space with a shrinking identified set:
\begin{align*}\Theta_{n} :=\left\{ (\mu,\sigma^{2},\mu^{*})\mid \mu=0+\frac{h_{1}}{\sqrt{n}},\sigma^{2}=\sigma_{0}^{2}+\frac{h_{2}}{\sqrt{n}},\mu^{*}\in I_n\left(\frac{h_{1}}{\sqrt{n}}\right)\right\},
\end{align*}
where
${I_{n}}\left(\frac{h_{1}}{\sqrt{n}}\right)  =\left[\frac{h_{1}}{\sqrt{n}}-k_{n},\frac{h_{1}}{\sqrt{n}}+{k_n}\right]$
for some $k_{n}$ such that $\sqrt{n}k_{n}=C>0$.
Now, consider the scaled  risk function of a  decision rule $d$ based on data $\{Y_i\}_{i=1}^{n}$, which can be written as
\[
\sqrt{n}\mu^{*}\left(\mathbf{1}\left\{ \mu^{*}\geq0\right\} -\mathbb{E}_{(\mu,\sigma^{2})}[d((Y_i)_{i=1}^n)]\right).
\]
It follows that
\begin{align*}\sqrt{n}\mu^{*}  \in\sqrt{n}{I_{n}}\left(\frac{h_{1}}{\sqrt{n}}\right)=\sqrt{n}\left[\frac{h_{1}}{\sqrt{n}}-k_{n},\frac{h_{1}}{\sqrt{n}}+k_{n}\right] =\left[h_{1}-C,h_{1}+C\right].
\end{align*}
Furthermore, standard asymptotic arguments (e.g.,  Proposition 3.1 in \citealt{HiranoPorter2009}) yield that for any converging rule $d$, there is a limiting rule $d_{\infty}$ such that
\[
\mathbb{E}_{(\mu,\sigma^{2})}[d((Y_i)_{i=1}^n)]\rightarrow\mathbb{E}_{h}[d_{\infty}(\Delta)],
\]
where $\Delta\sim\mathcal{N}(h,\mathbf{I}_{0}^{-1})$. Thus, we have a stable limit game in which the researcher observes $\Delta\sim\mathcal{N}(h,\mathbf{I}_{0}^{-1})$, and the partially identified welfare contrast $\mu^{*}_{\infty}$ lies in a limit identified set $\left[h_{1}-C,h_{1}+C\right]$. The limit MMR problem becomes
\begin{equation}\label{eq:intro.limit.mmr}
\min_{d_{\infty}:\mathbb{R}^{2}\rightarrow[0,1]}\sup_{\mu_{\infty}^{*}\in[h_{1}-C,h_{1}+C],h\in\mathbb{R}^{2}}\mu_{\infty}^{*}\left[\mathbf{1}\left\{ \mu_{\infty}^{*}\geq0\right\} -\mathbb{E}_{h}[d_{\infty}(\Delta)]\right],
\end{equation}
recovering the stylized example; compare \eqref{eq:stylized}-\eqref{eq:illustrative.regret}. Letting  $w=(1,0)^{\top}$ and $\Sigma_{0}=(w)^{\top}\mathbf{I}_{0}^{-1}w$, we can apply existing results and find the solution of \eqref{eq:intro.limit.mmr} as follows:
If $C\leq\sqrt{\Sigma_{0}\pi/2}$, then $\mathbf{1}\{w^{\top}\Delta\geq0\}$ is MMR optimal; if $C>\sqrt{\Sigma_{0}\pi/2}$, one optimal rule is
$\Phi\left(w^{\top}\Delta\left(2C^{2}/\pi-\Sigma_{0}\right)^{-1/2}\right)$.
With this limit optimal rule, we may further find a matching feasible rule via the natural ``plug-in'' principle, i.e., replacing (i) $w^{\top}\Delta$ with $\sqrt{n}\hat{\mu}$, where $\hat{\mu}$ is a best regular estimator of $\mu$ (e.g., the MLE) and (ii) $\Sigma_{0}$ with a consistent estimator $\hat{\Sigma}$.  This is exactly what a researcher would use in practice based on the ``finite-sample'' results above if we view  $\hat{\mu}$ as $Y$ in \eqref{eq:illustrative.regret} and calibrate $C$  with  $\sqrt{n}k$.

Next, we formalize the above discussions in a general framework. We show that our idea of applying ``near point identification'' at a suitably defined hardest case allows one to find asymptotically MMR optimal rules (in a sense we also precisely define) for a large class of treatment decision problems with partially identified parameters, even beyond convexity and centrosymmetry.

\section{The Limiting Decision Problem} \label{sec:frame}

\subsection{Setup}

Suppose a DM's payoff when taking action $a\in[0,1]$ is $aW_{1}+(1-a)W_{0}$
for some scalars $W_{1}$ and $W_{0}$. Denote by $U^{*}:=W_{1}-W_{0}\in\mathbb{R}$
the \emph{welfare contrast} between actions $a=1$ and $a=0$.
If $U^{*}$ were known, the optimal action would be $\mathbf{1}\left\{ U^{*}\geq0\right\}$. The DM observes a random sample
$Y^{n}:=\left\{ Y_{i}\right\} _{i=1}^{n}\in\mathbf{Y}^{n}$ of size
$n$, where $Y_{i}\in\mathbf{Y}\subseteq\mathbb{R}^{d_{Y}}$, to learn
about $U^{*}$. We assume that $Y_{i}$ follows a parametric
model with a distribution $P_{\gamma}$ indexed by $\gamma\in M\subseteq\mathbb{R}^{d_{m}}$.
We require $\gamma$
to be point-identified in the usual sense, i.e., if $\gamma,\gamma^{\prime}\in M$ are such
that $\gamma\neq\gamma^{\prime}$, we have $P_{\gamma}\neq P_{\gamma^{\prime}}$.
We therefore treat the values of $\gamma\in M$ as the point-identified
reduced-form parameter.


As in \cite{christensen2022optimal}, our primary focus  is when $U^{*}$---the payoff-relevant parameter---is only partially identified, but the DM can use $\gamma$ to deduce restrictions on $U^*$. That is, if $\gamma$ were known, the most the DM could infer is that $U^{*}\in\mathcal{I}(\gamma)\subseteq\mathbb{R}$,
where $\mathcal{I}(\cdotp)$ is a set-valued mapping from $M$ to
$\mathbb{R}$ that may contain both positive and negative values. We refer to  $\mathcal{I}(\gamma)$ as the identified
set of $U^{*}$ given $\gamma\in M$.

\begin{assumption}\label{asm:1} The identified set $\mathcal{I}(\gamma)$
depends on $\gamma$ only via a known function $\mu(\cdot):\mathbb{R}^{d_{m}}\rightarrow\mathbb{R}^{d_{\mu}}$,
with $d_{\mu}\leq d_{m}$. \end{assumption}

We may think of $\mu(\cdot)$ as a \emph{decision-relevant} reduced-form
parameter and view the rest of $\gamma$ as some nuisance parameter
not directly relevant for the identification of $U^{*}$. For instance,
in the example of Section \ref{sec:eg.challenge}, $\gamma$ refers
to the mean and variance of a normal distribution, while the identified
set of $U^{*}$ only depends on the mean.

Under Assumption \ref{asm:1}, we can write $\mathcal{I}(\gamma):=I(\mu(\gamma))$,
where
\[
I(t):=\left\{ u\in\mathbb{R}\mid u\in\mathcal{I}(\gamma),\mu(\gamma)=t,\gamma\in M\right\}
\]
is the identified set of $U^{*}$ given a fixed value of $\mu(\cdotp)$.
Let
\[
M_{\mu}:=\left\{ t \in\mathbb{R}^{d_{\mu}}\mid\mu(\gamma)=t,\gamma\in M\right\}
\]
collect all the values that the function $\mu(\cdotp)$ can take as
$\gamma$ ranges over  $M$. For any $t\in M_{\mu}$, let
\[
\overline{I}(t):=\sup I(t);\quad\underline{I}(t):=\inf I(t)
\]
be the upper and lower bounds of $I(t)\subseteq\mathbb{R}$. Abusing  notation, sometimes we write $I(\mu)$, and pretend that
$\mu$ denotes a particular value of the function $\mu(\cdotp)$,
instead of the whole function. In the problems that we are interested
in, we assume there must be at least one point $t^{*}\in M_{\mu}$
for which $0\in I(t^{*})$ and $\underline{I}(t^{*})<0<\overline{I}(t^{*})$.
This means that at values of the decision-relevant reduced-form parameter
that equal $t^{*}$, the welfare contrast could be strictly positive
or strictly negative.

After observing data $Y^{n}$, the decision maker chooses a statistical
decision rule $d_{n}:\mathbf{Y}^{n}\rightarrow[0,1]$, interpreted
as the probability of treatment assignment or the fraction of the
target population to be treated. To  organize notation, it will be convenient to define
$\theta:=\left(\gamma^{\top},U^{*}\right)^{\top}$ and
\[
\Theta:=\left\{ \left(a^{\top},b\right)^{\top}\in\mathbb{R}^{d_{m}+1}\mid a\in M,b\in\mathcal{I}(a)\right\}.
\]
Note that the set $\Theta$ collects the values that $\gamma$ and $U^{*}$ can take jointly. We endow $\Theta$ with the standard subspace topology of $\mathbb{R}^{d_m+1}$. Throughout the rest of the paper we require that its topological interior, denoted as $\operatorname{int}(\Theta)$, is nonempty.

In order to connect with previous work in the literature---in particular, with the work of \cite{yata2021}---for each $\theta\in\Theta$, denote by $m(\theta)$ the mapping that
selects the first $d_{m}$ elements of $\theta$ (i.e., the point-identified $\gamma$), and by $U(\theta)$ the mapping
that selects the last element of $\theta$ (i.e., the partially identified payoff-relevant parameter $U^{*}$). By construction, both $m(\cdotp)$
and $U(\cdotp)$ are linear. By definition,  $m(\cdot)$ is clearly not injective: $m(\theta) = m(\theta')$ does not imply $\theta=\theta'$.\footnote{Even though the framework considered by \cite{yata2021} appears to be significantly more general than ours (because he allows for the possibility of an underlying infinite-dimensional parameter $\theta$), we argue that the treatment choice problems that he considers can be recast using our framework. To see this,
given some underlying $\tilde{\theta}\in\tilde{\Theta}$ that may
be infinitely dimensional, let $\tilde{m}(\cdotp):\tilde{\Theta}\rightarrow\mathbb{R}^{d_{\tilde{m}}}$
map $\theta$ to a finite-dimensional reduced-form parameter and $\tilde{U}(\cdotp):\tilde{\Theta}\rightarrow\mathbb{R}$
map $\theta$ to a partially identified payoff-relevant parameter. Since the
risk function depends on $\theta$ only via $\tilde{m}(\cdotp)$
and $\tilde{U}(\cdotp)$, one can always reparametrize so that
$\theta:=(\tilde{m}(\tilde{\theta})^{\top},\tilde{U}(\tilde{\theta}))^{\top}\in\Theta\subseteq\mathbb{R}^{1+d_{\tilde{m}}}$, where $\Theta:=\{(\tilde{m}(\tilde{\theta})^{\top},\tilde{U}(\tilde{\theta}))\in\mathbb{R}^{1+d_{\tilde{m}}}\mid\tilde{\theta}\in\tilde{\Theta}\}$
is a set in $\mathbb{R}^{1+d_{\tilde{m}}}$.
}

We evaluate the performance of a decision rule $d_{n}$ via expected
regret:
\begin{align*}
R(d_{n},\theta) =U(\theta)\left(\mathbf{1}\left\{ U(\theta)\geq0\right\} -\mathbb{E}_{m(\theta)}\left[d_{n}\left(Y^{n}\right)\right]\right),
\end{align*}
where $\mathbb{E}_{m(\theta)}\left[\cdotp\right]$ denotes expectation with respect
to $Y^{n}$, whose distribution only depends on $m(\theta)$. We omit the subscript whenever there is no risk of confusion. Note that we also omit the explicit dependence of the risk on the sample size, for the sake of notational simplicity. A rule is finite-sample MMR optimal if it solves
\begin{equation}
\min_{d_{n}}\sup_{\theta\in\Theta}R(d_{n},\theta).\label{eq:finite.sample.mmr}
\end{equation}

\textsc{Local Asymptotics.} To provide a simpler, large-sample representation of the
finite-sample MMR problem \eqref{eq:finite.sample.mmr}, we fix a parameter $\theta_0:=(\gamma_0^{\top},0)^{\top}\in \operatorname{int}(\Theta)$, such that $\gamma_0\in \operatorname{int}(M)$, $0\in \mathcal{I}(\gamma_0)$ and $\operatorname{int}(\mathcal{I}(\gamma_0))\neq \emptyset$. Then, we  consider a sequence of parameter values of the form:
\begin{equation}\label{eq:local.parameter.theta}
\theta_n=\theta_0+\frac{\theta_h}{\sqrt{n}},
\end{equation}
where $\theta_h:=(h^{\top},\mu^*)^{\top} \in \mathbb{R}^{d_m+1}$. Since $m(\cdot)$ and $U(\cdot)$ are linear,  we also have
\begin{align}
m(\theta_{n})=\gamma_{0}+\frac{h}{\sqrt{n}},\quad U(\theta_{n})=\frac{\mu^{*}}{\sqrt{n}}.\label{eq:local.parameter.1}
\end{align}
In words, we recentered the reduced-form parameter at $\gamma_{0}$
and the payoff-relevant but partially identified parameter at $0$.



\subsection{Choice of Localization for the Point-Identified Parameter}\label{sec:local.point}
In point-identified settings, a standard choice for the localization parameter $\gamma_0$ is any point such that the corresponding payoff-relevant parameter equals zero, which corresponds to a \emph{hardest case} where it is difficult to determine the best
treatment even with large sample sizes \citep[][p.1687]{HiranoPorter2009}.  In our setup, the payoff-relevant parameter $U(\theta)$ is partially identified. This means that the choice of $\gamma_0$ is more delicate,  as any  values of $\gamma_0$ such that $0\in I(\mu(\gamma_0))$  would be consistent with a situation in which $U(\theta)=0$, and hence it might be difficult to determine the best treatment even in large samples.

Let $\mu_0:=\mu(\gamma_0)$. We propose a notion of \emph{hardest case} that considers the population level MMR problem
\begin{equation}\label{eq:population.mmr}
\min_{a\in[0,1]}\sup_{u\in I(\mu_0)}u\left(\mathbf{1}\left\{  u\geq0\right\} -a\right)
\end{equation}
corresponding to oracle knowledge of  the decision-relevant reduced-form parameter.
Denote by $\overline{R}(\mu_0)$ the value of \eqref{eq:population.mmr}. Note $\overline{I}(\mu_0)=\underline{I}(\mu_0)$ corresponds to a point-identified problem, for which $\overline{R}(\mu_0)=0$. Otherwise, a MMR optimal  action for \eqref{eq:population.mmr} is \citep{manski2007identification}
\begin{equation}\label{eq:manski.rule.general}
\left[\frac{\overline{I}(\mu_0)}{\overline{I}(\mu_0)-\underline{I}(\mu_0)}\right]_{0}^{1},
\end{equation}
achieving a value of
\[
\overline{R}(\mu_0)=\max\left\{ \frac{-\overline{I}(\mu_0)\underline{I}(\mu_0)}{\overline{I}(\mu_0)-\underline{I}(\mu_0)},0\right\}.
\]
We already assumed the existence of a decision-relevant reduced-form parameter, $\mu_0$, such that
$0\in I(\mu_0)$ and $\overline{I}(\mu_0) > 0 > \underline{I}(\mu_0)$. We strengthen this requirement by assuming that there is
a point $\mu_0$ in the interior of $M_{\mu}$ that satisfies this property, and that yields the highest possible value for the
population level MMR problem.
\begin{assumption}\label{asm:2}
There exists $\theta_0\in \operatorname{int}(\Theta)$  such that $m(\theta_0)=\gamma_0\in \operatorname{int}(M)$, $\mu_{0}=\mu(\gamma_0)\in\operatorname{int}\left(M_{\mu}\right)$, where $\overline{I}(\mu_0)>0>\underline{I}(\mu_0)$, and $\overline{R}(\mu_{0})\geq\overline{R}(\mu^{\prime})$
for all $\mu^{\prime}\in M_{\mu}$. \end{assumption}
The value of the problem \eqref{eq:population.mmr} is maximized and strictly positive at $\mu_0$ satisfying Assumption \ref{asm:2}, which we interpret as $\mu_0$ being most difficult from the DM's perspective. We refer to such $\mu_0$ as a \emph{global
least favorable (decision-relevant reduced-form) point}. This means that $m(\theta_n)$ in \eqref{eq:local.parameter.1} is centered at any $\gamma_{0}\in \operatorname{int(M)}$
such that  $\mu(\gamma_{0})=\mu_{0}$. See Figure \ref{fig:lfp} for an illustrative example  of a global least favorable point.

We note that different values of the reduced-form parameter $\gamma_0$ can be associated with the
same decision-relevant reduced-form parameter $\mu_0$. This is  allowed in our theory,  as long as we have efficient estimators for $\gamma_0$ (as well as consistent estimators of the asymptotic variance). See analogous treatments in  \cite{HiranoPorter2009} with asymmetric welfare regret and \cite{kitagawa2026treatment} with nonlinear regret.

\begin{figure}
\centering
\begin{tikzpicture}[
    axis/.style={->, thick, black},
    grayline/.style={very thick, gray},
    redline/.style={very thick, red},
    myredline/.style={very thick, myred},
    blueline/.style={very thick, blue},
    greenline/.style={very thick, mygreen},
    tangent/.style={gray, ultra thick}
]

\begin{scope}
  \clip (-2.8,-2.8) rectangle (2.8,2.8);
  \fill[orange!30]
    (-2.8, {-2.8-1.5})
    -- (-2.8, {-2.8+1.5})
    -- (2.8, {2.8+1.5})
    -- (2.8, {2.8-1.5})
    -- cycle;
  \draw[grayline]  (-2.8, {-2.8+1.5}) -- (2.8, {2.8+1.5});
  \draw[grayline] (-2.8, {-2.8-1.5}) -- (2.8, {2.8-1.5});
\end{scope}

\draw[axis] (-2.8,0) -- (2.8,0) node[right] {$\mu$};
\draw[axis] (0,-2.8) -- (0,2.8) node[above] {$\mu^*$};

\draw[redline] (0,-1.5) -- (0,1.5);

\filldraw[orange] (-1.5, 0)  circle (1.5pt);
\filldraw[blue] (1.5,0)  circle (1.5pt);
\filldraw[red] (0, 0)   circle (1.5pt);

\node[red, font=\large\bfseries] at (0.4, 0.25) {$\mu_0$};

\node[below] at (1.5, 0) {$k$};
\node[above] at (-1.5 - 0.2, 0) {$-k$};

\end{tikzpicture}
\caption{A global least favorable point in the example of Section \ref{sec:eg.challenge}. On $\mu$-axis, every point to the left of the orange dot and to the right of the blue dot leads to $\overline{R}(\mu)=0$. Between these two points, $\overline{R}(\mu)$ is strictly positive and maximized at the red dot with a value of $k/2$.}
\label{fig:lfp}
\end{figure}


\subsection{Choice of Local Identified Set}

Let $\gamma_0$ and $\mu_0$ be defined as in Assumption \ref{asm:2}. We impose the following smoothness and regularity condition.
\begin{assumption}\label{asm:diff}
\begin{itemize}
\item[(i)] $\mu(\cdotp)$ is (fully) differentiable in a local neighborhood of $\gamma_{0}$.
\item[(ii)] In a local neighborhood of $\mu_{0}$, $\underline{I}(\cdotp)$ and $\overline{I}(\cdotp)$ are attained and (Hadamard) directionally
differentiable for all directions $\mathbf{v}\in\mathbb{R}^{d_{\mu}}$,  and $I(\cdot)$ is connected.
\end{itemize}
\end{assumption}
Under Assumption \ref{asm:diff}, let
\[
G_{0}:=\underset{\mathbb{R}^{d_{\mu}}\times\mathbb{R}^{d_{m}}}{\underbrace{\mu^{(1)}(\gamma_{0})}}
\]
be the gradient of $\mu(\cdotp)$ at $\gamma_{0}$.  Denote by $\overline{I}^{(1)}(\mu;\mathbf{v})$ the directional derivative
of $\overline{I}(\cdotp)$ at $\mu$ with respect to vector $\mathbf{v}\in\mathbb{R}^{d_{\mu}}$,
i.e.,
\[
\overline{I}^{(1)}(\mu;\mathbf{v}):=\lim_{\lambda\downarrow0,\mathbf{v'\rightarrow\mathbf{v}}}\frac{\overline{I}(\mu+\lambda\mathbf{v}')-\overline{I}(\mu)}{\lambda}.
\]
If $\overline{I}(\cdotp)$
is (fully) differentiable at $\mu$, we have
\[
\overline{I}^{(1)}(\mu;\mathbf{v})=\left(\overline{I}^{(1)}(\mu)\right)^{\top}\mathbf{v},
\]
where $\overline{I}^{(1)}(\mu)$ is the gradient of $\overline{I}(\cdotp)$
at $\mu$. The directional derivative and gradient of $\underline{I}(\cdotp)$
at $\mu$, written as $\underline{I}^{(1)}(\mu;\mathbf{v})$ and $\underline{I}^{(1)}(\mu)$
respectively, are defined analogously.

For each $h\in\mathbb{R}^{d_m}$, the point-identified parameter follows a sequence
$m(\theta_n)=\gamma_0+\frac{h}{\sqrt{n}}$. As $n\rightarrow\infty$,
the identified set of $U(\theta_n)$ will be
\begin{align}
I\left(\mu\left(\gamma_{0}+\frac{h}{\sqrt{n}}\right)\right) & \approx\left[\underline{I}(\mu_{0})+\frac{\underline{I}^{(1)}(\mu_{0};G_{0}h)}{\sqrt{n}},\overline{I}(\mu_{0})+\frac{\overline{I}^{(1)}(\mu_{0};G_{0}h)}{\sqrt{n}}\right]\nonumber\\
&\rightarrow\left[\underline{I}(\mu_{0}),\overline{I}(\mu_{0})\right],\text{ as }n\rightarrow\infty.\label{eq:fixed.id}
\end{align}
The limit identified set \eqref{eq:fixed.id} reveals two important  observations. First, in the limit $U(\theta_n)$ converges to zero (implying point identification); however, in the limit, \eqref{eq:fixed.id} is a set, which implies partial identification. We think this suggests an inconsistency in the analysis which ought to be reconciled. Second, \eqref{eq:fixed.id} does not depend on the local parameter $h$---in this limit,  the data became pure noise and the issue of partial identification overwhelms sampling uncertainty.


Given these two observations and also motivated by the local behavior of $I(\mu(\gamma_0+\frac{h}{\sqrt{n}}))$
with large $n$,  we will construct a particular sequence of local,  shrinking identified set in the form of \eqref{eq:local.id.U}-\eqref{eq:local.id.U.2} below, leading to a scaled limit identified set \eqref{eq:limit.ID}. We stress that our construction is not the only way in which one could approximate the decision problem of interest. As we explain in Appendix \ref{sec:global.ID}, there are other reasonable ways of constructing a shrinking identified set that may lead to a different geometry of the limit identified set.\footnote{For example, if the  identified set can be defined via a number of moment conditions, the analysis  in \cite{armstrong2021sensitivity} would yield a limiting identified set different from \eqref{eq:limit.ID}, although  the two approaches may coincide in simple cases, e.g., the stylized example in Section \ref{sec:illustrate}.}


Take  $\underline{I}(\mu_0)$, $\overline{I}(\mu_0)$, $\underline{I}^{(1)}(\mu_0;\cdot)$ and  $\overline{I}^{(1)}(\mu_0;\cdot)$ as given from the original identified set, we construct a new, local identified set  as follows:
\begin{equation}\label{eq:local.id.U}
U(\theta_n)\in\left[\underline{I}_{{n}}(\mu_{0})+\frac{\underline{I}^{(1)}(\mu_{0};G_{0}h)}{\sqrt{n}},\overline{I}_{{n}}(\mu_{0})+\frac{\overline{I}^{(1)}(\mu_{0};G_{0}h)}{\sqrt{n}}\right],
\end{equation}
where $\underline{I}_{{n}}(\mu_{0})\rightarrow0$ and $\overline{I}_{{n}}(\mu_{0})\rightarrow0$ as $n\rightarrow\infty$ and
\begin{equation}\label{eq:local.id.U.2}
\frac{\underline{I}_{{n}}(\mu_{0})}{\overline{I}_{{n}}(\mu_{0})}=\frac{\underline{I}(\mu_0)}{\overline{I}(\mu_0)},\quad\sqrt{n}\underline{I}_{{n}}(\mu_{0})=\underline{C}{},\quad\sqrt{n}\overline{I}_{{n}}(\mu_{0})=\overline{C},
\end{equation}
for some $\underline{C}<0,\overline{C}>0$. In particular,  the relative magnitude of  $\underline{I}_{{n}}(\mu_{0})$  versus $\overline{I}_{{n}}(\mu_{0})$  respects what we observe in the original identified set.  It follows that
\begin{align*}
\mu^*=\sqrt{n}U(\theta_n) & \in\left[\sqrt{n}\underline{I}_{{n}}(\mu_{0})+\underline{I}^{(1)}(\mu_{0};G_{0}h),\sqrt{n}\overline{I}_{{n}}(\mu_{0})+\overline{I}^{(1)}(\mu_{0};G_{0}h)\right]\\
 & =\left[\underline{C}+\underline{I}^{(1)}(\mu_{0};G_{0}h),\overline{C}+\overline{I}^{(1)}(\mu_{0};G_{0}h)\right].
\end{align*}
Therefore, we define the identified set
for $\mu^*$ in \eqref{eq:local.parameter.1} as
\begin{equation}
I_{\infty}(h):=[\underline{I}_{\infty}(h),\overline{I}_{\infty}(h)],\label{eq:limit.ID}
\end{equation}
where $\underline{I}_{\infty}(h):=\underline{C}+\underline{I}^{(1)}(\mu_{0};G_{0}h)$ and
$\overline{I}_{\infty}(h):=\overline{C}+\overline{I}^{(1)}(\mu_{0};G_{0}h)$. We refer to \eqref{eq:limit.ID} as a \emph{limiting (local) identified set}, since in the limit it only reflects the local information of the original identified set $I(\cdot)$ around $\mu_0$. This means that our limiting identified set enjoys a simpler geometry than the original identified set and allows for slightly more tractability. As \eqref{eq:limit.ID} should be nonempty, we define
\begin{equation}\label{eq:M.h}
M_h:=\{h\in\mathbb{R}^{d_m}\mid\overline{I}_{\infty}(h)\geq\underline{I}_{\infty}(h)\}.
\end{equation}

It is important to stress that,  by working with such ``near-point-identification'' asymptotics, we do not  mean that the identified set  is literally shrinking as sample size increases. As noted in \cite{armstrong2021sensitivity},  the usefulness of any asymptotic device ``\emph{should be judged by whether it yields accurate approximations to the finite-sample behavior}''. In this regard, we think our approach offers a unifying framework for researchers to judge a suitable course of action in treatment choice problems with partial identification.

\begin{itemize}
\item[(i)]  Our approach is most helpful if, given finite data and the  decision problem at hand, the researcher  believes that both sampling uncertainty and the severity of partial identification are  important considerations. In practice, they simply apply our theory below by replacing  $\overline{C}$ with $\sqrt{n}\overline{I}(\mu_0)$ and analogously replacing  $\underline{C}$ with $\sqrt{n}\underline{I}(\mu_0)$.

\item[(ii)] If the data are so abundant that the researcher believes sampling uncertainty is rather tiny compared to the degree of partial identification, then they are free to set $\overline{C}=-\underline{C}=\infty$.  In this case, we would simply recommend to apply \eqref{eq:manski.rule.general} with consistent estimators of the corresponding reduced-form parameters.

\item[(iii)] Finally, if the data is so scarce that the researcher believes the problem of lack of identification is not as important at all,  they may  simply apply our approach and evaluate the optimal rule by letting    $\overline{C}\downarrow0$ and $\underline{C}\uparrow0$.
\end{itemize}

\subsection{Local Asymptotic Optimality}

To complete characterizing the limit  game,  we impose a standard differentiability
in quadratic mean (DQM) condition for the statistical model  $P_{m(\theta)}$.
\begin{assumption}\label{asm:DQM} For an open set $\Gamma\subseteq M$
that contains $\gamma_{0}$, the model $\left\{ P_{m(\theta)}:m(\theta)\in\Gamma\right\} $
satisfies the DQM assumption when $m(\theta)=\gamma_{0}$, i.e., there
exists a function $s:\mathbf{Y}\rightarrow\mathbb{R}^{d_{m}}$ such
that
\begin{align*}
  \int\left[dP_{\gamma_{0}+h}^{1/2}(y)-dP_{\gamma_{0}}^{1/2}(y)-\frac{1}{2}h^{\top}s(y)dP_{\gamma_{0}}^{1/2}(y)\right]^{2}= o\left(\left\Vert h\right\Vert ^{2}\right),\text{ as }h\rightarrow0.
\end{align*}
Moreover, the Fisher information $\mathbf{I}_{0}=\mathbb{E}_{\gamma_{0}}\left[ss^{\prime}\right]$
is nonsingular. \end{assumption} \begin{lemma}\label{lem:DQM} Suppose
Assumptions \ref{asm:1}-\ref{asm:DQM} hold. For a sequence of
rules $d_{n}$, if $\mathbb{E}_{\gamma_{0}+\frac{h}{\sqrt{n}}}[d_{n}(Y^{n})]$
converges as $n\rightarrow\infty$ for every $h$, then there exists
some $d_\infty:\mathbb{R}^{d_m}\rightarrow[0,1]$ such that for every $h$,
\begin{equation}\label{eq:matching.rule}
\mathbb{E}_{\gamma_{0}+\frac{h}{\sqrt{n}}}[d_{n}(Y^{n})]\rightarrow\mathbb{E}[d_\infty(\Delta)]
\end{equation}
as $n\rightarrow\infty$, where $\Delta\sim\mathcal{N}(h,\mathbf{I}_{0}^{-1})$
follows a Gaussian distribution with mean $h$ and a covariance matrix
$\mathbf{I}_{0}^{-1}$. \end{lemma}

Lemma \ref{lem:DQM} simply restates Proposition 3.1 of \cite{HiranoPorter2009} in the context of our statistical model $P_{m(\theta)}$ and point of localization $\gamma_0$. Let
\begin{equation}\label{eq:theta.n}
\Theta_n: = \{ \theta \in \Theta \: | \: \theta = \theta_0 + \theta_h/\sqrt{n}, \quad \textrm{where } \theta_h=(h^{\top},\mu^*)^{\top}  \textrm{satisfies } h \in M_h, \mu^* \in I_{\infty}(h) \}.
\end{equation}
Together with
our choice of $\gamma_0$, $M_h$ and $I_\infty(h)$, observe that for each $\theta_n\in\Theta_n$, we have $m(\theta_n)=\gamma_0+\frac{h}{\sqrt{n}}$ and $U(\theta_n)=\frac{\mu^*}{\sqrt{n}}$, where $h\in M_h$ and $\mu^*\in I_\infty(h)$.
Therefore, for any $d_n$ that converges in the sense of Lemma \ref{lem:DQM}, we have
\begin{align*}
\sqrt{n}R(d_{n},\theta_n) & =\sqrt{n}\frac{\mu^{*}}{\sqrt{n}}\left(\mathbf{1}\left\{ \frac{\mu^{*}}{\sqrt{n}}\geq0\right\} -\mathbb{E}_{\gamma_{0}+\frac{h}{\sqrt{n}}}[d_{n}(Y^n)]\right)\\
 &\rightarrow \mu^{*}\left(\mathbf{1}\left\{ \mu^{*}\geq0\right\} -\mathbb{E}[d_{\infty}(\Delta)]\right),\text{as }n\rightarrow\infty.
\end{align*}
Then, the limit decision problem becomes the following:
The researcher observes $\Delta\sim\mathcal{N}(h,\mathbf{I}_{0}^{-1})$
and needs to come up with a rule $d_{\infty}:\mathbb{R}^{d_m}\rightarrow[0,1]$
that maps $\Delta$ to the unit interval. The limit payoff-relevant parameter is $\mu^{*}\in\mathbb{R}$,  partially identified with an
identified set $I_{\infty}(h)$.
In this limit game, a decision rule is MMR optimal if it achieves the value
\begin{equation}
R^*:=\min_{d_{\infty}}\sup_{h\in M_h,\mu^{*}\in I_{\infty}(h)}\mu^{*}\left(\mathbf{1}\left\{ \mu^{*}\geq0\right\} -\mathbb{E}[d_{\infty}(\Delta)]\right).\label{eq:limit.game.1}
\end{equation}
We are now ready to define our notion of local asymptotic MMR optimality.  For $\Theta_n$ that has the local parametrization \eqref{eq:theta.n},  let
\begin{equation}\label{eq:formal.local.space}
\Theta_{n}(J):=\left\{ \theta_n\in\Theta_n\mid \theta_n=\theta_0+\theta_{h}/\sqrt{n},\theta_h=(h^{\top},\mu^*)^{\top}, h\in J,\mu^{*}\in I_{\infty}(h)\right\},
\end{equation}
where $J$ is a finite subset of  $M_{h}$. A rule
$d_{n}$ is \emph{locally asymptotically MMR optimal} if\footnote{Let $\mathcal{D}$ be the set of all sequences of rules that converge in the sense of Lemma \ref{lem:DQM}. Our definition implies that $d_n$ is asymptotically MMR optimal in $\mathcal{D}$, i.e., the RHS of \eqref{eq:def.asymptotic.mmr} may be equivalently replaced with $\inf_{d'_n\in\mathcal{D}}\sup_{J}\liminf_{n\rightarrow\infty}\sup_{\theta_n\in\Theta_{n}(J)}\sqrt{n}R(d'_{n},\theta_n)$.}
\begin{equation}\label{eq:def.asymptotic.mmr}
\sup_{J}\liminf_{n\rightarrow\infty}\sup_{\theta_n\in\Theta_{n}(J)}\sqrt{n}R(d_{n},\theta_n)=R^{*}.
\end{equation}


\begin{theorem}\label{thm: general}
Suppose Assumptions \ref{asm:1}-\ref{asm:DQM} hold and let $d^*_{\infty}$ be a solution of  \eqref{eq:limit.game.1}. If a sequence of statistical decision rules $d_n:\mathbf{Y}^n\rightarrow[0,1]$ matches $d^*_{\infty}$ in the sense of \eqref{eq:matching.rule}, $d_n$ is locally asymptotically MMR optimal.
\end{theorem}
Theorem \ref{thm: general} is our first main result. It justifies a Gaussian experiment with a suitable local limit identified set  as a limit experiment for a large class of treatment choice problems with partial identification.
By finding a solution $d^*_\infty$ of the simpler Gaussian limit game \eqref{eq:limit.game.1}, one can discover an asymptotically optimal rule for the finite-sample problem \eqref{eq:finite.sample.mmr}. Specifically, Theorem \ref{thm: general} shows that any finite-sample rule $d_n$ such that $\mathbb{E}_{\gamma_{0}+h/\sqrt{n}}[d_{n}(Y^{n})]\rightarrow\mathbb{E}[d^*_\infty(\Delta)]$ for every $h$  as $n\rightarrow\infty$ is asymptotically MMR optimal in the sense of \eqref{eq:def.asymptotic.mmr}. In practice, once $d^*_\infty$ is found, $d_n$ can often be constructed by plugging in a best regular estimator for $\gamma_0$ as well as consistent estimators of other nuisance parameters in $d^*_\infty$.


\section{Solving the Limit Problem}\label{sec:results}

The illustrative example in Section \ref{sec:illustrate} corresponds to fully differentiable bounds and centrosymmetric parameter space. In general, however, the limit identified set in our characterization need not be centrosymmetric and may also display directionally differentiable bounds. In the next two results, we discover $d^*_\infty$ for two non-nested cases, namely by either relaxing centrosymmetry or full differentiability. These findings are also of interest as stand-alone results in the spirit of \citet{yata2021} or \citet{stoye2012minimax}.


\subsection{Fully Differentiable Bounds}
When $\underline{I}(\cdot)$
and $\overline{I}(\cdot)$ are fully differentiable at $\mu_{0}$, the identified set \eqref{eq:limit.ID} simplifies to
\[
I_{\infty}(h)=\left[\underline{C}+\left(\underline{I}^{(1)}(\mu_{0})\right)^{\top}G_{0}h,\overline{C}+\left(\overline{I}^{(1)}(\mu_{0})\right)^{\top}G_{0}h\right].
\]
In this case, the Gaussian limit game \eqref{eq:limit.game.1} can be solved by considering an even simpler, one-dimensional limit
game. Let $\kappa_{0}:=(\underline{I}(\mu_{0}))^{2}/(\overline{I}(\mu_{0}))^{2}>0$.  Lemma \ref{lem:CoMo} in Appendix \ref{sec:app.1}
establishes that
\begin{equation}\label{eq:co-mono}
\underline{I}^{(1)}(\mu_{0})=\overline{I}^{(1)}(\mu_{0})\cdotp\kappa_{0},
\end{equation}
i.e., the gradients of the lower and upper bounds of the original identified set at $\mu_0$ must share the same direction.\footnote{Intuitively, since $\mu_0$ is a global least favorable point in the sense of Assumption \ref{asm:2}, moving away from it cannot increase the value of $\overline{R}(\cdot)$. In the case of full differentiability, that pins down a first-order condition implying \eqref{eq:co-mono}.} It follows that $I_\infty(h)$ depends on $h$ only via $\upsilon:=\left(\overline{I}^{(1)}(\mu_{0})\right)^{\top}G_{0}h\in\mathbb{R}$. Therefore, consider the following one-dimensional game: a researcher observes a scalar-valued  signal
\[\hat{\upsilon}:=\left(\overline{I}^{(1)}(\mu_{0})\right)^{\top}G_0\Delta\sim\mathcal{N}\left(\upsilon,\sigma^{2}\right),\]
where $\sigma^{2}:=\left(\overline{I}^{(1)}(\mu_{0})\right)^{\top}G_{0}\mathbf{I}_{0}^{-1}G_{0}^{\top}\overline{I}^{(1)}(\mu_{0})$
is known and strictly positive;  a limit welfare contrast is $\upsilon^{*}\in\mathbb{R}$,
partially identified with an identified set
\[
I(\upsilon)=\left[\underline{C}+\kappa_{0}\upsilon,\overline{C}+\upsilon\right],
\]
for values of $\upsilon$ such that $\left(\kappa_{0}-1\right)\upsilon\leq\overline{C}-\underline{C}$.
In this simple, one-dimensional game, a rule $d:\mathbb{R}\rightarrow[0,1]$
is MMR optimal if it solves
\begin{equation}
\min_{d}\sup_{\upsilon\in\mathbb{R},\left(\kappa_{0}-1\right)\upsilon\leq\overline{C}-\underline{C},\upsilon^{*}\in I(\upsilon)}\upsilon^{*}\left(\mathbf{1}\left\{ \upsilon^{*}\geq0\right\} -\mathbb{E}[d(\hat{\upsilon})]\right)\label{eq:one.dim.limit.game.smooth}
\end{equation}

The shape of the parameter space in game \eqref{eq:one.dim.limit.game.smooth} depends on the value of $\kappa_0$ and can be one of the three cases illustrated in Figure \ref{fig:parameter.space.smooth}. The case of $\kappa_0=1$ nests a convex and centrosymmetric parameter space with fully differentiable bounds analyzed in the existing literature. Otherwise, the parameter space is still convex but not centrosymmetric in general.  We characterize analytically the solution of  \eqref{eq:one.dim.limit.game.smooth}, which we show is in fact a MMR solution of the limit game  \eqref{eq:limit.game.1}.

\begin{figure}[http]
\centering
\begin{tikzpicture}[
    axis/.style={->, thick, black},
    redline/.style={very thick, gray},
    blueline/.style={very thick, gray}
]

\begin{scope}[xshift=0cm]
  \begin{scope}
    \clip (-2.8,-2.8) rectangle (2.8,2.8);
    \fill[orange!30]
      (-1.67, -1.17)
      -- (2.8, {2.8+0.5})
      -- (2.8, {0.4*2.8-0.5})
      -- cycle;
    \draw[redline]  (-1.67,-1.17) -- (2.8, {2.8+0.5});
    \draw[blueline] (-1.67,-1.17) -- (2.8, {0.4*2.8-0.5});
  \end{scope}
  \draw[axis] (-2.8,0) -- (2.8,0) node[right] {$v$};
  \draw[axis] (0,-2.8) -- (0,2.8) node[above] {$v^*$};
\end{scope}

\begin{scope}[xshift=5.5cm]
  \begin{scope}
    \clip (-2.8,-2.8) rectangle (2.8,2.8);
    \fill[orange!30]
      (-2.8, {-2.8-0.5})
      -- (-2.8, {-2.8+0.5})
      -- (2.8, {2.8+0.5})
      -- (2.8, {2.8-0.5})
      -- cycle;
    \draw[redline]  (-2.8, {-2.8+0.5}) -- (2.8, {2.8+0.5});
    \draw[blueline] (-2.8, {-2.8-0.5}) -- (2.8, {2.8-0.5});
  \end{scope}
  \draw[axis] (-2.8,0) -- (2.8,0) node[right] {$v$};
  \draw[axis] (0,-2.8) -- (0,2.8) node[above] {$v^*$};
\end{scope}

\begin{scope}[xshift=11cm]
  \begin{scope}
    \clip (-2.8,-2.8) rectangle (2.8,2.8);
    \fill[orange!30]
      (1.5, 2.0)
      -- (-2.8, {-2.8+0.5})
      -- (-2.8, {3*-2.8-2.5})
      -- cycle;
    \draw[redline]  (-2.8, {-2.8+0.5})   -- (1.5, 2.0);
    \draw[blueline] (-2.8, {3*-2.8-2.5}) -- (1.5, 2.0);
  \end{scope}
  \draw[axis] (-2.8,0) -- (2.8,0) node[right] {$v$};
  \draw[axis] (0,-2.8) -- (0,2.8) node[above] {$v^*$};
\end{scope}


\end{tikzpicture}
\caption{Parameter space in limit game \eqref{eq:one.dim.limit.game.smooth} (from left to right: $\kappa_0<1$, $\kappa_0=1$ and $\kappa_0>1$)}
\label{fig:parameter.space.smooth}
\end{figure}
Let $\phi(\cdotp)$ and $\Phi(\cdotp)$ be the standard normal pdf and
cdf and  $\Phi^{-1}(\cdotp)$ the inverse of $\Phi(\cdotp)$. Define
\begin{align*}
\overline{\sigma} & :=-\frac{\overline{C}}{\underline{C}}\left(\overline{C}-\underline{C}\right)\phi\left(\Phi^{-1}\left(\frac{\overline{C}}{\overline{C}-\underline{C}}\right)\right),\quad
t^{*} :=-\Phi^{-1}\left(\frac{\overline{C}}{\overline{C}-\underline{C}}\right)\overline{\sigma}.
\end{align*}
Also, for each $t,\upsilon\in\mathbb{R}$, and each $\sigma>0$, write
\begin{align*}
R_{2}(t,\upsilon;\sigma)  :=(\upsilon+\overline{C})\Phi\left(\frac{t-\upsilon}{\sigma}\right),\quad
R_{1}(t,\upsilon;\sigma)  :=-(\kappa_{0}\upsilon+\underline{C})\Phi\left(\frac{\upsilon-t}{\sigma}\right).
\end{align*}






\begin{theorem}\label{thm:full-diff} Suppose
Assumptions \ref{asm:1}-\ref{asm:DQM} hold,  $\underline{I}(\cdot)$
and $\overline{I}(\cdot)$ are fully differentiable at $\mu_{0}$,  $G_{0}\mathbf{I}_{0}^{-1}G_{0}^{\top}$ is positive definite and $\sigma^{2}>0$.
The following results hold true for the limit game \eqref{eq:limit.game.1}.

\begin{itemize}
\item[(i)] If $\sigma<\overline{\sigma}$,
$
d_{\infty,RT}=\Phi\left(\frac{\hat{\upsilon}-t^{*}}{\sqrt{\overline{\sigma}^{2}-\sigma^{2}}}\right)
$
is MMR optimal.
\item[(ii)] If $\sigma\geq\overline{\sigma}$, then
$\mathbf{1}\left\{ \hat{\upsilon}\geq t_{0}\right\}$
is MMR optimal, where $t_{0}$ is the unique $t\in\mathbb{R}$ such
that
\begin{equation}\label{eq:fixed.point.program.simple}
\max_{\left\{ \upsilon\in\mathbb{R}:\left(\kappa_{0}-1\right)\upsilon\leq\overline{C}-\underline{C},\upsilon+\overline{C}\geq0\right\} }R_{2}(t,\upsilon;\sigma)=\max_{\left\{ \upsilon\in\mathbb{R}:\left(\kappa_{0}-1\right)\upsilon\leq\overline{C}-\underline{C},\kappa_{0}\upsilon+\underline{C}\leq0\right\} }R_{1}(t,\upsilon;\sigma).
\end{equation}
\end{itemize}
\end{theorem}

Theorem \ref{thm:full-diff} demonstrates two interesting features of the multivariate-signal game \eqref{eq:limit.game.1} with fully differentiable bounds at $\mu_0$. First, there exists a MMR optimal rule that depends on data $\Delta$ only via an efficient  linear combination $\hat{\upsilon}$. Along this direction, the limit MMR problem becomes effectively one-dimensional. Second, the key insights of the MMR optimal rule for the symmetric case \citep{stoye2012minimax,yata2021,MQS2023decision} extend to a non-centrosymmetric parameter space. In particular, $\overline{\sigma}$ depends only on $\overline{C}$ and $\underline{C}$ and can be interpreted as a measure of the severity of the identification problem, while $\sigma$ signals the effective estimation difficulty. If estimation difficulty dominates the severity of partial identification (case (ii)), an MMR optimal rule is a threshold rule, although the threshold is no longer zero in general and must be found via a (simple one-dimensional) fixed point  program \eqref{eq:fixed.point.program.simple}.\footnote{The idea of the fixed point program is analogous to that in \cite{tetenov2012statistical}, who focuses on MMR optimal rules in point-identified situations with asymmetric welfare regret. In this regard, we uncover an interesting connection between a point-identified problem with asymmetric regret and a partially identified problem with symmetric regret but non-centrosymmetric parameter space.}  If, however, the severity of partial identification  dominates estimation difficulty (case (i)), the  MMR optimal rule randomizes, in which case we find one randomized rule $d_{\infty,RT}$. We are confident that, in case (i), there are  infinitely many MMR optimal rules, although pinning down a piecewise linear rule in the spirit of \cite{MQS2023decision}  appears to be extremely algebraically involved. Theorem \ref{thm:full-diff} informs the following asymptotic results.

\begin{theorem}\label{thm:full-diff-asymptotic}

Suppose the conditions of Theorem \ref{thm:full-diff} hold.  Moreover, let $\hat{\gamma}$
be a best regular estimator of $\gamma_{0}$ such that
\[
\sqrt{n}(\hat{\gamma}-\gamma_{0}-\frac{h}{\sqrt{n}})\overset{h}{\rightsquigarrow}N(\mathbf{0},\mathbf{I}_{0}^{-1}),\text{ for all }h\in\mathbb{R}^{d_m}.
\]
Let $\hat{\sigma}$ be a consistent estimator of $\sigma$ when $h=\mathbf{0}$,
$\hat{t}$ be the unique $t\in\mathbb{R}$ such that
\eqref{eq:fixed.point.program.simple} holds with $\sigma$ replaced by $\hat{\sigma}$.  Then, the following feasible  rule is locally asymptotically
MMR optimal:
\[
d_{F}:=\mathbf{1}\left\{ \hat{\sigma}\geq\overline{\sigma}\right\} \cdot d_{F,\hat{t}}+\mathbf{1}\left\{ \hat{\sigma}<\overline{\sigma}\right\} \cdot d_{F,RT},
\]
where
\begin{align*}
d_{F,\hat{t}} & :=\mathbf{1}\left\{ \sqrt{n}\left(\overline{I}^{(1)}(\mu_{0})\right)^{\top}(\mu(\hat{\gamma})-\mu_0)\geq\hat{t}\right\} ,\quad
d_{F,RT}:  =\Phi\left(\frac{\sqrt{n}\left(\overline{I}^{(1)}(\mu_{0})\right)^{\top}(\mu(\hat{\gamma})-\mu_0)-t^{*}}{\sqrt{\overline{\sigma}^{2}-\hat{\sigma}^{2}}}\right).
\end{align*}
\end{theorem}
In our partially identified setting,
the shape of the limit optimal rule found in Theorem \ref{thm:full-diff} depends on the value of $\sigma$, which is unknown and must be estimated. As a result, establishing asymptotic optimality of $d_F$
 is more challenging than in \cite{HiranoPorter2009}. We show the discontinuities caused by the regime change do not affect the asymptotic validity of $d_F$. In particular, irrespective of the true value of $\sigma$, $d_F$ is always matched  with the right limiting optimal rule found in Theorem \ref{thm:full-diff}.

The form of our feasible  rule  explicitly depends on the location of $\mu_0$. Our asymptotic theory applies  to any least favorable point that satisfies Assumption \ref{asm:2} and does not require uniqueness of $\mu_0$. However, if there are multiple least favorable points, they may lead to different decision rules. When this happens, we think researchers should have the freedom to judge judiciously  which point of localization is more pertinent to their analyses. Otherwise, we recommend to report the one associated with a higher calculated value of $R^*$ defined in \eqref{eq:limit.game.1}.


\subsection{Directionally Differentiable Bounds}\label{sec:centro}
Many partially identified models feature only directionally differentiable bounds. When $\underline{I}(\cdot)$ and $\overline{I}(\cdot)$
 are only directionally differentiable at $\mu_{0}$, the geometry of our limit game becomes more complicated.
To gain tractability, let
\[
\Theta_{\mathcal{M}}:=\left\{ \left(\mu(m(\theta))^{\top},U(\theta)\right)^{\top}\in\mathbb{R}^{d_{\mu}+1}\mid\theta\in\Theta\right\}
\]
be the set of values the decision-relevant parameter $\mu(m(\theta))$
and the payoff-relevant parameter $U(\theta)$ can take for the original problem \eqref{eq:finite.sample.mmr}.
We maintain
the following:
\begin{assumption}\label{asm:centro} \begin{itemize}
\item[(i)] $\Theta_{\mathcal{M}}$ is
convex and centrosymmetric.
\item[(ii)] There exists some $\mu\in M_{\mu}$
such that $\overline{I}(\mu)>\overline{I}(\mathbf{0})$.
\end{itemize}
\end{assumption}
With the additional Assumption \ref{asm:centro}(i), $\overline{I}(\cdot)$ is concave.\footnote{For example, one may apply \citet[][Lemma B.4]{yata2021} or \citet[][Lemma C.2]{MQS2023decision} to $\Theta_{\mathcal{M}}$.} Moreover, Lemma \ref{lem:global.lfp} shows that  $\mathbf{0}\in\mathbb{R}^{d_{\mu}}$ is a global least favorable point, at which we recenter our reduced-form parameter. As $h$ only affects the limit identified set via $G_{0}h$, we may still simplify the problem by  redefining a limit decision-relevant reduced-form parameter $\mu:=G_{0}h$\footnote{We slightly abuse notation here because $\mu$ is not the same as $\mu\in M_{\mu}$ in the original parameter space.} and
searching among rules that depend on $\Delta$ only via
\[
G_0\Delta\sim\mathcal{N}(\mu,\Sigma_{0}),
\]
where $\Sigma_{0}:=G_{0}\mathbf{I}_{0}^{-1}G_{0}^{\top}$ is assumed to be positive definite. The identified set for the limit payoff-relevant parameter then simplifies to
\begin{equation}
 I_{\infty}(\mu)=[\underline{I}_{\infty}(\mu),\overline{I}_{\infty}(\mu)],\label{eq:limit.id.set.simple}
\end{equation}
where
$\underline{I}_{\infty}(\mu)=\underline{C}+\underline{I}^{(1)}(\mathbf{0};\mu)$,
$\overline{I}_{\infty}(\mu)=\overline{C}+\overline{I}^{(1)}(\mathbf{0};\mu)$, and $\mu$ can take values in the set
\[
M_{\mu,\infty}:=\{\mu\in\mathbb{R}^{d_{\mu}}\mid \overline{I}_{\infty}(\mu)\geq\underline{I}_{\infty}(\mu)\}.
\]
Thus, one may consider finding a rule $d:\mathbb{R}^{d_\mu}\rightarrow[0,1]$ that solves  the following problem
\begin{equation}
\min_{d}\sup_{\mu^{*}\in[\underline{I}_{\infty}(\mu),\overline{I}_{\infty}(\mu)],\mu\in M_{\mu,\infty}}\mu^{*}\left[\mathbf{1}\left\{ \mu^{*}\geq0\right\} -\mathbb{E}[d(G_0\Delta)]\right].\label{pf:R.star}
\end{equation}
Note the parameter space of $(\mu,\mu^*)$ in \eqref{pf:R.star} is still convex and centrosymmetric, but data $G_0\Delta$ becomes Gaussian. This characterization endorses  the finite-sample results considered by \citet{yata2021} as in fact a \emph{bona fide} limiting decision problem, with a suitable adjustment of the  original identified set to its localized version. Therefore,  the main results in \cite{yata2021} apply to \eqref{pf:R.star}; in particular, an optimal MMR rule utilizes an efficient linear combination of $G_0\Delta$ and can be found by applying \citet[][Theorem 1]{yata2021}. Moreover, when $\overline{C}$ is sufficiently large, the results in \cite{MQS2023decision} imply there are infinitely many MMR optimal rules. Next, we contribute to the literature by showing that the efficient linear combination can  be alternatively found by solving either a simple quadratic program \eqref{eq:quadratic.program}  or a nonlinear optimization program \eqref{eq:non.benign} below.

Denote by $\partial\overline{I}(\mathbf{0})$ the superdifferential
of $\overline{I}(\cdotp)$ at $\mathbf{0}$.  By \citet[][p.217-218, Section 23]{rockafellarconvex}, we have
\[
\overline{I}^{(1)}(\mathbf{0};\mu)=\min_{p\in\partial\overline{I}(\mathbf{0})}p^{\top}\mu,\quad \forall  \mu \in \mathbb{R}^{d_\mu},
\]
where $\partial\overline{I}(\mathbf{0})$ is nonempty, compact, and
convex and $\overline{I}^{(1)}(\mathbf{0};\cdotp)$ is a finite positively
homogeneous concave function (i.e., $\overline{I}^{(1)}(\mathbf{0};\alpha\mu)=\alpha\overline{I}^{(1)}(\mathbf{0};\mu)$ for all $\alpha>0$). Moreover, Assumption \ref{asm:centro}(ii) implies
that $\mathbf{0}\notin\partial\overline{I}(\mathbf{0})$.
Therefore, denote  by  $\overline{p}:=\overline{p}(\Sigma_0)$ the
unique solution of
\begin{equation}\label{eq:quadratic.program}
\min_{p\in\partial\overline{I}(\mathbf{0})}\left\{ p^{\top}\Sigma_{0}p\right\}^{1/2},
\end{equation}
and write
\begin{align*}
\overline{t} & :=\frac{2\overline{C}}{\underline{I}^{(1)}(\mathbf{0};\Sigma_{0}\overline{p})-\overline{I}^{(1)}(\mathbf{0};\Sigma_{0}\overline{p})}\in(0,\infty],
\end{align*}
with the understanding that $\overline{t}=\infty$ if $\underline{I}^{(1)}(\mathbf{0};\Sigma_{0}\overline{p})-\overline{I}^{(1)}(\mathbf{0};\Sigma_{0}\overline{p})=0$.

\begin{theorem}\label{thm:non-diff}

Suppose Assumptions \ref{asm:1}-\ref{asm:centro} hold and $\Sigma_0$ is positive definite. Consider the limit game \eqref{eq:limit.game.1} with $\mu_0=\mathbf{0}$. The following statements hold true.
\begin{itemize}
\item[(i)] If $\left( \frac{\pi}{2}\overline{p}^{\top}\Sigma_{0}\overline{p}\right) ^{1/2}<\overline{C}$,
\[
d_{\infty,RT}:=\Phi\left(\frac{\overline{p}^{\top}G_0\Delta}{\sqrt{\frac{2}{\pi}\overline{C}^2-\overline{p}^{\top}\Sigma_{0}\overline{p}}}\right)
\]
is MMR optimal.
\item[(ii)]  If $\left( \frac{\pi}{2}\overline{p}^{\top}\Sigma_{0}\overline{p}\right) ^{1/2}\geq\overline{C}$,
and
\begin{equation}\label{eq:benign.non.benign.boundary}
\overline{x}:=\overline{x}(\Sigma_0) :=\arg\max_{x\in[0,\infty)}(\overline{C}+x)\Phi\left(-\frac{x}{\left(\overline{p}^{\top}\Sigma_{0}\overline{p}\right)^{1/2}}\right)\leq\overline{t}\overline{I}^{(1)}(\mathbf{0};\Sigma_{0}\overline{p}),
\end{equation}
$d_{\infty,\overline{p}}:=\mathbf{1}\left\{ \overline{p}^{\top}G_0\Delta\geq0\right\}$
is MMR optimal.
\item[(iii)] Otherwise, $d_{\infty,\mu^{o}}:=\mathbf{1}\left\{ \left(\mathbf{\mu}^{o}\right)^{\top}\Sigma_{0}^{-1}G_0\Delta\geq0\right\}$ is MMR optimal, where
$\mathbf{\mu}^{o}:=\mathbf{\mu}^{o}(\Sigma_0)$ uniquely solves
\begin{equation}\label{eq:non.benign}
\max_{\{\mu\in M_{\mu,\infty}:\overline{I}^{(1)}(\mathbf{0};\mu)\geq0\}}r(\mu;\Sigma_0),
 \end{equation}
and  $r(\mu;\Sigma_0):=\left(\overline{C}+\overline{I}^{(1)}(\mathrm{\mathbf{0}};\mu)\right)\Phi\left(-\sqrt{\mu^{\top}\Sigma_0^{-1}\mu}\right)$.
\end{itemize}
\end{theorem}


The intuition behind Theorem \ref{thm:non-diff} is as follows. Focus on the limit game among rules that depend on $\Delta$ only via $G_0\Delta$. Consistent with the existing results in the literature, a least favorable prior randomizes evenly between two symmetric points around $\mathbf{0}$, i.e., between  $(-\mu,\underline{I}_\infty(-\mu))$ and $(\mu,\overline{I}_\infty(\mu))$ where $\overline{I}_\infty(\mu)> 0$.  In ``benign cases'' where $\Sigma_0$ is not too large (statements (i) and (ii) above), the direction of $\mu$ is found via \eqref{eq:quadratic.program}, which determines an efficient direction along $\Sigma_0\overline{p}$. Along this direction, the game again essentially becomes one-dimensional. In cases where $\left( \frac{\pi}{2}\overline{p}^{\top}\Sigma_{0}\overline{p}\right) ^{1/2}\leq\overline{C}$,  one has $\mu=0$, and a limit MMR optimal rule randomizes whenever the inequality is strict. In fact, applying \citet[][Theorem 1]{MQS2023decision} to case (i), we can conclude that infinitely many MMR optimal rules exist, among them the piecewise linear rule
\begin{equation*}\label{eq:d.linear}
d^{*}_{\text{linear}}:=\left[\frac{\overline{p}^{\top}G_0\Delta+\rho^{*}}{2\rho^{*}}\right]_0^1,
\end{equation*}
where $\rho^*$ is the unique strictly positive solution of
\begin{equation*}
\left(\frac{\rho^{*}}{2\overline{C}}\right) -\frac{1}{2} + \Phi\left(-\frac{\rho^{*}}{\left(\overline{p}^{\top}\Sigma_{0}\overline{p}\right)^{1/2}}\right) =0.
\end{equation*}
In case (ii), $\mu$ has a strictly  positive length (i.e., $\mu=t\Sigma_0\overline{p}$ for some $t>0$) whenever the inequality is strict, and $d_{\infty,\overline{p}}$ is the (a.e.) unique Bayes response. A caveat is that the described  equilibrium can only be sustained if the length of $\mu$ is not too large,  as $\mu$ must also be in  $M_{\mu,\infty}$, which may be bounded in direction $\Sigma_0\overline{p}$. Condition \eqref{eq:benign.non.benign.boundary} gives this boundary condition. If it is  violated, we must search for a possibly alternative optimal direction via the nonlinear program \eqref{eq:non.benign}. In all cases, we also provide a suitable two-point symmetric prior in the space of $(h,\mu^*)$ that supports our decision rules, confirming that, although they depend only on $G_0\Delta$, they are MMR optimal for problem \eqref{eq:limit.game.1}. With Theorem \ref{thm:non-diff}, we next provide  a feasible plug-in rule that will be asymptotically MMR optimal.

\begin{theorem}\label{thm:non-diff-asymptotic}

Suppose the conditions of Theorem \ref{thm:non-diff} hold. Moreover, let $\hat{\gamma}$ be a best regular estimator of $\gamma_{0}$ such that \[ \sqrt{n}(\hat{\gamma}-\gamma_{0}-\frac{h}{\sqrt{n}})\overset{h}{\rightsquigarrow}N(\mathbf{0},\mathbf{I}_{0}^{-1}),\text{ for all }h\in\mathbb{R}^{d_m}. \]
Let $\hat{\Sigma}$ be a positive definite and consistent estimator of $\Sigma_0$ when $h=\mathbf{0}$,  $\hat{p}:=\overline{p}(\hat{\Sigma})$, $\hat{x}:=\overline{x}(\hat{\Sigma})$ and let $\mathbf{\hat{\mu}}^{o}$ solve \eqref{eq:non.benign} with $\Sigma_0$ replaced by $\hat{\Sigma}$. The following feasible rule is locally asymptotically MMR optimal:
\begin{align*}
d_{F} & :=\mathbf{1}\left\{ \left(\frac{\pi}{2}\hat{p}^{\top}\hat{\Sigma}\hat{p}\right)^{1/2}<\overline{C}\right\} d_{F,RT}\\
 & +\mathbf{1}\left\{ \left(\frac{\pi}{2}\hat{p}^{\top}\hat{\Sigma}\hat{p}\right)^{1/2}\geq\overline{C},\hat{x}\leq\frac{2\overline{C}\overline{I}^{(1)}(\mathbf{0};\hat{\Sigma}\hat{p})}{\underline{I}^{(1)}(\mathbf{0};\hat{\Sigma}\hat{p})-\overline{I}^{(1)}(\mathbf{0};\hat{\Sigma}\hat{p})}\right\} d_{F,\hat{p}}\\
 & +\mathbf{1}\left\{ \left(\frac{\pi}{2}\hat{p}^{\top}\hat{\Sigma}\hat{p}\right)^{1/2}\geq\overline{C},\hat{x}>\frac{2\overline{C}\overline{I}^{(1)}(\mathbf{0};\hat{\Sigma}\hat{p})}{\underline{I}^{(1)}(\mathbf{0};\hat{\Sigma}\hat{p})-\overline{I}^{(1)}(\mathbf{0};\hat{\Sigma}\hat{p})}\right\} d_{F,\hat{\mu}^{o}},
\end{align*}
where
\begin{align*}
d_{F,RT} & :=\Phi\left(\frac{\sqrt{n}\hat{p}^{\top}\mu(\hat{\gamma})}{\sqrt{\frac{2}{\pi}\overline{C}^{2}-\hat{p}^{\top}\hat{\Sigma}\hat{p}}}\right),d_{F,\hat{p}}:=\mathbf{1}\left\{ \hat{p}^{\top}\mu(\hat{\gamma})\geq0\right\} ,
d_{F,\hat{\mu}^{o}} :=\mathbf{1}\left\{ \left(\mathbf{\hat{\mu}}^{o}\right)^{\top}\hat{\Sigma}^{-1}\mu(\hat{\gamma})\geq0\right\}.
\end{align*}

\end{theorem}

\section{Applications}\label{sec:applications}

\subsection{Contaminated and Corrupted Data}\label{sec:contaminate}
This example corresponds to a parameter space that can be recast as  convex and centrosymmetric with fully differentiable bounds; see \cite{qiu2026statistical} for a related example in regression discontinuity design.  Suppose a DM must choose between two treatments based on RCT data, but some outcomes are contaminated. For each unit $i$, let  $D_i\in\{0,1\}$ be its treatment status, where $D_i=1$ denotes treatment and $D_i=0$ denotes control. The realized outcome for each unit is  $Y_i=D_iY_i(1)+(1-D_i)Y_i(0)$, where $Y_i(1),Y_i(0)\in\{0,1\}$ are binary potential outcomes (success/no success) under treatment and control, respectively. The DM also observes $S_i\in\{0,1\}$, where $S_i=1$ means the realized outcome $Y_i$ is contaminated or missing and therefore cannot be used for decision making. The payoff-relevant parameter is the average treatment effect $\mu^*=\mathbb{E}[Y(1)]-\mathbb{E}[Y(0)]$. \cite{horowitz1995identification} derive the sharp identified set of $\mu^*$:
\begin{equation}\label{eq:HM-bound}
[(1-p_1)\gamma_1-(1-p_0)\gamma_0-p_0,(1-p_1)\gamma_1-(1-p_0)\gamma_0+p_1],
\end{equation}
where $\gamma_1=\mathbb{E}[Y_i\mid D_i=1,S_i=0]$, $\gamma_0=\mathbb{E}[Y_i\mid D_i=0,S_i=0]$, $p_1=\mathbb{E}[S_i=1\mid D_i=1]$, and $p_0=\mathbb{E}[S_i=1\mid D_i=0]$. Suppose the contamination rates $p_1,p_0\in(0,1)$ are known, e.g., from RCT experts or through validations from past experiments. This fits  our framework with a convex and centrosymmetric parameter space by writing $\theta=(\mu^*,\gamma_1,\gamma_0)$, $m(\theta)=(\gamma_1,\gamma_0)$, and $U(\theta)=\mu^*$. The identified set depends on $\gamma_1$ and $\gamma_0$ only via $\mu:=(1-p_1)\gamma_1-(1-p_0)\gamma_0-(p_0-p_1)/2$, so that we can rewrite the identified set more compactly as $I(\mu)=[\underline{I}(\mu),\overline{I}(\mu)]$, where
\begin{equation}\label{eq:HM-bound-centro}
 \underline{I}(\mu)=\mu-k, \overline{I}(\mu)=\mu+k
\end{equation}
and $k=(p_1+p_0)/2\in(0,1)$. The global least favorable point has $\mu_0=0$,  $\overline{I}^{(1)}(\mu_0)=1$, and $\kappa_0=1$.  Theorem \ref{thm:full-diff-asymptotic} directly applies.

If $p_1$ and $p_0$ are also unknown, they become part of $m(\theta)$. In this case, $k$ in \eqref{eq:HM-bound-centro} is unknown as well. The global least favorable point  occurs on the boundary violating  Assumption \ref{asm:2}, namely at $p_1=p_0=1$, i.e.,  no realized outcome is observed and the identified set becomes completely uninformative (also resembling an issue faced in \cite{song2014point}). Our result does not apply to this irregular case.   In Appendix \ref{sec:non.standard}, we slightly modify our asymptotic framework to accommodate this situation so that one may use  estimated versions of $p_1$ and $p_0$ to construct optimal rules. See also \cite{xu2026asymptotic} for an alternative treatment.

\subsection{Robust Welfare Analyses}
This example is inspired by \cite{kang2025robustness} and corresponds to a scenario with fully differentiable bounds but non-centrosymmetric parameter space. A policy maker contemplates levying an \emph{ad valorem}  tax $\tau$ on a good and using the tax revenue to generate some known social surplus $G$. Suppose the demand curve does not shift before and after the tax levy.  The status-quo price and quantity, denoted as $p_0$ and $q_0$, are known. The new price after the levy, $p_1=p_0(1+\tau)$, is also known.  The policy maker observes some individual-level experimental data that point-identify the new quantity $q_1:=q_1(\gamma)$, which itself can be a smooth transformation of some finite-dimensional parameter $\gamma$ from   structural or reduced-form modeling. Even if $q_1$ is perfectly observed, we still do not know the true knowledge of the demand curve between $p_1$ and $p_0$. As a result, the loss in consumer surplus   $\Delta CS$ is only partially identified. In this example, $U(\theta)=G-\Delta CS$, while $m(\theta)=\gamma$.
\begin{figure}[http]
\centering
\begin{tikzpicture}[axis/.style={->, thick, black}]

\pgfmathsetmacro{\xuint}{4.75}
\pgfmathsetmacro{\xlint}{\xuint * 0.2}
\pgfmathsetmacro{\slopeU}{2.0/\xuint}
\pgfmathsetmacro{\slopeL}{4.0/\xuint}

\pgfmathsetmacro{\lowerAtRight}{0.8 - \slopeL*\xuint}
\pgfmathsetmacro{\upperAtLeft} {2.0 - \slopeU*\xlint}

\coordinate (UL) at (\xlint, \upperAtLeft);
\coordinate (UR) at (\xuint, 0);
\coordinate (LR) at (\xuint, \lowerAtRight);
\coordinate (LL) at (\xlint, 0);

\fill[orange!25] (UL) -- (UR) -- (LR) -- (LL) -- cycle;

\draw[very thick, gray]
    (-0.3, {2.0 - \slopeU*(-0.3)}) -- ({\xuint+0.8}, {2.0 - \slopeU*(\xuint+0.8)});
\draw[very thick, gray]
    (-0.3, {0.8 - \slopeL*(-0.3)}) -- ({\xuint+0.8}, {0.8 - \slopeL*(\xuint+0.8)});

\draw[axis] (-0.5, 0) -- ({\xuint+1.2}, 0) node[right] {$q_1$};
\draw[axis] (0, -4.0) -- (0, 3.5)          node[above] {$G - \Delta CS$};

\draw[thick, myred, dashed] (\xlint,  3.2) -- (\xlint, -3.8);
\draw[thick, myred, dashed] (\xuint,  3.2) -- (\xuint, -3.8);

\end{tikzpicture}
    \caption{Identified set of the net social surplus in a stylized example from \cite{kang2025robustness}. The area between the dotted red lines is where the decision problem is nontrivial and where the global least favorable point is located.}
    \label{fig:robust.welfare}
\end{figure}

Assuming the gradient of the demand curve at any price between $p_1$ and $p_0$ is bounded between
\[
\left[\frac{q_1-q_0}{p_1-p_0}\frac{1}{1-r},\frac{q_1-q_0}{p_1-p_0}(1-r)\right]
\]
for some given $r\in(0,1)$, \cite{kang2025robustness} find the sharp identified set of $G-\Delta CS$ as $I(q_1):=[\underline{I}(q_1),\overline{I}(q_1)]$, where
\[
\overline{I}(q_1)= G-\frac{(p_{1}-p_{0})(1-r)q_{0}}{2-r}-\frac{p_{1}-p_{0}}{2-r}q_{1},\quad \underline{I}(q_1)= G-\frac{(p_{1}-p_{0})q_{0}}{2-r}-\frac{\left(p_{1}-p_{0}\right)\left(1-r\right)}{2-r}q_{1},
\]
which  depend on $\gamma$ only through $q_1$ and are affine in $q_1$.  See Figure \ref{fig:robust.welfare} for an illustration.  The problem is only interesting when $G$ is in the relevant range so that there exists some $q_1$ such that $\overline{I}(q_1)>0>\underline{I}(q_1)$. In that case, the global least favorable location $q_1^0$ has a closed-form solution. Once  $q_1^0$ is found, we can set $\overline{C}=\sqrt{n}\overline{I}(q_1^0)$ and $\underline{C}=\sqrt{n}\underline{I}(q_1^0)$.  Furthermore, note $\overline{I}^{(1)}(q_1^0)=-\left(p_{1}-p_{0}\right)/(2-r)$ and $\kappa_0=1-r<1$, implying the limit identified set corresponds to a case in the left panel of Figure \ref{fig:parameter.space.smooth}. Then, Theorem \ref{thm:full-diff-asymptotic} can be applied with an asymptotically efficient estimator for $q_1$ (e.g., by applying the transformation $q_1(\hat{\gamma})$ where $\hat{\gamma}$ is the MLE for $\gamma$) and  a consistent estimator of its asymptotic variance.



\subsection{Evidence Aggregation}\label{sec:IK}
We revisit an evidence aggregation example in \citet{ishihara2021}.
Suppose a DM is interested in implementing a new policy in a target
country and observes a random sample $\left\{ Y_{i},S_{i}\right\} _{i=1}^{n}$
collected from (for simplicity) two nearby countries, where $Y_{i}$
denotes the outcome of interest when the policy is applied to unit
$i$ and $S_{i}$ is a binary indicator for the country of origin
of each unit ($S_{i}=0$ means unit $i$ is from country 1 and $S_{i}=1$
means unit $i$ is from country 2). The DM specifies a parametric
model $P_{\vartheta}$ for $\{Y_{i},S_{i}\}$. The policy effects
of the two nearby countries, denoted as $\mu_{1}:=\mu_{1}(\vartheta)$
and $\mu_{2}:=\mu_{2}(\vartheta)$ respectively, are modeled as known
smooth functionals of $\vartheta$. The status quo policy effect is
known and normalized to zero. The DM is willing to extrapolate the
policy effect in the target country based on those from nearby countries:
\[
\left\{ \mu_{0}\in\mathbb{R}:\left|\mu_{0}-\mu_{j}\right|\leq C_{j},j=1,2\right\}
\]
for some $C_{j}>0,j=1,2$. In this example, we have $\theta=\left(\mu_{0},\vartheta\right)\in\mathbb{R}^{1+d_{\vartheta}}$,
$U(\theta)=\mu_{0}$, $m(\theta)=\vartheta\in\mathbb{R}^{d_{\vartheta}}$,
and the identified set of $\mu_{0}$ depends on $m(\theta)$ only
via $\mu_{1}$ and $\mu_{2}$.
Furthermore, the identified set for $\mu_{0}$ can be written as intersection
bounds:
\begin{equation}\label{eq:IK.original.bound}
I(\mu_{1},\mu_{2})=\left[\underline{I}(\mu_{1},\mu_{2}),\overline{I}(\mu_{1},\mu_{2})\right],
\end{equation}
where $\underline{I}(\mu_{1},\mu_{2})=\max\left\{ \mu_{1}-C_{1},\mu_{2}-C_{2}\right\} $
and $\overline{I}(\mu_{1},\mu_{2})=\min\left\{ \mu_{1}+C_{1},\mu_{2}+C_{2}\right\} $.
The space of $\left(\mu_{0},\mu_{1},\mu_{2}\right)$ is convex and
centrosymmetric, so $\mu_{1}=\mu_{2}=0$ is a global least favorable
point.

Suppose there is one unique nearest neighbor, e.g., $0<C_{1}<C_{2}$.
Then, note at $\left(0,0\right)$, $\overline{I}(\mu_{1},\mu_{2})$
is fully differentiable with $\overline{I}^{(1)}(0,0)=(1,0)$. As
a result, our Theorems \ref{thm:full-diff} and \ref{thm:full-diff-asymptotic}
apply. In particular, the limit identified set in the sense of (\ref{eq:limit.id.set.simple})
would be
\begin{equation}\label{eq:ik.limit.bound.1}
\bar{I}_{\infty}(\mu_{1},\mu_{2})=\mu_{1}+\overline{C},~~\underline{I}_{\infty}(\mu_{1},\mu_{2})=\mu_{1}-\overline{C},
\end{equation}
where $\overline{C}=\sqrt{n}\overline{I}(0,0)=\sqrt{n}C_{1}$. Note
how the geometry of \eqref{eq:ik.limit.bound.1} differs from the original identified set \eqref{eq:IK.original.bound},  precisely due to the way we constructed our shrinking identified set in \eqref{eq:local.id.U}-\eqref{eq:local.id.U.2}
that only incorporates the local information of $I(\cdotp)$ around $\left(0,0\right)$.\footnote{This is not the only reasonable way to approximate the decision problem. In Appendix \ref{sec:global.ID}, we discuss an alternative approach to construct the shrinking identified set, which allows incorporating more features of the original identified set. }  Let $\hat{\vartheta}$ be the MLE for $\vartheta$ and define $\hat{\mu}:=(\hat{\mu}_{1},\hat{\mu}_{2}):=(\mu_{1}(\hat{\vartheta}),\mu_{2}(\hat{\vartheta}))$.
Suppose
\[
\Sigma_{0}:=\left(\begin{array}{cc}
\sigma_{1}^{2} & \rho\sigma_{1}\sigma_{2}\\
\rho\sigma_{1}\sigma_{2} & \sigma_{2}^{2}
\end{array}\right)
\]
is positive definite with $\sigma_{1},\sigma_{2}>0$ and $\rho\in(-1,1)$.
Let
\[
\hat{\Sigma}:=\left(\begin{array}{cc}
\hat{\sigma}_{1}^{2} & \hat{\rho}\hat{\sigma}_{1}\hat{\sigma}_{2}\\
\hat{\rho}\hat{\sigma}_{1}\hat{\sigma}_{2} & \hat{\sigma}_{2}^{2}
\end{array}\right)
\]
be a positive definite and consistent estimator of $\Sigma_{0}$.
Theorems \ref{thm:full-diff} and \ref{thm:full-diff-asymptotic}
imply that the asymptotically optimal decision rule would be either
$\mathbf{1}\left\{ \hat{\mu}_{1}\geq0\right\} $ or its randomized
version depending on the magnitude of $2\overline{C}_{1}^{2}/\pi$
versus  $\hat{\sigma}_{1}^{2}$.

Now, consider the case when $C_{1}=C_{2}=C>0$. Then, at
$\left(0,0\right)$, $\overline{I}(\mu_{1},\mu_{2})$ is only directionally
differentiable and our Theorems \ref{thm:non-diff} and \ref{thm:non-diff-asymptotic}
apply. We set $\overline{C}=-\underline{C}=\sqrt{n}C$, and can calculate
that
\[
\partial\overline{I}(0,0)=\left\{ \left(e,1-e\right)^{\top}\in\mathbb{R}^{2}\mid e\in[0,1]\right\} .
\]
For each $\mathbf{v}=(\mathbf{v}_{1},\mathbf{v}_{2})$ with unit length,
we have $\overline{I}^{(1)}((0,0);\mathbf{v})=\min\left\{ \mathbf{v}_{1},\mathbf{v}_{2}\right\} $,
which is nonnegative for all $\mathbf{v}$ in the first quadrant (i.e.,
both $\mathbf{v}_{1}$ and $\mathbf{v}_{2}$ are nonnegative). Analogously,
$\underline{I}^{(1)}((0,0);\mathbf{v})=\max\left\{ \mathbf{v}_{1},\mathbf{v}_{2}\right\} $.
Thus, in this example, the identified set in the limit game becomes
\[
\bar{I}_{\infty}(\mu)=\min\left\{ \mu_{1},\mu_{2}\right\} +\overline{C},~~\underline{I}_{\infty}(\mu)=\max\left\{ \mu_{1},\mu_{2}\right\} -\overline{C}.
\]
Moreover, along each direction $\mathbf{v}$ in the first quadrant,
the maximum possible length that $\mu$ can take is
\begin{align*}
\overline{t}(\mathbf{v}) & =\frac{2\overline{C}}{\max\left\{ \mathbf{v}_{1},\mathbf{v}_{2}\right\} -\min\left\{ \mathbf{v}_{1},\mathbf{v}_{2}\right\} }=\begin{cases}
\frac{2\overline{C}}{\mathbf{v}_{1}-\mathbf{v}_{2}}, & \mathbf{v}_{1}>\mathbf{v}_{2},\\
\infty, & \mathbf{v}_{1}=\mathbf{v}_{2},\\
\frac{2\overline{C}}{\mathbf{v}_{2}-\mathbf{v}_{1}}, & \mathbf{v}_{1}<\mathbf{v}_{2}.
\end{cases}
\end{align*}
See also Figure \ref{fig:ik.example} for an illustration of $M_{\mu,\infty}$
and $\left\{ \mu\in M_{\mu,\infty}:\overline{I}^{(1)}(\mathbf{0};\mu)\geq0\right\} $
in this example.
\begin{figure}
\centering

\begin{tikzpicture}[
    axis/.style={->, thick, black},
    grayline/.style={very thick, gray},
    redline/.style={very thick, red},
    blueline/.style={very thick, blue}
]

\begin{scope}[xshift=0cm]
          
  \begin{scope}
    \clip (-2.8,-2.8) rectangle (2.8,2.8);
    \begin{scope}
      \clip (0,0) rectangle (2.8,2.8);
      \fill[orange!30]
        (-2.8, {-2.8-1.5})
        -- (-2.8, {-2.8+1.5})
        -- (2.8, {2.8+1.5})
        -- (2.8, {2.8-1.5})
        -- cycle;
    \end{scope}
    \draw[grayline]  (-2.8, {-2.8+1.5}) -- (2.8, {2.8+1.5});
    \draw[grayline] (-2.8, {-2.8-1.5}) -- (2.8, {2.8-1.5});
    \begin{scope}
      \clip (0,0) rectangle (2.8,2.8);
      \draw[gray, thin]  (-2.8, {-2.8+1.5}) -- (2.8, {2.8+1.5});
      \draw[gray, thin] (-2.8, {-2.8-1.5}) -- (2.8, {2.8-1.5});
    \end{scope}
    \draw[very thick, magenta] (-2.8,-2.8) -- (2.8,2.8);
    \draw[->, very thick, magenta] (2.8-0.01,2.8-0.01) -- (2.8,2.8);
  \end{scope}

  \draw[axis] (-2.8,0) -- (2.8,0) node[right] {$\mu_1$};
  \draw[axis] (0,-2.8) -- (0,2.8) node[above] {$\mu_2$};

  \node[left]  at (0, 1.5+0.1)   {$2\overline{C}$};
  \node[right] at (0, -1.5)      {$-2\overline{C}$};
  \node[below] at (1.5+0.1, 0)   {$2\overline{C}$};
  \node[above] at (-1.5-0.2, 0)  {$-2\overline{C}$};

  \node[magenta, font=\small] at (2.3, 1.62)
    {$\dfrac{\Sigma_0 \overline{p}}{\|\Sigma_0 \overline{p}\|}$};
\end{scope}

\begin{scope}[xshift=8cm]
  \tdplotsetmaincoords{80}{-20}
  \begin{scope}[tdplot_main_coords, scale=0.9, xshift=0cm, yshift=-0.75cm]
                    
    \fill[blue!20, opacity=0.75]
      ({0-1.2}, {0+1.2}, {0*0.43})
      -- ({3.5-1.2}, {3.5+1.2}, {3.5*0.43})
      -- ({3.5}, {3.5}, {3.5*0.43-1.5})
      -- ({0}, {0}, {0*0.43-1.5})
      -- cycle;
    \draw[blue!60, thick]
      ({0-1.2}, {0+1.2}, {0*0.43})
      -- ({0}, {0}, {0*0.43-1.5});
    \draw[blue!60, thick]
      ({3.5-1.2}, {3.5+1.2}, {3.5*0.43})
      -- ({3.5}, {3.5}, {3.5*0.43-1.5});
    \draw[blue!60, thick]
      ({0}, {0}, {0*0.43-1.5})
      -- ({3.5}, {3.5}, {3.5*0.43-1.5});

    \fill[blue!20, opacity=0.75]
      ({0+1.2}, {0-1.2}, {0*0.43})
      -- ({3.5+1.2}, {3.5-1.2}, {3.5*0.43})
      -- ({3.5}, {3.5}, {3.5*0.43-1.5})
      -- ({0}, {0}, {0*0.43-1.5})
      -- cycle;
    \draw[blue!60, thick]
      ({0+1.2}, {0-1.2}, {0*0.43})
      -- ({0}, {0}, {0*0.43-1.5});
    \draw[blue!60, thick]
      ({3.5+1.2}, {3.5-1.2}, {3.5*0.43})
      -- ({3.5}, {3.5}, {3.5*0.43-1.5});
    \draw[blue!60, thick]
      ({0}, {0}, {0*0.43-1.5})
      -- ({3.5}, {3.5}, {3.5*0.43-1.5});





    \draw[very thick, magenta, ->]
      (-2.5, -2.5, 0) -- (6.5, 6.5, 0);

    \node[magenta, font=\small] at (-3.5, -3.5, 0){};

    \fill[magenta!20, opacity=0.75]
      ({0-1.2}, {0+1.2}, {0*0.43})
      -- ({3.5-1.2}, {3.5+1.2}, {3.5*0.43})
      -- ({3.5}, {3.5}, {3.5*0.43+1.5})
      -- ({0}, {0}, {0*0.43+1.5})
      -- cycle;
    \draw[magenta!60, thick]
      ({0-1.2}, {0+1.2}, {0*0.43})
      -- ({0}, {0}, {0*0.43+1.5});
    \draw[magenta!60, thick]
      ({3.5-1.2}, {3.5+1.2}, {3.5*0.43})
      -- ({3.5}, {3.5}, {3.5*0.43+1.5});
    \draw[magenta!60, thick]
      ({0}, {0}, {0*0.43+1.5})
      -- ({3.5}, {3.5}, {3.5*0.43+1.5});

    \fill[magenta!20, opacity=0.75]
      ({0+1.2}, {0-1.2}, {0*0.43})
      -- ({3.5+1.2}, {3.5-1.2}, {3.5*0.43})
      -- ({3.5}, {3.5}, {3.5*0.43+1.5})
      -- ({0}, {0}, {0*0.43+1.5})
      -- cycle;
    \draw[magenta!60, thick]
      ({0+1.2}, {0-1.2}, {0*0.43})
      -- ({0}, {0}, {0*0.43+1.5});
    \draw[magenta!60, thick]
      ({3.5+1.2}, {3.5-1.2}, {3.5*0.43})
      -- ({3.5}, {3.5}, {3.5*0.43+1.5});
    \draw[magenta!60, thick]
      ({0}, {0}, {0*0.43+1.5})
      -- ({3.5}, {3.5}, {3.5*0.43+1.5});

    \draw[thick, magenta!60]
      ({0+1.2}, {0-1.2}, {0*0.43})
      -- ({3.5+1.2}, {3.5-1.2}, {3.5*0.43});
    \draw[thick, magenta!60]
      ({0-1.2}, {0+1.2}, {0*0.43})
      -- ({3.5-1.2}, {3.5+1.2}, {3.5*0.43});

    \draw[thick, ->] (0,0,-2.25) -- (0,0,4)   node[anchor=south]{$\mu_0$};
    \draw[thick, ->] (0,-6.5,0)  -- (0,6.5,0)  node[anchor=south]{$\mu_2$};
    \draw[thick, ->] (-3.5,0,0)  -- (4.5,0,0)  node[anchor=north]{$\mu_1$};
  \end{scope}
\end{scope}

\end{tikzpicture}

\caption{Left: Illustration of $M_{\mu,\infty}$ in the evidence aggregation example when $C_1=C_2=C$. The
highlighted area represents $\{ \mu\in M_{\mu,\infty}:\overline{I}^{(1)}(\mathbf{0};\mu)\geq0\}$. Nature's equilibrium choice of $(\mu_1,\mu_2)$ in the limit game is along the red line whenever $\rho\leq\text{min}\{ \sigma_{1}/{\sigma_{2}},\sigma_{2}/\sigma_{1}\}$. Right: A 3D illustration of the parameter space in the same example.
\label{fig:ik.example}}
\end{figure}
Given the simple structure of $\partial\overline{I}(0,0)$, the unique
solution of the empirical analog of \eqref{eq:quadratic.program}
is $\hat{p}=\left(\hat{e},1-\hat{e}\right)^{\top}$, where
\[
\hat{e}=\left[\frac{\hat{\sigma}_{2}^{2}-\hat{\rho}\hat{\sigma}_{1}\hat{\sigma}_{2}}{\hat{\sigma}_{1}^{2}+\hat{\sigma}_{2}^{2}-2\hat{\rho}\hat{\sigma}_{1}\hat{\sigma}_{2}}\right]_{0}^{1}.
\]
Moreover,
\begin{align*}
\underline{I}^{(1)}(\mathbf{0};\hat{\Sigma}\hat{p})-\overline{I}^{(1)}(\mathbf{0};\hat{\Sigma}\hat{p}) & =\max\left\{ \hat{\Sigma}\hat{p}\right\} -\min\left\{ \hat{\Sigma}\hat{p}\right\} =\begin{cases}
0, & \hat{\rho}\leq\text{min}\left\{ \hat{\sigma}_{1}/\hat{\sigma}_{2},\hat{\sigma}_{2}/\hat{\sigma}_{1}\right\} ,\\
\hat{\rho}\hat{\sigma}_{1}\hat{\sigma}_{2}-\hat{\sigma}_{2}^{2}, & \hat{\rho}>\hat{\sigma}_{2}/\hat{\sigma}_{1},\\
\hat{\rho}\hat{\sigma}_{1}\hat{\sigma}_{2}-\hat{\sigma}_{1}^{2}, & \hat{\rho}>\hat{\sigma}_{1}/\hat{\sigma}_{2}.
\end{cases}
\end{align*}
Theorem \ref{thm:non-diff-asymptotic} can be directly applied. In
particular, if $\hat{\rho}\leq\min\left\{ \hat{\sigma}_{1}\big/\hat{\sigma}_{2},\hat{\sigma}_{2}\big/\hat{\sigma}_{1}\right\} $,
$\hat{e}\in[0,1]$ and information from both countries will always
be used if the inequality is strict. The asymptotically optimal decision
is either a randomized one based on $\hat{p}^{\top}\hat{\mu}$ or
a threshold rule $\mathbf{1}\{\hat{p}^{\top}\hat{\mu}\geq0\}$, depending
on the magnitude of $(\pi/2)\hat{p}^{\top}\hat{\Sigma}\hat{p}$ versus
$\overline{C}^{2}$. This characterization agrees with those in \citet[Section 3.2]{ishihara2021}
and \citet[Example 1]{yata2021} that focused on $\rho=0$.\footnote{\citet{ishihara2021} characterize the efficient weight numerically
for a class of non-randomized threshold rules. \citet{yata2021} characterizes
the optimal decision with techniques analogous to \citet{armstrong2018optimal}.
To be clear, their general results also apply to $\rho\neq0$. } The cases for $\hat{\rho}>\hat{\sigma}_{2}/\hat{\sigma}_{1}$ or
$\hat{\rho}>\hat{\sigma}_{1}/\hat{\sigma}_{2}$ are more involved.
For example, suppose $\hat{\rho}>\hat{\sigma}_{2}/\hat{\sigma}_{1}$.
Then, $\hat{p}=(0,1)^{\top}$. If $\left(\pi/2\right)^{1/2}\hat{\sigma}_{2}<\overline{C}$,
an asymptotic optimal rule is a randomized one based on $\hat{\mu}_{2}$
only. Let $\hat{x}$ solve $\max_{x\in[0,\infty)}(\overline{C}+x)\Phi\left(-x/\hat{\sigma}_{2}\right)$.
If $\left(\pi/2\right)^{1/2}\hat{\sigma}_{2}\geq\overline{C}$ and
$\hat{x}\leq(2\overline{C}\hat{\sigma}_{2})/(\hat{\rho}\hat{\sigma}_{1}-\hat{\sigma}_{2})$,
our optimal decision is simply $\mathbf{1}\left\{ \hat{\mu}_{2}\geq0\right\} $.
Otherwise, case (iii) of Theorem \ref{thm:non-diff} would apply.
The optimal weight of aggregating information from the two countries
must be computed by finding $\hat{\mu}^{o}=\arg\max_{\{\mu\in M_{\mu,\infty}:\min\left\{ \mu_{1},\mu_{2}\right\} \geq0\}}r(\mu;\hat{\Sigma})$,
and the optimal decision $\mathbf{1}\bigl\{\left(\mathbf{\hat{\mu}}^{o}\right)^{\top}\hat{\Sigma}^{-1}\hat{\mu}\geq0\bigr\}$
may use information from both countries again.



\section{Conclusion}\label{sec:conclude}
In this paper, we propose a new asymptotic framework to study treatment assignment problems with partially identified parameters. Our proposal has two novel features: (i) we consider a sequence of drifting parameters such that the extent of partial identification vanishes at the same rate as sampling uncertainty; (ii) we localize the reduced-form parameter at what we call a \emph{global least favorable point}, which we argue is a natural counterpart of the  \emph{hardest case} considered by \cite{HiranoPorter2009} for point-identified settings. Our approach characterizes the limit treatment choice problem in a Gaussian location shift model with a limit local identified set; to show this, we not only establish a rigorous link between finite-sample and limit problems but also generalize earlier analyses of the latter. Similarly, we show how to handle a large class of non-centrosymmetric models. Finally, our approach is computationally inexpensive, providing a simple recipe for applied researchers to find approximately optimal rules for a wide range of problems with partially identified parameters.

Several extensions and open questions remain. First, it may be possible to solve the limit game with directional differentiability of the bounds even with centrosymmetry dropped (e.g., by utilizing Lemma \ref{lem:CoMo}(i)), although we suspect that this will involve more technicalities; we leave it for future research. Second, we also do not consider reduced-form parameters that are semiparametric in nature. It remains to be investigated to what extent our results for parametric models are relevant for their semiparametric counterparts.
Third, it may also be interesting to investigate whether one can combine our approach with recent ones by \citet{BenTal}, \citet{eisenhauer}, or \citet{andrews2025certified}, who take the confidence sets for identified quantities as constraining  the state space and solve the decision problem subject to those restrictions.