EconBase
← Back to paper

Locally Asymptotically Minimax Statistical Treatment Rules Under Partial Identification

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

94,479 characters · 16 sections · 54 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Locally Asymptotically Minimax Statistical Treatment Rules Under Partial Identification

abstractPolicymakers often desire a statistical treatment rule (STR) that determines a treatment assignment rule deployed in a future population from available data. With the true knowledge of the data generating process, the average treatment effect (ATE) is the key quantity characterizing the optimal treatment rule. Unfortunately, the ATE is often not point identified but partially identified. Presuming the partial identification of the ATE, this study conducts a local asymptotic analysis and develops the locally asymptotically minimax (LAM) STR. The analysis does not assume the full differentiability but the directional differentiability of the boundary functions of the identification region of the ATE. Accordingly, the study shows that the LAM STR differs from the plug-in STR. A simulation study also demonstrates that the LAM STR outperforms the plug-in STR.\\ Keywords: Statistical Treatment Rules, Partial Identification, Hadamard Directional Differentiability, Local Asymptotic Minimaxity

Introduction

Consider the problem of choosing a future treatment for the population of interest from two alternatives, say “treat” and “do not treat”. The objective is to maximize the population's mean outcome, such as average income and cure rate from a specific disease. In making the decision, data that is somewhat informative about the performance of each treatment is available. A data-dependent procedure that effectively chooses the future treatment based on the obtained information from the data is, thus, desirable. Manski2004 calls such data-dependent procedures statistical treatment rules (STRs) and proposes to analyze this problem within the framework of statistical decision theory to derive the optimal STR. Since then, studies on STRs have developed under various settings Stoye2009,Hirano2009,Tetenov2012,Kitagawa2018,Mbakop2021,Athey2021. A common presumption in these studies is the point identification of the (conditional) average treatment effect (ATE) in the target population. In other words, if the true data generating process (dgp), the distribution from which the data is drawn, was known, the target population's ATE could be uniquely determined. It would then be optimal to choose the future treatment depending on the sign of the ATE. Although the true dgp is not known in practice, the preceding argument suggests that one of the most intuitive STRs is the empirical success rule that chooses “treat” if and only if the estimated ATE is positive Manski2004. This rule has several desirable properties. Specifically, Stoye2009 shows that this STR is a nearly finite-sample minimax regret STR under several circumstances. Moreover, Hirano2009 shows that it is asymptotically minimax optimal. However, in practice, the ATE is frequently not point identified but partially identified. For example, one might not assume the random assignment of the treatments in an observational study. Even in experimental studies, the experiment participants may not be the random sample from the population of interest. Moreover, non-compliance with assigned treatment may be allowed. In these cases, the ATE of the target population cannot be point identified but partially identified. Namely, even the full knowledge of the true dgp cannot uniquely pin down the value of the ATE. Rather, the knowledge only gives the identification region, the set to which the ATE belongs. The identification region can take various forms depending on the identifying assumption researchers impose on dgps. Still, the set is not generally a singleton. When the ATE is partially identified, it is troubling for two reasons. First, as the identification region depends on the dgp, it must be estimated from the data. Thus, there is the problem of statistical uncertainty. This is the fundamental feature any statistical decision problem has in common. Second, superior treatment might be ambiguous even with the true knowledge of the dgp. Suppose the true identification region contains positive and negative values. That is, “treat” may be strictly better than “do not treat” in the target population, and vice versa. Hence, it is unclear which treatment should be implemented in the population, even with the true knowledge of the identification region. This is the challenge inherent in the treatment decision problem under partial identification of the ATE. Although the partial identification of the ATE adds ambiguity to the treatment decision problem, the framework of statistical decision theory still applies. In particular, presuming the partial identification of the ATE, several studies have derived finite-sample minimax regret STRs by restricting the class of dgps to some analytically tractable classes or focusing on a specific identifying assumption. Specifically, Ishihara2021,Stoye2012,Yata2021 solve the statistical decision problem exactly by restricting the class of dgps to Gaussian experiments, where the data comes from a normal distribution with unknown mean and known covariance matrix. Manski2007 also solves the finite-sample problem exactly when the partial identification is conducted with empirical evidence alone. Unfortunately, it is challenging to solve the statistical treatment decision problem for a general class of dgps induced from arbitrary identification assumptions. Instead, in the same spirit as Hirano2009, this study conducts a local asymptotic analysis to derive asymptotically optimal STRs. A brief discussion of the setup and results are as follows. The setup comprises semi- or non-parametric models. Suppose the lower and upper bounds of the identification region can be written as $\tau_L(\theta)$ and $\tau_U(\theta)$. Here, $\theta$ is the possibly infinite-dimensional intermediate parameter that summarizes some features of the dgp relevant to the region. Denote by $\theta_0$ the value of the intermediate parameter at the true dgp. The boundary functions $\tau_L$ and $\tau_U$ are known real-valued functions from theory. Assume that these functions are Hadamard directionally differentiable, a version of directional differentiability for maps between two normed spaces. This is the key assumption that determines the scope of this study. Do not assume full differentiability because $\tau_L$ and $\tau_U$ are often not fully differentiable. For instance, $\max$ and $\min$ are not fully differentiable but directionally differentiable, although they often appear in the bounds Manski1990,Manski2000a,Balke1997. Allow for nondeterministic treatment rules: “treat” and “do not treat” can be implemented with positive probabilities in the target population. As noted, when the ATE is partially identified, the optimal treatment cannot be uniquely determined even if the true dgp is known. Thus, to address the ambiguity arising from the partial identification of the ATE, use the minimax regret criterion to uniquely determine the optimal treatment rule that should be implemented if the true dgp were known. Manski2007,Manski2009 have shown that the optimal treatment rule is generally nondeterministic. Specifically, if $\theta_0$ were known, then the optimal treatment rule would be

equation[equation omitted — 152 chars of source]

where $f^+(\cdot) = \max\dbrace{f(\cdot),0}$ and $f^-(\cdot) = \max\dbrace{-f(\cdot),0}$ for arbitrary real-valued function $f$. Concretely, “treat” would be implemented with probability $\kappa(\theta_0)$, while “do not treat” would be implemented with probability $1-\kappa(\theta_0)$. In particular, if the identification region contains positive and negative values, this rule is certainly nondeterministic. When the boundary functions, $\tau_L$ and $\tau_U$, are only directionally differentiable, the optimal treatment rule is also only directionally differentiable. The study seeks an asymptotically optimal STR. Formally, an STR is a function that maps data to a treatment rule. The asymptotic optimality criterion investigated here is the local asymptotic minimaxity (LAM). Given a risk function that measures the loss from the repeated use of an STR for a fixed dgp, the local asymptotic maximum risk of the STR is the worst-case risk over neighborhoods of the true dgp. Then, the LAM STR is the STR that minimizes the local asymptotic maximum risk. The main result of this study implies that the LAM STR takes the form of

equation*[equation* omitted — 86 chars of source]

where $\hat{\theta}_n$ is the efficient estimator for $\theta_0$, and $w_{\theta_0}^*$ are the optimal adjustment terms at $\theta_0$. The adjustment term $w_{\theta_0}^*$ is not necessarily zero. Hence, the plug-in STR (i.e., $\kappa(\hat{\theta}_n)$) is not generally LAM. The optimal adjustment term generally depends on $\theta_0$; hence, it is unknown in practice unless $\theta_0$ is known. Instead, this study employs a procedure to determine a data-dependent adjustment term $\hat{w}_n$ and show that the feasible STR $\kappa(\hat{\theta}_n + \hat{w}_n/\sqrt{n})$ is LAM. A simulation study exemplifies the proposed STR relative to the plug-in STR. \paragraph{Related Literature.}{ This study mainly contributes to a line of research that studies the statistical treatment choice problem with partially identified ATE. As noted, Manski2007,Stoye2012,Yata2021,Ishihara2021 have derived exact finite-sample minimax regret STRs by restricting the class of data generating process such that the finite-sample problem becomes analytically tractable. In a similar split of this study, Christensen2022 have also conducted the local asymptotic analysis to derive asymptotically optimal STR under partial identification of the ATE. The major difference from their analysis is the adopted optimality criterion: while they analyze the asymptotic average risk optimality, this study analyzes the asymptotic minimax optimality. Further, they restrict themselves to deterministic treatment rules, while this study allows for nondeterministic treatment rules. Several studies examine individualized treatment assignment problems using observable covariates Adjaho2022,Kido2022,DAdamo2022,Pu2021,Mo2021,Zhao2019,Si2020,Siforthcoming,Kallus2018. However, such studies focus on deterministic treatment rules; this study posits that the treatment assignment in the future population can be probabilistic. Furthermore, the result of this study applies to such individualized treatment assignment problems as long as it considers local asymptotics pointwise in covariates, and there is no constraint on feasible individualized treatment rules. This study is also closely related to studies of efficient estimation of parameters expressed as $f(\theta_0)$ for some known function $f$. When $f$ is Hadamard differentiable, the plug-in estimator $f(\hat{\theta}_n)$ is the efficient estimator Vaart1991b. When $f$ is not Hadamard differentiable but Hadamard directionally differentiable, then regular estimators, which are robust to small perturbations of dgps, do not exist Hirano2012,Vaart1991. Thus, the classical notion of efficiency based on best regularity breaks down for such nondifferentiable functions. However, the efficiency based on LAM still works Song2014a,Fang2016,Ponomarev2022. The relevant assume the (Hadamard) directional differentiability of $f$ and derive the LAM point estimators for $f(\theta)$. In such studies, the LAM point estimators also require some adjustment terms. In particular, Ponomarev2022 shows that the LAM point estimators generally take the form of $f(\hat{\theta}_n + w_{\theta_0}^\dagger/\sqrt{n}) + v_{\theta_0}^\dagger/\sqrt{n}$. Accordingly, it is possible to construct the LAM point estimator for $\kappa(\theta_0)$. However, the optimal adjustment terms $w_{\theta_0}^*$ and $w_{\theta_0}^\dagger$ generally differ because the objective functions differ in the estimation problem and treatment assignment problem; the point estimation problem aims at minimizing the symmetric loss function such as mean squared error, while the treatment assignment problem aims at minimizing the asymmetric loss function such as welfare regret loss. The study will exemplify this difference using an example. The Hadamard directional differentiability has also been adopted in Fang2019,Hong2018 to develop asymptotically valid inference procedures on $f(\theta_0)$. } \paragraph{Structure of the Paper.}{ (ref) formally sets up the statistical treatment decision problem under the partial identification of the ATE. It also introduces the precise requirement for the identification region of the ATE with some motivating examples. (ref) discusses the framework of the local asymptotic analysis employed in this study. (ref) derives the LAM STRs, and (ref) concludes the paper. All proofs are relegated to (ref). }

Setup

Statistical Treatment Choice Problem

Suppose there is a binary treatment encoded as 0 and 1. Let $Y(a) \in \mathcal{Y}$ denote the potential outcome that would be realized if an individual was exposed to treatment $a \in \{0,1\}$. The set $\mathcal{Y}$ is assumed to be a bounded subset of $\mathbb{R}$, and let $y_L \coloneqq \inf \mathcal{Y}$ and $y_U \coloneqq \sup \mathcal{Y}$. Consider a social planner who wants to assign treatment 0 or 1 to each individual of a population of interest. The planner is allowed to implement probabilistic assignments independently of potential outcomes. Specifically, the planner can assign treatment 1 to an individual with probability $\delta \in [0,1]$, which will be called a treatment rule. When the planner adopts a treatment rule $\delta$, the welfare at the population of interest is given by

equation[equation omitted — 205 chars of source]

where the expectation is taken with respect to the target population's joint distribution of $(Y(1),Y(0))$. Suppose that higher welfare is more desirable. Then, if the planner knew the true marginal distributions of potential outcomes, she would search for the treatment rule that maximizes the welfare. The right-hand side of (ref) shows that the important quantity for her optimization is the average treatment effect (ATE), $\operatorname{\mathbb{E}} \dbrack{Y(1) - Y(0)}$. The optimal treatment rules assign treatment 1 with probability one (zero) if and only if the ATE is positive (negative). Therefore, if she knew the signs of ATE, she could obtain an optimal treatment rule. Unfortunately, this rule is generally infeasible, as the ATE is unknown. Suppose instead a sample $\mathbf{Z}_n = (Z_1, \cdots, Z_n)$ is available to the planner. It is assumed that each observation $Z_i$ belongs to a complete separable metric space $\mathcal{Z}$ equipped with the Borel $\sigma$-algebra $\mathcal{B}(\mathcal{Z})$. Further, it is independent and identically distributed (i.i.d.) according to distribution $P_0 \in \mathcal{P}$, where $\mathcal{P}$ is a known class of Borel probability measures.\footnote{The completeness and separability of the space $\mathcal{Z}$ is put to ensure the separability of $L_2(P)$ space for every $P \in \mathcal{P}$. Indeed, every Borel probability measure $P$ is Radon Bogachev2007. This then implies that every $P$ is separable Bogachev2007. The separability of $L_2(P)$ follows from the separability of $P$ \Parencite[][Exercise 4.7.63]{Bogachev2007}.} Given this i.i.d. assumption, the sample is generated following $\mathbf{Z}_n \sim P_0^n$, which is the $n$-fold product of $P_0$. It is common to often refer each $P \in \mathcal{P}$ to a data generating process (dgp) and also refer $P_0$ to the true dgp. If an infinite number of observations of $Z_i$ were available, the dgp $P_0$ would be learnable with arbitrary precision. This study assumes that the ATE cannot be point identified but partially identified. Formally, even with knowledge of the true dgp $P_0$, it is only possible to conclude that the ATE belongs to the identification region $\tau(P_0)$, which is a subset of $[y_L - y_U,y_U-y_L]$ by the boundedness of $\mathcal{Y}$. Note that $\tau(P_0)$ may or may not contain positive and negative values, in which case the better treatment cannot be determined. The study places some restrictions on $\tau(P_0)$, which will be discussed in (ref). Given the availability of a sample $\mathbf{Z}_n$, it is natural that the planner decides a treatment rule, utilizing the knowledge obtained from $\mathbf{Z}_n$. From this perspective, what the planner needs is the procedure that outputs a good treatment rule for each possible realization of $\mathbf{Z}_n$. The planner's problem can then be interpreted as one of the statistical decision problems Manski2004. A statistical decision problem mainly comprises two components: a statistical decision rule and a risk. In the current context, a statistical decision rule is a mapping $\delta_n:\mathcal{Z}^n \to [0,1]$, which is called a statistical treatment rule (STR). Here, $\delta_n(\mathbf{z}_n) \in [0,1]$ is interpreted as the probability of assigning treatment 1 in the target population when the observed sample is $\mathbf{z}_n$. For the risk function, it is important to first specify a loss function that penalizes an arbitrary treatment rule $\delta$ when ATE is $\tau$. Following the literature, this study focuses on the welfare regret loss given by

align*[align* omitted — 324 chars of source]

This measures the loss the planner incurs by not knowing $\tau$. Given this loss function, for every $P \in \mathcal{P}$ and $\tau \in \tau(P)$, the risk of a STR $\delta_n$ is defined by $\operatorname{\mathbb{E}}_{P^n} \dbrack*{ L \dparen*{ \delta_n \dparen*{\mathbf{Z}_n}, \tau } }$. The definition clearly shows that the risk depends on the true dgp $P_0$ and ATE $\tau$. As neither $P_0$ nor $\tau$ are known with a sample of finite size, the planner would seek the STR that minimizes the risk for every possible $P \in \mathcal{P}$ and $\tau \in \tau(P)$. However, such STRs do not necessarily exist. To alleviate this difficulty, the literature on statistical decision theory has considered several criteria. Among those criteria, this study employs the minimax regret criterion, as suggested by Manski2004. Specifically, the minimax regret STR minimizes

equation[equation omitted — 223 chars of source]

From here, suppress the dependence of an STR $\delta_n(\mathbf{Z}_n)$ on the data $\mathbf{Z}_n$ to simplify the notation.

Identification Region of ATE

Among the features of the identification region $\tau(P)$, the important quantities in the current context are its lower and upper bounds. Indeed, when one fixes $P \in \mathcal{P}$ and consider the inner maximization problem in (ref), the direct calculation yields

align*[align* omitted — 289 chars of source]

where for any real value $x \in \mathbb{R}$, $(x)^+ = \max\{x,0\}$, and $(x)^- = \max\{-x,0\}$. Note that the boundedness of the outcome space implies that the bounds of $\tau(P)$ are bounded by a constant uniformly in $P$. The equation reveals that whatever the form of $\tau(P)$ is, the value of the inner maximization in (ref) depends only on $\inf\tau(P)$ and $\sup\tau(P)$. This study requires that the bounds, $\inf\tau(P)$ and $\sup\tau(P)$, are sufficiently smooth in $P$.

Motivating Examples

To motivate the concrete smoothness conditions, the study first presents at some examples of partial identification of ATE.

exampleImagine an observational study in which treatments are not randomly assigned. Formally, suppose a researcher draws $n$ i.i.d. copies of $(Y(1),Y(0),D)$ from a population of interest and observes $Z_i = (Y_i,D_i),~i=1,\cdots,n$. Here, $D_i \in \{0,1\}$ denotes the treatment $i$-th subject received, and $Y_i$ denotes the $i$-th subject's observed outcome satisfying $Y_i = Y_i(D_i)$. Let $p = P(D = 1)$ and $\mu_d = \operatorname{\mathbb{E}}_P[Y|D=d]$ for $d \in \{0,1\}$. With this empirical evidence alone, the ATE is not point identified but partially identified. Specifically, Manski1990 gives the sharp bounds as \begin{align} \inf\tau(P) = (\mu_1-y_U)p + (y_L - \mu_0)(1-p), \quad \sup\tau(P) = (\mu_1-y_L)p + (y_U - \mu_0)(1-p). \end{align} The width of $\tau(P)$ is always one irrespective of $P$. This implies that $\tau(P)$ always contains zero, whence the sign of ATE is never point identified. \qed
exampleImagine an experimental study in which an experimenter recruits participants non-randomly. More concretely, suppose a researcher has access to $n$ i.i.d. copies of $(Y(1),Y(0),D,S)$ from a population of interest and observes $Z_i = (S_iY_i,S_iD_i,S_i)$, where $Y_i$ and $D_i$ are defined in the same way as (ref), and $S_i \in \{0,1\}$ denotes the $i$-th subject's participation in the experiment ($S_i = 1$ if and only if subject $i$ participates in the experiment). Thus, for subjects with $S_i = 0$, the experimenter only observes that these subjects do not participate. However, among subjects with $S_i = 1$, the experimenter randomly assigns treatments with known probability, where one can assume $(Y(1),Y(0)) \perp D |S=1$. This setting has often been studied in the literature on external validity Hotz2005,Cole2010. Let $p = P(S_i = 1)$ and $\mu_d = \operatorname{\mathbb{E}}_P\dbrack{Y_i | D_i = d, S_i = 1}$ for $d \in \{0,1\}$. Without any further assumptions such as $(Y(1),Y(0))\perp S$, the ATE is partially identified. In particular, Manski1989's result implies \begin{align} \inf\tau(P) = (\mu_1 - \mu_0)p - (y_U - y_L)(1-p), \quad \sup\tau(P) = (\mu_1 - \mu_0)p + (y_U - y_L)(1-p). \end{align} The width of $\tau(P)$ is $2(1-p)$, which depends on the dgp $P$. Moreover, the set $\tau(P)$ may or may not contain zero depending on $P$. \qed
exampleIn (ref), suppose the researcher has access to a binary instrument $V \in \{v_0,v_1\}$. More precisely, suppose the researcher draws $n$ i.i.d. copies of $(Y(1),Y(0),D,V)$ from a population of interest and observes $Z_i = (Y_i,D_i,V_i),~i=1,\cdots,n$, where $Y_i$ and $D_i$ are defined in the same way as (ref). Here, suppose the instrument satisfies the exclusion restriction, which rules out the possibility of the direct effect of the instrument on outcomes. Assume the instrument $V$ is mean independent of potential outcomes, that is, \begin{equation*} \operatorname{\mathbb{E}}[Y(d)|V=v_j] = \operatorname{\mathbb{E}}[Y(d)]\quadfor every\quad (d,j) \in \{0,1\} \times \{0,1\}. \end{equation*} Let $p_j = P(D = 1|V = v_j)$ for $j \in \{0,1\}$ and $\mu_{d,j} = \operatorname{\mathbb{E}}_P[Y|D = d, V = v_j]$ for each $(d,j)$. Then, Manski1990's result gives the sharp bounds as \begin{align} \begin{aligned} \inf\tau(P) &= \max\dbrace*{ \begin{aligned} & \mu_{1,1}p_1 + y_L(1-p_1) - \mu_{0,1}(1-p_1) - y_Up_1\\ & \mu_{1,1}p_1 + y_L(1-p_1) - \mu_{0,0}(1-p_0) - y_Up_0\\ & \mu_{1,0}p_0 + y_L(1-p_0) - \mu_{0,1}(1-p_1) - y_Up_1\\ & \mu_{1,0}p_0 + y_L(1-p_0) - \mu_{0,0}(1-p_0) - y_Up_0 \end{aligned} },\\ \sup\tau(P) &= \min\dbrace*{ \begin{aligned} & \mu_{1,1}p_1 + y_U(1-p_1) - \mu_{0,1}(1-p_1) - y_Lp_1\\ & \mu_{1,1}p_1 + y_U(1-p_1) - \mu_{0,0}(1-p_0) - y_Lp_0\\ & \mu_{1,0}p_0 + y_U(1-p_0) - \mu_{0,1}(1-p_1) - y_Lp_1\\ & \mu_{1,0}p_0 + y_U(1-p_0) - \mu_{0,0}(1-p_0) - y_Lp_0 \end{aligned} }. \end{aligned} \end{align} In the binary outcome setting (i.e., $\mathcal{Y} = \{0,1\}$), Balke1997 showed that the bounds given in (ref) can be tightened with the assumption, $(Y(1),Y(0)) \perp V$, which is stronger than the mean-independence. Although the exact form of their bounds is not shown to save space, note that they still contain $\max$ and $\min$ functions. The set $\tau(P)$ may or may not contain zero depending on $P$. \qed

Directional Differentiability of the Bounds

The inspection of (ref) reveals that the lower and upper bounds of $\tau(P)$ can often be written as a composition of maps. The relevant features of the data-generating process $P$ are first summarized by intermediate parameters $\theta$. Then, $\theta$ is mapped by certain functions, say $\tau_L$ and $\tau_U$, to the lower and upper bounds of $\tau(P)$. For the smoothness of $\tau_L$ and $\tau_U$, we must deal with the possible nondifferentiability of these functions. For instance, max and min functions, which appear in (ref), are not differentiable at some points. Rather, the max and min functions are directionally differentiable. Thus, to cover such instances, it is adequate to assume the directional differentiability of $\tau_L$ and $\tau_U$ in some strength. In the following, this study formalizes these features as the restrictions on the identification region. Suppose there exist $\theta:\mathcal{P} \to \Theta \subset \mathbb{B}$ for a Banach space $\mathbb{B}$ with norm $\normB{\cdot}$ and known functions $\tau_j:\Theta \to \mathbb{R}$ for $j \in \{L,U\}$ such that

equation[equation omitted — 145 chars of source]

One can take $\mathbb{B} = \mathbb{R}^k$ for an integer $k$ in most applications, but the formulation allows infinite-dimensional parameters, such as $\mathbb{B} = L_2(P)$. The intermediate parameter $\theta(P)$ is required to be pathwise differentiable, which will be discussed in (ref). On the other hand, the functions $\tau_L(\theta)$ and $\tau_U(\theta)$ are assumed to be directionally differentiable in a certain strength, which is clarified in (ref) below. In that statement, let $P_0 \in \mathcal{P}$ be the true data generating process and let $\theta_0 \coloneqq \theta(P_0)$ be the value of the intermediate parameter under true dgp.

assumptionThe functions $\tau_L:\Theta \to \mathbb{R}$ and $\tau_U:\Theta \to \mathbb{R}$ are continuous in $\theta$ and satisfies $\tau_L(\theta) \leq \tau_U(\theta)$ for all $\theta \in \Theta$ and $\tau_L(\theta_0) < \tau_U(\theta_0)$ at $\theta_0 \in \Theta$. Moreover, the following properties hold. \begin{assumptionenum} • For each $j \in \{L,U\}$, the function $\tau_j:\Theta \to \mathbb{R}$ is Hadamard directionally differentiable at $\theta_0 \in \Theta$. That is, there is a continuous map $\tau_{j,\theta_0}':\mathbb{B} \to \mathbb{R}$ such that \begin{equation*} \lim_{n\to\infty} \abs*{ \frac{\tau_j(\theta_0 + t_nb_n) - \tau_j(\theta_0)}{t_n} - \tau_{j,\theta_0}'(b) } = 0 \end{equation*} for all sequences $\{b_n\} \subset \mathbb{B}$ and $\{t_n\} \subset \mathbb{R}_+$ such that $t_n \downarrow 0$, $b_n \to b \in \mathbb{B}$ as $n \to \infty$ and $\theta_0 + t_nb_n \in \Theta$ for all $n$. • For each $j \in \{L,U\}$, the Hadamard directional derivative $\tau_{j,\theta_0}':\mathbb{B} \to \mathbb{R}$ is Lipschitz continuous; that is, there exists $C_{\tau_{j,\theta_0}'} > 0$ such that \begin{equation*} \abs*{\tau_{j,\theta_0}'(b) - \tau_{j,\theta_0}'(\tilde{b})} \leq C_{\tau_{j,\theta_0}'}\normB*{b - \tilde{b}} \end{equation*} for all $b,\tilde{b} \in \mathbb{B}$. • For every $v \in \mathbb{R}$, there exists $\tilde{v} \in \mathbb{B}$ such that $\tau_{j,\theta_0}'(b) + v = \tau_{j,\theta_0}'(b + \tilde{v})$ for all $b \in \mathbb{B}$ and $j \in \{L,U\}$. \end{assumptionenum}

(ref) gives the precise strength of directional differentiability needed in this paper. It is important to note that the Hadamard directional differentiability does not demand the linearity of the derivative $\tau_{j,\theta_0}':\mathbb{B} \to \mathbb{R}$. When it is linear, the function $\tau_{j}:\mathbb{B} \to \mathbb{R}$ is said to be Hadamard differentiable at $\theta_0 \in \Theta$ Fang2019. Conventionally, Hadamard differentiability has often been utilized to derive the asymptotic distribution of statistics by applying the delta method. Concretely, if $\theta_n$ is an estimator for $\theta_0$ such that $\sqrt{n}(\theta_n - \theta_0) \leadsto \mathbb{G}$ and $\tau_j$ is Hadamard differentiable at $\theta_0$, then

equation[equation omitted — 142 chars of source]

where $\leadsto$ indicates the weak convergence. Unfortunately, the Hadamard differentiability is too strong to cover max and min functions that often appear in the bounds of identification region as in (ref). Instead, this study only assumes the directional differentiability. Among the several notions of directional differentiability Shapiro1990, the Hadamard directional differentiability is the most suitable for this study. Especially, as emphasized in Fang2019, this directional differentiability allows for the direct generalization of the standard delta method: as long as $\tau_{j}(\theta)$ is Hadamard directionally differentiable at $\theta_0$, the weak convergence in (ref) continues to hold. Based on this generalized delta-method, Fang2016,Ponomarev2022 developed the locally asymptotically minimax estimator for $\tau_j(\theta_0)$ itself. This study utilizes the generalized delta method to develop the LAM STR. In addition to the generalization of the delta method, the Hadamard directional differentiability is also useful, as it admits the chain rule Shapiro1990. (ref) requires that the Hadamard directional derivatives are Lipschitz continuous. (ref) is imposed to simplify the analysis. When $\mathbb{B} = \mathbb{R}^k$, (ref) is satisfied if $\tau_{j,\theta_0}'$ is translation equivariant. In this case, $\tau_{j,\theta_0}'(b + v\mathbf{1}) = \tau_{j,\theta_0}'(b) + v$ for every $b$ and $v$, where $\mathbf{1} = (1,\cdots,1) \in \mathbb{R}^k$. Particularly, $\max$ and $\min$ functions, which often appear in the bounds of partial identification, satisfy the translation equivariance.

examplecontinued{ex:identification-region-with-empirical-evidence-alone} In this example, let $\theta(P) = (\mu_1,\mu_0,p)^\intercal \in \Theta = [y_L,y_U]^2\times[0,1] \subset \mathbb{R}^3$, and let $\tau_L(\theta)$ and $\tau_U(\theta)$ be specified in the right-hand sides of (ref). As long as $\theta_0$ is an interior point of $\Theta$, (ref) is satisfied. In particular, $\tau_L(\theta)$ and $\tau_U(\theta)$ are Hadamard differentiable, and the corresponding derivative can be calculated from the usual partial derivatives. For instance, \begin{equation*} \tau_{L,\theta_0}'(b) = \dparen*{\frac{\partial \tau_L}{\partial \theta}\Bigr|_{\theta = \theta_0}}^\intercal b \end{equation*} for $b \in \mathbb{R}^3$.
examplecontinued{ex:experimental-study-with-non-random-sampling} Let $\theta(P) = (\mu_1,\mu_0,p)^\intercal \in \Theta = [y_L,y_U]^2\times[0,1] \subset \mathbb{R}^3$, and let $\tau_L(\theta)$ and $\tau_U(\theta)$ be given in the right-hand side of (ref). Again, (ref) is satisfied as long as $\theta_0$ is an interior point. The functions $\tau_L(\theta)$ and $\tau_U(\theta)$ are Hadamard differentiable with the derivatives calculated from the partial derivative as in (ref).
examplecontinued{ex:identification-region-with-mean-independent-IV} Let $\theta(P) = (\mu_{1,1},\mu_{1,0},\mu_{0,1},\mu_{0,0},p_1,p_0)^\intercal \in \Theta = [y_L,y_U]^4\times[0,1]^2 \subset \mathbb{R}^6$. Moreover, specify $\tau_L(\theta)$ and $\tau_U(\theta)$ by the right-hand sides of (ref). Then, (ref) holds as long as $\theta_0$ is in the interior of $\Theta$. To see the Hadamard directionally differentiability of $\tau_L(\theta)$, it is convenient to express the function as the composition of the max function and \begin{equation*} \psi_L(\theta) = (\psi_{L,1}(\theta),\cdots,\psi_{L,4}(\theta))^\intercal = \begin{pmatrix} \mu_{1,1}p_1 + y_L(1-p_1) - \mu_{0,1}(1-p_1) - y_Up_1\\ \mu_{1,1}p_1 + y_L(1-p_1) - \mu_{0,0}(1-p_0) - y_Up_0\\ \mu_{1,0}p_0 + y_L(1-p_0) - \mu_{0,1}(1-p_1) - y_Up_1\\ \mu_{1,0}p_0 + y_L(1-p_0) - \mu_{0,0}(1-p_0) - y_Up_0 \end{pmatrix} \in\mathbb{R}^4. \end{equation*} As in (ref), one can easily observe that $\psi_L(\theta)$ is Hadamard differentiable with the derivative given by $\psi_{L,\theta_0}'(b) = \dparen*{\frac{\partial \psi_L}{\partial \theta^\intercal }|_{\theta = \theta_0}}^\intercal b$ for $b = (b_1,\cdots,b_6) \in \mathbb{R}^6$. Moreover, the $\max$ function is Hadamard directionally differentiable; the directional derivative of the $\max$ function at $\psi_L(\theta_0)$ is given by \begin{equation*} \max_{j \in B_L(\theta_0)} \psi_j \quadfor every\quad (\psi_1,\dots,\psi_4) \in \mathbb{R}^4, \end{equation*} where $B_L(\theta_0) = \operatorname*{arg\,max}_{j = 1,\cdots,4} \psi_{L,j}(\theta_0)$. By the chain rule for the Hadamard directionally differentiable maps, $\tau_L(\theta)$ is Hadamard directionally differentiable with derivative \begin{equation*} \tau_{L,\theta_0}'(b) = \max_{j \in B_L(\theta_0)} \psi_{L,\theta_0}'(b) \end{equation*} Similarly, $\tau_U(\theta)$ is also Hadamard directionally differentiable.

Optimal Treatment Rule with Knowledge of True DGP

If the social planner knew the true dgp $P_0$, there would be no need to consider statistical uncertainty, and hence, they would solve $\inf_{\delta \in [0,1]}\sup_{\tau \in \tau(P_0)} L(\delta,\tau)$. As has been observed in Manski2007,Manski2009, the minimax regret value is

equation[equation omitted — 240 chars of source]

and the optimal treatment rule is given by (ref). Depending on the true intermediate parameter $\theta_0$, the future assignment given the optimal treatment rule may (not) be deterministic, and the corresponding minimax value may (not) be greater than zero. Following Manski2009, this study introduces a concept of error because of selecting an inferior treatment to interpret the regret value and rule. Specifically, refer to choosing treatment 0 when the ATE $\tau$ is positive as a Type 0 error. Conversely, refer to choosing treatment 1 when $\tau$ is negative as a Type 1 error. Further, recall that the welfare regret loss is given by $\tau(1-\delta)\cdot \operatorname{1}\dbrace{\tau > 0} - \tau\delta\cdot \operatorname{1}\dbrace{\tau \leq 0}$. The regret loss measures the loss from Type 0 or 1 errors: $\tau(1-\delta)$ corresponds to the loss from the Type 0 error, while $-\tau\delta$ corresponds to the loss given the Type 1 error. First, if $\tau_L(\theta_0) < \tau_U(\theta_0) \leq 0$, then $\kappa(\theta_0) = 0$, and hence, treatment 1 is assigned with probability 0. In this case, the ATE is certainly non-positive, and what matters is the Type 1 error. From the knowledge of the identification region $\tau(P_0)$, treatment 0 is never worse than treatment 1. Accordingly, the superior treatment can be chosen for sure, and the loss from the Type 1 error can be made zero. Similarly, if $0 \leq \tau_L(\theta_0) < \tau_U(\theta_0)$, then $\kappa(\theta_0) = 1$, and the minimax regret is again zero. This case is also interpreted similarly. Finally, if $\tau_L(\theta_0) < 0 < \tau_U(\theta_0)$, $\kappa(\theta_0) = \tau_U(\theta_0)/\dparen{\tau_U(\theta_0) - \tau_L(\theta_0)} \in (0,1)$, and the minimax regret value is greater than zero. In this case, positive and negative ATEs are possible, and hence, Type 0 and Type 1 errors matter. In particular, the worst-case loss from Type 0 error is $\tau_U(\theta_0)(1-\delta)$, and the worst-case loss from Type 1 error is $-\tau_L(\theta_0)\delta$. As $\delta$ increases from 0 to 1, the former loss decreases from $\tau_U(\theta_0)$ to $0$, while the latter loss increases from $0$ to $-\tau_L(\theta_0)$. As the welfare regret loss equally weighs the losses from both errors, the optimal treatment rule balances the worst-case losses from Type 0 and 1 errors; that is, $\tau_U(\theta_0)(1-\kappa(\theta_0)) = -\tau_L(\theta_0)\kappa(\theta_0)$. Under (ref), the optimal treatment rule $\kappa(\theta)$ inherits the properties of $\tau_L$ and $\tau_U$.

lemmaIf (ref) is in force, the following hold. \begin{lemmaenum} • The optimal treatment rule $\kappa: \Theta \to \mathbb{R}$ is Hadamard directionally differentiable at $\theta_0$ with the derivative given by \begin{equation} \kappa_{\theta_0}'(b) \coloneqq \frac{\tau_L^-(\theta_0)\cdot m_{\tau_U(\theta_0)}'\dparen*{\tau_{U,\theta_0}'(b)} - \tau_U^+(\theta_0)\cdot m_{-\tau_L(\theta_0)}'\dparen*{-\tau_{L,\theta_0}'(b)}}{\dparen*{\tau_U^+(\theta_0) + \tau_L^-(\theta_0)}^2} \end{equation} for all $b \in \mathbb{B}$. Here, $m_y':\mathbb{R}\to\mathbb{R}$ given by \begin{equation} m_{y}'(x) \coloneqq x \cdot \operatorname{1}\dbrace{y > 0} + \max\dbrace{x,0}\cdot \operatorname{1}\dbrace{y = 0} + 0 \cdot \operatorname{1}\dbrace{y < 0} \end{equation} denotes the Hadamard directional derivative of $x \mapsto \max\{x,0\}$ at $y \in \mathbb{R}$. • The Hadamard directional derivative $\kappa_{\theta_0}':\mathbb{B} \to \mathbb{R}$ is Lipschitz continuous. • For every $v \in \mathbb{R}$, there exists $\tilde{v} \in \mathbb{B}$ such that $\kappa_{\theta_0}'(b + \tilde{v}) = \kappa_{\theta_0}'(b) + v$ for all $b \in \mathbb{B}$. \end{lemmaenum}

Importantly, $\kappa(\theta)$ is also Hadamard directionally differentiable at $\theta_0$. The Hadamard directional derivative is derived by applying the chain rule for Hadamard directionally differentiable maps. Note that $\kappa(\theta)$ may not be differentiable at $\theta_0$ even if $\tau_L(\theta)$ and $\tau_U(\theta)$ are Hadamard differentiable at $\theta_0$. For instance, even if $\tau_U(\theta)$ is Hadamard differentiable at $\theta_0$, $\max\{\tau_U(\theta_0),0\}$ is not Hadamard differentiable when $\tau_U(\theta_0) = 0$; so is $\kappa(\theta_0)$. See also (ref)

examplecontinued{ex:identification-region-with-empirical-evidence-alone} This example always have $\tau_L(\theta) \leq 0$ and $\tau_U(\theta) \geq 0$ for all $\theta \in \Theta$. Hence, the optimal treatment rule can be simplified as $\kappa(\theta_0) = \tau_U(\theta_0)/(y_U - y_L)$. As is clear from the simplified expression, $\kappa(\theta_0)$ is Hadamard differentiable with derivative $\kappa_{\theta_0}'(b) = \tau_{U,\theta_0}'(b)/(y_U - y_L)$.
examplecontinued{ex:experimental-study-with-non-random-sampling} We separate the arguments into cases. If $\tau_L(\theta_0) < 0 < \tau_U(\theta_0)$, $\tau_U(\theta_0)$ and $\tau_L(\theta_0)$ are Hadamard differentiable at $\theta_0$, and so are $\tau_U^+(\theta_0)$ and $\tau_L^-(\theta_0)$. Hence, $\kappa(\theta_0)$ is also Hadamard differentiable, and $\kappa_{\theta_0}'(b)$ can be obtained from (ref). If $\tau_L(\theta_0) = 0 < \tau_U(\theta_0)$, $\tau_L^-(\theta_0)$ is only Hadamard directionally differentiable at $\theta_0$, and so is $\kappa(\theta_0)$. In particular, (ref) imply \begin{equation*} \kappa_{\theta_0}'(b) = \frac{1}{\tau_U(\theta_0)}\min\dbrace*{ \tau_{L,\theta_0}'(b), 0 }. \end{equation*} Finally, if $0 < \tau_L(\theta_0) < \tau_U(\theta_0)$, $\kappa_{\theta_0}'(b) = 0$. For other cases, it is possible to calculate $\kappa_{\theta_0}'(b)$ similarly.

Framework of Local Asymptotic Analysis

It is often challenging to exactly solve the finite-sample problem (ref) for general data-generating processes. Instead, in the same spirit of Hirano2009, this study conducts the local asymptotic analysis and then derives asymptotically minimax STRs. This section discusses the framework of the local asymptotic analysis. Specifically, (ref) setups local statistical experiments that are tailored to i.i.d. observations. (ref) gives the precise smoothness required for the intermediate parameter. Finally, (ref) formally defines the notion of asymptotic minimax optimality employed in this study.

Local Statistical Experiments

The study considers a sequence of local statistical experiments in the following asymptotic analysis. Although the original statistical experiment involves all of the possible dgps in $\mathcal{P}$, the fixed true dgp $P_0$ is learnable with arbitrary precision as the sample size approaches infinity. Thus, from the perspective of asymptotic analysis, the asymptotic analysis, considering dgps that are too distant from $P_0$ may be meaningless. Instead, the local asymptotic analysis focuses only on dgps that are statistically difficult to distinguish from the true dgp $P_0$ to approximate the finite-sample situation well. The next concrete construction of the sequence of local experiments closely follows Hirano2009,Vaart1991a. Given the fixed true dgp $P_0 \in \mathcal{P}$, let $\mathcal{P}(P_0)$ be a collection of paths $(0, \epsilon) \ni t \mapsto P_{t} \in \mathcal{P}$ such that

equation[equation omitted — 162 chars of source]

for some measurable function $h: \mathcal{Z} \to \mathbb{R}$, which is often called the score function. It is possible to view each path $t \mapsto P_t$ as a one-dimensional parametric submodel $\dbrace{P_t : t \in (0,\epsilon)} \subset \mathcal{P}$, in which $t$ is treated as the only parameter. Gathering all score functions associated with paths in $\mathcal{P}(P_0)$ induces the tangent set $T(P_0)$, which is the collection of score functions. It is known that $\int h dP_0 = 0$ and $\int h^2 dP_0 < \infty$ for any score function satisfying (ref). Thus, regarding $P_0$-almost surely equal score functions as an equivalence class, $T(P_0)$ can be viewed as a subset of $L_2(P_0)$, which is the separable Hilbert space with the inner product $\dangle{h,g} \coloneqq \int hg dP_0$ and norm $\norm{h}_{2,P_0} \coloneqq \sqrt{\dangle{h,h}}$. To simplify the exposition, this study posits that the tangent set is closed with respect to addition and scalar multiplication.\footnote{The result of this paper can be extended to the case where the tangent set is a convex cone with some appropriate modifications of proofs. See, e.g., Vaart1989,Ponomarev2022.}

assumptionThe tangent set $T(P_0)$ is a linear subspace of the separable Hilbert space $L_2(P_0)$.

We introduce some notations to be used hereafter. For each path $t\mapsto P_t$ with score function $h \in T(P_0)$, let $P_{1/\sqrt{n},h} \in \mathcal{P}$ denote the value of the path evaluated at $t = 1/\sqrt{n}$. Moreover, let $\mathbb{P}_{n,h} = P_{1/\sqrt{n},h}^n$ and let $\mathbb{P}_{n,0} = P_0^n$. Correspondingly, the expectation with respect to $\mathbb{P}_{n,h}$ is denoted by $\operatorname{\mathbb{E}}_{n,h}[\cdot]$. Further, for every path with score function $h$, write $\theta_n(h) = \theta(P_{1/\sqrt{n},h})$. An important consequence of the property (ref) is the local asymptotic normality of the statistical experiments $\mathcal{E}_n = (\mathcal{Z}^n,\mathcal{B}(\mathcal{Z}^n),\mathbb{P}_{n,h}:h \in T(P_0))$. Specifically, for each $h \in T(P_0)$, the log-likelihood ratio admits a linear expansion such that

equation[equation omitted — 188 chars of source]

where $\Delta_{n,h} = \sum_{i=1}^n h(Z_i)/\sqrt{n}$ converges weakly to $N(0,\norm{h}_{2,P_0}^2)$ under $\mathbb{P}_{n,0}$. This local asymptotic normality further implies the convergence of experiments $\mathcal{E}_n$ to a certain limit experiment $\mathcal{E}$. Under (ref), a useful characterization of $\mathcal{E}$ is available. (ref) implies that $\overline{T(P_0)}$ itself is one of separable Hilbert spaces. Then, there exists a complete orthonormal basis $\{h_1,h_2,\cdots\} \subset T(P_0)$, and every $h \in T(P_0)$ can be written as $h = \sum_{j=1}^\infty \dangle{h,h_j}h_j$. In this case, the following convergence of experiments holds:

equation*[equation* omitted — 166 chars of source]

The latter experiment corresponds to observe a single draw $\Xi = (\Xi_1,\Xi_2,\cdots) \in \mathbb{R}^\infty$ such that $\Xi_1,\Xi_2,\cdots$ are independent and $\Xi_j \sim N(\dangle{h,h_j},1)$ for each $j \in \mathbb{N}$. For ease of notation, write $\mathbb{P}_h$ for $N_\infty((\dangle{h,h_j}),\mathbf{1})$ for every $h \in T(P_0)$. Accordingly, the expectation with respect to $\mathbb{P}_h$ is denoted by $\operatorname{\mathbb{E}}_{h}[\cdot]$. As one expects, this limit experiment $\mathcal{E}$ is more tractable relative to the original experiment $\mathcal{E}_n$. Finally, the convergence of experiments culminates in the following asymptotic representation theorem.

lemma[{Vaart1991a}] Suppose a sequence of experiments $\mathcal{E}_n$ converges to a dominated experiment $\mathcal{E}$. Let $T_n:\mathcal{Z}^n \to \mathbb{D}$ be a statistic with values in a metric space $\mathbb{D}$. Suppose the sequence $T_n \overset{h}{\leadsto} Q_h$ for every $h \in T(P_0)$ and that there exists a complete separable subset $\mathbb{D}_0 \subset \mathbb{D}$ such that $Q_h(\mathbb{D}_0) = 1$ for every $h$. Then there exists a randomized statistic $T$ in $\mathcal{E}$ such that $T_n \overset{h}{\leadsto} T$ for every $h$.

In the statement of (ref), the randomized statistic $T$ refers to a measurable map $T:\mathbb{R}^\infty \times [0,1] \ni (\Xi,U) \mapsto T(\Xi,U) \in \mathbb{D}$, where $\Xi \sim \mathbb{P}_h$ and $U \sim \mathrm{Unif}([0,1])$ independently of $\Xi$. Moreover, $\overset{h}{\leadsto}$ means the weak convergence under $\mathbb{P}_{n,h}$ for $h \in T(P_0)$. (ref) allows for guessing a statistic with a certain asymptotic property by the following procedures; (i) focus on the limit experiment $\mathcal{E}$; (ii) obtain the statistic $T$ with the desired property in $\mathcal{E}$; (iii) construct a sequence of statistics $T_n$ that converges weakly to $T$.

Regularity of Intermediate Parameter

For the intermediate parameter $\theta:\mathcal{P} \to \mathbb{B}$ introduced in (ref), assume its pathwise differentiability.

assumptionThe parameter $\theta:\mathcal{P} \to \mathbb{B}$ is pathwise differentiable at $P_0$ relative to the tangent set $T(P_0)$. That is, there exists a continuous linear map $\Dot{\theta}_0:T(P_0) \to \mathbb{B}$ such that \begin{equation*} \frac{\theta(P_{t}) - \theta(P_0)}{t} \to \Dot{\theta}_0(h) \quadas\quad t \downarrow 0 \end{equation*} for each path $t \mapsto P_t$ in $\mathcal{P}(P_0)$ with score function $h \in T(P_0)$.

The pathwise differentiability especially implies $\sqrt{n}\dparen{\theta_n(h) - \theta_0} \to \Dot{\theta}_0(h)$ as $n \to \infty$. The pathwise derivative $\Dot{\theta}_0(h)$ must be linear maps. When $\mathbb{B}$ is the Euclidean space, the Riesz representation theorem gives a useful characterization of the pathwise derivative. Specifically, as $\overline{T(P_0)}$ is a Hilbert space, the theorem implies that there exists $\tilde{\theta}_0 = (\tilde{\theta}_{0,1},\cdots,\tilde{\theta}_{0,k})^\top \in \overline{T(P_0)}^k$ such that $\Dot{\theta}_{0,j}(h) = \dangle{\tilde{\theta}_{0,j},h}$ for all $h \in \overline{T(P_0)}$, where $\Dot{\theta}_{0,j}$ denotes the $j$-th element of $\Dot{\theta}_0$. The function $\tilde{\theta}_0$, thus constructed, is often called efficient influence function. More generally, if $\mathbb{B}$ is an arbitrary Banach space, an adjoint map can characterize the pathwise derivative. In particular, there exists $\Dot{\theta}_0^*:\mathbb{B}^* \to \overline{T(P_0)}$ determined by $\dangle{\Dot{\theta}_0^*b^*, h} = b^*(\Dot{\theta}_0(h))$ for every $h \in T(P_0)$ and $b^* \in \mathbb{B}^*$. The pathwise differentiability assumption is often employed in the study of efficient estimators Vaart1998,Vaart1996. Among the results obtained in the study, one of the most important is the convolution theorem that characterizes the limit law of regular estimators. An estimator $\theta_n:\mathcal{Z}^n \to \Theta \subset \mathbb{B}$ is said to be regular at $P_0$ for estimating $\theta_0$ if

equation*[equation* omitted — 126 chars of source]

where $G$ is a tight random element in $\mathbb{B}$ such that its law does not depend on $h$. In words, the limit law of the standardized statistics $\sqrt{n}\dparen{\theta_n - \theta_n(h)}$ are not affected by small perturbations of the dgp. Given the pathwise differentiability of $\theta(P)$, the convolution theorem Vaart1996 ensures that the limit law of any regular estimator $\theta_n$ can be expressed as in

equation*[equation* omitted — 68 chars of source]

where $\mathbb{G}_0$ and $W$ are independent, tight, Borel measurable random elements in $\mathbb{B}$ such that $b^*\mathbb{G}_0 \sim N(0,\norm{\Dot{\theta}_0^*b^*}_{2,P_0}^2)$ for $b^* \in \mathbb{B}^*$, and the support of $\mathbb{G}_0$ is the closure of $\Dot{\theta}_0(T(P_0))$. The random element $W$ can be interpreted as an independent noise. Since adding an independent noise solely increases the variance, the noise injection is usually undesirable. Thus, the best regular estimator $\hat{\theta}_n:\mathcal{Z}^n \to \Theta$ is defined as the regular estimator whose limit law equals $\mathbb{G}_0$. That is, the limit law of best regular estimators is maximally concentrated around zero. Hence, the best regularity is often considered one of the desirable properties for an estimator. In parametric models, typical examples of the best regular estimator include the maximum likelihood estimator and Bayesian posterior mean Ibragimov1981. The availability of a best regular estimator $\hat{\theta}_n$ for $\theta_0$ immediately provides a best regular estimator for $f(\theta_0)$ as far as $f: \Theta \to \mathbb{R}$ is an arbitrary Hadamard differentiable map. In this case, the plug-in estimator (i.e., $f(\hat{\theta}_n)$) is a best regular estimator for $f(\theta_0)$. Indeed, it is easy to show that the limit law of any regular estimator $f_n$ for $f(\theta_0)$ can be written as

equation*[equation* omitted — 165 chars of source]

Vaart1996. Again, $\tilde{W}$ is a noise term independent of $\mathbb{G}_0$, and hence the limit law of best regular estimator for $f(\theta_0)$ is $f_{\theta_0}'(\mathbb{G}_0)$. Using the standard delta method, $f(\hat{\theta}_n)$ is the best regular. When $f$ is not Hadamard differentiable but directionally differentiable, the plug-in estimator is no longer the best regular estimator. On the contrary, if $f$ is not Hadamard differentiable, there is no best regular estimator for $f(\theta_0)$ Hirano2012,Vaart1991. Therefore, in the current setting, there exists no estimators for $\tau_L(\theta_0)$, $\tau_U(\theta_0)$, and $\kappa(\theta_0)$ that are efficient in terms of best regularity.

examplecontinued{ex:identification-region-with-empirical-evidence-alone} In this example, the tangent set is equal to $L_2^0(P_0) \coloneqq \dbrace{ h \in L_2(P_0) : \int hdP_0 = 0 }$, and $\theta(P_t)$ is pathwise differentiable at $P_0$ relative to this tangent set. The best regular estimator for $\theta_0$ is \begin{equation*} \hat{\theta}_n = \dparen*{ \hat{\mu}_{1,n},\hat{\mu}_{0,n},\hat{p}_n } = \dparen*{ \frac{\sum_{i=1}^n Y_iD_i}{\sum_{i=1}^n D_i}, \frac{\sum_{i=1}^nY_i(1-D_i)}{\sum_{i=1}^n (1-D_i)}, \frac{\sum_{i=1}^n D_i}{n} }, \end{equation*} and the limit law of $\sqrt{n}\dparen{\hat{\theta}_n - \theta_0}$ is the mean-zero normal distribution whose covariance matrix is the diagonal matrix with its diagonal elements given by $\sigma_{Y|D=1}^2/p_0$, $\sigma_{Y|D=0}^2/(1-p_0)$, and $p_0(1-p_0)$, where $\sigma_{Y|D=d}^2 = \operatorname{\mathbb{E}}_{P_0}[\dparen{Y - \mu_{d,0}}^2|D=d]$ for $d \in \{0,1\}$.
examplecontinued{ex:experimental-study-with-non-random-sampling} As in (ref), the tangent set is $L_2^0(P_0)$, and $\theta(P_t)$ is pathwise differentiable relative to this tangent set. The best regular estimator is \begin{equation*} \hat{\theta}_n = (\hat{\mu}_{1,n},\hat{\mu}_{0,n},\hat{p}_n) = \dparen*{ \frac{\sum_{i=1}^n Y_iD_iS_i}{\sum_{i=1}^nD_iS_i}, \frac{\sum_{i=1}^nY_i(1-D_i)S_i}{\sum_{i=1}^n (1-D_i)S_i},\frac{\sum_{i=1}^n S_i}{n} }, \end{equation*} and $\mathbb{G}_0$ is distributed according to the mean-zero normal distribution whose covariance matrix is the diagonal matrix with its diagonal elements given by $\sigma_{Y|D=1,S=1}^2/(\pi_0p_0)$, $\sigma_{Y|D=0,S=1}^2/((1-\pi_0)p_0)$, and $p_0(1-p_0)$, where $\sigma_{Y|D=d,S=1}^2 = \operatorname{\mathbb{E}}_{P_0}[(Y-\mu_{d,0})^2|D=d,S=1]$ for $d \in \{0,1\}$ and $\pi_0 = P_0(D=1|S=1)$.

Local Asymptotic Minimaxity

Here, we precisely define our asymptotic minimax criterion (i.e., the LAM). First, define the criterion for a general risk function. For an STR $\delta_n$ and a dgp $P \in \mathcal{P}$, let $R_n(\delta_n,P) \geq 0$ be the risk function that quantifies the loss from adopting the STR $\delta_n$ when the dgp is $P$. The (local) asymptotic maximum risk of $\delta_n$ with respect to $R_n$ is defined by

equation[equation omitted — 173 chars of source]

where the first supremum is taken with respect to arbitrary finite subsets $I$ of $T(P_0)$. This notion of local asymptotic maximum risk has been adopted in Vaart1996,Vaart1998,Vaart1991a,Hirano2009,Fang2016,Ponomarev2022. For the motivation of this notion, refer to Fang2016. The STR $\delta_n$ is said to be locally asymptotically minimax (LAM) with respect to the risk function $R_n$ if it minimizes (ref). Note that the LAM property depends on the choice of the risk function. Now, we specify the risk function employed. For an STR $\delta_n$ and $h \in T(P_0)$, its identifiable maximum risk (IMR) is given by\footnote{This name originally comes from Song2014a.}

equation[equation omitted — 227 chars of source]

where $\operatorname{\mathbb{E}}_{n,h*}[\cdot]$ denotes the inner integral with respect to $\mathbb{P}_{n,h}$. The right-hand side of (ref) measures the worst-case welfare regret incurred by not knowing the ATE even with the knowledge of the dgp $P_{1/\sqrt{n},h}$, and it has the same structure with that of (ref). Note that the IMR can be expressed as

equation[equation omitted — 297 chars of source]

The LAM based on the IMR is unsatisfactory. It is easy to show that any consistent estimator for $\kappa(\theta_0)$ is LAM with respect to the IMR. This includes not only the plug-in STR; that is,

equation[equation omitted — 121 chars of source]

but also the STR, $\kappa(\hat{\theta}_n + a_n)$, with arbitrary adjustment term $a \in \mathbb{B}$ such that $a = o_p(1)$. That is, it is not possible to rank consistent estimators. To overcome the drawback of the asymptotic minimaxity based on the IMR, it is necessary to employ other risk functions. The choice here is the regret concerning the IMR: for an STR $\delta_n$ and $h \in T(P_0)$, its regret of identifiable maximum risk (RIMR) is given by

equation[equation omitted — 178 chars of source]

This measures the loss in terms of the IMR because of not knowing the true state $h$. The RIMR is scaled by $\sqrt{n}$ in the analysis. A similar approach has also been adopted in Christensen2022 and Song2014. From (ref), observe that

equation*[equation* omitted — 169 chars of source]

where the minimum is attained by $\delta_n = \kappa(\theta_n(h))$. Substituting the above equation and (ref) in (ref), the RIMR can be expressed as

align[align omitted — 414 chars of source]

For the following technical reason, it is challenging to deal directly with $\mathrm{RIMR}_{n}(\delta_n,h)$ in the asymptotic analysis, whence this study employs a modified version. In the process of obtaining a lower bound of the asymptotic maximum risk for a broad class of statistics, it is often necessary to evaluate the asymptotic lower bound of $\operatorname{\mathbb{E}}_{n,h*}[\ell(\sqrt{n}(\delta_n-\kappa(\theta_n(h)))]$ for some function $\ell:\mathbb{R}\to\mathbb{R}$. When one is interested in the point estimation of $\kappa(\theta_0)$, $\ell$ is usually chosen to be a subconvex loss function, such as squared loss $\ell(x) = x^2$. As the subconvex loss functions are bounded from below and satisfy lower semicontinuity, one can apply the Portmanteau lemma to obtain the desired bound. In the current problem, however, $\ell$ is the identity function that is not bounded; hence, one cannot employ the usual strategy. To circumvent this challenge, this study modifies $\operatorname{RIMR}_{n}(\delta_n,h)$ by replacing $\ell$ with its truncated version. Specifically, it employs $\ell_M(x) \coloneqq \max\{-M,\min\{M,x\}\}$ for $M \in \mathbb{N}$ instead of $\ell$. Correspondingly, the modified $\operatorname{RIMR}_n(\delta_n,h)$ becomes

align*[align* omitted — 466 chars of source]

Asymptotically Minimax Optimal Statistical Treatment Rules

Lower Bound of Asymptotic Maximum RIMR

We first obtain the lower bound of the asymptotic maximum RIMR for a broad class of STRs in (ref).

theoremSuppose (ref) hold. Let $\delta_n$ be an arbitrary STR such that $\sqrt{n}\dparen{\delta_n - \kappa(\theta_0)}$ is asymptotically tight and asymptotically measurable under $P_0^n$. Then, there exists $M_0 \in \mathbb{N}$ such that for every $M \geq M_0$, \begin{align} \MoveEqLeft \adjustlimits{\sup}_{I}{\liminf}_{n \to \infty}\sup_{h \in I} \sqrt{n}\mathrm{RIMR}_{M,n}(\delta_n,h)\nonumber\\ \geq & \adjustlimits{\inf}_{w \in \mathbb{B}}{\sup}_{h \in T(P_0)} \max\dbrace*{ \begin{aligned} &\tau_L^-(\theta_0)\operatorname{\mathbb{E}}\dbrack*{ \kappa_{\theta_0}'\dparen*{\mathbb{G}_0 + w + \Dot{\theta}_0(h) } - \kappa_{\theta_0}'\dparen*{ \Dot{\theta}_0(h) } },\\ -&\tau_U^+(\theta_0)\operatorname{\mathbb{E}}\dbrack*{ \kappa_{\theta_0}'\dparen*{\mathbb{G}_0 + w + \Dot{\theta}_0(h) } - \kappa_{\theta_0}'\dparen*{ \Dot{\theta}_0(h) } } \end{aligned} }. \end{align}

(ref) gives the lower bound of the asymptotic maximum RIMR. This lower bound holds for a broad class of STRs: the only requirement for STRs is the asymptotic tightness and measurability of the standardized statistic $\sqrt{n}(\delta_n - \kappa(\theta_0))$. To see the benefit of this lower bound, it is important to notice that the integrand in the right-hand side of (ref) is the weak limit of a certain class of statistics. Specifically, noting that

equation*[equation* omitted — 136 chars of source]

for any $w$ and $h$, apply the generalized delta method to obtain

align*[align* omitted — 427 chars of source]

This implies that if there exists an adjustment term $w_{\theta_0}^* \in \mathbb{B}$ that solves the minimization in the right-hand side of (ref), the LAM STR can be constructed by

equation[equation omitted — 119 chars of source]

If $w_{\theta_0}^* \neq 0$, this implies that the plug-in STR is not LAM. Instead, it is essential to find an optimal adjustment term $w_{\theta_0}^*$ to obtain the LAM STR. A special case is when $\kappa(\theta)$ is Hadamard differentiable at $\theta_0$. In this case, the optimal adjustment term $w_{\theta_0}^*$ always equals zero. Indeed, the full differentiability implies the linearity of $\kappa_{\theta_0}'$, and hence

align*[align* omitted — 738 chars of source]

Moreover, the linearity of $\kappa_{\theta_0}'$ implies that $\kappa_{\theta_0}'$ is an element of the dual space $\mathbb{B}^*$, where $\kappa_{\theta_0}'(\mathbb{G}_0) \sim N(0, \norm{\Dot{\theta}_0^*\kappa_{\theta_0}'}_{2,P_0}^2)$. Thus, noting that the right-hand side of (ref) is not smaller than zero, the choice, $w_{\theta_0}^* = 0$, is optimal.

remarkThe lower bound in (ref) has an equivalent expression. Specifically, let $s = \Dot{\theta}_0(h) \in \mathbb{B}$ for $h \in T(P_0)$. Then, it is possible to replace the supremum for $h \in T(P_0)$ with the supremum for $s \in \overline{\Dot{\theta}_0(T(P_0))}$, as the corresponding objective function is continuous in $s$. Therefore, the right-hand side of (ref) equals \begin{align} \begin{aligned} \adjustlimits{\inf}_{w\in\mathbb{B}}{\sup}_{s \in S(\mathbb{G}_0)} \max\dbrace*{ \begin{aligned} &\tau_L^-(\theta_0)\operatorname{\mathbb{E}}\dbrack*{ \kappa_{\theta_0}'\dparen*{ \mathbb{G}_0 + w + s } - \kappa_{\theta_0}'(s) },\\ -&\tau_U^+(\theta_0)\operatorname{\mathbb{E}}\dbrack*{ \kappa_{\theta_0}'\dparen*{ \mathbb{G}_0 + w + s} - \kappa_{\theta_0}'(s) } \end{aligned} }, \end{aligned} \end{align} where $S(\mathbb{G}_0)$ denotes the support of $\mathbb{G}_0$. Especially when $\mathbb{B} = \mathbb{R}^k$, the support is equal to the linear span of the column vectors of the efficiency bound of estimators for $\theta_0$, $\Sigma_{\theta_0} = \operatorname{\mathbb{E}}_{P_0}[\tilde{\theta}_0^{~}\tilde{\theta}_0^\top]$. Moreover, if $\Sigma_{\theta_0}$ is nonsigular, then it holds that $S(\mathbb{G}_0) = \mathbb{R}^k$. This expression will be used for constructing the optimal STR that is LAM with respect to RIMR in (ref).
remarkPonomarev2022,Song2014a,Fang2016 have derived the lower bounds of the asymptotic maximum mean squared error (MSE) to develop the LAM point estimators for parameters characterized by nondifferentiable functionals. When one views STRs as estimators for $\kappa(\theta)$, Ponomarev2022's result, combined with the conditions in (ref), implies that the asymptotic maximum MSE is bounded from below as in \begin{align} \MoveEqLeft \adjustlimits\sup_I\liminf_{n\to\infty}\sup_{h \in I} \mathbb{E}_{n,h*}\dbrack*{\dparen*{\sqrt{n}(\delta_n - \kappa(\theta_n(h)))}^2}\nonumber\\ \geq & \adjustlimits\inf_{w \in \mathbb{B}}\sup_{s \in S(\mathbb{G}_0)}\mathbb{E}\dbrack*{ \dparen*{ \kappa_{\theta_0}'\dparen*{ \mathbb{G}_0 + w + s } - \kappa_{\theta_0}'\dparen*{ s } }^2 }. \end{align} This lower bound suggests that the LAM point estimator also takes the form of $\kappa(\hat{\theta}_n + w_{\theta_0}^\dagger)$, where $w_{\theta_0}^\dagger$ is the optimal adjustment term that solves the minimization problem in the right-hand side of the preceding inequality. As the objective functions in (ref) differ, the optimal adjustment terms, $w_{\theta_0}^*$ and $w_{\theta_0}^\dagger$, generally differ as well. This is exemplified using (ref) below.
examplecontinued{ex:identification-region-with-empirical-evidence-alone} As $\kappa(\theta)$ is Hadamard differentiable at $\theta_0$, as long as $\theta_0$ is an interior point, the above discussion implies that the lower bound in (ref) can always be made zero by choosing $w_{\theta_0}^* = 0$. Thus, the plug-in STR is LAM in this example.
examplecontinued{ex:experimental-study-with-non-random-sampling} As noted above, if neither $\tau_L(\theta_0) = 0$ nor $\tau_U(\theta_0) = 0$, then $\kappa(\theta)$ is Hadamard differentiable at $\theta_0$. Thus, the above argument implies that the optimal adjustment term $w_{\theta_0}^*$ equals zero. If $\tau_L(\theta_0) = 0 < \tau_U(\theta_0)$, $\kappa(\theta)$ is Hadamard directionally differentiable at $\theta_0$ with derivative $\kappa_{\theta_0}'(b) = \min\dbrace{\tau_{L,\theta_0}'(b),0}/\tau_U(\theta_0)$. With the linearity of $\tau_{L,\theta_0}'$, (ref) implies that the lower bound is the infimum of \begin{equation*} \sup_{s\in\mathbb{R}^3}\max\dbrace*{0, \min\dbrace*{\tau_{L,\theta_0}'(s),0} - \operatorname{\mathbb{E}}\dbrack*{ \min\dbrace*{\tau_{L,\theta_0}'(\mathbb{G}_0) + \tau_{L,\theta_0}'(w) + \tau_{L,\theta_0}'(s), 0 } } }. \end{equation*} The direct argument based on the closed-form expression of $\operatorname{\mathbb{E}}\dbrack{ \min\dbrace{\tau_{L,\theta_0}'(\mathbb{G}_0 + w + s), 0 } }$ shows that the supremum in the preceding display always attained at $s$ such that $\tau_{L,\theta_0}'(s) = 0$. That is, the object in the preceding display is equal to \begin{equation*} \max\dbrace*{0,-\mathbb{E}\dbrack*{\min\dbrace*{\tau_{L,\theta_0}'(\mathbb{G}_0) + \tau_{L,\theta_0}'(w),0}}}. \end{equation*} Then, it is clear that the above object monotonically approaches zero from above by moving $w$ such that $\tau_{L,\theta_0}'(w)\to\infty$. Hence, the lower bound remains zero, although the optimal adjustment term $w_{\theta_0}^*$ does not exist. This argument implies that the plug-in STR $\kappa(\hat{\theta}_n)$ is not LAM STR. The plug-in STR is asymptotically inferior to the STR $\kappa(\hat{\theta}_n + w /\sqrt{n})$ for any $w$ satisfying $\tau_{L,\theta_0}'(w) > 0$. The LAM point estimator that minimizes the asymptotic maximum MSE is obtained by solving \begin{equation*} \adjustlimits\inf_{w \in \mathbb{R}^3}\sup_{s \in \mathbb{R}^3} \mathbb{E}\dbrack*{ \dparen*{ \min\dbrace*{ \tau_{L,\theta_0}'(\mathbb{G}_0 + w + s), 0 } - \min\dbrace*{ \tau_{L,\theta_0}'(s), 0 } }^2 }. \end{equation*} The direct calculation shows that the optimal adjustment term that solves the above infimum is zero (i.e., $w_{\theta_0}^\dagger = 0$). Therefore, the LAM point estimator is the plug-in estimator $\kappa(\hat{\theta}_n)$. The argument in the preceding paragraph implies that the LAM point estimator is also inferior to the STR $\kappa(\hat{\theta}_n + w/\sqrt{n})$ in terms of the RIMR when $\tau_{L,\theta_0}'(w) > 0$.

Feasible Construction of LAM STR

If the value $\theta_0$ of the intermediate parameter at the true data generating process $P_0$ were known, one could solve the minimax problem in (ref) to obtain the optimal adjustment term $w_{\theta_0}^*$. The value $\theta_0$ is unknown in practice; hence, one cannot solve the minimax problem because of the objects that depend on $\theta_0$. Specifically, five unknown objects exist given the unknown identity of $\theta_0$. First, the limit law $\mathbb{G}_0$ of the best regular estimator $\hat{\theta}_n$ is unknown. Correspondingly, the support $S(\mathbb{G}_0)$ of $\mathbb{G}_0$ is not known as well. The Hadamard directional derivatives $\tau_{L,\theta_0}'$ and $\tau_{U,\theta_0}'$ are also unknown, and so is $\kappa_{\theta_0}'$. Finally, $\tau_L^-(\theta_0)$ and $\tau_U^+(\theta_0)$ are unknown quantities. To alleviate this challenge, this study appropriately estimates those unknown objects and solves the sample analog of (ref), where the unknown objects are replaced with the estimators, to obtain the estimate $\hat{w}_n \in \mathbb{B}$ for $w_{\theta_0}^*$. If $\hat{w}_n$ is a consistent estimator for $w_{\theta_0}^*$, then the STR given by $\kappa(\hat{\theta}_n + \hat{w}_n/\sqrt{n})$ is LAM. To ensure the consistency of $\hat{w}_n$, the estimators of unknown quantities must satisfy the appropriate conditions. In what follows, the study states the required conditions, defines our proposed STR formally, and then shows that the STR is LAM with respect to RIMR. First, estimate the law of $\mathbb{G}_0$ by that of the estimator $(\mathbf{Z}_n,\mathbf{W}_n) \mapsto \hat{\mathbb{G}}_n \in \mathbb{B}$, where $\mathbf{W}_n$ is independent of $\mathbf{Z}_n$. This estimator is particularly constructed via the bootstrap or simulation. For the bootstrap, let $\mathbf{W}_n = (W_{1},\dots,W_{n})$ be the nonparametric bootstrap weights independent of $\mathbf{Z}_n$, and let $(\mathbf{Z}_n,\mathbf{W}_n) \mapsto \hat{\theta}_n^*$ be the bootstrap version of the best regular estimator $\hat{\theta}_n$. Then, use $\hat{\mathbb{G}}_n = \sqrt{n}(\hat{\theta}_n^* - \hat{\theta}_n)$ to approximate the law of $\mathbb{G}_0$. For the simulation, presume the law of $\mathbb{G}_0$ belongs to a known class of distributions. For example, $\mathbb{G}_0 \sim N(0,\Sigma_{\theta_0})$ when $\mathbb{B} = \mathbb{R}^{d_\theta}$. Then, given a consistent estimator $\hat{\Sigma}_n$ for $\Sigma_{\theta_0}$, use $ \hat{\mathbb{G}}_n = \mathbf{W}_n \sim N(0,\hat{\Sigma}_n)$ to approximate the law of $\mathbb{G}_0$. In general, the estimator $\hat{\mathbb{G}}_n$ must satisfy (ref) below. Given this assumption, $\hat{\mathbb{G}}_n \leadsto \mathbb{G}_0$ conditionally on $Z^n$; and hence, it is reasonable to approximate the law of $\mathbb{G}_0$ by $\hat{\mathbb{G}}_n$ Vaart1996.

assumptionThe estimator $(\mathbf{Z}_n,\mathbf{W}_n) \mapsto \hat{\mathbb{G}}_n \in \mathbb{B}$ with $\mathbf{W}_n$ independent of $\mathbf{Z}_n$ has the following properties: (i) $\sup_{f\in \mathrm{BL}_1(\mathbb{B})} \abs{ \operatorname{\mathbb{E}} \dbrack[]{f\dparen{\hat{\mathbb{G}}_n}|\mathbf{Z}_n} - \operatorname{\mathbb{E}} \dbrack[]{f(\mathbb{G}_0)} } = o_p(1)$, where $\mathrm{BL}_1(\mathbb{B})$ is the space of functions $f:\mathbb{B} \to \mathbb{R}$ such that $\sup_{b \in \mathbb{B}}\abs*{f(b)} \leq 1$ and $\abs*{f(b_1) - f(b_2)} \leq \normB{b_1 - b_2}$ for all $b_1,b_2 \in \mathbb{B}$; (ii) $\hat{\mathbb{G}}_n$ is asymptotically measurable jointly in $(\mathbf{Z}_n,\mathbf{W}_n)$.

Second, approximate the support $S(\mathbb{G}_0)$ by an expanding sequence of compact sets. The rationale of this strategy is based on the tightness of $\mathbb{G}_0$. Specifically, $\mathbb{G}_0$ has a tight Borel law if and only if its support is a $\sigma$-compact set (i.e., a countable union of compact sets). Thus, the support of $\mathbb{G}_0$ can be approximated by an expanding sequence of compact sets $S_n \subset \mathbb{B}$. For the property of this set sequence, assume the following.

assumptionThere exists an expanding sequence of compact sets $S_n \subset \mathbb{B}$ such that (i) for every $s \in S(\mathbb{G}_0)$ and for any $\epsilon > 0$, there is $s_n \in S_n$ such that $\norm{s_n - s}_{\mathbb{B}} \leq \epsilon$ for sufficiently large $n$; (ii) $\sup_{s \in S_n} \normB{s} = o(\sqrt{n})$. The sequence $\{S_n\}$ are known or estimated by $\{\hat{S}_n\}$ satisfying $d_H(\hat{S}_n,S_n) = o_p(1)$, where $d_H(\hat{S}_n,S_n)$ denotes the Hausdorff distance between the sets $\hat{S}_n$ and $S_n$.

As stated in (ref), the expanding sequence $\{S_n\}$ is not necessary to be known ex-ante. Rather, it is sufficient to have an estimator $\hat{S}_n$ whose Hausdorff distance from $S_n$ goes to zero in probability. A typical candidate of $\hat{S}_n$ and $S_n$ for the case when $\mathbb{B} = \mathbb{R}^{d_\theta}$ is constructed as follows. Let $\hat{\Sigma}_n$ be a $\sqrt{n}$-consistent estimator for the efficiency bound $\Sigma_{\theta_0}$ of estimators for $\theta_0$. Moreover, let $\hat{\sigma}_j$ and $\sigma_j$ be the $j$-th column vector of $\hat{\Sigma}_n$ and $\Sigma_{\theta_0}$, respectively. Then, the following expanding sequence of compact sets usually fulfills the requirement:

equation*[equation* omitted — 199 chars of source]

where $\lambda_n = o(\sqrt{n})$. Furthermore, in some examples such as (ref), $\Sigma_{\theta_0} = \mathbb{R}^{d_\theta}$. Then, it suffices to take as $S_n$ the closed ball with center $0$ and radius $\lambda_n$. For the choice of $S_n$ and $\hat{S}_n$ when $\mathbb{B}$ is a Banach space, see Section 5.2.2 in Ponomarev2022. Third, this study postulates that the estimators $\hat{\tau}_{U,n}'$ and $\hat{\tau}_{L,n}'$ for $\tau_{U,\theta_0}'$ and $\tau_{L,\theta_0}'$ satisfying the following assumption are available.

assumptionFor each $j \in \{U,L\}$, the estimator $\hat{\tau}_{j,n}':\mathbb{B} \to \mathbb{R}$ is a function of $\mathbf{Z}_n$ and has the following properties: (i) for every $\delta > 0$, $\sup_{b \in K_n^\delta} \abs{ \hat{\tau}_{j,n}'(b) - \tau_{j,\theta_0}'(b) } = o_p(1)$ for arbitrary expanding sequence of compact sets $K_n \subset \mathbb{B}$ with $\sup_{b \in K_n}\normB{b} = o(\sqrt{n})$; (ii) there exists an asymptotically tight random variable $C_{\hat{\tau}_{j,n}'} \in \mathbb{R}$ such that $\abs{\hat{\tau}_{j,n}'(b_1) - \hat{\tau}_{j,n}'(b_2)} \leq C_{\hat{\tau}_{j,n}'}\norm{b_1-b_2}_{\mathbb{B}}$ holds outer almost surely for every $(b_1,b_2) \in \mathbb{B}\times\mathbb{B}$.

(ref).(i) requires some uniform consistency of $\hat{\tau}_{j,n}'$. In many applications, estimators with stronger properties such as $\hat{\tau}_{j,n}(b) = \tau_{j,\theta_0}'(b)$ with probability approaching 1 are often available. See Remarks 3.3 and 3.4 and Section S.4 in the supplementary material of Fang2019 for more details. Usually, one can construct such estimators from the analytical expression of $\tau_{j,\theta_0}'$ or numerical method proposed in Hong2018. It is now appropriate to define the proposed STR. Given the estimators $\hat{\tau}_{U,n}'$ and $\hat{\tau}_{L,n}'$ for $\tau_{U,\theta_0}'$ and $\tau_{L,\theta_0}'$, define the estimator $\hat{\kappa}_n'$ for $\kappa_{\theta_0}'$ by

equation[equation omitted — 346 chars of source]

for $b \in \mathbb{B}$. Here, $\hat{m}_{y,n}$ is the estimator of the directional derivative $m_y'$ of $\max\{x,0\}$ at $y \in \mathbb{R}$, and, for any $x \in \mathbb{R}$, it is given by

equation*[equation* omitted — 213 chars of source]

with $\epsilon_n > 0$ such that $\epsilon_n \downarrow 0$ and $\sqrt{n}\epsilon_n \uparrow \infty$. Let $\{K_m\}$ be an expanding sequence of compact sets in $\mathbb{B}$. For a sufficiently large constant $M > 0$, let

equation[equation omitted — 642 chars of source]

Then, define the STR by

equation[equation omitted — 137 chars of source]

where $\hat{\theta}_n$ is the best regular estimator for $\theta_0$. The following theorem gives an upper bound of local asymptotic maximum RIMR of $\delta_n^{\text{LAM}}$.

theoremSuppose (ref) hold. Then, \begin{align} \begin{aligned} \MoveEqLeft \adjustlimits\limsup_{m \to \infty}\limsup_{M \to \infty}\adjustlimits\sup_{I}\liminf_{n\to\infty}\sup_{h \in I} \sqrt{n}\operatorname{RIMR}_{M,n}(\delta_n^{\operatorname{LAM}},h)\\ \leq & \adjustlimits{\inf}_{w \in \mathbb{B}}{\sup}_{h \in T(P_0)} \max\dbrace*{ \begin{aligned} &\tau_L^-(\theta_0)\operatorname{\mathbb{E}}\dbrack*{ \kappa_{\theta_0}'\dparen*{\mathbb{G}_0 + w + \Dot{\theta}_0(h) } - \kappa_{\theta_0}'\dparen*{ \Dot{\theta}_0(h) }},\\ -&\tau_U^+(\theta_0)\operatorname{\mathbb{E}}\dbrack*{ \kappa_{\theta_0}'\dparen*{\mathbb{G}_0 + w + \Dot{\theta}_0(h) } - \kappa_{\theta_0}'\dparen*{ \Dot{\theta}_0(h) }} \end{aligned} }. \end{aligned} \end{align}

The right-hand sides of (ref) and (ref) coincide. Therefore, combining (ref), $\delta_n^{\operatorname{LAM}}$ is LAM STR.

Simulation Study

This simulation study examines the finite-sample performance of the proposed LAM STR in the context of (ref). In the setup in (ref), consider a binary outcome (i.e., $\mathcal{Y} = \{0,1\}$) and set the true parameter to

equation*[equation* omitted — 110 chars of source]

In this case, $\tau_L(\theta_0) = 0$ and $\tau_U(\theta_0) = 1/2$. Although the boundary functions $\tau_L(\theta)$ and $\tau_U(\theta)$ are differentiable at $\theta_0$, the optimal treatment rule $\kappa(\theta)$ is only directionally differentiable at $\theta_0$. Focus on a specific set of perturbed data generating processes characterized by

equation*[equation* omitted — 152 chars of source]

for $h \in I \coloneqq \{-2,-1.95,-1.90,\dots,2\}$. When $h = 0$, $\theta_n(h)$ corresponds to the true parameter value $\theta_0$. As $|h|$ gets larger, $\theta_n(h)$ deviates more from the nondifferentiable point. The goal here is to compare the maximum RIMR,

equation*[equation* omitted — 70 chars of source]

between the proposed LAM STR and plug-in STR. This study describes details for constructing the LAM STR. Given a sample $\{(S_iY_i,S_iD_i,S_i)\}_{i=1}^n$, $\theta$ is estimated by the best regular estimator $\hat{\theta}_n$ described in (ref) on page (ref). This estimator satisfies $\sqrt{n}(\hat{\theta}_n - \theta_n(h)) \overset{h}{\leadsto} \mathbb{G}_0 \sim N(0,\Sigma_{\theta_0})$ for any $h$. Given a $\sqrt{n}$ consistent estimator $\hat{\Sigma}_n$ for $\Sigma_{\theta_0}$, estimate the law of $\mathbb{G}_0$ by drawing $G_l,l=1,\dots,L$ independently from $N(0,\hat{\Sigma}_n)$, where $L = 1000$. The Hadamard derivatives, $\tau_{L,\theta_0}'(b)$ and $\tau_{U,\theta_0}'(b)$, are estimated by $\hat{\tau}_{L,n}'(b) = \dparen{ \frac{\partial \tau_L}{\partial \theta}|_{\theta = \hat{\theta}_n} }^\intercal b$ and $\hat{\tau}_{U,n}'(b) = \dparen{ \frac{\partial \tau_U}{\partial \theta}|_{\theta = \hat{\theta}_n} }^\intercal b$, respectively. In the estimator for the directional derivative of $\max\{x,0\}$, set $\epsilon_n = n^{1/3}$. This study sets the truncation constant $M$ for the loss function to be $1000$. As the efficiency bound $\Sigma_{\theta_0}$ is nonsingular, the support of $\mathbb{G}_0$ corresponds to $\mathbb{R}^3$. It implies that the support can be approximated by expanding cubes. Hence, set $\hat{S}_n = S_n = [-\lambda_n,\lambda_n]^3$, where $\lambda_n = n^{1/3}$. The compact set $K_m$ over which the optimal adjustment term is searched is set to $[-2,2]^3$. Given these specifications, solve the next minimization problem to obtain the data-dependent adjustment term $\hat{w}_n$:

equation*[equation* omitted — 389 chars of source]

For each $h \in H$, draw $J = 5000$ independent samples of data $\mathbf{Z}_n^{(j)},j = 1,\dots,J$. Each observation $Z_i^{(j)}$ in data $\mathbf{Z}_n^{(j)} = (Z_1^{(j)},\dots,Z_{n}^{(j)})$ comprises $(S_iY_i,S_iD_i,S_i)$ drawn from the data generating process characterized by $\theta_n(h)$ and $\Pr(D=1|S=1) = 1/2$. Based on this data, obtain the best regular estimate $\hat{\theta}_n^{(j)}$ and data-dependent adjustment term $\hat{w}_n^{(j)}$. Given the collection, $\hat{\theta}_n^{(1)},\dots,\hat{\theta}_n^{(J)}$ and $\hat{w}_n^{(1)},\dots,\hat{w}_n^{(J)}$, estimate $\mathbb{E}_{n,h}[\delta_n^{\operatorname{LAM}}]$ by $J^{-1}\sum_{j=1}^J \kappa(\hat{\theta}_n^{(j)} + \hat{w}_n^{(j)}/\sqrt{n})$. Finally, the RIMR is calculated following equation (ref). Similarly, calculate the RIMR of the plug-in STR by ignoring the data-dependent adjustment term. The size $n$ of a sample is set to 300 or 1000. (ref) plots the relation between the local parameter $h$ and RIMR by the size of a sample. For both STRs, the RIMR is maximized at the dgp such that the optimal treatment rule $\kappa(\theta)$ is nondifferentiable (i.e., at $h = 0$). Importantly, the maximized RIMR of the LAM STR is lower than that of the plug-in STR. The LAM STR is designed to minimize the (asymptotic) maximum RIMR, and thus this result is as expected. Conversely when $h \neq 0$, $\kappa(\theta)$ is differentiable at $\theta_n(h)$. Nevertheless, the RIMR is relatively large in a neighborhood of zero. The LAM STR outperforms the plug-in STR in such a neighborhood.

figure[figure omitted — 563 chars of source]

Conclusion

This study analyzed the statistical treatment choice problem under the partial identification of the ATE. Solving the finite-sample problem exactly is difficult for general data generating processes. Instead, this study conducted the local asymptotic analysis and derived the LAM STRs. Specifically, the LAM STR takes the form of

equation*[equation* omitted — 88 chars of source]

where $\hat{\theta}_n$ is the efficient estimator for the true value $\theta_0$ of parameters, $w_{\theta_0}^*$ is the optimal adjustment term that can depend on $\theta_0$, and $\kappa$ is the function that outputs the optimal treatment rule given a value of the parameter. The adjustment term $w_{\theta_0}^*$ generally differs from zero. This implies that the LAM STR differs from the plug-in STR $\kappa(\hat{\theta}_n)$. In practice, the true parameter value $\theta_0$ is unknown; hence, so is the optimal adjustment term $w_{\theta_0}^*$. To alleviate this challenge, the study developed the data-dependent adjustment term $\hat{w}_n$, showing that the STR $\kappa(\hat{\theta}_n + \hat{w}_n/\sqrt{n})$ remained LAM.