Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
94,479 characters · 16 sections · 54 citation commands
Locally Asymptotically Minimax Statistical Treatment Rules Under Partial Identification
Consider the problem of choosing a future treatment for the population of interest from two alternatives, say “treat” and “do not treat”. The objective is to maximize the population's mean outcome, such as average income and cure rate from a specific disease. In making the decision, data that is somewhat informative about the performance of each treatment is available. A data-dependent procedure that effectively chooses the future treatment based on the obtained information from the data is, thus, desirable. Manski2004 calls such data-dependent procedures statistical treatment rules (STRs) and proposes to analyze this problem within the framework of statistical decision theory to derive the optimal STR. Since then, studies on STRs have developed under various settings Stoye2009,Hirano2009,Tetenov2012,Kitagawa2018,Mbakop2021,Athey2021. A common presumption in these studies is the point identification of the (conditional) average treatment effect (ATE) in the target population. In other words, if the true data generating process (dgp), the distribution from which the data is drawn, was known, the target population's ATE could be uniquely determined. It would then be optimal to choose the future treatment depending on the sign of the ATE. Although the true dgp is not known in practice, the preceding argument suggests that one of the most intuitive STRs is the empirical success rule that chooses “treat” if and only if the estimated ATE is positive Manski2004. This rule has several desirable properties. Specifically, Stoye2009 shows that this STR is a nearly finite-sample minimax regret STR under several circumstances. Moreover, Hirano2009 shows that it is asymptotically minimax optimal. However, in practice, the ATE is frequently not point identified but partially identified. For example, one might not assume the random assignment of the treatments in an observational study. Even in experimental studies, the experiment participants may not be the random sample from the population of interest. Moreover, non-compliance with assigned treatment may be allowed. In these cases, the ATE of the target population cannot be point identified but partially identified. Namely, even the full knowledge of the true dgp cannot uniquely pin down the value of the ATE. Rather, the knowledge only gives the identification region, the set to which the ATE belongs. The identification region can take various forms depending on the identifying assumption researchers impose on dgps. Still, the set is not generally a singleton. When the ATE is partially identified, it is troubling for two reasons. First, as the identification region depends on the dgp, it must be estimated from the data. Thus, there is the problem of statistical uncertainty. This is the fundamental feature any statistical decision problem has in common. Second, superior treatment might be ambiguous even with the true knowledge of the dgp. Suppose the true identification region contains positive and negative values. That is, “treat” may be strictly better than “do not treat” in the target population, and vice versa. Hence, it is unclear which treatment should be implemented in the population, even with the true knowledge of the identification region. This is the challenge inherent in the treatment decision problem under partial identification of the ATE. Although the partial identification of the ATE adds ambiguity to the treatment decision problem, the framework of statistical decision theory still applies. In particular, presuming the partial identification of the ATE, several studies have derived finite-sample minimax regret STRs by restricting the class of dgps to some analytically tractable classes or focusing on a specific identifying assumption. Specifically, Ishihara2021,Stoye2012,Yata2021 solve the statistical decision problem exactly by restricting the class of dgps to Gaussian experiments, where the data comes from a normal distribution with unknown mean and known covariance matrix. Manski2007 also solves the finite-sample problem exactly when the partial identification is conducted with empirical evidence alone. Unfortunately, it is challenging to solve the statistical treatment decision problem for a general class of dgps induced from arbitrary identification assumptions. Instead, in the same spirit as Hirano2009, this study conducts a local asymptotic analysis to derive asymptotically optimal STRs. A brief discussion of the setup and results are as follows. The setup comprises semi- or non-parametric models. Suppose the lower and upper bounds of the identification region can be written as $\tau_L(\theta)$ and $\tau_U(\theta)$. Here, $\theta$ is the possibly infinite-dimensional intermediate parameter that summarizes some features of the dgp relevant to the region. Denote by $\theta_0$ the value of the intermediate parameter at the true dgp. The boundary functions $\tau_L$ and $\tau_U$ are known real-valued functions from theory. Assume that these functions are Hadamard directionally differentiable, a version of directional differentiability for maps between two normed spaces. This is the key assumption that determines the scope of this study. Do not assume full differentiability because $\tau_L$ and $\tau_U$ are often not fully differentiable. For instance, $\max$ and $\min$ are not fully differentiable but directionally differentiable, although they often appear in the bounds Manski1990,Manski2000a,Balke1997. Allow for nondeterministic treatment rules: “treat” and “do not treat” can be implemented with positive probabilities in the target population. As noted, when the ATE is partially identified, the optimal treatment cannot be uniquely determined even if the true dgp is known. Thus, to address the ambiguity arising from the partial identification of the ATE, use the minimax regret criterion to uniquely determine the optimal treatment rule that should be implemented if the true dgp were known. Manski2007,Manski2009 have shown that the optimal treatment rule is generally nondeterministic. Specifically, if $\theta_0$ were known, then the optimal treatment rule would be
where $f^+(\cdot) = \max\dbrace{f(\cdot),0}$ and $f^-(\cdot) = \max\dbrace{-f(\cdot),0}$ for arbitrary real-valued function $f$. Concretely, “treat” would be implemented with probability $\kappa(\theta_0)$, while “do not treat” would be implemented with probability $1-\kappa(\theta_0)$. In particular, if the identification region contains positive and negative values, this rule is certainly nondeterministic. When the boundary functions, $\tau_L$ and $\tau_U$, are only directionally differentiable, the optimal treatment rule is also only directionally differentiable. The study seeks an asymptotically optimal STR. Formally, an STR is a function that maps data to a treatment rule. The asymptotic optimality criterion investigated here is the local asymptotic minimaxity (LAM). Given a risk function that measures the loss from the repeated use of an STR for a fixed dgp, the local asymptotic maximum risk of the STR is the worst-case risk over neighborhoods of the true dgp. Then, the LAM STR is the STR that minimizes the local asymptotic maximum risk. The main result of this study implies that the LAM STR takes the form of
where $\hat{\theta}_n$ is the efficient estimator for $\theta_0$, and $w_{\theta_0}^*$ are the optimal adjustment terms at $\theta_0$. The adjustment term $w_{\theta_0}^*$ is not necessarily zero. Hence, the plug-in STR (i.e., $\kappa(\hat{\theta}_n)$) is not generally LAM. The optimal adjustment term generally depends on $\theta_0$; hence, it is unknown in practice unless $\theta_0$ is known. Instead, this study employs a procedure to determine a data-dependent adjustment term $\hat{w}_n$ and show that the feasible STR $\kappa(\hat{\theta}_n + \hat{w}_n/\sqrt{n})$ is LAM. A simulation study exemplifies the proposed STR relative to the plug-in STR. \paragraph{Related Literature.}{ This study mainly contributes to a line of research that studies the statistical treatment choice problem with partially identified ATE. As noted, Manski2007,Stoye2012,Yata2021,Ishihara2021 have derived exact finite-sample minimax regret STRs by restricting the class of data generating process such that the finite-sample problem becomes analytically tractable. In a similar split of this study, Christensen2022 have also conducted the local asymptotic analysis to derive asymptotically optimal STR under partial identification of the ATE. The major difference from their analysis is the adopted optimality criterion: while they analyze the asymptotic average risk optimality, this study analyzes the asymptotic minimax optimality. Further, they restrict themselves to deterministic treatment rules, while this study allows for nondeterministic treatment rules. Several studies examine individualized treatment assignment problems using observable covariates Adjaho2022,Kido2022,DAdamo2022,Pu2021,Mo2021,Zhao2019,Si2020,Siforthcoming,Kallus2018. However, such studies focus on deterministic treatment rules; this study posits that the treatment assignment in the future population can be probabilistic. Furthermore, the result of this study applies to such individualized treatment assignment problems as long as it considers local asymptotics pointwise in covariates, and there is no constraint on feasible individualized treatment rules. This study is also closely related to studies of efficient estimation of parameters expressed as $f(\theta_0)$ for some known function $f$. When $f$ is Hadamard differentiable, the plug-in estimator $f(\hat{\theta}_n)$ is the efficient estimator Vaart1991b. When $f$ is not Hadamard differentiable but Hadamard directionally differentiable, then regular estimators, which are robust to small perturbations of dgps, do not exist Hirano2012,Vaart1991. Thus, the classical notion of efficiency based on best regularity breaks down for such nondifferentiable functions. However, the efficiency based on LAM still works Song2014a,Fang2016,Ponomarev2022. The relevant assume the (Hadamard) directional differentiability of $f$ and derive the LAM point estimators for $f(\theta)$. In such studies, the LAM point estimators also require some adjustment terms. In particular, Ponomarev2022 shows that the LAM point estimators generally take the form of $f(\hat{\theta}_n + w_{\theta_0}^\dagger/\sqrt{n}) + v_{\theta_0}^\dagger/\sqrt{n}$. Accordingly, it is possible to construct the LAM point estimator for $\kappa(\theta_0)$. However, the optimal adjustment terms $w_{\theta_0}^*$ and $w_{\theta_0}^\dagger$ generally differ because the objective functions differ in the estimation problem and treatment assignment problem; the point estimation problem aims at minimizing the symmetric loss function such as mean squared error, while the treatment assignment problem aims at minimizing the asymmetric loss function such as welfare regret loss. The study will exemplify this difference using an example. The Hadamard directional differentiability has also been adopted in Fang2019,Hong2018 to develop asymptotically valid inference procedures on $f(\theta_0)$. } \paragraph{Structure of the Paper.}{ (ref) formally sets up the statistical treatment decision problem under the partial identification of the ATE. It also introduces the precise requirement for the identification region of the ATE with some motivating examples. (ref) discusses the framework of the local asymptotic analysis employed in this study. (ref) derives the LAM STRs, and (ref) concludes the paper. All proofs are relegated to (ref). }
Suppose there is a binary treatment encoded as 0 and 1. Let $Y(a) \in \mathcal{Y}$ denote the potential outcome that would be realized if an individual was exposed to treatment $a \in \{0,1\}$. The set $\mathcal{Y}$ is assumed to be a bounded subset of $\mathbb{R}$, and let $y_L \coloneqq \inf \mathcal{Y}$ and $y_U \coloneqq \sup \mathcal{Y}$. Consider a social planner who wants to assign treatment 0 or 1 to each individual of a population of interest. The planner is allowed to implement probabilistic assignments independently of potential outcomes. Specifically, the planner can assign treatment 1 to an individual with probability $\delta \in [0,1]$, which will be called a treatment rule. When the planner adopts a treatment rule $\delta$, the welfare at the population of interest is given by
where the expectation is taken with respect to the target population's joint distribution of $(Y(1),Y(0))$. Suppose that higher welfare is more desirable. Then, if the planner knew the true marginal distributions of potential outcomes, she would search for the treatment rule that maximizes the welfare. The right-hand side of (ref) shows that the important quantity for her optimization is the average treatment effect (ATE), $\operatorname{\mathbb{E}} \dbrack{Y(1) - Y(0)}$. The optimal treatment rules assign treatment 1 with probability one (zero) if and only if the ATE is positive (negative). Therefore, if she knew the signs of ATE, she could obtain an optimal treatment rule. Unfortunately, this rule is generally infeasible, as the ATE is unknown. Suppose instead a sample $\mathbf{Z}_n = (Z_1, \cdots, Z_n)$ is available to the planner. It is assumed that each observation $Z_i$ belongs to a complete separable metric space $\mathcal{Z}$ equipped with the Borel $\sigma$-algebra $\mathcal{B}(\mathcal{Z})$. Further, it is independent and identically distributed (i.i.d.) according to distribution $P_0 \in \mathcal{P}$, where $\mathcal{P}$ is a known class of Borel probability measures.\footnote{The completeness and separability of the space $\mathcal{Z}$ is put to ensure the separability of $L_2(P)$ space for every $P \in \mathcal{P}$. Indeed, every Borel probability measure $P$ is Radon Bogachev2007. This then implies that every $P$ is separable Bogachev2007. The separability of $L_2(P)$ follows from the separability of $P$ \Parencite[][Exercise 4.7.63]{Bogachev2007}.} Given this i.i.d. assumption, the sample is generated following $\mathbf{Z}_n \sim P_0^n$, which is the $n$-fold product of $P_0$. It is common to often refer each $P \in \mathcal{P}$ to a data generating process (dgp) and also refer $P_0$ to the true dgp. If an infinite number of observations of $Z_i$ were available, the dgp $P_0$ would be learnable with arbitrary precision. This study assumes that the ATE cannot be point identified but partially identified. Formally, even with knowledge of the true dgp $P_0$, it is only possible to conclude that the ATE belongs to the identification region $\tau(P_0)$, which is a subset of $[y_L - y_U,y_U-y_L]$ by the boundedness of $\mathcal{Y}$. Note that $\tau(P_0)$ may or may not contain positive and negative values, in which case the better treatment cannot be determined. The study places some restrictions on $\tau(P_0)$, which will be discussed in (ref). Given the availability of a sample $\mathbf{Z}_n$, it is natural that the planner decides a treatment rule, utilizing the knowledge obtained from $\mathbf{Z}_n$. From this perspective, what the planner needs is the procedure that outputs a good treatment rule for each possible realization of $\mathbf{Z}_n$. The planner's problem can then be interpreted as one of the statistical decision problems Manski2004. A statistical decision problem mainly comprises two components: a statistical decision rule and a risk. In the current context, a statistical decision rule is a mapping $\delta_n:\mathcal{Z}^n \to [0,1]$, which is called a statistical treatment rule (STR). Here, $\delta_n(\mathbf{z}_n) \in [0,1]$ is interpreted as the probability of assigning treatment 1 in the target population when the observed sample is $\mathbf{z}_n$. For the risk function, it is important to first specify a loss function that penalizes an arbitrary treatment rule $\delta$ when ATE is $\tau$. Following the literature, this study focuses on the welfare regret loss given by
This measures the loss the planner incurs by not knowing $\tau$. Given this loss function, for every $P \in \mathcal{P}$ and $\tau \in \tau(P)$, the risk of a STR $\delta_n$ is defined by $\operatorname{\mathbb{E}}_{P^n} \dbrack*{ L \dparen*{ \delta_n \dparen*{\mathbf{Z}_n}, \tau } }$. The definition clearly shows that the risk depends on the true dgp $P_0$ and ATE $\tau$. As neither $P_0$ nor $\tau$ are known with a sample of finite size, the planner would seek the STR that minimizes the risk for every possible $P \in \mathcal{P}$ and $\tau \in \tau(P)$. However, such STRs do not necessarily exist. To alleviate this difficulty, the literature on statistical decision theory has considered several criteria. Among those criteria, this study employs the minimax regret criterion, as suggested by Manski2004. Specifically, the minimax regret STR minimizes
From here, suppress the dependence of an STR $\delta_n(\mathbf{Z}_n)$ on the data $\mathbf{Z}_n$ to simplify the notation.
Among the features of the identification region $\tau(P)$, the important quantities in the current context are its lower and upper bounds. Indeed, when one fixes $P \in \mathcal{P}$ and consider the inner maximization problem in (ref), the direct calculation yields
where for any real value $x \in \mathbb{R}$, $(x)^+ = \max\{x,0\}$, and $(x)^- = \max\{-x,0\}$. Note that the boundedness of the outcome space implies that the bounds of $\tau(P)$ are bounded by a constant uniformly in $P$. The equation reveals that whatever the form of $\tau(P)$ is, the value of the inner maximization in (ref) depends only on $\inf\tau(P)$ and $\sup\tau(P)$. This study requires that the bounds, $\inf\tau(P)$ and $\sup\tau(P)$, are sufficiently smooth in $P$.
To motivate the concrete smoothness conditions, the study first presents at some examples of partial identification of ATE.
The inspection of (ref) reveals that the lower and upper bounds of $\tau(P)$ can often be written as a composition of maps. The relevant features of the data-generating process $P$ are first summarized by intermediate parameters $\theta$. Then, $\theta$ is mapped by certain functions, say $\tau_L$ and $\tau_U$, to the lower and upper bounds of $\tau(P)$. For the smoothness of $\tau_L$ and $\tau_U$, we must deal with the possible nondifferentiability of these functions. For instance, max and min functions, which appear in (ref), are not differentiable at some points. Rather, the max and min functions are directionally differentiable. Thus, to cover such instances, it is adequate to assume the directional differentiability of $\tau_L$ and $\tau_U$ in some strength. In the following, this study formalizes these features as the restrictions on the identification region. Suppose there exist $\theta:\mathcal{P} \to \Theta \subset \mathbb{B}$ for a Banach space $\mathbb{B}$ with norm $\normB{\cdot}$ and known functions $\tau_j:\Theta \to \mathbb{R}$ for $j \in \{L,U\}$ such that
One can take $\mathbb{B} = \mathbb{R}^k$ for an integer $k$ in most applications, but the formulation allows infinite-dimensional parameters, such as $\mathbb{B} = L_2(P)$. The intermediate parameter $\theta(P)$ is required to be pathwise differentiable, which will be discussed in (ref). On the other hand, the functions $\tau_L(\theta)$ and $\tau_U(\theta)$ are assumed to be directionally differentiable in a certain strength, which is clarified in (ref) below. In that statement, let $P_0 \in \mathcal{P}$ be the true data generating process and let $\theta_0 \coloneqq \theta(P_0)$ be the value of the intermediate parameter under true dgp.
(ref) gives the precise strength of directional differentiability needed in this paper. It is important to note that the Hadamard directional differentiability does not demand the linearity of the derivative $\tau_{j,\theta_0}':\mathbb{B} \to \mathbb{R}$. When it is linear, the function $\tau_{j}:\mathbb{B} \to \mathbb{R}$ is said to be Hadamard differentiable at $\theta_0 \in \Theta$ Fang2019. Conventionally, Hadamard differentiability has often been utilized to derive the asymptotic distribution of statistics by applying the delta method. Concretely, if $\theta_n$ is an estimator for $\theta_0$ such that $\sqrt{n}(\theta_n - \theta_0) \leadsto \mathbb{G}$ and $\tau_j$ is Hadamard differentiable at $\theta_0$, then
where $\leadsto$ indicates the weak convergence. Unfortunately, the Hadamard differentiability is too strong to cover max and min functions that often appear in the bounds of identification region as in (ref). Instead, this study only assumes the directional differentiability. Among the several notions of directional differentiability Shapiro1990, the Hadamard directional differentiability is the most suitable for this study. Especially, as emphasized in Fang2019, this directional differentiability allows for the direct generalization of the standard delta method: as long as $\tau_{j}(\theta)$ is Hadamard directionally differentiable at $\theta_0$, the weak convergence in (ref) continues to hold. Based on this generalized delta-method, Fang2016,Ponomarev2022 developed the locally asymptotically minimax estimator for $\tau_j(\theta_0)$ itself. This study utilizes the generalized delta method to develop the LAM STR. In addition to the generalization of the delta method, the Hadamard directional differentiability is also useful, as it admits the chain rule Shapiro1990. (ref) requires that the Hadamard directional derivatives are Lipschitz continuous. (ref) is imposed to simplify the analysis. When $\mathbb{B} = \mathbb{R}^k$, (ref) is satisfied if $\tau_{j,\theta_0}'$ is translation equivariant. In this case, $\tau_{j,\theta_0}'(b + v\mathbf{1}) = \tau_{j,\theta_0}'(b) + v$ for every $b$ and $v$, where $\mathbf{1} = (1,\cdots,1) \in \mathbb{R}^k$. Particularly, $\max$ and $\min$ functions, which often appear in the bounds of partial identification, satisfy the translation equivariance.
If the social planner knew the true dgp $P_0$, there would be no need to consider statistical uncertainty, and hence, they would solve $\inf_{\delta \in [0,1]}\sup_{\tau \in \tau(P_0)} L(\delta,\tau)$. As has been observed in Manski2007,Manski2009, the minimax regret value is
and the optimal treatment rule is given by (ref). Depending on the true intermediate parameter $\theta_0$, the future assignment given the optimal treatment rule may (not) be deterministic, and the corresponding minimax value may (not) be greater than zero. Following Manski2009, this study introduces a concept of error because of selecting an inferior treatment to interpret the regret value and rule. Specifically, refer to choosing treatment 0 when the ATE $\tau$ is positive as a Type 0 error. Conversely, refer to choosing treatment 1 when $\tau$ is negative as a Type 1 error. Further, recall that the welfare regret loss is given by $\tau(1-\delta)\cdot \operatorname{1}\dbrace{\tau > 0} - \tau\delta\cdot \operatorname{1}\dbrace{\tau \leq 0}$. The regret loss measures the loss from Type 0 or 1 errors: $\tau(1-\delta)$ corresponds to the loss from the Type 0 error, while $-\tau\delta$ corresponds to the loss given the Type 1 error. First, if $\tau_L(\theta_0) < \tau_U(\theta_0) \leq 0$, then $\kappa(\theta_0) = 0$, and hence, treatment 1 is assigned with probability 0. In this case, the ATE is certainly non-positive, and what matters is the Type 1 error. From the knowledge of the identification region $\tau(P_0)$, treatment 0 is never worse than treatment 1. Accordingly, the superior treatment can be chosen for sure, and the loss from the Type 1 error can be made zero. Similarly, if $0 \leq \tau_L(\theta_0) < \tau_U(\theta_0)$, then $\kappa(\theta_0) = 1$, and the minimax regret is again zero. This case is also interpreted similarly. Finally, if $\tau_L(\theta_0) < 0 < \tau_U(\theta_0)$, $\kappa(\theta_0) = \tau_U(\theta_0)/\dparen{\tau_U(\theta_0) - \tau_L(\theta_0)} \in (0,1)$, and the minimax regret value is greater than zero. In this case, positive and negative ATEs are possible, and hence, Type 0 and Type 1 errors matter. In particular, the worst-case loss from Type 0 error is $\tau_U(\theta_0)(1-\delta)$, and the worst-case loss from Type 1 error is $-\tau_L(\theta_0)\delta$. As $\delta$ increases from 0 to 1, the former loss decreases from $\tau_U(\theta_0)$ to $0$, while the latter loss increases from $0$ to $-\tau_L(\theta_0)$. As the welfare regret loss equally weighs the losses from both errors, the optimal treatment rule balances the worst-case losses from Type 0 and 1 errors; that is, $\tau_U(\theta_0)(1-\kappa(\theta_0)) = -\tau_L(\theta_0)\kappa(\theta_0)$. Under (ref), the optimal treatment rule $\kappa(\theta)$ inherits the properties of $\tau_L$ and $\tau_U$.
Importantly, $\kappa(\theta)$ is also Hadamard directionally differentiable at $\theta_0$. The Hadamard directional derivative is derived by applying the chain rule for Hadamard directionally differentiable maps. Note that $\kappa(\theta)$ may not be differentiable at $\theta_0$ even if $\tau_L(\theta)$ and $\tau_U(\theta)$ are Hadamard differentiable at $\theta_0$. For instance, even if $\tau_U(\theta)$ is Hadamard differentiable at $\theta_0$, $\max\{\tau_U(\theta_0),0\}$ is not Hadamard differentiable when $\tau_U(\theta_0) = 0$; so is $\kappa(\theta_0)$. See also (ref)
It is often challenging to exactly solve the finite-sample problem (ref) for general data-generating processes. Instead, in the same spirit of Hirano2009, this study conducts the local asymptotic analysis and then derives asymptotically minimax STRs. This section discusses the framework of the local asymptotic analysis. Specifically, (ref) setups local statistical experiments that are tailored to i.i.d. observations. (ref) gives the precise smoothness required for the intermediate parameter. Finally, (ref) formally defines the notion of asymptotic minimax optimality employed in this study.
The study considers a sequence of local statistical experiments in the following asymptotic analysis. Although the original statistical experiment involves all of the possible dgps in $\mathcal{P}$, the fixed true dgp $P_0$ is learnable with arbitrary precision as the sample size approaches infinity. Thus, from the perspective of asymptotic analysis, the asymptotic analysis, considering dgps that are too distant from $P_0$ may be meaningless. Instead, the local asymptotic analysis focuses only on dgps that are statistically difficult to distinguish from the true dgp $P_0$ to approximate the finite-sample situation well. The next concrete construction of the sequence of local experiments closely follows Hirano2009,Vaart1991a. Given the fixed true dgp $P_0 \in \mathcal{P}$, let $\mathcal{P}(P_0)$ be a collection of paths $(0, \epsilon) \ni t \mapsto P_{t} \in \mathcal{P}$ such that
for some measurable function $h: \mathcal{Z} \to \mathbb{R}$, which is often called the score function. It is possible to view each path $t \mapsto P_t$ as a one-dimensional parametric submodel $\dbrace{P_t : t \in (0,\epsilon)} \subset \mathcal{P}$, in which $t$ is treated as the only parameter. Gathering all score functions associated with paths in $\mathcal{P}(P_0)$ induces the tangent set $T(P_0)$, which is the collection of score functions. It is known that $\int h dP_0 = 0$ and $\int h^2 dP_0 < \infty$ for any score function satisfying (ref). Thus, regarding $P_0$-almost surely equal score functions as an equivalence class, $T(P_0)$ can be viewed as a subset of $L_2(P_0)$, which is the separable Hilbert space with the inner product $\dangle{h,g} \coloneqq \int hg dP_0$ and norm $\norm{h}_{2,P_0} \coloneqq \sqrt{\dangle{h,h}}$. To simplify the exposition, this study posits that the tangent set is closed with respect to addition and scalar multiplication.\footnote{The result of this paper can be extended to the case where the tangent set is a convex cone with some appropriate modifications of proofs. See, e.g., Vaart1989,Ponomarev2022.}
We introduce some notations to be used hereafter. For each path $t\mapsto P_t$ with score function $h \in T(P_0)$, let $P_{1/\sqrt{n},h} \in \mathcal{P}$ denote the value of the path evaluated at $t = 1/\sqrt{n}$. Moreover, let $\mathbb{P}_{n,h} = P_{1/\sqrt{n},h}^n$ and let $\mathbb{P}_{n,0} = P_0^n$. Correspondingly, the expectation with respect to $\mathbb{P}_{n,h}$ is denoted by $\operatorname{\mathbb{E}}_{n,h}[\cdot]$. Further, for every path with score function $h$, write $\theta_n(h) = \theta(P_{1/\sqrt{n},h})$. An important consequence of the property (ref) is the local asymptotic normality of the statistical experiments $\mathcal{E}_n = (\mathcal{Z}^n,\mathcal{B}(\mathcal{Z}^n),\mathbb{P}_{n,h}:h \in T(P_0))$. Specifically, for each $h \in T(P_0)$, the log-likelihood ratio admits a linear expansion such that
where $\Delta_{n,h} = \sum_{i=1}^n h(Z_i)/\sqrt{n}$ converges weakly to $N(0,\norm{h}_{2,P_0}^2)$ under $\mathbb{P}_{n,0}$. This local asymptotic normality further implies the convergence of experiments $\mathcal{E}_n$ to a certain limit experiment $\mathcal{E}$. Under (ref), a useful characterization of $\mathcal{E}$ is available. (ref) implies that $\overline{T(P_0)}$ itself is one of separable Hilbert spaces. Then, there exists a complete orthonormal basis $\{h_1,h_2,\cdots\} \subset T(P_0)$, and every $h \in T(P_0)$ can be written as $h = \sum_{j=1}^\infty \dangle{h,h_j}h_j$. In this case, the following convergence of experiments holds:
The latter experiment corresponds to observe a single draw $\Xi = (\Xi_1,\Xi_2,\cdots) \in \mathbb{R}^\infty$ such that $\Xi_1,\Xi_2,\cdots$ are independent and $\Xi_j \sim N(\dangle{h,h_j},1)$ for each $j \in \mathbb{N}$. For ease of notation, write $\mathbb{P}_h$ for $N_\infty((\dangle{h,h_j}),\mathbf{1})$ for every $h \in T(P_0)$. Accordingly, the expectation with respect to $\mathbb{P}_h$ is denoted by $\operatorname{\mathbb{E}}_{h}[\cdot]$. As one expects, this limit experiment $\mathcal{E}$ is more tractable relative to the original experiment $\mathcal{E}_n$. Finally, the convergence of experiments culminates in the following asymptotic representation theorem.
In the statement of (ref), the randomized statistic $T$ refers to a measurable map $T:\mathbb{R}^\infty \times [0,1] \ni (\Xi,U) \mapsto T(\Xi,U) \in \mathbb{D}$, where $\Xi \sim \mathbb{P}_h$ and $U \sim \mathrm{Unif}([0,1])$ independently of $\Xi$. Moreover, $\overset{h}{\leadsto}$ means the weak convergence under $\mathbb{P}_{n,h}$ for $h \in T(P_0)$. (ref) allows for guessing a statistic with a certain asymptotic property by the following procedures; (i) focus on the limit experiment $\mathcal{E}$; (ii) obtain the statistic $T$ with the desired property in $\mathcal{E}$; (iii) construct a sequence of statistics $T_n$ that converges weakly to $T$.
For the intermediate parameter $\theta:\mathcal{P} \to \mathbb{B}$ introduced in (ref), assume its pathwise differentiability.
The pathwise differentiability especially implies $\sqrt{n}\dparen{\theta_n(h) - \theta_0} \to \Dot{\theta}_0(h)$ as $n \to \infty$. The pathwise derivative $\Dot{\theta}_0(h)$ must be linear maps. When $\mathbb{B}$ is the Euclidean space, the Riesz representation theorem gives a useful characterization of the pathwise derivative. Specifically, as $\overline{T(P_0)}$ is a Hilbert space, the theorem implies that there exists $\tilde{\theta}_0 = (\tilde{\theta}_{0,1},\cdots,\tilde{\theta}_{0,k})^\top \in \overline{T(P_0)}^k$ such that $\Dot{\theta}_{0,j}(h) = \dangle{\tilde{\theta}_{0,j},h}$ for all $h \in \overline{T(P_0)}$, where $\Dot{\theta}_{0,j}$ denotes the $j$-th element of $\Dot{\theta}_0$. The function $\tilde{\theta}_0$, thus constructed, is often called efficient influence function. More generally, if $\mathbb{B}$ is an arbitrary Banach space, an adjoint map can characterize the pathwise derivative. In particular, there exists $\Dot{\theta}_0^*:\mathbb{B}^* \to \overline{T(P_0)}$ determined by $\dangle{\Dot{\theta}_0^*b^*, h} = b^*(\Dot{\theta}_0(h))$ for every $h \in T(P_0)$ and $b^* \in \mathbb{B}^*$. The pathwise differentiability assumption is often employed in the study of efficient estimators Vaart1998,Vaart1996. Among the results obtained in the study, one of the most important is the convolution theorem that characterizes the limit law of regular estimators. An estimator $\theta_n:\mathcal{Z}^n \to \Theta \subset \mathbb{B}$ is said to be regular at $P_0$ for estimating $\theta_0$ if
where $G$ is a tight random element in $\mathbb{B}$ such that its law does not depend on $h$. In words, the limit law of the standardized statistics $\sqrt{n}\dparen{\theta_n - \theta_n(h)}$ are not affected by small perturbations of the dgp. Given the pathwise differentiability of $\theta(P)$, the convolution theorem Vaart1996 ensures that the limit law of any regular estimator $\theta_n$ can be expressed as in
where $\mathbb{G}_0$ and $W$ are independent, tight, Borel measurable random elements in $\mathbb{B}$ such that $b^*\mathbb{G}_0 \sim N(0,\norm{\Dot{\theta}_0^*b^*}_{2,P_0}^2)$ for $b^* \in \mathbb{B}^*$, and the support of $\mathbb{G}_0$ is the closure of $\Dot{\theta}_0(T(P_0))$. The random element $W$ can be interpreted as an independent noise. Since adding an independent noise solely increases the variance, the noise injection is usually undesirable. Thus, the best regular estimator $\hat{\theta}_n:\mathcal{Z}^n \to \Theta$ is defined as the regular estimator whose limit law equals $\mathbb{G}_0$. That is, the limit law of best regular estimators is maximally concentrated around zero. Hence, the best regularity is often considered one of the desirable properties for an estimator. In parametric models, typical examples of the best regular estimator include the maximum likelihood estimator and Bayesian posterior mean Ibragimov1981. The availability of a best regular estimator $\hat{\theta}_n$ for $\theta_0$ immediately provides a best regular estimator for $f(\theta_0)$ as far as $f: \Theta \to \mathbb{R}$ is an arbitrary Hadamard differentiable map. In this case, the plug-in estimator (i.e., $f(\hat{\theta}_n)$) is a best regular estimator for $f(\theta_0)$. Indeed, it is easy to show that the limit law of any regular estimator $f_n$ for $f(\theta_0)$ can be written as
Vaart1996. Again, $\tilde{W}$ is a noise term independent of $\mathbb{G}_0$, and hence the limit law of best regular estimator for $f(\theta_0)$ is $f_{\theta_0}'(\mathbb{G}_0)$. Using the standard delta method, $f(\hat{\theta}_n)$ is the best regular. When $f$ is not Hadamard differentiable but directionally differentiable, the plug-in estimator is no longer the best regular estimator. On the contrary, if $f$ is not Hadamard differentiable, there is no best regular estimator for $f(\theta_0)$ Hirano2012,Vaart1991. Therefore, in the current setting, there exists no estimators for $\tau_L(\theta_0)$, $\tau_U(\theta_0)$, and $\kappa(\theta_0)$ that are efficient in terms of best regularity.
Here, we precisely define our asymptotic minimax criterion (i.e., the LAM). First, define the criterion for a general risk function. For an STR $\delta_n$ and a dgp $P \in \mathcal{P}$, let $R_n(\delta_n,P) \geq 0$ be the risk function that quantifies the loss from adopting the STR $\delta_n$ when the dgp is $P$. The (local) asymptotic maximum risk of $\delta_n$ with respect to $R_n$ is defined by
where the first supremum is taken with respect to arbitrary finite subsets $I$ of $T(P_0)$. This notion of local asymptotic maximum risk has been adopted in Vaart1996,Vaart1998,Vaart1991a,Hirano2009,Fang2016,Ponomarev2022. For the motivation of this notion, refer to Fang2016. The STR $\delta_n$ is said to be locally asymptotically minimax (LAM) with respect to the risk function $R_n$ if it minimizes (ref). Note that the LAM property depends on the choice of the risk function. Now, we specify the risk function employed. For an STR $\delta_n$ and $h \in T(P_0)$, its identifiable maximum risk (IMR) is given by\footnote{This name originally comes from Song2014a.}
where $\operatorname{\mathbb{E}}_{n,h*}[\cdot]$ denotes the inner integral with respect to $\mathbb{P}_{n,h}$. The right-hand side of (ref) measures the worst-case welfare regret incurred by not knowing the ATE even with the knowledge of the dgp $P_{1/\sqrt{n},h}$, and it has the same structure with that of (ref). Note that the IMR can be expressed as
The LAM based on the IMR is unsatisfactory. It is easy to show that any consistent estimator for $\kappa(\theta_0)$ is LAM with respect to the IMR. This includes not only the plug-in STR; that is,
but also the STR, $\kappa(\hat{\theta}_n + a_n)$, with arbitrary adjustment term $a \in \mathbb{B}$ such that $a = o_p(1)$. That is, it is not possible to rank consistent estimators. To overcome the drawback of the asymptotic minimaxity based on the IMR, it is necessary to employ other risk functions. The choice here is the regret concerning the IMR: for an STR $\delta_n$ and $h \in T(P_0)$, its regret of identifiable maximum risk (RIMR) is given by
This measures the loss in terms of the IMR because of not knowing the true state $h$. The RIMR is scaled by $\sqrt{n}$ in the analysis. A similar approach has also been adopted in Christensen2022 and Song2014. From (ref), observe that
where the minimum is attained by $\delta_n = \kappa(\theta_n(h))$. Substituting the above equation and (ref) in (ref), the RIMR can be expressed as
For the following technical reason, it is challenging to deal directly with $\mathrm{RIMR}_{n}(\delta_n,h)$ in the asymptotic analysis, whence this study employs a modified version. In the process of obtaining a lower bound of the asymptotic maximum risk for a broad class of statistics, it is often necessary to evaluate the asymptotic lower bound of $\operatorname{\mathbb{E}}_{n,h*}[\ell(\sqrt{n}(\delta_n-\kappa(\theta_n(h)))]$ for some function $\ell:\mathbb{R}\to\mathbb{R}$. When one is interested in the point estimation of $\kappa(\theta_0)$, $\ell$ is usually chosen to be a subconvex loss function, such as squared loss $\ell(x) = x^2$. As the subconvex loss functions are bounded from below and satisfy lower semicontinuity, one can apply the Portmanteau lemma to obtain the desired bound. In the current problem, however, $\ell$ is the identity function that is not bounded; hence, one cannot employ the usual strategy. To circumvent this challenge, this study modifies $\operatorname{RIMR}_{n}(\delta_n,h)$ by replacing $\ell$ with its truncated version. Specifically, it employs $\ell_M(x) \coloneqq \max\{-M,\min\{M,x\}\}$ for $M \in \mathbb{N}$ instead of $\ell$. Correspondingly, the modified $\operatorname{RIMR}_n(\delta_n,h)$ becomes
We first obtain the lower bound of the asymptotic maximum RIMR for a broad class of STRs in (ref).
(ref) gives the lower bound of the asymptotic maximum RIMR. This lower bound holds for a broad class of STRs: the only requirement for STRs is the asymptotic tightness and measurability of the standardized statistic $\sqrt{n}(\delta_n - \kappa(\theta_0))$. To see the benefit of this lower bound, it is important to notice that the integrand in the right-hand side of (ref) is the weak limit of a certain class of statistics. Specifically, noting that
for any $w$ and $h$, apply the generalized delta method to obtain
This implies that if there exists an adjustment term $w_{\theta_0}^* \in \mathbb{B}$ that solves the minimization in the right-hand side of (ref), the LAM STR can be constructed by
If $w_{\theta_0}^* \neq 0$, this implies that the plug-in STR is not LAM. Instead, it is essential to find an optimal adjustment term $w_{\theta_0}^*$ to obtain the LAM STR. A special case is when $\kappa(\theta)$ is Hadamard differentiable at $\theta_0$. In this case, the optimal adjustment term $w_{\theta_0}^*$ always equals zero. Indeed, the full differentiability implies the linearity of $\kappa_{\theta_0}'$, and hence
Moreover, the linearity of $\kappa_{\theta_0}'$ implies that $\kappa_{\theta_0}'$ is an element of the dual space $\mathbb{B}^*$, where $\kappa_{\theta_0}'(\mathbb{G}_0) \sim N(0, \norm{\Dot{\theta}_0^*\kappa_{\theta_0}'}_{2,P_0}^2)$. Thus, noting that the right-hand side of (ref) is not smaller than zero, the choice, $w_{\theta_0}^* = 0$, is optimal.
If the value $\theta_0$ of the intermediate parameter at the true data generating process $P_0$ were known, one could solve the minimax problem in (ref) to obtain the optimal adjustment term $w_{\theta_0}^*$. The value $\theta_0$ is unknown in practice; hence, one cannot solve the minimax problem because of the objects that depend on $\theta_0$. Specifically, five unknown objects exist given the unknown identity of $\theta_0$. First, the limit law $\mathbb{G}_0$ of the best regular estimator $\hat{\theta}_n$ is unknown. Correspondingly, the support $S(\mathbb{G}_0)$ of $\mathbb{G}_0$ is not known as well. The Hadamard directional derivatives $\tau_{L,\theta_0}'$ and $\tau_{U,\theta_0}'$ are also unknown, and so is $\kappa_{\theta_0}'$. Finally, $\tau_L^-(\theta_0)$ and $\tau_U^+(\theta_0)$ are unknown quantities. To alleviate this challenge, this study appropriately estimates those unknown objects and solves the sample analog of (ref), where the unknown objects are replaced with the estimators, to obtain the estimate $\hat{w}_n \in \mathbb{B}$ for $w_{\theta_0}^*$. If $\hat{w}_n$ is a consistent estimator for $w_{\theta_0}^*$, then the STR given by $\kappa(\hat{\theta}_n + \hat{w}_n/\sqrt{n})$ is LAM. To ensure the consistency of $\hat{w}_n$, the estimators of unknown quantities must satisfy the appropriate conditions. In what follows, the study states the required conditions, defines our proposed STR formally, and then shows that the STR is LAM with respect to RIMR. First, estimate the law of $\mathbb{G}_0$ by that of the estimator $(\mathbf{Z}_n,\mathbf{W}_n) \mapsto \hat{\mathbb{G}}_n \in \mathbb{B}$, where $\mathbf{W}_n$ is independent of $\mathbf{Z}_n$. This estimator is particularly constructed via the bootstrap or simulation. For the bootstrap, let $\mathbf{W}_n = (W_{1},\dots,W_{n})$ be the nonparametric bootstrap weights independent of $\mathbf{Z}_n$, and let $(\mathbf{Z}_n,\mathbf{W}_n) \mapsto \hat{\theta}_n^*$ be the bootstrap version of the best regular estimator $\hat{\theta}_n$. Then, use $\hat{\mathbb{G}}_n = \sqrt{n}(\hat{\theta}_n^* - \hat{\theta}_n)$ to approximate the law of $\mathbb{G}_0$. For the simulation, presume the law of $\mathbb{G}_0$ belongs to a known class of distributions. For example, $\mathbb{G}_0 \sim N(0,\Sigma_{\theta_0})$ when $\mathbb{B} = \mathbb{R}^{d_\theta}$. Then, given a consistent estimator $\hat{\Sigma}_n$ for $\Sigma_{\theta_0}$, use $ \hat{\mathbb{G}}_n = \mathbf{W}_n \sim N(0,\hat{\Sigma}_n)$ to approximate the law of $\mathbb{G}_0$. In general, the estimator $\hat{\mathbb{G}}_n$ must satisfy (ref) below. Given this assumption, $\hat{\mathbb{G}}_n \leadsto \mathbb{G}_0$ conditionally on $Z^n$; and hence, it is reasonable to approximate the law of $\mathbb{G}_0$ by $\hat{\mathbb{G}}_n$ Vaart1996.
Second, approximate the support $S(\mathbb{G}_0)$ by an expanding sequence of compact sets. The rationale of this strategy is based on the tightness of $\mathbb{G}_0$. Specifically, $\mathbb{G}_0$ has a tight Borel law if and only if its support is a $\sigma$-compact set (i.e., a countable union of compact sets). Thus, the support of $\mathbb{G}_0$ can be approximated by an expanding sequence of compact sets $S_n \subset \mathbb{B}$. For the property of this set sequence, assume the following.
As stated in (ref), the expanding sequence $\{S_n\}$ is not necessary to be known ex-ante. Rather, it is sufficient to have an estimator $\hat{S}_n$ whose Hausdorff distance from $S_n$ goes to zero in probability. A typical candidate of $\hat{S}_n$ and $S_n$ for the case when $\mathbb{B} = \mathbb{R}^{d_\theta}$ is constructed as follows. Let $\hat{\Sigma}_n$ be a $\sqrt{n}$-consistent estimator for the efficiency bound $\Sigma_{\theta_0}$ of estimators for $\theta_0$. Moreover, let $\hat{\sigma}_j$ and $\sigma_j$ be the $j$-th column vector of $\hat{\Sigma}_n$ and $\Sigma_{\theta_0}$, respectively. Then, the following expanding sequence of compact sets usually fulfills the requirement:
where $\lambda_n = o(\sqrt{n})$. Furthermore, in some examples such as (ref), $\Sigma_{\theta_0} = \mathbb{R}^{d_\theta}$. Then, it suffices to take as $S_n$ the closed ball with center $0$ and radius $\lambda_n$. For the choice of $S_n$ and $\hat{S}_n$ when $\mathbb{B}$ is a Banach space, see Section 5.2.2 in Ponomarev2022. Third, this study postulates that the estimators $\hat{\tau}_{U,n}'$ and $\hat{\tau}_{L,n}'$ for $\tau_{U,\theta_0}'$ and $\tau_{L,\theta_0}'$ satisfying the following assumption are available.
(ref).(i) requires some uniform consistency of $\hat{\tau}_{j,n}'$. In many applications, estimators with stronger properties such as $\hat{\tau}_{j,n}(b) = \tau_{j,\theta_0}'(b)$ with probability approaching 1 are often available. See Remarks 3.3 and 3.4 and Section S.4 in the supplementary material of Fang2019 for more details. Usually, one can construct such estimators from the analytical expression of $\tau_{j,\theta_0}'$ or numerical method proposed in Hong2018. It is now appropriate to define the proposed STR. Given the estimators $\hat{\tau}_{U,n}'$ and $\hat{\tau}_{L,n}'$ for $\tau_{U,\theta_0}'$ and $\tau_{L,\theta_0}'$, define the estimator $\hat{\kappa}_n'$ for $\kappa_{\theta_0}'$ by
for $b \in \mathbb{B}$. Here, $\hat{m}_{y,n}$ is the estimator of the directional derivative $m_y'$ of $\max\{x,0\}$ at $y \in \mathbb{R}$, and, for any $x \in \mathbb{R}$, it is given by
with $\epsilon_n > 0$ such that $\epsilon_n \downarrow 0$ and $\sqrt{n}\epsilon_n \uparrow \infty$. Let $\{K_m\}$ be an expanding sequence of compact sets in $\mathbb{B}$. For a sufficiently large constant $M > 0$, let
Then, define the STR by
where $\hat{\theta}_n$ is the best regular estimator for $\theta_0$. The following theorem gives an upper bound of local asymptotic maximum RIMR of $\delta_n^{\text{LAM}}$.
The right-hand sides of (ref) and (ref) coincide. Therefore, combining (ref), $\delta_n^{\operatorname{LAM}}$ is LAM STR.
This simulation study examines the finite-sample performance of the proposed LAM STR in the context of (ref). In the setup in (ref), consider a binary outcome (i.e., $\mathcal{Y} = \{0,1\}$) and set the true parameter to
In this case, $\tau_L(\theta_0) = 0$ and $\tau_U(\theta_0) = 1/2$. Although the boundary functions $\tau_L(\theta)$ and $\tau_U(\theta)$ are differentiable at $\theta_0$, the optimal treatment rule $\kappa(\theta)$ is only directionally differentiable at $\theta_0$. Focus on a specific set of perturbed data generating processes characterized by
for $h \in I \coloneqq \{-2,-1.95,-1.90,\dots,2\}$. When $h = 0$, $\theta_n(h)$ corresponds to the true parameter value $\theta_0$. As $|h|$ gets larger, $\theta_n(h)$ deviates more from the nondifferentiable point. The goal here is to compare the maximum RIMR,
between the proposed LAM STR and plug-in STR. This study describes details for constructing the LAM STR. Given a sample $\{(S_iY_i,S_iD_i,S_i)\}_{i=1}^n$, $\theta$ is estimated by the best regular estimator $\hat{\theta}_n$ described in (ref) on page (ref). This estimator satisfies $\sqrt{n}(\hat{\theta}_n - \theta_n(h)) \overset{h}{\leadsto} \mathbb{G}_0 \sim N(0,\Sigma_{\theta_0})$ for any $h$. Given a $\sqrt{n}$ consistent estimator $\hat{\Sigma}_n$ for $\Sigma_{\theta_0}$, estimate the law of $\mathbb{G}_0$ by drawing $G_l,l=1,\dots,L$ independently from $N(0,\hat{\Sigma}_n)$, where $L = 1000$. The Hadamard derivatives, $\tau_{L,\theta_0}'(b)$ and $\tau_{U,\theta_0}'(b)$, are estimated by $\hat{\tau}_{L,n}'(b) = \dparen{ \frac{\partial \tau_L}{\partial \theta}|_{\theta = \hat{\theta}_n} }^\intercal b$ and $\hat{\tau}_{U,n}'(b) = \dparen{ \frac{\partial \tau_U}{\partial \theta}|_{\theta = \hat{\theta}_n} }^\intercal b$, respectively. In the estimator for the directional derivative of $\max\{x,0\}$, set $\epsilon_n = n^{1/3}$. This study sets the truncation constant $M$ for the loss function to be $1000$. As the efficiency bound $\Sigma_{\theta_0}$ is nonsingular, the support of $\mathbb{G}_0$ corresponds to $\mathbb{R}^3$. It implies that the support can be approximated by expanding cubes. Hence, set $\hat{S}_n = S_n = [-\lambda_n,\lambda_n]^3$, where $\lambda_n = n^{1/3}$. The compact set $K_m$ over which the optimal adjustment term is searched is set to $[-2,2]^3$. Given these specifications, solve the next minimization problem to obtain the data-dependent adjustment term $\hat{w}_n$:
For each $h \in H$, draw $J = 5000$ independent samples of data $\mathbf{Z}_n^{(j)},j = 1,\dots,J$. Each observation $Z_i^{(j)}$ in data $\mathbf{Z}_n^{(j)} = (Z_1^{(j)},\dots,Z_{n}^{(j)})$ comprises $(S_iY_i,S_iD_i,S_i)$ drawn from the data generating process characterized by $\theta_n(h)$ and $\Pr(D=1|S=1) = 1/2$. Based on this data, obtain the best regular estimate $\hat{\theta}_n^{(j)}$ and data-dependent adjustment term $\hat{w}_n^{(j)}$. Given the collection, $\hat{\theta}_n^{(1)},\dots,\hat{\theta}_n^{(J)}$ and $\hat{w}_n^{(1)},\dots,\hat{w}_n^{(J)}$, estimate $\mathbb{E}_{n,h}[\delta_n^{\operatorname{LAM}}]$ by $J^{-1}\sum_{j=1}^J \kappa(\hat{\theta}_n^{(j)} + \hat{w}_n^{(j)}/\sqrt{n})$. Finally, the RIMR is calculated following equation (ref). Similarly, calculate the RIMR of the plug-in STR by ignoring the data-dependent adjustment term. The size $n$ of a sample is set to 300 or 1000. (ref) plots the relation between the local parameter $h$ and RIMR by the size of a sample. For both STRs, the RIMR is maximized at the dgp such that the optimal treatment rule $\kappa(\theta)$ is nondifferentiable (i.e., at $h = 0$). Importantly, the maximized RIMR of the LAM STR is lower than that of the plug-in STR. The LAM STR is designed to minimize the (asymptotic) maximum RIMR, and thus this result is as expected. Conversely when $h \neq 0$, $\kappa(\theta)$ is differentiable at $\theta_n(h)$. Nevertheless, the RIMR is relatively large in a neighborhood of zero. The LAM STR outperforms the plug-in STR in such a neighborhood.
This study analyzed the statistical treatment choice problem under the partial identification of the ATE. Solving the finite-sample problem exactly is difficult for general data generating processes. Instead, this study conducted the local asymptotic analysis and derived the LAM STRs. Specifically, the LAM STR takes the form of
where $\hat{\theta}_n$ is the efficient estimator for the true value $\theta_0$ of parameters, $w_{\theta_0}^*$ is the optimal adjustment term that can depend on $\theta_0$, and $\kappa$ is the function that outputs the optimal treatment rule given a value of the parameter. The adjustment term $w_{\theta_0}^*$ generally differs from zero. This implies that the LAM STR differs from the plug-in STR $\kappa(\hat{\theta}_n)$. In practice, the true parameter value $\theta_0$ is unknown; hence, so is the optimal adjustment term $w_{\theta_0}^*$. To alleviate this challenge, the study developed the data-dependent adjustment term $\hat{w}_n$, showing that the STR $\kappa(\hat{\theta}_n + \hat{w}_n/\sqrt{n})$ remained LAM.