EconBase
← Back to paper

Orthogonal Policy Learning Under Ambiguity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

129,280 characters · 11 sections · 148 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Orthogonal Policy Learning Under Ambiguity

}

center[center omitted — 146 chars of source]
abstractThis paper studies the problem of estimating individualized treatment rules when treatment effects are partially identified, as it is often the case with observational data. By drawing connections between the treatment assignment problem and classical decision theory, we characterize several notions of optimal treatment policies in the presence of partial identification. The proposed framework allows to incorporate user-defined constraints on the policies, such as restrictions for transparency or interpretability, while also ensuring computational feasibility. We show that partial identification leads to a novel statistical learning problem with risk directionally -- but not fully -- differentiable with respect to an infinite-dimensional nuisance component. We propose an estimation procedure that ensures Neyman-orthogonality with respect to the nuisance component and provide statistical guarantees that depend on the amount of concentration around the points of non-differentiability in the data-generating process. The proposed method is illustrated using data from the Job Partnership Training Act study.

Introduction

The problem of choosing an optimal treatment assignment based on data is ubiquitous in economics and other fields, including medicine and marketing. Individuals often display heterogeneous responses to the same treatment. Decision-makers in policy and industry are therefore interested in leveraging the growing availability of rich granular data to tailor treatment assignment to individuals based on their characteristics. As a result, a fast-growing literature has emerged focused on developing procedures for estimation of individualized treatment rules. While a variety of approaches have been recently established, these typically assume that the available data allow to provide credible point estimates for the effect of the treatment, that is treatment effects are point identified. While of important stylized value, this assumption is often hard to justify in many empirical settings. For example, economists have long been aware that popular quasi-experimental and observational research designs, such as instrumental variables (IV), allow to point identify treatment effects only for specific sub-populations imbensangrist1994. Even in randomized control trials, point identification of the treatment effects is often precluded due to non-random attrition, e.g.\@ when participants dropout from a program or the researcher is denied information on the outcome variable lee2009. In such settings, the data may only provide partial knowledge about the treatment response in the form of credible bounds, i.e.\@ the treatment effects are partially identified. As a result, the decision-maker may have ambiguous evidence on whether a candidate policy should be preferred to another, so that only a partial ordering of policies can be deduced in general. While informative from a scientific perspective, a partial ordering of policies is unsatisfying when the ultimate goal of the analysis is to select a single policy to be implemented in the real world. In this scenario, a decision-maker has to confront two sources of ambiguity. The first source concerns ambiguous knowledge of the treatment response $\tau$ conditional on knowledge of distribution of the data $P$, due to partial identification. The second source is the lack of knowledge of the distribution $P$, which must be estimated from the data.

In this paper, we develop methods to handle both sources of ambiguity within the framework of “empirical welfare maximization" kitagawatetenov2018, also referred to as “policy learning" atheywager2021. This approach considers treatment policies that are exogenously constrained to have low complexity in terms of Vapnik-Chervonenkis (VC) dimension. This encompasses many practical settings of interest, as policies often have to satisfy requirements imposed for institutional or practical reasons, such as fairness, budget or interpretability. The empirical welfare maximization (EWM) method selects the optimal policy as the maximizer of the empirical analogue of the population welfare, formulated as the average of the individual outcomes in the target population. The EWM estimation procedure has the convenient structure of an empirical risk minimization problem, which is exploited by kitagawatetenov2018 and atheywager2021 to study its statistical properties.

We extend the EWM framework to settings with partial identification by making several contributions. First, we study the problem of assigning treatment under partial identification at the population level (i.e.\@ where the distribution of the data $P$ is known) from a general perspective. In particular, we show how classic optimality criteria for decision under ambiguity, such as minimax risk and minimax regret, can be applied in the context of welfare maximization. Our unified framework accommodates different attitudes towards ambiguity and a wide range of popular identification assumptions, including manski1990 and manskiPepper bounds. Our analysis delivers several notions of optimal treatment policies, which we refer to as ambiguity-robust: they are “robust" in the sense that each of them delivers a notion of single optimal policy in the presence of partial identification, while they all reduce to the same optimal treatment assignment in the special case of point-identification. As part of this analysis, we establish general conditions on the identification sets under which the treatment assignment problem can be expressed in a simplified form, leading to computationally tractable sample analogues. In particular, we show that all ambiguity-robust policies can be represented as maximizers of a “surrogate" welfare, in which identification bounds are combined to form a proxy for the partially identified CATE. The surrogate welfare depends on several nuisance components, and its specific form is determined by the identification assumptions and attitude towards ambiguity held by the decision-maker.

We then propose an algorithm for computing the estimated ambiguity-robust policy and provide statistical guarantees on its performance in terms of the regret convergence of the surrogate welfare. Similarly to atheywager2021 and fosterSyrgkanis, our procedure leverages insights from the literature on double/de-biased machine learning chernozhukov2018 by making use of Neyman-orthogonalized estimates of the surrogate welfare. This, coupled with sample-splitting, allows us to guarantee fast rates of convergence for the estimated ambiguity-robust optimal policy while imposing minimal requirements on the estimation of the nuisance components. One unique feature of the partially identified setting studied in this paper is the restricted degree of smoothness enjoyed by the welfare criterion. In particular, we show that popular choices of identification assumptions and optimality criteria for choice under ambiguity lead to surrogate welfare criteria that are only directionally differentiable with respect to the data-generating process. We highlight the importance of this feature for the problem at hand and develop new theoretical results showing how the extent of non-differentiability in the data-generating process affects the statistical properties of the learning procedure. To the best of our knowledge, we are the first to investigate the role of non-differentiabilities in the context of semiparametric statistical learning problems. Our results are therefore of independent interest and may be relevant beyond the treatment assignment problem of this paper.

Finally, we apply the proposed method to experimental data from the Job Training Partnership Act study, a dataset that has been extensively used to study the effect of subsidized job training on labor market outcomes. We study the optimal participation of workers into the job training programme based on their education and previous earnings, and show that allowing for partial identification delivers substantially different programme participation policies compared to existing methods that assume point-identification.

Related literature

The results of this paper contribute to the recent literature on EWM methods, e.g.\@ kitagawatetenov2018, atheywager2021, mbakopTabord, viviano2021, sun2021, and more broadly to the literature studying statistical treatment choice, including manski2004, dehejia2005, hiranoPorter2009ECMA, stoye2009, chamberlain2011, christensenMoonSchor2020, kitagawaLeeQiu2022.\footnote{See also hiranoPorter2020HoE and referenes therein.}

kitagawatetenov2018 introduced the EWM method and provided theoretical results showing its optimality when implemented with experimental data. atheywager2021 leverage insights from the recent literature on orthogonal machine learning chernozhukov2018 and propose doubly-robust estimation of the treatment effect which leads to optimal learning rates even with observational data. We build on their work by adopting Neyman-orthogonal estimates while we relax the fundamental assumption that treatment effects are point identified. cuiTchetgen2021 also develop procedures for learning optimal treatments rules with instrumental variables but consider unconstrained policy classes. Similarly to atheywager2021, they ensure point-identification of treatment response by restricting their analysis to the effect on compliers.

kasy2016, han2019 and Byambadalai2022 provide methods for comparing policies in the presence of covariates and partial identification of treatment effects. The focus of their work is on characterizing the partial ordering of policies in terms of their associated welfare rather than resolving the ambiguity and estimating an optimal treatment rule.

In a series of papers, manski2009 (manski2009, manski2010, manski2011) studies the problem of a social planner who must choose treatment for a population under partial knowledge of the treatment response in the absence of covariates. He shows that when the sign of the treatment effect is ambiguous, the minimax regret criterion leads to policies that randomize treatment in the population. While our study of the population problem is inspired by Manski's work in this area, the focus of our paper is on deterministic rules assigning individualized treatment, i.e.\@ based on (potentially continuous) covariates. STOYE2012138, ishiharaKitagawa and yata consider treatment assignment under partial identification from a finite-sample minimax perspective, while christensenMoonSchor2020 adopt a local-asymptotic approach. However, these works do not consider individualization of the treatment assignment.

More closely related to our work is kallus2018, who extend the EWM framework to learn an optimal policy in the presence of partially identified treatment effects under violations of unconfoundedness. In particular, they target welfare improvement with respect to a baseline pre-existing policy and consider partial-identification of the welfare criterion within Rosenbaum's sensitivity model rosenbaum1987. christensenAdjaho and kido2022 examine policies with maximin welfare guarantees when the target population lies in a Wasserstein neighborhood of the experimental population. The identification assumptions (and associated estimation procedures) considered in these papers are distinct and do not nest those covered by our framework. As a result, our contributions are complementary to these works.

russell2020 considers estimation of the optimal policy under partial identification within a “probably approximately correct" learning framework valiant. His proposed procedure has the advantage of side-stepping direct estimation of the identified set, and can be applied in the context of incomplete models for which the identification bounds cannot typically be obtained in closed form. However, in the context of the identification assumptions considered in this paper (e.g.\@ Manski bounds), the theoretical results in russell2020 require that the covariates have discrete support. On the other hand, our proposed procedure requires computation of the identification bounds in closed form but accommodates continuous covariates.

In independent work, Pu_2021 study policy learning under ambiguity from a classification perspective and derive an optimal policy which coincides with one notion of ambiguity-robust policy studied in this paper. However, our estimation procedure crucially differs from theirs for the use of Neyman-orthogonalization which, combined with a refined proof-strategy that accounts for the lack of full-differentiability in the welfare criterion, allows us to guarantee considerably faster rates of convergence. In this sense, our results extend and improve on those in Pu_2021.

Finally, we contribute to a body of literature dealing with estimation and inference for directionally-differentiable functionals. hiranoPorter2012ECMA show that if a target estimand is not differentiable in the parameters of the data distributions, then no asymptotically unbiased or regular estimator exists. ponomarev studies efficient estimation of directionally differentiable functionals from a local minimax perspective. fangSantos2019 and kitagawaEtAl2020 provide inference results for directionally differentiable functions from a frequentist and Bayesian perspective, respectively. Also motivated by partial identification, christensenMoonSchor2020 consider estimation of treatment rules when the welfare criterion is only directionally differentiable with respect to a finite-dimensional nuisance component. Our framework instead involves infinite-dimensional nuisance components and therefore our analysis must account for the lack of differentiability with novel theoretical results that substantially differ from christensenMoonSchor2020.

The rest of this paper is organized as follows. Section (ref) introduces the setup. Section (ref) presents several notions of ambiguity-robust optimal policies. Section (ref) presents the proposed estimation procedure for the ambiguity-robust optimal policy. Section (ref) provides statistical guarantees for the estimated optimal policy. Section (ref) presents an empirical illustration based on the Job Training Partnership Act Study. Section (ref) concludes the paper. Proofs and extensions are given in the Appendix.

\\ \noindentNotation. Throughout the paper, for $d\in \mathbb{N}$, let $\mathbb{R}^d$ denote the Euclidean space, with $\|\cdot \|_p$ and $\langle \cdot, \cdot \rangle$ being the usual $\ell_p$-norm and inner product, respectively. For two vectors $x \in \mathbb{R}^p$ and $y \in \mathbb{R}^q$, $x\subset y$ means that $x$ is a sub-vector of $y$. For a symmetric matrix $A$, $\lambda_{{\rm max}}(A)$ denotes its largest eigenvalue. Unless otherwise stated, the expectation $\mathbb{E}[\cdot ]$, probability $\mathbb{P}(\cdot)$, and variance ${\rm Var}(\cdot)$ operators will be taken with respect to the underlying distribution of observables $P$. Given a random variable $Z\in \mathcal{Z}$ with $\mathcal{Z}\subseteq \mathbb{R}^d$, the associated probability measure $P_Z$, and a function $f\,:\, \mathcal{Z} \to \mathcal{W}$ with $\mathcal{W}\subseteq\mathbb{R}^q$, we define $\norm{f}_{Lp(P_Z)} = \left( \mathbb{E}_{P_Z} \left[ \|f(Z)\|^p_p \right]\right)^{1/p}$ for $p\in(0\,\infty)$. We extend this definition to $p=\infty$ in the natural way. For a sequence of real numbers $x_n$ and $y_n$, $x_n=o(y_n)$ and $x_n=O(y_n)$ mean, respectively, that $x_n/y_n\to 0$ and $x_n\leq C y_n$ for some constant $C$ as $n\to \infty$. For real numbers $a,b$, $a\lesssim b$ means that there exists a constant $C$ such that $a\leq C b$. For a positive real number $a$, $\lfloor a \rfloor$ denotes its nearest smallest integer. The notation $\rightarrow_p$ denotes convergence in probability.

Setup

Let $Y_i\in \mathbb{R}$ be an outcome measuring utility, $D_i\in\{0,1\}$ a binary treatment, $X_i \in \mathcal{X}\subseteq\mathbb{R}^{k_x}$ a set of pre-treatment covariates for an individual $i$ from an i.i.d.\@ population of interest. We use standard notation to define the potential outcomes $Y_i(0), Y_i(1)$. The conditional average treatment effect (CATE) $\tau: \mathcal{X}\to \mathbb{R}$ is then defined as

align*[align* omitted — 88 chars of source]

where the expectation is taken with respect to the distribution of the population, and we will henceforth suppress the $i$-subscript for convenience. The decision-maker (DM) is interested in choosing a deterministic treatment assignment rule (or policy) $\pi:\mathcal{X}\to \{0,1\}$, which maps from the support of individual pre-treatment covariates to the binary decision “treat" ($\pi(x)=1$) or “do not treat" ($\pi(x)=0$). Following manski2004, we define the utilitarian social welfare associated with a policy $\pi$ and a given configuration of the expected potential outcomes $y_0(\cdot),y_1(\cdot)$ as

equation[equation omitted — 261 chars of source]

where $I_\tau(\pi)$ represents the average impact of policy $\pi$. The optimal policy for a given configuration of the CATE function is the one that maximizes the associated welfare:

align[align omitted — 156 chars of source]

where $\Pi$ is a family of candidate policies.\footnote{Throughout the paper, we will assume that the maximization problem in (ref) has at least one solution. If multiple solutions exist, the DM is assumed to arbitrarily pick $\pi^*$ from the set of maximizers.} The DM has knowledge of the CATE through the distribution $P\in \mathcal{P}$ of observable random variables $W$, where $(Y,D,X)\subseteq W$. In particular, we denote $\mathcal{T}(P)$ the set of plausible CATE functions associated with a certain distribution of observables. When the DM has perfect knowledge of $P$ and $\mathcal{T}(P)$ is a singleton, i.e.\@ $\tau$ is point-identified, she can obtain $\pi^*$ by solving (ref).

Suppose now that $\mathcal{T}(P)$ is a non-singleton set, i.e.\@ $\tau$ is partially identified. In that case, even under perfect knowledge of $P$, there exists a set of plausible values for the impact $I_\tau(\pi)$ of a candidate policy $\pi$. Notice that partial identification of the CATE does not necessarily imply that the DM cannot obtain the optimal policy $\pi^*$. In particular, it is easy to see that under point-identification of the CATE one has $\pi^*(x)=\mathbbm{1}\{\tau(x)\geq 0\}$ when the class of candidate policies $\Pi$ is unrestricted, so that identification of the sign of the CATE is sufficient to obtain the optimal policy.\footnote{cuiTchetgen2021 study a case in which sole point-identification of the sign of the CATE via an instrumental variable allows to obtain the optimal policy.} However, the unrestricted policy class has limited relevance in many practical settings. For example, the policy space $\Pi$ may be exogenously constrained for institutional reasons, e.g.\@ as policies may be required to satisfy specific requirements for budget, fairness or interpretability. While the DM may still hope that his specification of $\Pi$ contains the first-best policy $\mathbbm{1}\{\tau(X)\geq 0\} $, it is useful to interpret $\pi^*$ as the “best-in-class" policy for the chosen class $\Pi$, when this does not contain the first-best. When $\Pi$ is constrained, the DM is not able to obtain $\pi^*$ in general without full knowledge of the CATE, although a partial ordering of policies can still be deduced kasy2016,han2019,Byambadalai2022.

Under partial identification, the DM therefore faces two sources of ambiguity. First, she does not know the distribution $P$. However, we assume that she has access to a random sample $(W_i)_{i=1,\dots,n}$ from which she can learn about $P$. Second, she does not have knowledge about $\tau$ within the identified-set $\mathcal{T}(P)$, even under perfect knowledge of $P$. The broad objective of this paper is to provide a framework that allows the DM to handle both sources of ambiguity. We will approach the problem in two steps. First, we will study the decision problem faced by the DM under perfect knowledge of $P$. In particular, we will handle the ambiguity arising from partial identification of $\tau$ using well-known optimality criteria for decision under ambiguity. Each of the optimality criteria we consider will deliver a corresponding notion of optimal policy, which we call “ambiguity-robust". The ambiguity-robust optimal policy is a unique treatment assignment rule that is preferred to all other policies in $\Pi$ according to preferences of the DM, and that coincides with the usual notion of optimal policy $\pi^*$ in (ref) in the special case of point identification of the CATE. In the next section, we study several notions of ambiguity-robust optimal policy.

In the second part of our analysis, we study how to handle the ambiguity in $P$ by showing how the random sample $(W_i)_{i=1,\dots,n}$ can be used to obtain an estimate $\widehat \pi_n$ for the ambiguity-robust optimal policy. The estimation procedure and the associated statistical guarantees are presented in Section (ref) and (ref), respectively.

remarkUnrestricted policy classes may also be precluded for practical reasons related to the estimation of the optimal policy. For example, the researcher may need to condition on a large number of covariates $X$ for identification of the treatment effects, but only be interested in assigning treatment based on a restricted set of the covariates $\widetilde X \subset X$ (e.g.\@ because she may not observe the full set of covariates when assigning treatment to new individuals from the population). In that case, a practical way to side-step computation of an estimate for the lower-dimensional CATE, $\mathbb{E}[Y_i(1)|\widetilde{X}_i=\widetilde{x}]-\mathbb{E}[Y_i(0)|\widetilde{X}_i=\widetilde{x} ]$, is to impose restrictions directly on the policy class $\Pi$ and estimate the optimal policy based on the estimated higher-dimensional CATE via the sample analogue of (ref).
commentSuppose now that the CATE function $\tau(\cdot)$ is not point identified but only identified to be in a set $\mathcal{T}$. In that case, the welfare associated with a single policy is also not point identified and the notion of optimality in (ref) only allows to characterize a partial ordering of the policies kasy2016,han2019,Byambadalai2022, meaning that only a set of potentially optimal policies can be identified. Partial identification of the CATE function arises in many contexts of practical relevance such as missing outcomes and, most notably, instrumental variables estimation. atheywager2021 and cuiTchetgen2021 consider the problem of policy learning with instrumental variables, but rule out the presence of unobservable treatment effect heterogeneity in order to guarantee point-identification of the CATE function. While one might be unwilling to make such strong an assumption, identifying a set of optimal policies might also not be satisfying when the researcher's objective is to recommend a single policy to be implemented in a real-world setting. Achieving this objective requires adopting a criterion of optimality that delivers the notion of a single optimal policy under the ambiguity that arises from partial identification of the CATE. The analysis of next subsection introduces several notions of such policies, which we call ambiguity-robust optimal policies. Let $Y_i\in \mathbb{R}$ be an outcome measuring utility, $D_i\in\{0,1\}$ a binary treatment, $X_i \in \mathcal{X}\subseteq\mathbb{R}^{d_x}$ a set of observable pre-treatment covariates for an individual $i$, and denote $Y_i(0), Y_i(1)$ the usual potential outcomes. The conditional average treatment effect (CATE) $\tau: \mathcal{X}\to \mathbb{R}$ at $x\in \mathcal{X}$ is defined as \begin{align*} \tau(x)= y_1(x)-y_0(x),\quad y_d(x)=\mathbb{E}[Y_i(d)|X_i=x ],\, d=0,1, \end{align*} where the expectation is take with respect to the distribution of an i.i.d.\@ population of interest, and we will henceforth suppress the $i$-subscript for convenience. We assume that a policy-maker can choose $\pi:\mathcal{X}\to \{0,1\}$ a deterministic treatment assignment rule (“policy") that maps from the space of observed covariates to the binary decision “treat" ($\pi(x)=1$) or “do not treat" ($\pi(x)=0$). Following manski2004, we define the utilitarian social welfare associated with a policy $\pi$ and a given configuration of the expected potential outcomes $y_0(\cdot),y_1(\cdot)$ as \begin{equation} \begin{aligned} W_{y_0,y_1}(\pi)&= \mathbb{E}_{P_X}\left[y_1(X)\cdot\pi(X) +y_0(X)\cdot(1-\pi(X))\right]\\ &=\underbrace{\mathbb{E}_{P_X}\left[\pi(X)\cdot\tau(X)\right]}_{=:I_\tau(\pi)} + E_{P_X}\left[y_0(X)\right]. \end{aligned} \end{equation} where $I_\tau(\pi)$ represents the impact of policy $\pi$. The optimal policy for a given configuration of the CATE function is the one that maximizes the associated welfare: \begin{align} \pi^*= \operatorname*{argmax}_{\pi \in \Pi }W_{y_0,y_1}(\pi)=\operatorname*{argmax}_{\pi \in \Pi }I_{\tau}(\pi), \end{align} where $\Pi$ is a family of candidate policies. Suppose now that the CATE function $\tau(\cdot)$ is not point identified but only identified to be in a set $\mathcal{T}$. In that case, the welfare associated with a single policy is also not point identified and the notion of optimality in (ref) only allows to characterize a partial ordering of the policies kasy2016,han2019,Byambadalai2022, meaning that only a set of potentially optimal policies can be identified. Partial identification of the CATE function arises in many contexts of practical relevance such as missing outcomes and, most notably, instrumental variables estimation. atheywager2021 and cuiTchetgen2021 consider the problem of policy learning with instrumental variables, but rule out the presence of unobservable treatment effect heterogeneity in order to guarantee point-identification of the CATE function. While one might be unwilling to make such strong an assumption, identifying a set of optimal policies might also not be satisfying when the researcher's objective is to recommend a single policy to be implemented in a real-world setting. Achieving this objective requires adopting a criterion of optimality that delivers the notion of a single optimal policy under the ambiguity that arises from partial identification of the CATE. The analysis of next subsection introduces several notions of such policies, which we call ambiguity-robust optimal policies.
commentFollowing kitagawatetenov2018, we define the population welfare associated with a policy $\pi$ and a given configuration of the CATEs as \begin{align} W_\tau(\pi)=\mathbb{E}_{P_X}\left[\pi(X)\cdot\tau(X)\right]. \end{align} and the associate optimal policy as its maximizer \begin{align} \pi^*= \operatorname*{argmax}_{\pi \in \Pi }W_\tau(\pi), \end{align} where $\Pi$ is a family of policies of finite VC-dimension. If the CATE function $\tau(x)$ is not point identified but only identified to be in a set $\mathcal{T}$ instead, then the welfare is also not point identified and the notion of optimality in ((ref)) leads to a partial ordering of the policies (see, e.g., kasy2016, kasy2016, and Byambadalai2022, Byambadalai2022). Partial identification of the CATE function arises in many contexts of practical relevance such as missing outcomes and, most notably, instrumental variables estimation. atheywager2021 and cuiTchetgen2021 consider the problem of policy learning with instrumental variables, but rule out the presence of unobservable treatment effect heterogeneity in order to guarantee point-identification of the CATE function. While such strong an assumption might be unrealistic in practical settings, identifying a set of optimal policies might also not be satisfying when the researcher's objective is to recommend a single policy to be implemented in a real-world setting. Achieving this objective therefore requires adopting a criterion of optimality that delivers a notion of single optimal policy in the presence of ambiguity in the treatment effects.

Ambiguity-robust optimal policies

The study of decision under ambiguity has a long tradition in decision theory and has received considerable attention in the context of treatment assignment problems (see manski2011, manski2011, for a review). In this section we review some classical optimality criteria for decision under ambiguity and study how they can be applied in the context of the treatment assignment problem at hand, leading to several notions of ambiguity-robust optimal policy.

A well-known optimality criterion for decision under ambiguity is minimax risk wald_1950. In the context of our treatment assignment problem we can interpret welfare as negative risk, and this criterion leads to the optimal maximin welfare policy

align[align omitted — 160 chars of source]

where $\mathcal{Y}(P)$ is the ambiguity set for $\left(y_0(\cdot),y_1(\cdot)\right)$ identified from the distribution $P$ of observables random variables. The optimal maximin welfare policy maximizes the lowest possible welfare under any configuration of the expected potential outcome functions in the identified set $\mathcal{Y}(P)$. An alternative application of minimax risk optimality in the context of treatment assignment is maximin impact, leading to the optimal policy

align[align omitted — 139 chars of source]

where $\mathcal{T}(P)$ denotes the ambiguity set for the CATE function. The optimal maximin impact policy maximizes the lowest possible impact under any configuration of the CATE in the identified set $\mathcal{T}(P)$. Notice that the minimax welfare criterion reflects an extreme degree of pessimism with regards to outcomes associated with both treatment and non-treatment scenarios; on the other hand, the minimax impact criterion reflects an extreme degree of pessimism with regards to the impact of the policy, thus directly raising the threshold for treatment.\footnote{In the empirical application of Section (ref), both minimax welfare and minimax impact criteria result in $\pi(x)=0$ for the entire population.} Despite its intuitive appeal, minimax optimality has been criticised for being too conservative and often delivering decisions that are especially sensitive to changes in the ambiguity set.\footnote{In his classic textbook, Berger goes as far as saying that “In actually making decisions, the use of the minimax principle is definitely suspect." Berger.}

An alternative criterion that alleviates some of these concerns is minimax regret, with corresponding optimal policy

equation[equation omitted — 447 chars of source]

The minimax regret criterion delivers a policy that minimizes the largest possible distance between attained welfare and the highest level of welfare attainable by the “oracle" treatment rule $\pi^*=\mathbbm{I}\left\{\tau(x)\geq0\right\}$ that has knowledge of the true $\tau$. Minimax regret optimality has been advocated by manski2004 for its balanced consideration of the possible states of nature and for delivering more “reasonable" decisions rules in practice, compared to minimax risk approaches.

remarkAn alternative version of the minimax regret criterion is minimax regret with respect to the welfare attained by the best-in-class policy in $\Pi$, resulting in the objective \begin{align} \pi^*_{MMR2} = \operatorname*{argmin}_{\pi \in \Pi } \max_{\tau \in \mathcal{T}(P) }\left[\left( \max_{\pi \in \Pi}I_\tau(\pi)\right) - I_\tau(\pi)\right]. \end{align} While these two versions of the minimax regret criterion can be expected to enjoy similar properties, the first version we have considered is considerably more tractable. In fact, the innermost maximization in ((ref)) has the closed-form solution $\max_{\pi\, : \, \mathcal{X}\to\{0,1\}}W_\tau(\pi)= \mathbb{E}_{P_X}\left[\max\left\{\tau(X),0\right\}\right]$. As we show in Proposition (ref) below, this allows to more explicitly characterize the properties of the optimization problem and the resulting optimal policy, as well as reduce the computational burden in solving the empirical analogue of the problem. For this reason we will focus on the version in ((ref)) of the criterion. We also note that whenever the class $\Pi$ is “well-specified", in the sense that $\mathbbm{I}\left\{\tau(x)\geq0\right\}\in \Pi$ for all $\tau\in\mathcal{T}(P)$, the two optimality criteria are equivalent.

One critical drawback in the application of the optimality criteria just presented to the treatment assignment problem of this paper is that the optimal policies cannot be obtained in closed form. This is due to the form of (ref), (ref) and ((ref)) involving several nested optimizations whose solutions cannot be easily characterized at the current level of generality when $X$ includes continuously distributed covariates and $\Pi$ may be arbitrarily restricted, which are both primary cases of interest of this paper. To make progress, we impose the following restrictions on the ambiguity sets for the expected potential outcomes and CATE.

assumption[Rectangular identified set for $(y_0,y_1)$] The identified set for $(y_0,y_1)$ is rectangular, that is, $\mathcal{Y}$ is of the form \begin{equation*} \mathcal{Y} = \{ \left(y_0(\cdot),y_1(\cdot)\right) : \left(y_0(x),y_1(x)\right) \in \mathcal{Y}(x)\}, \end{equation*} where $\mathcal{Y}(x)$ is a compact subset of $\mathbb{R}^2$.
assumption[Rectangular identified set for $\tau$] The identified set for $\tau$ is rectangular, that is, $\mathcal{T}$ is of the form \begin{align*} \mathcal{T} = \{ \tau(\cdot) : \tau(x) \in [\tau(x), \overline{\tau}(x) ]\}, \end{align*} where $|\overline{\tau}(x)|<\infty$, $|\underline{\tau}(x)|<\infty$ for all $x\in\mathcal{X}$.

Assumptions (ref) and (ref) impose separation of the identified sets for the expected potential outcomes and CATE across the support of the covariates $\mathcal{X}$.\footnote{Notice that Assumption (ref) implies Assumption (ref), but not viceversa.} They are typically satisfied by identification schemes that do not impose shape restrictions on counterfactual outcomes with respect to the covariates $X_i$. These assumptions are widely adopted in the partial identification literature, and we refer the reader to Appendix B in kasy2016 for an extensive review of identification schemes that result in rectangular identified sets. Below we present three examples of identification schemes for the CATE that satisfy this assumption.

commentA well-known optimality criterion for decision under ambiguity is minimax loss. In the context of our treatment assignment problem, this criterion leads to the optimal maximin welfare policy \begin{align} \pi^*_{MMW} = \operatorname*{argmax}_{\pi \in \Pi } \min_{\tau \in \mathcal{T}(P) }W_\tau(\pi), \end{align} where $\mathcal{T}(P)$ is identified from the distribution $P$ of observables random variables and is assumed to be a subset of the space of bounded functions $B(\mathcal{X})$. The optimal maximin welfare policy maximizes the lowest possible welfare under any configuration of the CATE in the identified set $\mathcal{T}$. Despite its intuitive appeal, the minimax optimality criterion has been criticised for being too conservative and often delivering policies that are especially sensitive to changes in the ambiguity set $\mathcal{T}$.\footnote{In his classic textbook, Berger goes as far as saying that “In actually making decisions, the use of the minimax principle is definitely suspect." Berger. } An alternative criterion that alleviates some of these concerns is minimax welfare regret, with corresponding optimal policy \begin{align} \pi^*_{MMR}=\operatorname*{argmin}_{\pi \in \Pi } \max_{\tau \in \mathcal{T} }\left[ \left(\max_{\pi\, : \, \mathcal{X}\to\{0,1\}}W_\tau(\pi)\right) - W_\tau(\pi)\right]. \end{align} The minimax regret criterion delivers a policy that minimizes the largest possible distance between attained welfare and the highest possible welfare attained by the “oracle" treatment rule $\pi^*=\mathbbm{I}\left\{\tau(x)\geq0\right\}$ under knowledge of the true CATE configuration $\tau$. The minimax criterion has been advocated by manski2004 for its balanced consideration of the possible states of nature and for delivering “reasonable" decisions rules in practice. \begin{remark} An alternative version of the minimax regret criterion is minimax regret with respect to the welfare attained by the best-in-class policy in $\Pi$, resulting in the objective \begin{align} \pi^*_{MMR2} = \operatorname*{argmin}_{\pi \in \Pi } \max_{\tau \in \mathcal{T} }\left[\left( \max_{\pi \in \Pi}W_\tau(\pi)\right) - W_\tau(\pi)\right]. \end{align} Both versions have previously appeared in the literature and can be expected to enjoy similar properties. However the first version we have considered is considerably more tractable, as the innermost maximization in ((ref)) has closed form solution $\max_{\pi\, : \, \mathcal{X}\to\{0,1\}}W_\tau(\pi)= \mathbb{E}_{P_X}\left[\max\left\{\tau(X),0\right\}\right]$. As we show in Proposition (ref) below, this allows to more explicitly characterize the properties of the problem and its solution, as well as reduce the computational burden in solving the empirical analogue of the problem. For this reason we will focus on the version in ((ref)) of the criterion. We also note that whenever the class $\Pi$ contains the first-best oracle policy, i.e.\@ $\mathbbm{I}\left\{\tau(x)\geq0\right\}\in \Pi$, the two optimality criteria are equivalent. \end{remark} One critical drawback in the application of the two popular criteria just presented in the setting of this paper is that the optimal policies cannot be obtained in closed form. This is due to the typical form of ((ref)) and ((ref)) involving several nested optimizations whose solutions cannot be easily characterized at current level of generality. In order to make progress, we impose a restriction on the set of possible CATEs. \begin{assumption}[Rectangular identified set for $\tau$] The identified set for $\tau$ is rectangular, that is, $\mathcal{T}$ is of the form \begin{align*} \mathcal{T} = \{ \tau(\cdot) : \tau(x) \in [\tau(x), \overline{\tau}(x) ]\}. \end{align*} \end{assumption} Assumption (ref) is widely adopted in the partial identification literature, and we refer the reader to Appendix B in kasy2016 for an extensive review of identification schemes that result in rectangular identified sets. Below we present three examples of identification schemes for the CATE that satisfy this assumption.
example[Manski bounds] Suppose there exists a binary instrument $Z_i \in \{0,1\}$ that satisfies the well know exogeneity and exclusion restrictions $Y_i(0),Y_i(1),D_i(0),D_i(1) \perp Z_i | X_i$, where $Y_i(d)$ and $D_i(z)$ denote the counterfactual outcome and treatment functions, respectively. If the instrument $Z_i$ also satisfies the overlap condition \begin{align*} \eta \leq \mathbb{P}(Z_i =1 | X_i) \leq 1-\eta , \quad \eta>0, \end{align*} and the monotonicity condition (also known as no-defiers condition): \begin{align*} \mathbb{P} \big( D_i(1) \leq D_i(0) | X_i\big) =1\quad {\rm or} \quad \mathbb{P} \big( D_i(1) \geq D_i(0) | X_i\big) =1, \end{align*} then seminal work by imbensangrist1994 shows point-identification of the conditional local average treatment effect (LATE): \begin{align*}\mathbb{E}[Y_i(1) - Y_i(0) \,| \, D_i(1) \neq D_i(0), \,X_i=x].\end{align*} Let us now assume that $Y\in [Y_L,Y_U]$, i.e.\@\ the outcome is bounded, and define \begin{align*} &h(z,x) = \mathbb{E}[Y_i | Z_i=z, X_i=x],\\ &m(d,z,x) = \mathbb{E}[Y_i |D_i=d, Z_i=z, X_i=x],\\ &p(z,x)=\mathbb{P}(D_i=1 | Z_i=z, X_i=x),\\ &z(x)= \mathbb{P}(Z_i=1 | X_i=x). \end{align*} The identified sets for the expected potential outcomes $y_0(x)$ and $y_1(x)$ are contained within the bounds \begin{align*} \overline{y}_0(x)& = \min_{z\in\{0,1\}}\big\{ m(0,z,x)\cdot (1-p(z,x)) +Y_U \cdot p(z,x)\big\}, \\ y_0(x)& = \max_{z\in\{0,1\}}\big\{ m(0,z,x)\cdot (1-p(z,x)) +Y_L \cdot p(z,x)\big\}, \end{align*} and \begin{align*} \overline{y}_1(x)& = \min_{z\in\{0,1\}}\big\{ m(1,z,x)\cdot p(z,x) +Y_U \cdot (1-p(z,x))\big\}, \\ y_1(x)& = \max_{z\in\{0,1\}}\big\{ m(1,z,x)\cdot p(z,x) +Y_L \cdot (1-p(z,x))\big\}. \end{align*} The identified set for the CATE is then contained within the bounds \begin{align*} \overline{\tau}(x)& = \overline{y}_1(x) - y_0(x), \\ \tau(x) &=y_1(x) -\overline{y}_0(x) . \end{align*} If no further functional form assumption on the distribution of potential outcomes is made, these bounds are sharp heckman2001instrumental and the sharp identified sets for the average potential outcomes and CATE respectively satisfy Assumption (ref) and Assumption (ref).
example[Balke-Pearl] Suppose that the same assumptions as in Example (ref) hold, and additionally the monoticity assumption is strengthened to \begin{align*} \mathbb{P} \big( D_i(1) \geq D_i(0) | X_i\big) =1, \end{align*} that is, the direction of the monotonicity is known and positive. The bounds for the potential outcomes simplify to \begin{align*} \overline{y}_0(x)& = m(0,0,x)\cdot (1-p(0,x)) +Y_U \cdot p(0,x), \\ y_0(x)& = m(0,0,x)\cdot (1-p(0,x)) +Y_L \cdot p(0,x),\\ \overline{y}_1(x)& = m(1,1,x)\cdot p(1,x) +Y_U \cdot (1-p(1,x), \\ y_1(x)& = m(1,1,x)\cdot p(1,x) +Y_L \cdot (1-p(1,x)), \end{align*} and the CATE is contained within the bounds \begin{align*} \overline{\tau}(x)& = h(1,x) - h(0,x) + p(0,x)\cdot\big( m(1,0, x) - Y_L\big) + (1-p(1,x))\cdot\big(Y_U - m(0,1, x) \big), \\ \tau(x) &= h(1,x) - h(0,x) + p(0,x)\cdot\big( m(1,0, x) - Y_U\big) + (1-p(1,x))\cdot\big(Y_L - m(0,1, x) \big) . \end{align*} where $p(0,x)$ and $1-p(1,x)$ identify the proportions of always-takers and never-takers at $X_i=x$, respectively. If no further functional form assumption on the distribution of outcomes for non-compliant populations is made, these bounds are sharp balkePearl2007 and the sharp identified sets for the average potential outcomes and CATE respectively satisfy Assumption (ref) and Assumption (ref).
example[Manski-Pepper bounds] Suppose that instead of full exogeneity, the instrumental variable $Z_i$ satisfies the weaker “monotone IV" condition \begin{align} \mathbb{E}[Y_i(d) | Z_i=0 , X_i ] \leq \mathbb{E}[Y_i(d) | Z_i=1 , X_i ] , \quad d=0,1. \end{align} manskiPepper show that when the outcome is bounded one has \begin{align*} \sum_{z=0,1}\mathbb{P}\left(Z_i=z|X_i \right)\cdot \max_{z_1\leq z} &\left\{ m(d,z_1,X_i)\cdot \mathbb{P}(D_i=d | Z_i=z_1, X_i) + Y_L\cdot \mathbb{P}(D_i=1-d|Z_i=z_1,X_i) \right\} \\ &\qquad \qquad \qquad \leq \,\mathbb{E}[Y_i(d) | X_i] \,\leq\\ \sum_{z=0,1}\mathbb{P}\left(Z_i=z|X_i \right)\cdot \min_{z_2\geq z}&\left\{ m(d,z_2,X_i)\cdot \mathbb{P}(D_i=d | Z_i=z_2, X_i) + Y_L\cdot \mathbb{P}(D_i=1-d|Z_i=z_2,X_i)\right\}. \end{align*} Upper (lower) bounds for the CATE are obtained by combining upper (lower) bounds for $\mathbb{E}[Y_i(1)|X_i=x]$ with the lower (upper) bound for $\mathbb{E}[Y_i(0) | X_i=x]$: \begin{align*} \overline{\tau}(x)& = z(x)\cdot \psi_{1,1}(x;Y_U) + (1-z(x))\cdot \min\left\{\psi_{0,1}(x;Y_U),\psi_{1,1}(x;Y_U)\right\} \\ &-z(x)\cdot\max\left\{\psi_{0,0}(x;Y_L),\psi_{1,0}(x;Y_L)\right\} -(1-z(x))\cdot \psi_{0,0}(x;Y_L), \\ \tau(x) &= z(x)\cdot \max\left\{\psi_{0,1}(x;Y_L),\psi_{1,1}(x;Y_L)\right\} + (1-z(x))\cdot \psi_{0,1}(x; Y_L)\\ &-z(x)\cdot \psi_{1,1}(x;Y_U) -(1-z(x))\cdot \min\left\{\psi_{0,0}(x;Y_U),\psi_{1,0}(x;Y_U)\right\}, . \end{align*} where \begin{align*} \psi_{z,d}\left(x;Y_{(\cdot)}\right)=m(d,z,x)\cdot (d\cdot p(z,x) +(1-d)\cdot (1-p(z,x)) + Y_{(\cdot)} \cdot (d\cdot (1-p(z,x)) +(1-d)\cdot p(z,x)). \end{align*} Under no further assumption on the distribution of potential outcomes, these bounds are sharp manskiPepper and satisfy Assumptions (ref) and (ref).

Having restricted the identified sets $\mathcal{Y}$ and $\mathcal{T}$ as in Assumptions (ref)-(ref), we are now able to provide a simpler characterization of the maximin welfare and maximin impact policies.

propositionDefine $\underline{y}_d(x)=\min_{y_d(x) \in \mathcal{Y}(x)} y_d(x)$ and $\overline{y}_d(x)=\max_{y_d(x) \in \mathcal{Y}(x)} y_d(x)$. Under Assumption (ref) the optimal maximin welfare policy is \begin{align} \pi^*_{MMW} = \operatorname*{argmax}_{\pi \in \Pi } \mathbb{E}_{P_X}\Bigg[ \left( 2\pi(X)-1\right) \cdot (y_1(X) -y_0(X)) \Bigg]. \end{align} Furthermore, under Assumption (ref) the optimal maximin impact policy is \begin{align} \pi^*_{MMI} = \operatorname*{argmax}_{\pi \in \Pi } \mathbb{E}_{P_X}\Bigg[ \left( 2\pi(X)-1\right) \cdot \underline{\tau}(X) \Bigg]. \end{align}

Proposition (ref) shows that the optimal maximin welfare and maximin impact policies maximize surrogate versions of the welfare substituting the unidentified CATE with the difference in the lower bounds of potential outcomes $\underline{y}_1(x)-\underline{y}_0(x)$ and the lower bound for CATE $\underline{\tau}(x)$ at every point in the covariate space, respectively. Notice that when $\mathcal{Y}(x)$ is rectangular with respect to the two potential outcomes, i.e.\@ $y_0(x)\in \mathcal{Y}_0(x), y_1(x)\in \mathcal{Y}_1(x)$ and $\mathcal{Y}(x)=\mathcal{Y}_0(x)\times\mathcal{Y}_1(x)$, we have $\underline{\tau}(x)=\underline{y}_1(x)-\overline{y}_0(x)$, thus highlighting the “pessimistic" nature of the maximin impact criterion.

remarkThe maximin welfare and maximin impact optimal policies coincide when $y_0(\cdot)$ is point-identified. This case is relevant when $y_0(x)$ represents the (conditional) average outcome under the status-quo in the entire population and is typically point-identified from observational data.

The simplification of these two maximin problems into single maximisation problems has important advantages for the study of the optimal policies and their estimation from the data. In fact, the sample analogues of optimizations (ref) and (ref) are amenable to standard computation procedures for a variety of policy classes $\Pi$. Furthermore, their solution can be studied using tools for empirical risk minimisation problems, as discussed in Section (ref).

Despite the involvement of an additional maximization problem compared to maximin welfare and maximin impact, Assumption (ref) allows to provide a simpler characterization also for the minimax regret optimal policy.

propositionUnder Assumption (ref) the optimal minimax welfare regret policy is \begin{align} \pi^*_{MMR} = \operatorname*{argmax}_{\pi \in \Pi } \mathbb{E}_{P_X}\Bigg[ \left( 2\pi(X)-1\right) \cdot \widetilde \tau(X) \Bigg] \end{align} where \begin{align} \widetilde \tau(x)=\overline{\tau}(x)\cdot\mathbbm{1}\big\{\overline{\tau}(x) \geq0 \big\} + \tau(x)\cdot \mathbbm{1}\big\{\tau(x) \leq0 \big\} \end{align}

This simpler characterization of the minimax regret problem as a single maximization sheds light on the properties of its associated optimal policy. In particular, we see that the objective function symmetrically treats individuals whose expected treatment effect sign is identified by assigning as surrogate for the CATE their outer bound, i.e.\@\ the CATE upper (lower) bound for individuals with identified positive (negative) sign for CATE. Individuals for which the sign of the treatment effect is ambiguous are assigned an intermediate point within their respective CATE bounds. The location of this intermediate point depends on the extent to which the identified set lies in the positive or negative region. Intuitively, the criterion prioritizes correct treatment allocation to individuals who unambiguously benefit from (or are harmed by) the treatment and down-weights the importance of individuals for which the sign of the treatment response is ambiguous within the treatment allocation problem. As an extreme case, individuals with CATE bounds exactly symmetric around 0 (i.e.\@ $\overline{\tau}(x)=-\underline{\tau}(x)$) are given no consideration in the solution of the treatment allocation problem. This intuition can be further supported by noticing that the original welfare maximization under point-identification in ((ref)) can be re-casted as the weighted classification problem

align*[align* omitted — 163 chars of source]

of which the minimax welfare regret optimal policy in (ref) turns out to solve the minimax version under Assumption (ref):

align*[align* omitted — 206 chars of source]

It is from this minimax classification risk perspective that Pu_2021 obtain and study the minimax regret policy, which they call the “IV-optimal policy".

An alternative version of minimax regret optimality which has been used in the context of treatment choice is minimax regret with respect to a baseline policy. kallus2018 assume the existence of a fixed policy $\pi_{\texttt{B}}$ from which the DM does not want to unnecessarily deviate. They define the optimal policy as minimizing regret with respect to this baseline policy:

align*[align* omitted — 422 chars of source]

where the second equality uses Assumption (ref). While potentially appealing in certain settings, e.g.\@ when $\pi_{\texttt{B}}$ represents the existing standard of care in a medical setting, this optimality criterion suffers the potential drawback of requiring the DM to specify (and motivate) the baseline policy for it to be operational. Adopting the never-treat baseline policy, i.e.\@ $\pi_{\texttt{B}}(x)=0, \, \forall x\in\mathcal{X}$, could be seen as an appealing “agnostic" choice, which however makes this criterion default to maximin impact and thus inherit its potentially undesirable properties.

The last notion of ambiguity-robust optimal policy that we present in this section is based on the Hurwicz criterion hurwicz, arguably one of the most widely used in decision-making under ambiguity. In the context of the treatment assignment problem at hand, the Hurwicz criterion leads to the ambiguity-robust policy

align*[align* omitted — 316 chars of source]

where $\delta_1\in[0,1]$ and $\delta_0\in[0,1]$ are user-defined weights reflecting the degree of optimism with respect to the outcomes under treatment and non-treatment, respectively. It is easy to see that the maximin welfare criterion in (ref) corresponds to the choice $\delta_1=0,\delta_0=0$. An analogous notion of optimality focused on impact rather than welfare, leads to the optimal policy

align*[align* omitted — 221 chars of source]

where $\delta\in[0,1]$ controls the degree of optimism with respect to the effect of treatment, with the maximin impact optimal policy corresponding to the choice $\delta=0$. Under Assumption (ref) and $\mathcal{Y}(x)=\mathcal{Y}_0(x)\times\mathcal{Y}_1(x)$, the Hurwicz Impact criterion is nested into the Hurwicz Welfare for the choice of parameters $\delta=\delta_0=1-\delta_1$; unlike Hurwicz Welfare, however, the Hurwicz Impact criterion is still well-defined under the weaker Assumption (ref). Interestingly, when $\delta=1/2$ and $\Pi$ is well-specified, in the sense that it contains the first-best assignment $\mathbbm{1}\{\overline{\tau}(x) + \underline{\tau}(x)\geq 0 \}$, we have

align*[align* omitted — 177 chars of source]

Therefore the minimax regret and Hurwicz impact optimal policies coincide under correct specification of $\Pi$, as they assign treatment based on the middle point between the upper and lower CATE bounds. When $\Pi$ is not well-specified, however, minimax regret optimality is not nested into any of the Hurwicz-type criteria just presented, thus highlighting the radically different attitude towards ambiguity implied by minimax regret compared to maximin welfare/impact. In particular, minimax regret is the only criterion of those presented (along with Hurwicz impact under $\delta=1/2$) that treats symmetrically individuals with identified sets symmetric around 0, in the sense that $\widetilde\tau_1(x)= -\widetilde\tau_2(x)$ whenever $\overline{\tau}_1(x)=-{\underline{\tau}}_2(x)$ and $\underline{\tau}_1(x)=-{\overline{\tau}}_2(x)$. For this reason, minimax regret does not reflect an optimistic/pessimistic attitude towards ambiguity but rather an “opportunistic" one, in light of its prioritization of correct treatment assignment to individuals whose CATE sign is unambiguously identified.

figure[figure omitted — 1,051 chars of source]

A common framework

While accommodating a wide range of attitudes towards ambiguity, the notions of optimality presented in Section (ref) share a common structure. In fact, by virtue of Assumptions (ref) and (ref), the corresponding optimal policies can all be written as

align[align omitted — 180 chars of source]

for a specific score function\footnote{The term `score function' is borrowed from atheywager2021.} $\Gamma(P;\,\cdot\,)$, where we have highlighted the dependence of the score on the distribution $P$. The specific dependence on $P$ is determined by the optimality criterion (as summarized in Table (ref)) as well as the identification assumptions (e.g.\@ Balke-Pearl, Manski-Pepper etc.\@). This common structure also nests the point identified setting as the special case $\Gamma(P;X)=\tau(X)$ and thus suggests that existing estimation procedures for this special case can be extended to the partially identified setting.

However, one peculiar feature of the partially identified setting is the restricted degree of smoothness enjoyed by the objective function, in particular the differentiability of the scores with respect to $P$. Under point-identification of the CATE via standard unconfoundedness assumptions, one has $\Gamma(P;x)=\mathbb{E}[Y| D=1,X=x]-\mathbb{E}[Y|D=1,X=x]$ and the full differentiability of the score with respect to the expectation $\mathbb{E}[Y|D,X]$ is immediately apparent. However, for the minimax regret criterion we notice that the score is directionally differentiable\footnote{ Let $P\in\mathcal{P}$ be a probability distribution on which the function $f:\mathcal{P}\to \mathbb{R}$ depends. We say that $f$ is directionally differentiable at $P_0$ if the limit

align*[align* omitted — 86 chars of source]

exists for every $h\in \mathcal{P}$, in which case $\dot{f}_{P_0}[\cdot]$ denotes the directional derivative of $f$ at $P_0$. If it exists, the directional derivative $\dot{f}_{P_0}[\cdot]$ is positively homogeneous of degree one but not necessarily linear. If $\dot{f}_{P_0}[\cdot]$ is linear then f is fully differentiable at $P_0$.} with respect to $P$ at $\overline{\tau}(x)=0$ or $\underline{\tau}(x)=0$. Even when $\Gamma(P;x)$ depends smoothly on expected outcomes/CATE bounds, lack of full differentiability of the scores can arise through a lack of differentiability of the expected outcomes and CATE bounds themselves. In fact, many popular identification assumptions, including the Manski and Manski-Pepper bounds from Examples (ref) and (ref), deliver bounds that are only directionally differentiable with respect to identified parameters due to the presence of $\min/\max$ operators chernozhukovIntersectionBounds. Whether a consequence of the optimality criterion or the identification assumptions, lack of full differentiability of the scores is a unique and pervasive feature of the treatment assignment problem under partial identification, one that has not been explicitly acknowledged in the most recent contributions in this area.\footnote{The only exception is christensenMoonSchor2020, who deal with estimation of optimal treatment decisions in the absence of individualization.} A major contribution of this paper is to account for the role played by the lack of full differentiability when we establish procedures for estimating ambiguity-robust optimal policies in Section (ref).

comment\subsubsection{Neyman orthogonal moment conditions for IV CATE bounds} In this section we provide explicit expression for Neyman-orthogonal identifying moment conditions of CATE bounds $\left( \overline{\tau}(x), \underline{\tau}(x)\right)$. These will later be used in the optimal policy estimation. Collecting the observables in the vector $W_i=\left(Y_i,D_i,Z_i,X_i\right)$ and having defined the quantities \begin{align*} &h(z,x) = \mathbb{E}[Y_i | Z_i=z, X_i=x]\\ &m(d,z,x) = \mathbb{E}[Y_i |D=d, Z_i=z, X_i=x]\\ &e_{d,z}(x)=\mathbb{P}(D_i=d, Z_i=z| X_i=x),\\ &g(Z_i, X_i) = \frac{Z_i - z(X_i)}{z(X_i)(1-z(X_i)}, \qquad z(x)=\mathbb{P}(Z_i=1 | X_i=x), \end{align*} we can re-write the CATE bounds as \begin{align*} &\overline{\tau}(X_i) = h(1,X_i) - h(0,X_i) + \frac{e_{10}(X_i)}{1-z(X_i)}\cdot\big( m(1,0, X_i) - Y_L\big) + \frac{e_{01}(X_i)}{z(X_i)}\cdot\big(Y_U - m(0,1, X_i) \big), \\ &\tau(X_i) = h(1,X_i) - h(0,X_i) + \frac{e_{10}(X_i)}{1-z(X_i)}\cdot\big( m(1,0, X_i) - Y_U\big) + \frac{e_{01}(X_i)}{z(X_i)}\cdot\big(Y_L - m(0,1, X_i) \big) . \end{align*} Following ichimura2015 we compute the influence functions for the CATE upper bound { \begin{align*} &\phi_U(W_i) =\underbrace{g(Z_i, X_i)\big(Y_i - h(Z_i,X_i)\big)}_{\phi^U_1}\\ & + \underbrace{\frac{m(1,0, X_i)}{1-z(X_i)}\cdot \big(D_i(1-Z_i) - e_{10}(X_i)\big) + \frac{D_i(1-Z_i)}{1-z(X_i)}\cdot \big( Y_i -m(D_i,Z_i,X_i)\big) + \frac{m(1,0, X_i)}{\big(1-z(X_i)\big)^2}\cdot \big(Z_i - z(X_i))\big)}_{\phi^U_2} \&\underbrace{- Y_L\cdot \frac{1}{1-z(X_i)}\cdot \big(D_i(1-Z_i) - e_{10}(X_i)\big) - Y_L\cdot \frac{e_{10}(X_i)}{\big(1-z(X_i)\big)^2}\cdot \big(Z_i - z(X_i))\big)}_{\phi^U_3}\\ & - \underbrace{\frac{m(0,1, X_i)}{z(X_i)}\cdot \big((1-D_i)Z_i - e_{01}(X_i)\big) - \frac{(1-D_i)Z_i}{z(X_i)}\cdot \big( Y_i -m(D_i,Z_i,X_i)\big) - \frac{m(0,1, X_i)}{z(X_i)^2}\cdot \big(Z_i - z(X_i))\big)}_{\phi^U_3} \&\underbrace{+ Y_U\cdot \frac{1}{z(X_i)}\cdot \big((1-D_i)Z_i - e_{01}(X_i)\big) - Y_U\cdot \frac{e_{10}(X_i)}{\big(1-z(X_i)\big)^2}\cdot \big(Z_i - z(X_i))\big)}_{\phi^U_4}, \end{align*}} and the analogous $\phi_L(W_i)$ for the lower bound. Neyman-orthogonal moment functions for the bounds are then formed as \begin{align*} &\overline{\tau}^{NO}_i =\overline{\tau}(X_i) + \phi_U(W_i), \\ &\tau^{NO}_i =\tau(X_i) + \phi_L(W_i). \end{align*}
table[table omitted — 1,040 chars of source]

Estimation

In this section we present the statistical framework underlying the problem of estimation of optimal treatment rules under partial identification. We will discuss heuristics underlying several features of the estimation problem, and then present our proposed estimation procedure.

We work in a learning setting where the estimand $\pi^*(P)$ is as in ((ref)), and we observe an i.i.d. sample $(W_i)_{i=1,\dots,n}$ of size $n$ from the unknown distribution $P$ of the observed random variables $W\in \mathcal{W}$, $X\subset W$. To retain generality of the framework, we do not specify the exact dependence of the functional $\Gamma(P;x)$ on $P$, which will depend on the choice of optimality criterion for the resolution of ambiguity (maximin welfare, minimax regret etc.) and identification assumptions determining the identification sets $\mathcal{Y}(P),\mathcal{T}(P)$. However, we will assume that the scores depend on $P$ only through a vector of nuisance functions $g:\mathcal{V}\to \mathbb{R}^J$ specified by the moment equations

align[align omitted — 68 chars of source]

where $U$ and $V$ are random vectors with $U\subseteq W$ and $X\subseteq V\subset W$. Furthermore, we will stipulate that the dependence of $\Gamma(g;x)$ on the nuisance functions $g$ from the possibly infinite-dimensional space $\mathcal{G}$ can be reduced as

align*[align* omitted — 46 chars of source]

where, for a fixed $x$, the parameter $\theta(x)\in \Theta_x \subseteq\mathbb{R}^M$ is a finite-dimensional vector of conditional moments of $U$ deduced from $g$. This latter restriction rules out scores $\Gamma(g;X)$ that at a single point in the covariate space depend on exhaustive evaluations of the nuisance functions $g$ over continuous supports. This is the case, for example, in versions of the CATE bounds from Examples (ref)-(ref) featuring instruments with continuous support $\mathcal{Z}$. In those settings, the CATE bounds depend on objects such as $\sup_{z\in \mathcal{Z}}\mathbb{E}[Y \mid Z=z, \,X=x ]$, and are therefore not covered by the results of this paper. Finally, we will assume that that $\Gamma(\theta(x);x)$ can be expressed as

comment\begin{equation} \begin{aligned} &\Gamma(\theta(x);x) =\\ &\varphi_0(\theta(x);x) + \sum_{\ell=1}^L\varphi_\ell(\theta(x);x)\cdot \mathbbm{1}\left\{\varphi_\ell}(\theta(x);x)\geq 0 \right\} +\sum_{m=1}^M\Psi^{(m)}(\theta(x);x)\cdot \mathbbm{1}\left\{\Psi^{(m)}(\theta(x);x)\leq 0 \right\}, \end{aligned} \end{equation}
equation[equation omitted — 264 chars of source]

where the functions $\varphi_\ell(\theta(x);x): \Theta_x\times \mathcal{X} \to \mathbb{R}$ are fully differentiable with respect to $\theta(x)$ for all $x\in\mathcal{X}$. While seemingly ad-hoc, this restriction is sufficiently general to accommodate a wide range of popular partial identification assumptions for the CATE as well as optimality criteria for the resolution of ambiguity. In particular, formulation ((ref)) accommodates linear combinations of $\min$/$\max$ operators, which typically feature in many identification bounds for the CATE with discrete instruments. In fact, our framework can be shown to be applicable to any combination of the optimality criteria discussed in Section (ref) and the identification schemes contained in the recent survey paper by Swanson et al. (2018).\footnote{Albeit not directly accommodated by formulation (ref), our framework and theoretical results also apply to scores that feature a finite number of nested linear combinations of $\min/\max$ operators. We discuss this extension in Appendix (ref).}

example[continues=balkePearl] Under the identification assumptions of the Balke-Pearl bounds and resolution of ambiguity via Minimax Regret, we have \begin{align*}g&=(h, m, p),\\ \theta(x)&=(h(1,x), h(0,x), m(1,0,x), m(0,1,x), p(1,x), p(0,x)),\end{align*} and \begin{align*} \Gamma(g;x)= \varphi_1(\theta(x);x)\cdot \mathbbm{1}\left\{\varphi_1(\theta(x);x)\geq 0 \right\} - \varphi_2(\theta(x);x)\cdot \mathbbm{1}\left\{\varphi_2(\theta(x);x)\geq 0 \right\}, \end{align*} where $\varphi_1(\theta(x);x)=\overline{\tau}(\theta(x);x)$, $\varphi_2(\theta(x);x)=-\underline{\tau}(\theta(x);x)$ are differentiable with respect to $\theta(x)$.

In this framework, a natural approach for estimation is via the so-called “empirical risk minimisation" (ERM) principle vapnik1998, in which the estimate for the optimal policy is obtained as the maximiser of a sample analogue of the population objective $Q$:

align[align omitted — 218 chars of source]

where $\widehat{\Gamma}_i$ is some suitable estimate for $\Gamma(g;X_i)$. The ERM approach is a cornerstone of statistical learning theory and is at the foundation of many traditional and modern estimation methods in statistics, econometrics and machine learning. The ERM principle has also guided much of the recent literature on individualized treatment rules, where different variations have been applied under the names of “outcome-weighted learning" kosorokOWL2012 and “empirical welfare maximization" kitagawatetenov2018. A major challenge in the implementation of (ref) comes from the presence of the nuisance functions $g$, which are typically unknown and thus need to be estimated. Assuming that we have access to appropriate algorithms/nonparametric procedures for estimation the nuisance functions, one simple approach would be to use the sample $(W_i)_{i=1,\dots,n}$ to obtain the estimates $\widehat g$ and then form plug-in estimates for the score as $\widehat \Gamma_i=\Gamma(\widehat g; X_i)$. While seemingly natural, this “naive plug-in" approach has undesirable properties. In particular, policies estimated via the naive plug-in approach can typically only be shown to converge at sub-optimal rates to their population counterparts, unless very restrictive assumptions are imposed on first-stage estimators for the nuisance functions $g$ fosterSyrgkanis.

One crucial reason underlying the undesirable statistical properties of the naive plug-in approach is that the resulting objective function estimate $\widehat Q_n$ is overtly sensitive to error in estimating the nuisance functions $g$. In order to gain intuition, it is useful to consider the following expansion of the population objective function $Q(g;\pi)=\mathbb{E}_{P_X}\left[(2\pi(X)-1)\cdot\Gamma(g;X)\right]$,

align[align omitted — 225 chars of source]

where

align*[align* omitted — 282 chars of source]

Von Mises expansions like the above are at the heart of the theory of orthogonal machine learning chernozhukov2018. In our setting, it allows to describe the impact of a small deviation from $g$ in the direction $\widetilde g -g$ as consisting of three terms. The first term is the so-called “pathwise derivative" of $Q(g;\pi)$ and typically scales with $\|\widetilde g - g\|_{L_1(P)}$. The second term $\Delta(\widetilde g, g; \pi)$ is due the presence of the type of non-differentiabilities arising under partial identification, and is unique to the framework of this paper. This term accounts for the bias that arises from misclassifying whether the component functions $\varphi_\ell$ are above or below 0, as we move away from $g$ in the direction $\widetilde g - g$. The third term is a second-order remainder scaling with the mean-square distance between $\widetilde g$ and $g$. A central feature of our proposed estimation procedure is the construction of a new objective function, called Neyman-orthogonal, with reduced sensitivity to local perturbations away from $g$. For this purpose, we will assume that there exists functionals $\alpha_\ell(\{g,f\}; V)$ such that for every $\widetilde g \in \mathcal{G}$

align*[align* omitted — 156 chars of source]

where $f\in \mathcal{F}$ is a vector of additional nuisance functions defined analogously to $g$ .\footnote{We will also assume that for the $j$-th entry of the Riesz-representer we have $\alpha^{(j)}_\ell(\{(\widetilde g_{-j},\widetilde g_{j}),\widetilde f\}, x)=\alpha^{(j)}_\ell(\{\widetilde g_{-j},\widetilde f\}, x)$, where $\widetilde g_{-j}$ denotes the exclusion of the $j$-th entry $\widetilde g_{j}$ from the vector of nuisance functions $\widetilde g$. This restriction is sufficiently general to accommodate component functions $\varphi_\ell(\theta(x);x)$ that feature linear combinations of products of the parameters $\theta(x)$, thus encompassing all the discussed identification schemes, including Examples (ref)-(ref).} We then construct Neyman-orthogonal formulations for the component functions as

align*[align* omitted — 177 chars of source]

The functionals $\alpha_\ell$ are the Riesz-representers of $\varphi_\ell$, while the functionals $\phi_\ell$ are their so-called influence function adjustments. We refer the reader to ichimura2015 for their properties and general methods for their calculation, while we provide below their specific form for the Balke-Pearl CATE bounds of Example (ref).\footnote{See also kennedy2022 for a user-friendly discussion of methods for computation influence function adjustments.}

example[continues=balkePearl] Following ichimura2015, we compute the influence function adjustment $\phi_U(\{g,f\},W_i)$ for the CATE upper bound by taking the Gateaux derivative of $\overline{\tau}(g;X)$, which yields \begin{align*} \phi_U(\{g,f\};W_i) =&\underbrace{\left[\frac{Z_i}{z(X_i)} - \frac{1-Z_i}{1-z(X_i)}\right]}_{\alpha_U^{(1)}(\{g,f\},V_i)}\cdot\big(Y_i - h(Z_i,X_i)\big)\\ &+\underbrace{\left[\frac{D_i(1-Z_i)}{1-z(X_i)}+\frac{(1-D_i)Z_i}{z(X_i)}\right]}_{\alpha_U^{(2)}(\{g,f\},V_i)}\cdot \big( Y_i -m(D_i,Z_i,X_i)\big) \\ &+\underbrace{\left[(m(1,0, X_i)-Y_L)\cdot \frac{1-Z_i}{1-z(X_i)} - (Y_U-m(0,1, X_i))\cdot \frac{Z_i}{z(X_i)}\right]}_{\alpha_U^{(3)}(\{g,f\},V_i)}\cdot \big(D_i - p(Z_i,X_i)\big), \end{align*} where the associated Riesz-representer is $\alpha_U(\{g,f\},V_i)=(\alpha_U^{(1)},\alpha_U^{(2)},\alpha_U^{(3)})'$ with $f=z(x)$ and $V_i=(D_i,Z_i,X_i)'$. The influence function and Riesz-representer for the CATE lower bound are obtained by interchanging $Y_U$ and $Y_L$ in the expressions above.

Finally we construct Neyman-orthogonal formulations for the scores as

align*[align* omitted — 209 chars of source]

which are then used to form the Neyman-orthogonal objective function

align*[align* omitted — 151 chars of source]

Our construction of Neyman-orthogonal scores features the addition of the influence function adjustments $\phi_\ell$ to the component functions $\varphi_\ell$ outside the indicators, but crucially not inside. Heuristically, the influence function adjustments serve the purpose of reducing the bias induced by the evaluation of the component functions $\varphi_\ell(\cdot;x)$ away from $g$. Since the indicators vary discontinuously with $g$ it is not possible to linearly approximate the dependence of the indicators on the nuisance functions at the point of discontinuity $\varphi_\ell(g;x)=0$. As a result, it is not possible to reduce the bias induced by the presence of the indicators (represented by the term $\Delta(\widetilde g,g,\pi)$ in (ref)) by means of influence function adjustments, whose de-biasing properties implicitly rely on the validity of such linear approximation.\footnote{On the contrary, naively adding the influence function adjustments inside the indicators would lead to a bias increase, rather than a reduction.} Notice that $Q^{\texttt{NO}}(\{g,f\};\pi) = Q(g;\pi)$ by the mean-zero property of the influence function adjustments, so that orthogonalization of the objective does not change the notion of optimal policy $\pi^*(P)$. Nonetheless, for the orthogonalized objective we have that

align[align omitted — 245 chars of source]

Comparing the above with ((ref)), we see that the von Mises expansion for the orthogonalized objective does not feature the pathwise derivative term, implying that $Q^{\texttt{NO}}(\,\cdot\,;\pi)$ is less sensitive to deviations away from $g$ compared to the original objective $Q( \, \cdot \,;\pi)$. As shown in Section (ref), this property will generally translate in improved statistical guarantees for the estimated policy when the nuisance functions have to be learned from the data. It should however be noticed that the term $\Delta(\widetilde g,g;\pi)$ still appears in the relevant expansion after orthogonalization. The contribution of this term is quantified in Section (ref), where it is shown to be of first-order importance for the statistical properties of the estimation procedures.

The second key component of our approach is the use of sample-splitting, which is a commonly employed method in semiparametric inference chernozhukov2018 and statistical learning fosterSyrgkanis. The main purpose of sample-splitting is to reduce the risk of overfitting that generally arises from using the same data to estimate the nuisance functions as well as the optimal policy, as in the naive plug-in approach. Similarly to atheywager2021, we employ a particular form of sample-splitting known as K-fold cross-fitting (described below). This procedure ensures that in the estimate for $\Gamma^{\texttt{NO}}(\{g,f\}; W_i)$, the estimates for the nuisance functions $\{g,f\}$ are independent from the data-point $W_i$ for that same unit. This independence property is crucial for the theoretical guarantees of our proposed method.

commentA second fundamental ingredient to our proposed Following kitagawatetenov2018 and atheywager2021, we estimate the optimal policy by solving the sample analogue of the MMR population problem ((ref)): \begin{align} \widehat{\pi}_{MMR} = \operatorname*{argmax} \Bigg\{\frac{1}{n}\sum_{i=1}^n \big(2\pi(X_i) -1\big)\widehat{\Gamma}(X_i) \, : \, \pi \in \Pi \Bigg\}, \end{align} where $\widehat{\Gamma}(X_i), \, i=1,\dots,n$ are estimates for the scores to be obtained from the same sample. Assuming that we have access to algorithms/nonparametric procedures that deliver estimates $\big(\hat{h},\hat{m}, \hat{e}, \hat{z} \big)$ of the nuisance functions involved in the CATE bounds, we could then form plug-in estimates of the scores. This is the approach taken by Pu_2021.\footnote{Similarly to us, Pu_2021 also use cross-fitting.} We instead follow the insight of atheywager2021 and base estimation of the scores on Neyman-orthogonal estimates of the CATE bounds. Neyman-orthogonal estimates are obtained by adding the influence function with respect to nuisance parameters to the original estimating moment equation (see chernozhukov2018, chernozhukov2018). We refer the reader to ichimura2015 for the properties of the influence function and general methods for its calculation, and we provide below the specific form of the Neyman-orthogonal estimating equations for the Balke-Pearl bounds of Example (ref).

Our proposed estimation procedure is therefore as follows. We first randomly split the data into $K$ evenly-sized folds and for each fold $k=1,\dots, K$ we obtain estimates $\{\widehat g^{(-k)},\widehat{f}^{(-k)}\}$ using data from the remaining $K-1$ folds. These estimates are then used to form cross-fitted Neyman-orthogonal estimates for the scores

comment\begin{align} \widehat{\Gamma}^{NO}_i = \varphi^{(0), NO}(\hat{f}^{-k(i)};W_i) + \sum_{\ell=1}^L a_{\ell}\cdot\varphi^{NO}_\ell(\hat{f}^{-k(i)};W_i)\cdot \mathbbm{1}\left\{\varphi_\ell(\hat{f}^{-k(i)};X_i)\geq 0 \right\}, \end{align}
align[align omitted — 161 chars of source]

where $k(i)\in \{1,\dots, K\}$ denotes the fold containing the $i$-th observation. Finally, the estimated optimal policy rule $\widehat{\pi}_n$ is obtained via the optimization problem ((ref)).

Statistical guarantees for the estimated policy

Let $\widehat{\pi}_{n}$ be the estimated treatment policy defined in ((ref)), with estimated scores as in ((ref)). Following manski2004, we assess the performance of the estimated policy in terms of (statistical) regret with respect to population optimal policy. Let the population ambiguity-robust optimal policy be $\pi^*_{n}(P)\in \operatorname*{argmax}_{\pi \in \Pi_n }Q(P;\pi)$, where we have included the $n$-subscript to the policy class $\Pi_n$ to allow this to depend on the sample size for generality. The statistical regret of an estimated policy $\widehat{\pi}_n$ is defined as

align[align omitted — 124 chars of source]

where $\mathbb{E}_{P_n}$ is the expectation with respect to the i.i.d.\ sample of observable random variables $(W_i)_{i=1,\dots,n}$ used to estimate $\widehat{\pi}_n$. The next few subsections build up to a final result providing asymptotic convergence guarantees for $\widehat{\pi}_{n}$ to $\pi^{*}_{n}$ in terms of statistical regret.

Assumptions

We make the following assumptions.

assumption[VC-class] There exists constants $0\leq\nu<1/2$ and $N\geq1$ such that $\emph{VC}(\Pi_n)\lesssim n^\nu$ for all $n\geq N$.

Assumption (ref) restricts the policy class to have finite VC-dimension, which is a standard requirement for controlling the complexity of a policy class in the classification literature. The VC-dimension of the policy-class $\Pi$ is defined as the largest interger $m$ such that there exist points $x_1,\dots,x_m$ that are shattered by $\Pi$, i.e.\@ where the policy values $\pi(x_1),\dots,\pi(x_m)$ can take on all $2^m$ possible combinations in $\{0,1\}^m$ (for more on the VC-dimension, see wainwright_2019). Several practically relevant classes of treatment rules satisfy this requirement, including the linear-index and quadrant rules used in the empirical application of Section (ref). Our assumption allows the VC-dimension of the policy class to grow moderately with the sample size, thus allowing the treatment rule to depend on high-dimensional covariates.

comment\begin{assumption}[Overlap] There is a constant $\eta>0$ such that $z(x)\in[\eta,1-\eta]$ for all $x\in Supp(X)$. \end{assumption} The Overlap Assumption implies non-trivial identification for the CATE in the context of Example (ref). While not strictly necessary for the results of this section, it ensures the presence of nuisance functions in the problem, which is a key feature for the contributions of this paper.
assumption[Regularity conditions for data-generating process] \begin{itemize} • There exist constants $\mathcal{C}_{1,\varphi},\mathcal{C}_{1,\alpha}$ such that for all $\{\widetilde g, \widetilde f\}\in \mathcal{G}\times \mathcal{F}$ \begin{align*}\norm{\varphi_\ell(\widetilde g; X) - \varphi_\ell(g; X)}_{L_\infty(P_X)}&\leq \mathcal{C}_{1,\varphi} \cdot \norm{\widetilde g - g }_{L_\infty(P_V)},\\ \norm{\alpha_\ell(\{\widetilde g, \widetilde f\}; V) - \alpha_\ell(\{g, f\}; V)}_{L_\infty(P_V)}&\leq \mathcal{C}_{1,\alpha} \cdot \left(\norm{\widetilde g - g }_{L_\infty(P_V)} + \norm{\widetilde f - f }_{L_\infty(P_V)} \right), \end{align*} for $\ell=0,\dots, L$. • There exist constants $\mathcal{C}_{2,\varphi},\mathcal{C}_{2,\alpha}$ such that for all $\{\widetilde g, \widetilde f\}\in \mathcal{G}\times \mathcal{F}$ \begin{align*}\norm{\varphi_\ell(\widetilde g; X) - \varphi_\ell(g; X)}_{L_2(P_V)}&\leq \mathcal{C}_{2,\varphi} \cdot \norm{\widetilde g - g }_{L_2(P_V)},\\ \norm{\alpha_\ell(\{\widetilde g, \widetilde f\}; V) - \alpha_\ell(\{g, f\}; V)}_{L_2(P_V)}&\leq \mathcal{C}_{2,\alpha} \cdot \left(\norm{\widetilde g - g }_{L_2(P_V)} + \norm{\widetilde f - f }_{L_2(P_V)} \right), \end{align*} for $\ell=0,\dots, L$. • There exist constants $\mathcal{C}_{3,\varphi},\mathcal{C}_{3,\alpha}$ such that for all $\{\widetilde g, \widetilde f\}\in \mathcal{G}\times \mathcal{F}$ \begin{align*} \norm{\varphi_\ell(\widetilde g; X)}_{L_\infty(P_X)}&\leq\mathcal{C}_{3,\varphi},\\ \norm{\alpha_\ell(\{\widetilde f,\widetilde g\};V)}_{L_\infty(P_V)}&\leq\mathcal{C}_{3,\alpha}, \end{align*} for $\ell=0,\dots,L$. • The irreducible noise $\varepsilon_i \coloneqq U_i - g(V_i)$ is a sub-Gaussian vector conditional on $V_i$, with conditional variance ${\rm Var}(\varepsilon_i \mid V_i )= \Sigma(V_i)$ satisfying $\norm{\lambda_{\rm max}(\Sigma(V))}_{L_\infty(P_V)}\leq \overline{\lambda}< \infty$. \end{itemize}

Assumptions 5(i) and 5(ii) impose Lipschitz continuity of the component functions and Riesz-representers with respect to the nuisance component in the $L_\infty$ and $L_2$-norm, respectively. These requirements are typically met under mild conditions within the framework of this paper. For the Balke-Pearl bounds of Example (ref), these assumptions hold under the overlap condition whenever $\mathcal{G}$ and $\mathcal{F}$ are subsets of the space of bounded functions\footnote{That is, there exists a constant $B>0$ such that $\|\{\widetilde g, \widetilde f\}\|_{L_\infty(P_V)}\leq B, \, \forall \{\widetilde g, \widetilde f\} \in \mathcal{G}\times \mathcal{F} $}, which is automatically satisfied since $U_i=(Y_i, D_i, Z_i)'$ is a vector of random variables with bounded support. Assumption (ref)(iii) is a uniform bound on the component functions and Riesz-representers, the former implying uniform boundedness of the scores $\Gamma(g;\cdot)$. Assumption (ref)(iv) is a standard requirement in statistical learning theory restricting the tail behaviour of the statistical noise $\varepsilon_i$. It is automatically satisfied when $U_i$ has bounded support, as in the Balke-Pearl bounds, but also allows for outcomes with unbounded support whose conditional distributions have sufficiently thin tails. Together with Assumption (ref)(iii), this assumption implies sub-gaussianity of $\Gamma^{\texttt{NO}}(\{g,f\},W_i)$.

The next two assumptions impose requirements on the estimators for the nuisance components.

assumption[Regularity conditions for fist-step estimators] \begin{itemize} • The estimators of the nuisance functions $\{\widehat g_n, \widehat f_n\}$ belong to the function classes $\mathcal{G}\times \mathcal{F}$ with probability 1. • There exists a constant $\mathcal{C}_4>0$ such that \begin{align*} \norm{\widehat{g}_n - g}_{L_\infty(P_V)} \leq \mathcal{C}_4,\\ \norm{\widehat{f}_n - f}_{L_\infty(P_V)} \leq \mathcal{C}_4, \end{align*} with probability approaching 1 as $n\to \infty$. \end{itemize}

Part (i) of Assumption (ref) is needed to ensure the validity of the Lipschitz continuity requirements of Assumption (ref) for the component functions and Riesz-representers when evaluated at the first-stage estimates. In the context of the Balke-Pearl bounds, it is satisfied when ${\widehat g_n, \widehat f_n}$ are uniformly bounded and the estimated propensity score $\widehat z(X_i)$ is uniformly bounded away from 0 and 1, with probability one. The first condition is satisfied by virtually any estimation procedure when the outcomes $U_i$ have bounded support. The second requirement can be guaranteed under appropriate trimming of the estimated propensities. Part (ii) requires that estimation errors for the nuisance components are uniformly bounded, which is satisfied under Assumption (ref)(i) when $\mathcal{G}\times\mathcal{F}$ is a subset of the space of bounded functions. When $U_i$ has unbounded support and $\mathcal{G}\times\mathcal{F}$ includes unbounded functions, a more primitive condition for (ii) would be uniform consistency of the first stage estimates, that is $\|\{\widehat{g}_n, \widehat f_n \} - \{g,f\}\|_{L_\infty(P_V)} \rightarrow_p 0$.\footnote{However, it should be noted that the uniform consistency requirement is not completely innocuous when $\{\widehat g_n, \widehat f_n\}$ are machine learning estimators farrell2020.}

assumption[$L_2$ convergence rates] The estimators of the nuisance functions satisfy \begin{comment} \begin{align*} &\mathbb{E}\Big[\big(\hat{h}(Z,X) -h(Z,X)\big)^2\Big], \,\mathbb{E}\Big[\big( \hat{m}(D,Z,X) -m(D,Z,X)\big)^2 \Big], \,\sup_{d,z}\mathbb{E}\Big[\big(\hat{e}_{dz}(X) -e_{dz}(X)\big)^2 \Big]\leq \frac{a(n)}{n^{1/2}} \\ &\mathbb{E}\Big[\big(\hat{z}(X) -z(X)\big)^2 \Big],\, \mathbb{E}\Big[\big(\hat{g}(Z,X) -g(Z,X)\big)^2 \Big]\leq\frac{a(n)}{n^{1/2}}, \&\mathbb{E}\Big[\big(\hat{\overline{\tau}}(X) -\overline{\tau}(X)\big)^2 \Big],\,\mathbb{E}\Big[\big(\hat{\tau}(X) -\tau(X)\big)^2 \Big]\leq\frac{a(n)}{n^{1/2}},\\ \end{align*} \end{comment} \begin{align*} &\mathbb{E}_{P_n}\left[ \norm{ \widehat g_n - g }^2_{L_2(P_V)} \right] \leq \frac{r_n}{n^{1/2}},\\ &\mathbb{E}_{P_n}\left[ \norm{ \widehat f_n - f }^2_{L_2(P_V)} \right] \leq \frac{r_n}{n^{1/2}}, \end{align*} for some sequence $r_n=o(1)$.

The above requirement on the $L_2$-convergence rates for the learners of the nuisance functions is a standard assumption in the semiparametric inference literature (see, e.g., farrell2015, farrell2015, and chernozhukov2018, chernozhukov2018). It can be shown to provably hold for traditional nonparametric estimation methods such as sieve methods (chen2007, chen2007) as well as modern black-box machine learning algorithms including Lasso (see, e.g., farrell2015, farrell2015), deep neural networks (farrell2020, farrell2020), boosting and others, for which stronger guarantees such as Donsker-type properties are typically not available. The ability to invoke a mild $L_2$-convergence requirement is a virtue of the combined use of Neyman-orthogonalization and sample-splitting, a key insight brought forward by chernozhukov2018 for semiparametric GMM inference, and subsequently leveraged by atheywager2021 and fosterSyrgkanis in the context of statistical learning problems.\footnote{Unlike atheywager2021, our assumptions do not allow to trade-off accuracy in the estimation across the different nuisance functions. This is because our framework allows for $\varphi_\ell(g;x)$ to be a potentially non-linear functional of the nuisance functions $g$, as is the case in Examples (ref)-(ref), thus precluding such double-robustness property.}

Finally, we present an assumption that concerns the distribution of the component functions $\varphi_\ell$ at the population level.

assumption[Margin]There exist constants $\mathcal{C}_m>0$ and $\gamma\geq0$ such that \begin{align*} \mathbb{P}_X\Big( 0<|\varphi_\ell(g;X)|\leq t\ \Big) \leq \mathcal{C}_m t^\gamma , \quad \forall \,t>0. \end{align*} for $\ell=1,\dots , L$.

The above assumption restricts the extent to which the distribution of the component functions $\varphi_\ell(g;X)$ can concentrate around the point of non-differentiability 0 and it is a form of “margin assumption", first introduced by mammenTsybakov1999. Such an assumption has been widely used in statistics to obtain fast learning rates in classification problems (see, e.g., arlotBartlett2011, arlotBartlett2011). Notice that the above formulation for the margin assumption restricts the concentration of probability for the distribution of the components functions in a neighbourhood of 0, but still allows for arbitrary probability mass at 0.

example[$\gamma=1$] Suppose $X$ contains an absolutely continuous covariate $\mathrm{\tilde{x}}$ and $\varphi_\ell(g;X)\cdot\mathbbm{1}\{\mathrm{\tilde{x}}\neq 0\}$ is absolutely continuous with density bounded above by $\overline{f}$ for $\ell=1,\dots,L$. Then Assumption ((ref)) holds with $\gamma=1$ and $\mathcal{C}_m= 2\overline{f} $.
example[$\gamma=\infty$] Suppose there exists a $t_0>0$ such that $\mathbb{P}_X\left( 0<|\varphi_\ell(X)|<t_0 \right) = 0 $ for $\ell=1,\dots,L$. Then Assumption ((ref)) holds with $\gamma=\infty$ and some $\mathcal{C}_m>0$.

In the context of the Balke-Pearl bounds from Example (ref) with resolution of ambiguity via Minimax Regret, Assumption (ref) restricts the extent to which the CATE bounds $\overline \tau,\underline \tau$ can concentrate around 0 in the data-generating process. Under $\gamma=\infty$ the support of each CATE bound is required to be fully separated from 0, while $\gamma=1$ requires that each CATE bound has bounded density in a neighborhood of 0.

In the next section we present our theoretical results based on Assumptions (ref)-(ref).

Regret convergence rates

In this section, we provide asymptotic rates of convergence for the regret of the estimated policy $R_{n}(P;\widehat{\pi}_{n})$ as defined in (ref). In line with the existing literature, we study uniform regret bounds that are valid for all distributions $P\in\mathcal{P}$ satisfying Assumptions (ref)-(ref). All results in this section are thus intended to hold uniformly in the above sense, and we will drop the dependence on $P$ for notational convenience.

We begin by noticing that controlling the convergence of $\widehat{\pi}_n$ to the best-in-class policy $\pi^*_n$ intuitively requires accounting for: 1) the estimation error in the component functions $\varphi_\ell$ and influence function adjustments $\phi_\ell$ due to estimation of the nuisance components $\{g,f\}$, 2) the difference between the population Neyman-orthogonal score and true score\footnote{That is, we need to account for the fact that we have added the influence function adjustments to the component functions.}, and 3) the fact that we estimate our policy using a sample from the distribution of the covariates $X_i$ rather than their true distribution. We define the following quantities:

align*[align* omitted — 277 chars of source]

and formalize this intuition in the next proposition.

commentWe formalize these heuristics in the following proposition. To account for all of the above, we first define the quantities \begin{align*} {Q}_n(\pi)& = \frac{1}{n}\sum_{i=1}^n \big(2\pi(X_i) -1\big)\cdot\Gamma(g;X_i),\\ \widehat{Q}_n^{NO}(\pi)& = \frac{1}{n}\sum_{i=1}^n \big(2{\pi}(X_i) -1\big)\cdot \Gamma^{NO}(\{\widehat g, \widehat f\};W_i),\\ {Q}^{NO}_n(\pi)& = \frac{1}{n}\sum_{i=1}^n \big(2{\pi}(X_i) -1\big)\cdot\Gamma^{NO}(\{g,f\},W_i), \end{align*} and then follow the typical proof-structure of empirical risk minimisation methods by decomposing regret as follows: \begin{align*} Q(\pi_n^*) - Q(\widehat{\pi}_n) = {\Big[Q(\pi_n^*) - {Q}_n(\pi_n^*) \Big]}+ \Big[Q_n(\pi_n^*) - \widehat{Q}^{NO}_n(\widehat{\pi}_n)\Big] +\Big[\widehat{Q}^{NO}_n(\widehat{\pi}_n) - Q(\widehat{\pi}_n) \Big]. \end{align*} The first term is zero in expectation. The second term can be upper bounded as \begin{align*} \Big[Q_n(\pi^*_n) - \widehat{Q}^{\texttt{NO}}_n(\pi^*)\Big] +\underbrace{\Big[\widehat{Q}^{\texttt{NO}}_n(\pi^*) - \widehat{Q}^{\texttt{NO}}_n(\widehat{\pi}_n)\Big]}_{\leq 0} \leq \sup_{\pi \in \Pi_n}\left| Q_n(\pi) -Q^{\texttt{NO}}_n(\pi)\right| + \sup_{\pi \in \Pi_n}\left| Q^{\texttt{NO}}_n(\pi) - \widehat{Q}^{\texttt{NO}}_n(\pi) \right| . \end{align*} The third term can be further expanded and upper bounded as follows \begin{equation*} \widehat{Q}^{\texttt{NO}}_n(\widehat{\pi}_n) - Q(\widehat{\pi}_n) \leq \sup_{\pi \in \Pi_n}\left|\widehat{Q}^{\texttt{NO}}_n(\pi) - Q^{\texttt{NO}}_n(\pi)\right| +\sup_{\pi \in \Pi_n}\left|Q^{\texttt{NO}}_n(\pi) - Q_n(\pi)\right| +\sup_{\pi \in \Pi_n}\left| Q_n(\pi) - Q(\pi)\right| . \end{equation*}
propositionThe regret of $\widehat{\pi}_n$ obeys the following bound: \begin{align} R_n(\widehat{\pi}_n) \leq &2\mathbb{E}\left[\sup_{\pi \in \Pi_n}\left|\widehat{Q}^{NO}_n(\pi) - Q^{NO}_n(\pi)\right|\right] +\mathbb{E}\left[\sup_{\pi \in \Pi_n}\left|Q^{NO}_n(\pi) - Q(\pi)\right|\right]. \end{align}

The second term in the above bound accounts for points 2) and 3). $Q_n(\pi) - Q(\pi)$ is a centred (mean-zero) empirical process and therefore its uniform expectation can be shown to be $O\left(\sqrt{\text{VC}(\Pi_n)/n}\right)$ using symmetrization and chaining arguments wainwright_2019. Controlling the first term, which accounts for point 1), is particularly challenging and requires tailored arguments that deal with the particular form of the population scores in (ref), in particular their lack of full differentiability.

lemmaSuppose that Assumptions (ref)-(ref) hold and define $\kappa_n=\lfloor n(1-1/K) \rfloor$. Then we have \begin{align*} \mathbb{E}_{P_n}\left[\sup_{\pi \in \Pi_n}\left|\widehat{Q}^{NO}_n(\pi) - Q^{NO}_n(\pi)\right|\right]= O\left( \frac{r_{\kappa_n}}{\sqrt{n}} + \sqrt{\frac{VC(\Pi_n)}{n}} + \left( \frac{r_{\kappa_n}}{\sqrt{n}}\right) ^{\frac{\gamma+1}{\gamma+2}}\right). \end{align*}

Lemma (ref) is the central result of this paper. It provides an asymptotic rate of convergence to zero of the empirical process $\left|\widehat{Q}^{\texttt{NO}}_n(\pi) - Q^{\texttt{NO}}_n(\pi)\right|$ uniformly over the policy class $\Pi_n$, which depends on the VC-dimension of the class and the degree of concentration of the component functions $\varphi_\ell(g;X)$ around 0, as indexed by $\gamma$. In order to convey intuition on this result we provide a brief outline of the proof, which is based on the decomposition

comment\begin{align*} &\widehat{Q}^{NO}_n(\pi) - Q^{NO}_n(\pi) =\frac{1}{n}\sum_{i=1}^n (2\pi(X_i)-1)(\widehat{\Gamma}_i -\widetilde{\Gamma}_i) \\ &=\frac{1}{n}\sum_{i=1}^n (2\pi(X_i)-1) \cdot \Big(\hat{\overline{\tau}}^{NO}_i\cdot\mathbbm{1}\{\hat{\overline{\tau}}_i>0 \} - \overline{\tau}^{NO}_i\cdot\mathbbm{1}\{\overline{\tau}_i >0 \} \Big) + analogous terms for the lower bound \\ &= \underbrace{\frac{1}{n}\sum_{i=1}^n (2\pi(X_i)-1) \cdot\Big(\hat{\overline{\tau}}^{NO}_i - \overline{\tau}^{\texttt{NO}}_i\Big)\cdot\mathbbm{1}\left\{\hat{\overline{\tau}}_i >0 \right\}}_{C_1(\pi)} \\ &+\underbrace{\frac{1}{n}\sum_{i=1}^n (2\pi(X_i)-1)\cdot \phi_{U,i} \cdot\left( \mathbbm{1}\left\{\hat{\overline{\tau}}^{-k(i)}(X_i) \geq0\right\} - \mathbbm{1}\left\{\overline{\tau}(X_i) \geq0 \right\} \right)}_{ C_{2}(\pi)}\&+ \underbrace{\frac{1}{n}\sum_{i=1}^n (2\pi(X_i)-1) \cdot\left(\overline{\tau}_i \cdot\left( \mathbbm{1}\left\{\hat{\overline{\tau}}_i >0\right\} - \mathbbm{1}\left\{\overline{\tau}_i >0 \right\} \right)\right)}_{C_3(\pi)} +\quad \text{analogous terms for the lower bound}. \end{align*}
equation[equation omitted — 443 chars of source]

where

align*[align* omitted — 927 chars of source]

Terms $A_0(\pi)$ and $A_{1,\ell}(\pi)$ can be controlled using similar arguments to atheywager2021 and are responsible for the $O\left(r_{\kappa_n}/\sqrt{n}\right)$ term in the bound of Lemma (ref). The de-biasing properties of Neyman-orthogonalization combined with sample-splitting play a crucial role in this context, as they ensure that the error in estimating $\varphi_\ell(g;x)$ only has a second-order contribution. As a result, term $A_{1,\ell}(\pi)$ scales with the mean-squared estimation error in the nuisance functions and, under Assumption (ref), its expectation decays faster than $1/\sqrt{n}$ uniformly over $\Pi_n$.\footnote{Notice that, by virtue of sample-splitting, the presence of the indicator $\mathbbm{1}\{\varphi_\ell(\widehat g^{-k(i)};X_i)\geq0 \}$ is immaterial when controlling the expectation of $A_{1,\ell}(\pi)$ uniformly over $\Pi_n$.} If plug-in (non-orthogonalized) estimates for $\varphi_\ell$ are instead used to form the score estimates $\widehat{\Gamma}_i$, the estimation error in the nuisance functions has a first-order impact on term $A_{1,\ell}(\pi)$. As a result, its uniform expectation would scale with the $L_1$ estimation error which, under Assumption (ref), implies the much slower convergence $\mathbb{E}[\sup_{\pi \in \Pi_n}|A_{1,\ell }(\pi)|]= o(n^{1/4})$.

For term $A_{2,\ell}(\pi)$, the mean-zero property of the influence function adjustments together with sample-splitting ensures that this term is a centred empirical process and thus it is responsible for a $O\left(\sqrt{\text{VC}(\Pi_n)/n}\right)$ contribution again by symmetrization and chaining arguments.

Finally, for term $A_{3,\ell}(\pi)$ we show that

align*[align* omitted — 284 chars of source]

where the RHS can be recognized to be the classification loss of an estimator for the sign of $\varphi_\ell(g;x)$ based on thresholding $\varphi_\ell(\widehat g^{-k(i)};x)$. Rates of convergence in binary classification problems intuitively depend on the degree of separation of the true regression function from 0, as indexed by $\gamma$. We thus leverage results from the literature on classification audibertTsybakov2007 to quantify the contribution of $A_{3,\ell}$ in the bound of Lemma (ref) in terms of $\gamma$.

We are now ready to combine the rates of convergence for the three terms in Proposition (ref) to obtain a final regret bound for our proposed estimation procedure.

theoremSuppose Assumptions (ref)-(ref) hold. Then the regret obeys \begin{align*} R_n(\widehat{\pi}_n)=O\left( \sqrt{\frac{VC(\Pi_n)}{n} } \vee \left( \frac{r_{\kappa_n}}{\sqrt{n}}\right) ^{\frac{\gamma+1}{\gamma+2}}\right). \end{align*}

We see that regret convergence for our policy learning procedure happens at a rate corresponding to whichever is the leading term in the asymptotic expansion of Lemma (ref), which depends on $\nu$ and $\gamma$. When the policy class $\Pi_n$ has fixed VC-dimension ($\nu=0$), regret convergence happens at rates ranging from $o(n^{1/4})$ in the least favourable case ($\gamma=0$) to $O(\sqrt{\text{VC}(\Pi)/n})$ in the most favourable case ($\gamma=\infty$). The latter case is in line with existing results for policy learning with point-identified CATE, in which full-differentiability of the scores leads to $\sqrt{\text{VC}(\Pi_n)/n}$ learning rates kitagawatetenov2018, atheywager2021,fosterSyrgkanis. For the intermediate case $\gamma=1$ of Example (ref) our procedure guarantees regret convergence at rate $o(n^{1/3})$.

It is useful to compare the performance guarantees in this paper with Pu_2021, whose procedure involves the use of non-orthogonalized estimates for the scores with sample-splitting. They show that the regret of a policy estimated via the maximization $(\ref{estimationPolicy})$ based on cross-fitted non-orthogonalized scores is upper bounded by the $L_1$-norm of the estimation error in the nuisance functions. Under Assumption (ref), this implies $o(n^{1/4})$ convergence for the regret, which is strictly slower than our rates for all values of $\gamma>0$. The faster speed of convergence guaranteed by our procedure is not just due to a refined proof strategy but crucially depends on the use of Neyman-orthogonalization, as elucidated by our discussion of Lemma (ref).

remarkThe procedure of Pu_2021 also differs from ours in its final implementation, which in their case is carried out via support vector machines (SVM) with $\Pi_n$ assumed to be a reproducing kernel Hilbert space. While the use of surrogate losses (such as the hinge loss in SVM) to convexify problem ((ref)) can bring considerable computational benefits in terms of speed and scalability, it comes at the cost of even slower convergence guarantees than the $o(n^{-1/4})$ discussed above. We stress that our insights regarding the benefits of Neyman-orthogonalization in terms of faster learning rates apply irrespective of the final implementation. Notice also that the use of surrogate loss functions does not guarantee convergence of the estimated optimal policy to the best-in-class $\pi^*_n$ in general when the policy class $\Pi_n$ does not contain the “first-best" policy $\mathbbm{1}\{\Gamma(g;x)\geq 0\}$, as shown by the recent work of kitagawaSakaguchi2021.

Empirical application

In this section we apply the methods discussed in this paper to data from the National Job Training Partnership Act (JTPA) Study. This study randomly selected applicants to receive various training and services, including job-search assistance, for a period of 18 months. The study collected background information on applicants before random assignment and then recorded their earnings in the 30-month period following treatment assignment. kitagawatetenov2018 apply their EWM method to a sample of 9,223 adult JTPA applicants to estimate the optimal allocation of eligibility into the programme that maximizes individual earnings across the population. In particular, they take total individual earnings in the 30 months after assignment as the welfare outcome measure $Y_i$, and consider policies that allocate eligibility in the programme based on the individual's observable characteristics. kitagawatetenov2018's analysis is from an intent-to-treat perspective as they focus on the problem of deciding who should be given eligibility to participate in the programme. Since eligibility in the JTPA study is randomly assigned, the effect of eligibility on earnings is point identified from the data and methods for policy learning under point-identification can be applied in this setting. We depart from kitagawatetenov2018 and instead consider optimal assignment of actual participation in the training. This analysis would be of interest to a policy-maker that expects to achieve (close to) perfect compliance to her treatment decision, e.g. when participation is made a condition for receipt of a generous unemployment benefit.\footnote{No financial incentive had been put in place to promote compliance in the implementation of the JTPA study.} Compliance in the JTPA study is imperfect as roughly 23% of applicants' participation status $D_i=0,1$ deviates from their assigned eligibility status $Z_i=0,1$, as shown in Table (ref). As a result, random assignment of the eligibility instrument $Z_i$ is not sufficient to point-identify the effect of participation in the training, motivating the use of the methods proposed in this paper.

table[table omitted — 532 chars of source]

For partial identification of the CATE we consider the Balke-Pearl scheme of Example (ref), where bounds for the 30-month post-treatment earnings are $Y_L=\$0$ and $Y_U=\$59,640$.\footnote{The outcome upper bound corresponds to the 97.5th percentile of the earnings distribution rather than highest recorded value of $\$155,760$. Outcome bounds in Balke-Pearl bounds effectively impute unidentified expected earnings for never-takers and always-takers. Restricting expected earnings to be below such high quantile is in effect a mild requirement which brings considerable identification power.} We compare this with point-identification of the CATE as the conditional local average treatment effect (LATE), predicated under the assumption of no unobserved heterogeneity. We subtract \$1216 from both the CATE bounds and the conditional LATE; this is the average cost of services per actual treatment, estimated from Table 5 in bloom1997. Following kitagawatetenov2018, we condition treatment assignment on two pre-treatment variables: the individual's years of education and earnings in the year prior to assignment. Estimation of the optimal policy follows the procedure described in Section (ref), with $K=10$ evenly-sized data folds used to form cross-fitted Neyman-orthogonal estimates for the CATE bounds and conditional LATE functions. The nuisance functions are estimated via boosted regression trees, performed by the MATLAB function \verb+fitrensenmble+.\footnote{Tuning parameters have been chosen via cross-validation within each data-fold. For further details on the estimation procedure we refer to the MATLAB documentation for the command.}

Figure 1 demonstrates cross-fitted plug-in estimates for the CATE bounds (a) and the LATE/minimax regret scores (b), where the size of the dots indicates the number of individuals with different covariate values. We first notice that the estimated CATE lower bounds are negative for the whole sample, and thus the maximin impact optimal policy never assigns treatment in this application. We therefore focus our analysis on minimax regret (MMR).\footnote{Maximin welfare also results in no treatment for the whole population in this application.} Comparison of the conditional LATE and MMR scores highlights how partial identification leads to increased variation of the scores across different levels of education and pre-program earnings. In particular, the MMR scores are considerably higher (and positive) for individuals with fewer years of education and smaller pre-programmes. The conditional LATE estimates display less overall variation over the support of the covariates, compared to the MMR scores, but are lower for individuals with 0 pre-programme earnings.

figure[figure omitted — 422 chars of source]

We consider three alternative choices for the candidate policy class $\Pi$. The first is the class of quadrant treatment policies. To be assigned to treatment according to this policy, an individual's education and pre-program earnings have to be above (or below) some specific threshold. Figure (ref) illustrates the optimal quadrant treatment policies based on Neyman-orthogonalized cross-fitted MMR and LATE scores, where the colored shaded areas indicate individuals required to undertake job training by the respective policies. The optimal MMR policy (green) assigns treatment to individuals with education below 15 years and pre-treatment earnings below \$39,952. The optimal LATE policy (blue) selects the same threshold for education, but selects individuals with pre-treatment earnings above \$200 for treatment.

comment\begin{table}[H] \caption{Treatment proportions of alternative treatment assignment policies} \scalebox{1}{ \begin{threeparttable} \begin{tabular}{lcccccccc} \toprule &&&\shortstack{Share of Population\\ to be treated} & \shortstack{Share of Population receiving \\same treament under MMR and LATE}\\\midrule \multicolumn{4}{@c}{Quadrant Rule}& 0.68\\ Minimax Regret&&&$0.96$&--\\ LATE&&&$0.64$&--\\ \multicolumn{4}{@c}{Linear Index Rule}&0.70\\ Minimax Regret &&&$0.95$&--\\ LATE&&&$0.69$& --\\ \bottomrule \end{tabular} \begin{tablenotes} • The rows labeled “Minimax Regret" give information on estimated optimal minimax regret policies based on the scores in Equation (ref) with Balke-Pearl CATE bounds. The rows labeled “LATE" give information on optimal policies for scores that identify the conditional LATE. \end{tablenotes} \end{threeparttable}} \end{table}

While the two policies appear similar, they substantially differ in the proportion of population assigned to treatment (96% by the MMR policy versus 64% by the LATE policy), as shown in Table (ref). This is due to the large concentration of individuals with pre-treatment earnings close to (or equal) zero. As a result, 32% of individuals receive a different treatment assignment across the two policies. Figure (ref) also shows the optimal “na\"ive" MMR policy based on cross-fitted but non-orthogonalized scores (yellow), which recommends participation into the programme for the entire population. {

table[table omitted — 1,676 chars of source]

}

Second, we consider the class of linear treatment policies. This class consists of policies that assign treatment to an individual according to whether a linear index in his observable characteristics is above a certain threshold. Figure (ref) illustrates how the direction of treatment assignment as a function of prior earnings differs between the MMR and LATE policy in a similar fashion to the quadrant rules; contrary to the LATE policy, MMR prioritizes treatment assignment to individuals with lower pre-program earnings. Nonetheless, 70% of the population still receives the same treatment under the two different policies, in light of the relatively low concentration of individuals in the areas of the covariate space where the two policies differ. Similarly to the quadrant policy rule, the MMR policy assigns treatment to a larger share of the population (95%) compared to the LATE policy (69%). The na\"ive MMR policy is qualitatively similar to the one using Neyman-orthogonolization, but recommends programme participation to a larger share of individuals.

Finally, we consider linear treatment policies that additionally include quadratic and cubic terms for education. Figure (ref) shows how the additional flexibility in the policy class leads to rules that are less interpretable but maintain similar qualitative features compared to the more parsimonious classes previously considered.

comment\begin{figure}[H] \caption{Estimated optimal policies from the quadrant policy class} \end{figure} \begin{figure}[H] \caption{Estimated optimal policies from the linear-index policy class} \end{figure}
figure[figure omitted — 251 chars of source]
figure[figure omitted — 190 chars of source]
figure[figure omitted — 306 chars of source]
comment\begin{figure} \begin{subfigure}[b]{.48\linewidth} \caption{A mouse} \end{subfigure} \begin{subfigure}[b]{.48\linewidth} \caption{A gull} \end{subfigure} \begin{subfigure}[b]{.48\linewidth} \caption{A tiger} \end{subfigure} \caption{Picture of animals} \end{figure}

Conclusion

This paper develops a general policy learning framework for estimation of individualized treatment rules when treatment effects are partially identified. By drawing connections between the treatment assignment problem and classical decision theory, we have characterized several notions of optimal treatment policies in the presence of partial identification. We have shown how partial identification leads to a new policy learning problem where the risk is only directionally-differentiable with respect to a nuisance infinite-dimensional component. We have proposed an estimation procedure that ensures Neyman-orthogonality with respect to the nuisance components and we have provide statistical guarantees that depend on the amount of concentration around the points of non-differentiability in the data-generating process. Our proposed methods are illustrated with an application to the Job Training Partnership Act study, where we have shown that allowing for partial identification delivers substantially different programme participation policies compared to existing methods that assume point-identification.

There are several avenues for future research. First, it would be interesting to extend the theory of this paper to partial identification via instrumental variables with continuous support. Second, it would be useful to extend the methods to more general identification sets that incorporate smoothness restrictions on unobserved counterfactual quantities, such as those considered in simonLeeBounds2018. Finally, it would be interesting to assess the optimality of our proposed estimation procedure by deriving minimax lower bound rates for semiparametric statistical learning problems with directionally-differentiable risk.

commentThe ideal empirical application for this paper satisfies two main requirements. The first concerns the motivation for moving from ITT to ATE policies. In particular, the empirical setting should be such that imperfect compliance in the (quasi-)experimental research design for the available data is matched by virtually full control of treatment take-up in the policy implementation. The second requirement is that compliance needs to be high enough to guarantee sufficiently narrow bounds for the CATE (under reasonable identification assumptions). If the bounds are too wide then it is unlikely that a policy maker would want to base the implementation of a policy on that dataset. Furthermore, if the sign of the CATE is not identified at any point in the covariate support $\mathcal{X}$, then we automatically have that the optimal maximin welfare estimated policy is “treat nobody", irrespective of the policy class $\Pi$. A first candidate choice for the empirical application is the Job Partnership Training Act dataset. The instrument here would be `eligibility to get training' and the treatment is `actual training participation', with the outcome being earnings some time after the training. This is the empirical application in kitagawatetenov2018 who estimate the optimal ITT policy. While this application features high compliance levels (only 23% of individuals do not comply, mostly never-takers), it is hard to motivate an analysis based on assignment of the actual training in place of the ITT analysis of kitagawatetenov2018. One potential argument is that job training could be offered alongside a much more generous compensation than the one in the JPTA (e.g. as a condition for receipt of a generous unemployment benefit), which would result in very high compliance in practice. A second candidate empirical application is the Oregon Health Insurance Experiment (OHIE), in which a random sample of household was given eligibility to apply for free Medicaid insurance.