Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
82,086 characters · 14 sections · 49 citation commands
Distributionally Robust Policy Learning with Wasserstein Distance
The effects of treatments are often heterogeneous based on observable covariates. In the presence of multiple treatments, an important decision for a policymaker is to choose a rule that specifies a treatment for each value of covariates to maximize their objective function. Such a rule is often called an individualized treatment rule (ITR), and its significance has been acknowledged in many areas including healthcare Bertsimas2017, homeless services Kube2019, and energy conservation Ida2021. The typical decision-making process of a policymaker is as follows. They are interested in a specific population, called the target population. Population here refers to the joint distribution of potential outcomes and covariates. In determining an ITR, they can use experimental or observational data generated from another population, which is referred to as the source population. Note that they can identify they source covariate distribution and source distribution of potential outcome associated with each treatment conditional on a possible value of covariates from this data. Based on this data, they decide upon an ITR, and then, the resulting ITR is applied to the target population. Finally, the average outcome attained by the ITR is realized as the policymaker's reward. The average outcome is often called welfare; therefore, a policymaker's primal goal is to choose an optimal ITR that maximizes the target population's welfare. Most existing studies have developed efficient estimation methods of optimal ITRs under the implicit assumption that the target and source populations are essentially identical. Under this assumption, the target population's welfare attained by arbitrary ITRs can be point-identified. Consequently, the optimal ITR for the target population also becomes identifiable. Based on this fact, previous studies have proposed an estimation method of optimal ITR utilizing the inverse-probability weighting Kitagawa2018,Swaminathan2015,Zhao2012 and augmented inverse-probability weighting Athey2021,Dudik2011,Zhou2022. However, in practice, whether this assumption is valid is controversial. For example, imagine that a state government wants to decide whether or not to allow each individual to participate in an employment support program to improve average earnings in the state. The available data on the program come from a randomized controlled trial conducted in another state. In the framework above, two available treatments exist: one is to force an individual to participate in the program, and the other is to exclude the individual from the program. Associated with each treatment, the potential outcome is earnings that would be realized when the individual is assigned to the treatment. In addition, observable individual attributes, such as age and previous earnings, compose covariates. The joint distribution of potential outcomes and covariates in the former state corresponds to the target population, and that in the latter state corresponds to the source population. The critical question is whether the former and latter states can be considered the same. The demographics of the two states may be different. In addition, the response to the program may also be different because of differences in other economic circumstances. In such cases, it would be unrealistic to assume that the two populations are the same. The scenarios in which the assumption can be violated are not limited to the above example. For instance, when the state government wants to exploit a randomized controlled trial conducted earlier in its state to determine the current optimal ITR for the state, it must consider the validity of the assumption judiciously. When the source and target populations are different, it is challenging to estimate the optimal ITR via the straightforward application of existing methods. However, in some cases, a slight modification in existing methods is sufficient to estimate the optimal ITR for the target population. Suppose that a difference exists between the source and target populations only in terms of covariate distribution and the density ratio of the target and source covariate distributions is known or identifiable with additional covariates data from the target population. In this case, the target population's welfare attained by an ITR remains identifiable by reweighting the outcome with the density ratio of source and target covariate distributions. Kitagawa2018 and Uehara2020 discuss estimation methods that are tailored to this situation. Instead, if the distributions of covariates or distributions of potential outcomes associated with each treatment conditional on a value of covariates differ between the source and target populations in an unknown and unidentifiable way, it is impossible to estimate the optimal ITR for the target population. In the first place, the target population's welfare attained by an ITR cannot be point-identified, which is a contrast to the case where the two populations are the same. This implies that the optimal ITR for the target population cannot be identified; the estimation goal itself is unclear. Motivated by this identification problem in such a situation, one strand of recent literature has proposed replacing the unidentifiable goal with an identifiable one. In particular, Mo2021, Si2020,Si2021, and Zhao2019 consider utilizing the concept of distributionally robust optimization (DRO) to obtain a reasonable learning goal. DRO begins with constructing an ambiguity set, which is a set of populations, based on the available knowledge and additional assumptions regarding the relationship between the source and target populations. Then, for each feasible ITR, it evaluates the \emph{distributionally robust welfare}, which is the worst-case value of welfare over the ambiguity set. Finally, it chooses the ITR that maximizes the distributionally robust welfare. Following Mo2021, this study refers to such an ITR as a \emph{distributionally robust ITR (DR-ITR)}. By choosing the DR-ITR, it is guaranteed that the target population's welfare is at least tantamount to its distributionally robust welfare, as long as the target population belongs to the ambiguity set. From this perspective, the DR-ITR would be a reasonable goal. Although these attempts are the same, when they are applications of DRO, they vary based on underlying assumption. Zhao2019 and Mo2021 assume that the target population is absolutely continuous with respect to the source population and that a difference in the source and target populations exists only in the covariate distribution. Then, with the available data being the experimental or observational data generated from the source population, they construct the ambiguity set on the unknown target covariate distribution. Instead, Si2020,Si2021 include the case where the conditional distributions can also differ in the two populations, while assuming that the target population is absolutely continuous with respect to the source population. Using the experimental or observational data, which are derived from the source population, they construct the ambiguity set on the target population. As is clear from the current discussion, these attempts may not be appropriate when the assumed absolute continuity does not hold. This study also aims to define a reasonable ITR when the target population's welfare is not identifiable by exploiting the concept of DRO. Specifically, this study mainly deals with the situation where the policymaker has access to the experimental data from the source population and the covariate data from the target population. In other words, the policymaker has no knowledge about the target conditional distribution of potential outcomes. In contrast to the aforementioned attempts Mo2021,Zhao2019,Si2020,Si2021, this study does not assume that the target conditional distribution of potential outcomes is absolutely continuous with respect to that of the source population, while the absolute continuity of the covariate distributions is maintained. To construct the ambiguity set without absolute continuity, I employ \emph{Wasserstein distance} or order 1. This is a metric over probability distributions that does not require the Radon-Nykodim derivative, and thus, the proposed method can cover broader scenarios. In addition to this advantage, the proposed Wasserstein-style ambiguity set is computationally attractive. In particular, with this ambiguity set, one can obtain the analytical solution of the distributionally robust welfare. Owing to this analytical simplicity, one can draw a lot of intuition. For example, it can be shown that obtaining an ITR by replacing the unknown target conditional distribution of potential outcomes with that of the source population is already distributionally robust under special cases in the Wasserstein sense. This provides a justification for the assumption that the source and target conditional distributions of potential outcomes are the same. In addition, by utilizing the technique developed in Mo2021 or Zhao2019, one can easily extend the proposed method to a situation where no covariate data are available from the target population. For the estimation of the DR-ITR, the derived closed-from of the distributionally robust welfare implies a simple estimator. I assess its theoretical performance in terms of regret, which is in line with the literature on policy learning. Finally, I demonstrate the proposed DR-ITR using data from experimental evaluations of the changes in welfare-to-work programs in the early 1980s. The results show that the DR-ITR attains higher welfare in the target population than the ITR obtained by a naive approach. \paragraph{Related Literature} This study contributes to the literature on ITRs Kitagawa2018,Zhao2012,Swaminathan2015,Athey2021,Dudik2011,Zhou2022. The methods developed in existing studies are intended for use in situations where the source and target populations are the same. The policy learnings when the two populations are different are not well-explored. This study designs a DR-ITR that works appropriately in such situations. The fundamental problem that motivates this study is the so-called \emph{external validity} Campbell1957. The studies that explore external validity include Hotz2005,Cole2010,Stuart2011,Stuart2015,Stuart2018,Allcott2015,Andrews2019,Dehejia2021,Pritchett2013,Hartman2015,Gechter2019,Vivalt2020,Tipton2013,Tipton2014. The primary interest of the literature is to obtain a point identification of the average treatment effect for the target population. To achieve this goal, the literature often assumes that the conditional distributions of potential outcomes are the same between the source and target populations. By contrast, this study allows a difference in the conditional distributions, and DRO attempts to obtain the lower bound of the target population's welfare. As mentioned earlier, the idea of this study stems from DRO with a Wasserstein-style ambiguity set. For a comprehensive review of DRO, see Rahimian2019. A typical application of Wasserstein-DRO constructs an ambiguity set around the empirical distribution and lets its radius shrink to zero depending on the sample size to capture sampling uncertainty Esfahani2018,Blanchet2019,ShafieezadehAbadeh2015,ShafieezadehAbadeh2019,Gao2017a,Gao2020. By contrast, the ambiguity set considered in this study is not centered around empirical distribution, but is centered around the source population, which differs from empirical distribution. In addition, its radius is set at a fixed positive value and does not converge to zero. Adjaho2022 conducted a similar study.\footnote{This study and Adjaho2022 are independent contributions made public almost simultaneously in the arXiv repository.} They also apply Wasserstein-DRO to a similar situation to develop a method that works when the source and target populations are different. One apparent difference is a concrete form of the Wasserstein-style ambiguity set. Specifically, Adjaho2022 deals with a similar situation as that in this study, but the derived results are different due to the difference in the ambiguity set. Their result implies that the modifications discussed in Kitagawa2018 and Uehara2020 are already distributionally robust, while the result of this study implies that the distributional robustness of the modifications does not necessarily hold and this study's proposed approach outperforms the modifications in some aspects. In addition, Adjaho2022 extends Wasserstein-DRO to the case where both the target covariate distribution and target conditional distribution of potential outcomes are unknown and unidentifiable, without imposing the assumption of the absolute continuity of the covariate distributions. Their result implies that the resulting distributionally robust welfare remains unidentifiable without further assumptions. By contrast, this study extends the proposed method to such a case while maintaining the assumption of the absolute continuity of the covariate distributions, and adopts the DRO technique developed in Mo2021,Zhao2019. \paragraph{Structure of the paper} In (ref), I formally introduce the underlying model considered in this study. Then, in (ref), I discuss the application of Wasserstein-DRO to the model of (ref). (ref) develops the estimation method of the proposed DR-ITR and provides the theoretical guarantee of the estimator. Finally, (ref) demonstrates the proposed DR-ITR using experimental data. The additional information and proofs of the theoretical results are presented in the supplementary materials. \paragraph{Notation} For any metric space $(\mathcal{S},d_{\mathcal{S}})$, I use $\mathscr{B}(\mathcal{S})$ to denote Borel $\sigma$-algebra, which is the $\sigma$-algebra generated from the metric topology of $\mathcal{S}$. In addition, I write $\mathscr{P}(\mathcal{S})$ for the set of all probability measures on $(\mathcal{S},\mathscr{B}(\mathcal{S}))$. For a probability measure $\mathbb{P} \in \mathscr{P}(\mathcal{S})$, the support of the measure is denoted by $\operatorname{\mathrm{supp}}(\mathbb{P})$. Given a measure space $(\operatorname{\mathcal{S}},\mathscr{B}(\mathcal{S}),\mathbb{P})$, another measurable space $(\mathcal{T},\mathscr{B}(\mathcal{T}))$, and a measurable map $f:(\mathcal{S},\mathscr{B}(\mathcal{S}))\mapsto(\mathcal{T},\mathscr{B}(\mathcal{T}))$, I denote the induced probability measure on $\mathscr{B}(\mathcal{T})$ by $f_{\#}\mathbb{P}$; that is, $f_{\#}\mathbb{P}(A)=\mathbb{P}(f^{-1}(A))$ for all $A\in\mathscr{B}(\mathcal{T})$. In addition, if $(\mathcal{T},\mathscr{B}(\mathcal{T})) = (\mathbb{R},\mathscr{B}(\mathbb{R}))$, the expectation of $f$ with respect to the measure $\mathbb{P}$ is denoted by $\mathbf{E}_{\mathbb{P}}[f]$.
Here, I formally introduce the underlying model considered in this study ((ref)). After introducing the model, I explain the naive approach that assumes that the two populations differ only in the covariate distributions ((ref)). This approach often works as a bench mark. Before delving into details, a note of caution: (ref) exclude the consideration of identification and estimation and focus on the problem of interest in the population level. In particular, I often assume that a policymaker has the knowledge of a certain population, which is a joint distribution of potential outcomes and covariates, although the joint distribution of potential outcomes is never identified even with the experimental data. The identification and estimation will be discussed in (ref).
Let $\operatorname{\mathcal{A}}=\{1,\cdots,d\}$ be a finite set of possible treatments. Each individual in a population is characterized by a tuple $(Y_{1},\cdots,Y_{d},X)$, where $Y_{a}$ denotes the potential outcome that would be realized if one received treatment $a\in\operatorname{\mathcal{A}}$. I assume that for any $a\in\operatorname{\mathcal{A}}$, $Y_{a}$ takes a value in the outcome space $\operatorname{\mathcal{Y}}$, which is a closed subset of a Euclidean space. Then, a tuple $(Y_1,\cdots,Y_d)$ takes values in $\operatorname{\mathcal{Y}}^d$, which is assumed to be equipped with the $\ell_1$-distance; that is, for any pair $(y_1,\cdots,y_d),(y_1',\cdots,y_d') \in \operatorname{\mathcal{Y}}^d$, the distance is measured by $\sum_{a \in \operatorname{\mathcal{A}}} |y_a - y_a'|$. The last component $X$ denotes the individual's observable characteristics, which take value in a Polish space $\operatorname{\mathcal{X}}$. Let $\operatorname{\mathcal{Z}} := \operatorname{\mathcal{Y}}^d \times \operatorname{\mathcal{X}}$; thus, a tuple $(Y_1,\cdots,Y_d,X)$ takes values in $\operatorname{\mathcal{Z}}$. A particular population is characterized by a probability measure over $(\operatorname{\mathcal{Z}},\operatorname{\mathscr{B}}(\operatorname{\mathcal{Z}}))$. For an arbitrary population $\operatorname{\mathbb{U}} \in \operatorname{\mathscr{P}}(\operatorname{\mathcal{Z}})$, let $\operatorname{\mathbb{U}}_{(Y_1,\cdots,Y_d)\mid X=x} \in \operatorname{\mathscr{P}}(\operatorname{\mathcal{Y}}^d)$ denote the conditional distribution of potential outcomes, given by $X=x$, and let $\operatorname{\mathbb{U}}_X \in \operatorname{\mathscr{P}}(\operatorname{\mathcal{X}})$ denote the marginal distribution of $X$.\footnote{In this study, the conditional distributions of potential outcomes are understood as a regular conditional probability measure. The existence of a regular conditional probability measure is guaranteed by Bogachev2007 because the covariate space is Polish. Hence, any conditional distribution of potential outcomes can be regarded as a proper probability measure in the outcome space, as compared to the case where one defines a conditional probability as a Radon-Nikodym derivative.} Additionally, I define the conditional mean response (CMR) function as $m_{\operatorname{\mathbb{U}}}(x,a):=\mathbf{E}_{\operatorname{\mathbb{U}}_{(Y_{1},\cdots,Y_{d})|X=x}}[Y_{a}]$ for $a\in\operatorname{\mathcal{A}}$ and $x \in \operatorname{\mathrm{supp}}(\operatorname{\mathbb{U}}_X)$. An ITR is a measurable mapping from $\operatorname{\mathcal{X}}$ to $\operatorname{\mathcal{A}}$, and a set of candidate ITRs is denoted by $\operatorname{\mathcal{G}}$. As emphasized in Kitagawa2018 and often assumed in subsequent studies, the set $\operatorname{\mathcal{G}}$ may not necessarily contain all measurable mappings. For any ITR $g \in \operatorname{\mathcal{G}}$, the welfare of ITR $g$ in population $\operatorname{\mathbb{U}} \in \operatorname{\mathscr{P}}(\operatorname{\mathcal{Z}})$ is given by
where the last equality comes from the law of iterated expectations. Hence, the welfare of ITR $g$ can be calculated using the knowledge of the covariate distribution and CMR function of population $\operatorname{\mathbb{U}}$. Clearly, the welfare of ITR $g$ varies with the population under consideration. A policymaker focuses on the target population $\operatorname{\mathbb{T}}\in\operatorname{\mathscr{P}}(\operatorname{\mathcal{Z}})$, and aims to maximize the target population's welfare. Namely, given the target population $\operatorname{\mathbb{T}}$, their goal would be to obtain the optimal ITR $g_{\operatorname{\mathbb{T}}}^*$ for the target population as summarized in the following optimization problem:
It is obvious that the policymaker can achieve the goal if they have complete knowledge about $\operatorname{\mathbb{T}}$. More precisely, knowledge about the target covariate distribution ${\operatorname{\mathbb{T}}}_X$ and target CMR function $m_{\operatorname{\mathbb{T}}}(x,a)$ for all $a \in \operatorname{\mathcal{A}}$ and $x \in \operatorname{\mathrm{supp}}({\operatorname{\mathbb{T}}}_X)$ is necessary and sufficient to solve problem (ref). In other words, the policymaker cannot obtain the best ITR $g_{\operatorname{\mathbb{T}}}^*$ without such knowledge. In this study, I consider the situation where the policymaker has no knowledge on the target conditional distribution of potential outcomes. Contrarily, I assume that they have knowledge on the target covariate distribution $\operatorname{\mathbb{T}}_X$ and source population $\operatorname{\mathbb{S}} \in \operatorname{\mathscr{P}}(\operatorname{\mathcal{Z}})$, which differs from the target population. The two populations can differ in an arbitrary way as long as the following assumption is satisfied.
Similar assumptions to (ref) have been imposed in the literature on external validity Hotz2005. (ref) is satisfied when potential outcomes are $\operatorname{\mathbb{S}}$-almost surely bounded. In summary, the policymaker does not know anything about the target CMR function, but instead knows the source and target covariate distribution and source CMR function. This implies that they cannot calculate the target population's welfare $V(g;\operatorname{\mathbb{T}})$, whatever the ITR $g$ is. Consequently, they cannot obtain the optimal ITR $g_{\operatorname{\mathbb{T}}}^* \in \operatorname{\mathcal{G}}$ for the target population.
The reason the policymaker cannot obtain the optimal ITR for the target population is that the target CMR function $m_{\operatorname{\mathbb{T}}}(x,a)$ is unknown, in contrast to the source CMR function $m_{\operatorname{\mathbb{S}}}(x,a)$. Thus, a naive approach is to replace the target CMR function with the source CMR function and solve the optimization problem; that is,
Note that the problem (ref) is well-defined as an optimization problem due to (ref); otherwise, there exists a point $x \in \operatorname{\mathrm{supp}}(\operatorname{\mathbb{T}}_X)$ such that $m_{\operatorname{\mathbb{S}}}(x,a)$ is not uniquely determined. Hereafter, I refer to this approach as the naive approach and denote its optimal solution by $g_{\text{navie-ITR}}^*$, as indicated in (ref). When the target and source CMR functions are the same, the objective function in (ref) corresponds to the target population's welfare $V(g;\operatorname{\mathbb{T}})$, and hence, $g_{\text{naive-ITR}}^*$ equals $g_{\operatorname{\mathbb{T}}}^*$. This is the case where one assumes that the source and target conditional distributions of potential outcomes are the same, as often assumed in the literature on external validity Hotz2005. Unfortunately, however, one cannot verify this assumption from the available knowledge.
This section formally discusses the application of DRO to the current setting. First, I define the DR-ITR with an arbitrary ambiguity set, which is a set of populations. Generally, an ambiguity set is constructed based on available knowledge. I denote the ambiguity set by $\mathscr{U}(\operatorname{\mathbb{S}},{\operatorname{\mathbb{T}}}_X) \subset \mathscr{P}(\mathcal{Z})$. As a policymaker has knowledge of the source population and target covariate distribution, it is natural that the ambiguity set can depend on $\operatorname{\mathbb{S}}$ and $\operatorname{\mathbb{T}}_X$. Given the ambiguity set, DRO solves the following optimization problem:
where
The object $\underline{V}(g)$ is the distributionally robust welfare, which is the worst-case value of welfare over the ambiguity set. Thus, DRO's goal is to choose the DR-ITR $g_{\text{DR-ITR}}^{*}$ that maximizes the distributionally robust welfare.\footnote{One of the potential drawbacks of DRO is that the max-min criterion may be too conservative to generate a valid policy. This drawback can be alleviated by adopting the Hurwicz criterion, which allows a policymaker to decide the extent to which they become conservative Hurwicz1951. Then, the objective function can be defined as a convex combination of the worst-case and best-case values for welfare. All the results presented in this study are easily extensible with minor modifications.} Suppose that the target population $\operatorname{\mathbb{T}}$ is contained in the ambiguity set $\mathscr{U}(\operatorname{\mathbb{S}},\operatorname{\mathbb{T}}_X)$. In this case, the optimal ITR $g_{\operatorname{\mathbb{T}}}^*$ for the target population defined in (ref) performs better than the distributionally robust policy $g_{\text{DR-ITR}}^*$ in terms of the target population's welfare; that is, $V(g_{\operatorname{\mathbb{T}}}^*;\operatorname{\mathbb{T}}) \geq V(g_{\text{DR-ITR}}^*;\operatorname{\mathbb{T}})$. However, it is infeasible to obtain $g_{\operatorname{\mathbb{T}}}^*$ with the current knowledge. Instead, choosing the DR-ITR ensures that the ex-post target population's welfare is at least equal to its distributionally robust welfare; that is, $V(g_{\text{DR-ITR}}^*;\operatorname{\mathbb{T}}) \geq \underline{V}(g_{\text{DR-ITR}}^*)$. In the aforementioned DR-ITR formulation, several important issues remain unclear. First, I introduced the ambiguity set $\mathscr{U}(\operatorname{\mathbb{S}}, \operatorname{\mathbb{T}}_X)$ in an abstract manner. The choice of the ambiguity set is left to the policymaker's discretion, and they can construct the set based on the available information and their belief in the difference between the source and target populations. (ref) instantiate the DR-ITR with a Wasserstein-style ambiguity set. Second, given the specific choice of the ambiguity set, the DR-ITR must evaluate the distributionally robust welfare $\underline{V}(g)$ for each ITR $g \in \mathcal{G}$. However, this is deemed difficult because the minimization problem in (ref) generally involves an infinite number of distributions. Nevertheless, (ref) shows that the distributionally robust welfare can be calculated very efficiently. In addition to these issues, I discuss what kind of population is determined as the worst case by Wasserstein DRO ((ref)) and some equivalence of the DR-ITR and naive-ITR under special cases ((ref)).
In this study, the ambiguity set is constructed with the Wasserstein distance of order 1. Wasserstein distance is a kind of distance of probability measures that utilizes the structure of the underlying metric space. In the current setup, for a pair $(\operatorname{\mathbb{S}}_{(Y_1,\cdots,Y_d)\mid X=x},\operatorname{\mathbb{U}}_{(Y_1,\cdots,Y_d)\mid X=x})$ of probability measures in $\operatorname{\mathscr{P}}(\operatorname{\mathcal{Y}}^d)$, the Wasserstein distance of order 1 is defined by
where $\Pi$ denotes the set of all the joint distributions of $(Y_1,\cdots,Y_d,Y_1',\cdots,Y_d')$, such that $(Y_1,\cdots,Y_d) \sim \operatorname{\mathbb{S}}_{(Y_1,\cdots,Y_d)\mid X=x}$ and $(Y_1',\cdots,Y_d') \sim \operatorname{\mathbb{U}}_{(Y_1,\cdots,Y_d)\mid X=x}$. It is well-known that there exists $\pi \in \Pi$ that attains the infimum on the right-hand side of (ref). In addition, in our current application, it is important that the distance does not require Radon-Nikodym derivative. For a more detailed explanation, see Villani2009. Based on Wasserstein distance, I introduce the ambiguity set considered in this study. Concretely, given the source population $\operatorname{\mathbb{S}}$, the target covariate distribution $\operatorname{\mathbb{T}}_X$, and a real number $\delta \geq 0$, the ambiguity set is defined as
Note that the ambiguity set (ref) is well-defined under (ref) as the source conditional distribution of potential outcomes, $\operatorname{\mathbb{S}}_{(Y_1,\cdots,Y_d)\mid X=x}$, is unique for all $x \in \operatorname{\mathrm{supp}}(\operatorname{\mathbb{T}}_X)$. A population $\operatorname{\mathbb{U}}$ is contained in the ambiguity set (ref) only if it satisfies the following properties: (i) for each $x$ in support of the target covariate distribution $\operatorname{\mathbb{T}}_X$, its conditional distribution of potential outcomes at point $x$ differs from the counterpart of the source population by at most $\delta$ in terms of Wasserstein distance; (ii) its marginal distribution of covariates equals the target covariate distribution, $\operatorname{\mathbb{T}}_X$. That is, the ambiguity set point wisely constructs Wasserstein balls with the source conditional distribution of potential outcomes as the center and $\delta$ as the radius. Importantly, this allows the conditional distribution of potential outcomes to differ from that of the source population. The radius $\delta$ represents the ambiguity level of the target population, and its choice depends on the policymaker. On the one hand, the value should be large enough that the target population is contained in the ambiguity set; otherwise, it is not guaranteed that the distributionally robust welfare of the DR-ITR is a lower bound of its target population's welfare. On the other hand, the value should also be small. When the value is excessively large, the decision based on DRO can be too conservative to be helpful. Hence, the optimal value of $\delta$ would be $\delta=\sup_{x\in\operatorname{\mathrm{supp}}(\operatorname{\mathbb{T}}_{X})}W_{1}(\operatorname{\mathbb{S}}_{(Y_{1},\cdots,Y_{d})|X=x},\operatorname{\mathbb{T}}_{(Y_{1},\cdots,Y_{d})|X=x})$. Unfortunately, such a choice is infeasible in the current setup as knowledge on the target conditional distribution is lacking. Instead, (ref) shows the relationship between certain welfare-relevant parameters and the value of $\delta$, which guides the choice of $\delta$.
The first two items characterize the difference in the conditional distributions of potential outcomes for each possible $x$. In particular, the first one states that as long as a population is in the ambiguity set, its CMR function differs from the source CMR function by at most $\delta$. Similarly, the second one states that the difference in the CMR functions between any pairs of treatments differs from the counterpart of the source population by at most $\delta$. The latter two items characterize the difference in the marginal distributions of potential outcomes, while focusing on its mean. For the third item, the upper bound consists of the ambiguity level $\delta$ and an additional term that essentially measures the difference in the covariate distributions $\operatorname{\mathbb{T}}_X$ and $\operatorname{\mathbb{S}}_X$. A similar interpretation holds for the fourth item. Note that this result does not address the tightness of the aforementioned inequalities. In particular, the result does not imply the existence of a distribution $\operatorname{\mathbb{U}}\in\mathscr{U}(\operatorname{\mathbb{S}},\operatorname{\mathbb{T}}_X)$ with $|m_{\operatorname{\mathbb{U}}}(x,a)-m_{\operatorname{\mathbb{S}}}(x,a)|=\delta$ for some $x$ and $a$. The existence of such distributions are guaranteed in (ref), which will be presented later.
To solve problem (ref), one must evaluate the distributionally robust welfare $\underline{V}(g)$ in (ref). However, the infimum generally involves an infinite number of distributions and is not easy to calculate directly. In this section, I provide a tractable reformulation of the problem (ref) by exploiting the strong duality results from Blanchet2019a, which simplifies our analysis. First, the law of iterated expectations implies that the minimization problem (ref) can be solved component-wise. Specifically, it can be written as $\underline{V}(g) = \mathbf{E}_{\operatorname{\mathbb{T}}_X}[\underline{v}(X;g)]$, where
for each $x \in \operatorname{\mathrm{supp}}(\operatorname{\mathbb{T}}_X)$. The object $\underline{v}(x;g)$ can be interpreted as the worst-case conditional welfare of policy $g$ at covariate value $x$. With these notations, the evaluation must be reduced to $\underline{v}(x;g)$ for all $x\in\operatorname{\mathrm{supp}}(\operatorname{\mathbb{T}}_{X})$ and $g \in \mathcal{G}$. One can obtain the strong dual problem of (ref) by applying the results of Blanchet2019a. In \ifdefined(ref)\else(ref) in the supplementary materials\fi, I briefly review their result. The significant advantage of this result in our application is that the dual problem involves only the source conditional distribution and that the dual problem involves a one-dimensional concave programming with respect to the dual argument. By explicitly solving this dual problem, the following theorem is obtained.
As is clear from the first equation of (ref), the worst-case conditional welfare can be written in a simpler form, and its interpretation becomes simple as well. Consider an ITR $g \in \mathcal{G}$ that assigns treatment $a \in \mathcal{A}$ to individuals with the covariate $X=x$. The worst-case conditional welfare $\underline{v}(x;g)$ of the ITR for these individuals is obtained by translating the source CMR function of treatment $a$ by $-\delta$ (when such a translation is unrealistic owing to the lower bound of the outcome space, the worst-case value is set at the lower bound). Because of this simplification, the DR-ITR becomes tidy, as shown in (ref). This result provides an interesting connection with the naive approach in (ref), and also underlies the estimation method discussed in (ref). Note that the above result does not address the existence of a worst-case population that attains the infimum in (ref), which will be discussed in (ref).
It is natural to question whether there exists a worst-case population that attains the infimum in (ref), and if so, what kind of population it is. This section discusses the existence and characterization of such populations in the DR-ITR for specific outcome spaces. Again, the application of the results of Blanchet2019a helps determine the existence and characterization of the worst-case population, as shown in (ref).
(ref) shows the existence of the worst-case conditional distribution of potential outcomes, which attains the infimum in (ref) when the outcome space $\mathcal{Y}$ is a convex subset of $\mathbb{R}$. The result indicates that the worst-case conditional distribution of potential outcomes always exists when the outcome space is convex. Then, one can obtain the worst-case population by incorporating the worst-case conditional distribution with the target covariate distribution $\operatorname{\mathbb{T}}_X$. The corollary also gives examples of the worst-case conditional distribution. When the outcome space is convex and unbounded from below, one worst-case conditional distribution can be obtained by simply moving the source conditional distribution $\operatorname{\mathbb{S}}_{(Y_1,\cdots,Y_d)\mid X=x}$ by $-\delta$ along the $g(x)$-th dimension. Consequently, the worst-case conditional welfare is equal to $m_{\operatorname{\mathbb{S}}}(x,g(x)) - \delta$. When the outcome space is convex but bounded from below, the characterization depends on whether it is logically possible to decrease the source CMR by $\delta$. When $m_{\operatorname{\mathbb{S}}}(x,g(x))-\delta > \inf\operatorname{\mathcal{Y}}$, the worst-case distribution can be obtained by reducing the source CMR by $\delta$ and concentrating more mass around the mean. Conversely, if $m_{\operatorname{\mathbb{S}}}(x,g(x))-\delta\leq\inf\mathcal{Y}$, the worst-case conditional distribution concentrates on $\inf\mathcal{Y}$ for the $g(x)$-th dimension. Note that the distributions given in the corollary are merely examples of worst-case conditional distributions. For example, when $\mathcal{Y}=\mathbb{R}$, the worst-case distribution can also be characterized as an induced measure using a map
given $m_{\operatorname{\mathbb{S}}}(x,g(x))\neq0$. Finally, this result implies that the bound given in (ref) is tight.
This subsection discusses an interesting relationship between the DR-ITR and naive-ITR in (ref), which provides a justification for the naive approach from the perspective of Wasserstein DRO. Let $g_{\operatorname{\mathbb{S}}}^{\text{FB}}:\operatorname{\mathcal{X}} \to \operatorname{\mathcal{A}}$ be the first best ITR for the source population such that $g_{\operatorname{\mathbb{S}}}^{\text{FB}}(x) \in \operatorname*{arg\,max}_{a \in \operatorname{\mathcal{A}}} m_{\operatorname{\mathbb{S}}}(x,a)$ $\operatorname{\mathbb{S}}_X$-a.s.. I additionally impose the following assumption.
A simple example in which (ref) is satisfied is the case where $\operatorname{\mathcal{G}}$ consists of all measurable mappings from $\operatorname{\mathcal{X}}$ to $\operatorname{\mathcal{A}}$. (ref) informally states that the source CMR function $m_{\operatorname{\mathbb{S}}}(x,a)$ is sufficiently distant from the lower bound of the outcome space. One typical case is where the outcome space $\mathcal{Y}$ is unbounded from below; that is, $\inf\mathcal{Y}=-\infty$. In this case, the assumption holds regardless of the value of $\delta$ and the class $\mathcal{G}$ of the ITRs. Under (ref), one has the following theorem:
At first glance, this result seems unnatural. The naive approach imposes a somewhat strong assumption so that the target and source CMR functions become identical, and then searches for the best ITR. Contrarily, the DR-ITR considers populations with CMR functions that can deviate from the source CMR function more flexibly and optimizes against the worst-case over such populations. Hence, there appears to be a difference between the ITRs resulting from these two approaches. However, when the ambiguity set is specified using the 1-Wasserstein distance as in (ref) and (ref) holds, the naive-ITR is already distributionally robust. I explain the relation between (ref) in more depth. For (ref), notice that (ref) implies that the worst-case conditional welfare $\underline{v}(x;g)$ is always maximized by choosing the first best ITR $g_{\operatorname{\mathbb{S}}}^{\text{FB}}$. Thus, $g_{\operatorname{\mathbb{S}}}^{\text{FB}}$ also maximizes the worst-case welfare $\underline{V}(g)$. Therefore, as long as the first best ITR is in $\operatorname{\mathcal{G}}$, it is one of the DR-ITRs. However, when (ref) holds, (ref) implies that
for all $g \in \operatorname{\mathcal{G}}$. That is, the worst-case welfare is obtained by translating the objective function of the naive approach by $-\delta$. Thus, maximizing the distributionally robust welfare is equivalent to maximizing the objective function of the naive approach. Finally, I note that this consequence does not hold when (ref) is not satisfied as exemplified in (ref). In addition, such a relation does not necessarily hold when the ambiguity set is constructed using other metrics as demonstrated in (ref).
Thus far, it has been assumed that the target covariate distribution is known. However, what if it is unknown? Even in such a case, one can construct another ambiguity set on the target covariate distribution, and seek a DR-ITR that is distributionally robust against the unknown covariate shifts as well. In particular, if (ref) is satisfied, the ambiguity set Mo2021,Zhao2019 can be utilized. Here, I incorporate the above idea with the ambiguity set based on $\phi$-divergence. Suppose that a policymaker has no knowledge about the target population, but instead has knowledge on the source population. In addition, I assume that (ref) holds. Then, one can incorporate unknown covariate shifts by using the following ambiguity set:
for $\rho\geq 0$. With this ambiguity set, the DR-ITR is an ITR that maximizes the worst-case welfare over the ambiguity set; that is, $g_{\text{DR-ITR}}^* \in \operatorname*{arg\,max}_{g \in \operatorname{\mathcal{G}}} \inf_{\operatorname{\mathbb{U}} \in \mathscr{U}(\operatorname{\mathbb{S}})} V(g;\operatorname{\mathbb{U}})$. As the constraints on the conditional distribution of potential outcomes and the covariate distribution are independent, one can calculate the distributionally robust welfare in a two-step procedure: (i) obtain the worst-case conditional welfare $\underline{v}(x;g)$ and (ii) calculate the worst-case value of the marginalized worst-case conditional welfare over the ambiguity set. Thus, (ref) implies that the distributionally robust welfare is expressed as
As in (ref), (ref) provides an interesting result. Specifically, it can be shown that under the assumption it holds that
Thus, the DR-ITRs obtained by ignoring the possible difference in conditional distributions of potential outcomes are also distributionally robust. This result specifically reinforces the DR-ITR considered in Mo2021.
In this section, I discuss the identification and estimation of the DR-ITR $g_{\text{DR-ITR}}^{*}$. Then, I evaluate the theoretical performance of its estimator. The estimation method is based on (ref). Suppose that the experimental or observational data generated from the source population $\operatorname{\mathbb{S}}$ and the covariate data generated from the target population $\operatorname{\mathbb{T}}$ are available. Specifically, the data from the source population are $\{(Y_{i},A_{i},X_{i}^{s})\}_{i=1}^{n_{s}}$, where $X_{i}^{s}\in\mathcal{X}$ refers to the observable pre-treatment covariates of individual $i$, $A_{i}\in\mathcal{A}$ denotes the individual's treatment assignment, and $Y_{i}\in\mathcal{Y}$ is the post-treatment observed outcome. The sample is an i.i.d. draw from a joint distribution of $(Y_{1,i},\cdots,Y_{d,i},A_i,X_i)$. I assume that it satisfies unconfoundedness, the overlap condition, and consistency, and that its marginal distribution of $(Y_{1,i},\cdots,Y_{d,i},X_i)$ is the same as the source population $\operatorname{\mathbb{S}}$. However, the covariate data from the target population is $\{X_{j}^{t}\}_{j=1}^{n_{t}}$, where each $X_{j}^{t}$ is an i.i.d. draw from $\operatorname{\mathbb{T}}_{X}$. The following are identifiable from this data; the marginal distributions $\operatorname{\mathbb{S}}_{Y_a\mid X=x}$ of the source conditional distribution of potential outcomes for each $a \in \operatorname{\mathcal{A}}$ and $x \in \operatorname{\mathrm{supp}}(\operatorname{\mathbb{S}}_X)$, the source covariate distribution $\operatorname{\mathbb{S}}_X$, and the target covariate distribution $\operatorname{\mathbb{T}}_X$. Thus, it is obvious from (ref) that the distributionally robust welfare is identifiable, and thus, a natural estimator for $g_{\text{DR-ITR}}^*$ can be constructed as follows:
The object $\widehat{\underline{V}}(g)$ is an estimator of the distributionally robust welfare. The estimator is constructed by replacing the unknown source CMR function $m_{\operatorname{\mathbb{S}}}(x,a)$ with the estimated CMR function $\widehat{m}_{\operatorname{\mathbb{S}}}(x,a)$, and replacing the unknown target covariate distribution $\operatorname{\mathbb{T}}_X$ with the empirical distribution based on $\{X_j^t\}_{j=1}^t$. Arbitrary method can be used as the estimator for the source CMR function.
In line with the literature on policy learning, I evaluate the estimator's performance using a kind of regret. Usually, the regret of the ITR $g$ is defined as the difference in the target population's welfare between the optimal ITR $g_{\operatorname{\mathbb{T}}}^*$ for the target population and the ITR $g$; that is, the usual regret is defined as $V(g_{\operatorname{\mathbb{T}}}^*;\operatorname{\mathbb{T}})-V(g;\operatorname{\mathbb{T}})$. However, the goal of DRO is different from that of standard policy learning; therefore, it is natural to modify the definition of regret. Specifically, I focus on the distributionally robust regret $R_{\text{DRO}}(g)$ of the ITR $g$ defined by
In contrast to the usual regret, distributionally robust regret is defined as the difference in the distributionally robust welfare between the true DR-ITR and the ITR $g$. The same concept is also considered by, for example, Lee2018, Si2021, Tu2019. In the following, I derive a high probability bound on $R_{\mathrm{DRO}}(\widehat{g}_{\text{DR-ITR}})$. I begin by listing some assumptions necessary for the analysis. Thereafter, I denote the expectation with respect to the distribution of data by $\mathbf{E}[\cdot]$.
(ref) is the assumption on the data generating process. (ref) is included to control the complexity of the ITR class. The object $N_H(u,\operatorname{\mathcal{G}})$ in the assumption is the covering number of $\operatorname{\mathcal{G}}$, with $u$ as the radius and the Hamming distance as the distance. For its precise definition, see Definition 4 of Zhou2022. Then, under (ref), I introduce the entropy integral $\kappa(\operatorname{\mathcal{G}})$ by $\kappa(\mathcal{G}):=\int_{0}^{1}\sqrt{\log N_{H}(u^{2},\mathcal{G})}du < \infty$. Combining these assumptions, the next theorem follows.
(ref) provides a high-probability bound on the distributionally robust regret. The first term of the upper bound is a high probability bound on the error generated using the empirical distribution of $\{X_{j}^{t}\}_{j=1}^{n_{t}}$, instead of the true target covariate distribution $\operatorname{\mathbb{T}}_{X}$. Contrastingly, the second term is a high probability bound on the error generated by using the estimator $\widehat{m}_{\operatorname{\mathbb{S}}}(x,a)$, instead of the true source CMR function $m_{\operatorname{\mathbb{S}}}(x,a)$. Overall, the distributionally robust regret with high probability converges to zero at a rate of $O(n_{t}^{-1/2}\vee\psi_{n_{s}}^{-1/2})$.
I demonstrate the proposed DR-ITR using data from the experimental evaluations of changes in welfare-to-work programs during the early 1980s. The program changes were aimed to increase employment and decrease welfare receipt. The new programs required single-parent families who were receiving benefits from Aid to Families with Dependent Children (AFDC), to participate in employment support programs as a condition for receiving the full amount of their monthly welfare payments. For a more detailed explanation of the program changes, see Friedlander1995, Hotz2005, and the references therein. Several states conducted randomized controlled trials (RCTs) to evaluate the effects of the program changes. Among these states, I focus on two experiments conducted in Arkansas and San Diego, because these two states seem to be very different in terms of the covariate distribution and conditional distribution of potential outcomes. The new programs implemented and the corresponding experiments in Arkansas and San Diego differed slightly in terms of timing, central target family, and program components. The Arkansas WORK Program (WORK) targeted AFDC applicants and recipients with children aged at least three years and provided group job search and unpaid work experience. The corresponding experiment started in 1983 with 1,127 participants, of whom half were exposed to WORK and the rest were assigned to the control condition. Contrarily, the Saturation Work Initiative Model (SWIM) in San Diego targeted AFDC applicants and recipients with children aged at least six years and provided job search assistance, skills training, and unpaid work experience. The corresponding experiment started in 1985 with 3,211 participants, of whom half were exposed to the SWIM, and the remainder were assigned to the control condition. In both RCTs, the covariates related to personal characteristics and employment information were collected prior to the experiments. In addition, both experiments involved following up with participants at least two years after the start of the experiments, and collected earnings as outcome variables. For the illustration of the DR-ITR, I consider the following hypothetical situation: Suppose that the government of Arkansas desires a good ITR that specifies whether an individual should participate in the WORK or not. Assume that the data available to the government are limited only to the data from RCT on the SWIM and the data of covariates from RCT on the WORK. In other words, I assume that the experimental assignment variable and the outcome variable of the WORK are not observed. Thus, according to the terminology of this study, the target population comprises single-parent families in Arkansas, whereas the source population comprises single-parent families in San Diego.
The source and target populations actually differ in terms of the covariate distribution and conditional distribution of potential outcomes. First, for the difference in the covariate distribution, \ifdefined(ref) \else(ref) in the supplementary materials \fi lists the covariates used in the following estimation and provides the summary statistics of covariates by RCTs. In addition, it compares the means of each covariate in the two RCTs. The result implies that the two populations differ in terms of means of the covariates. To show that that conditional distributions of potential outcomes are also different, I focus on the difference in the conditional average treatment effect (CATE) functions of the two populations. Specifically, I estimate the CATE function of each location using S-learner with the random forest as the base learner Kuenzel2019,Jacob2021. Then, I evaluate the two estimated CATE functions at the observed values of covariates in Arkansas's RCT. Thus, the difference of the covariate distribution is essentially controlled, and any difference in the values of CATE derives from the difference in the conditional distribution of potential outcomes. (ref) visualizes the difference in the values of the CATE using a scatter plot. In the figure, each point represents an observation of covariates in Arkansas's RCT, and its $x$-coordinate and $y$-coordinate are an evaluation of the estimated CATE function for Arkansas and San Diego at that value of covariates. If a point is on the dashed line, the CATEs of the two locations are the same. In other words, the more distant the point is from this line, the more different the two CATEs are. From this figure, one can observe that the CATE functions of the two locations are substantively different. The maximum distance between the two CATEs are about \$4,618. In summary, the source and target populations are significantly different, and hence, the existing methods that assume two populations to be same are inappropriate. I apply the proposed approach to the current setting and estimate the DR-ITR. The outcome is the total earnings in two years after the start of the WORK or SWIM, and thus the outcome space is non-negative real numbers. The source CMR function is estimated by the random forest with tuned hyper parameters using R package randomForest. The ambiguity level $\delta$ in (ref) varies in $0,1,\cdots,10$. Note that the case of $\delta = 0$ corresponds to the naive approach explained in (ref), which is included for comparison. The ITR class consists of decision trees of depth 2 constructed by the covariates except sex and race information. These two variables are excluded because this information should not be used owing to ethical concerns. The optimization of the tree is conducted using R package policytree Zhou2022. After obtaining the ITRs, I evaluate their performance by estimating the total earnings that would be attained in two years if each ITR was implemented in Arkansas. Note that this corresponds to the target population's welfare. As the data from RCT conducted in Arkansas are actually available, it is possible to consistently estimate the welfare.
(ref) compares the performance of the DR-ITR with that of the naive-ITR. In the figure, the solid line with a circle represents the welfare of the DR-ITR, while the dashed line with a square represents that of the naive-ITR. Note that the naive-ITR does not depend on the value of $\delta$. It is clear that the DR-ITR outperforms the naive-ITR in terms of the point estimates of the target population's welfare by appropriately choosing the value of $\delta$. Specifically, the welfare due to the DR-ITR with positive ambiguity level $\delta$ is not less than that of the naive-ITR. Moreover, when $\delta = 9$, the t-test on the null hypothesis that the values of the target population's welfare due to the DR-ITR and naive-ITR are the same is rejected at the 5% confidence level.
This study examined how one should determine a reasonable ITR when the population in which the estimated ITR is implemented differs from the population from which the available experimental or observational data are generated. Specifically, this study focuses on the unknown and unidentifiable difference in the conditional distribution of potential outcomes between the source and target populations, and proposes obtaining a DR-ITR that maximizes the worst-case value of welfare over a certain set of populations. In combination with the Wasserstein distance-based ambiguity set, the DR-ITR offers a simple intuition and estimation method. Additionally, it provides a justification for the naive approach, which assumes that the conditional distributions of potential outcomes are the same between the source and target populations. Thus, this study can be regarded as complementing the existing studies on this topic. In addition, an estimator for the distributionally robust policy is developed, and the theoretical analysis shows that the estimator has regret converging to zero. Several interesting questions are left unanswered. First, a policy maker often has access to multiple sets of experimental data, each of which is generated from a different population. It is obscure how the proposed DR-ITR can be extended to such multiple-source population scenarios. Second, in the population level, the current problem discussed in (ref) can be understood as a two-person zero-sum game between a policymaker who chooses an ITR to maximize the target population's welfare and an adversary who chooses the target population to minimize the target population's welfare. However, this study does not consider whether a Nash equilibrium exists or what types of strategies constitute the equilibrium. I leave these questions for future work.
\ifdefined \subfile{supp} \fi