EconBase
← Back to paper

Locally Robust Policy Learning: Inequality, Inequality of Opportunity and Intergenerational Mobility

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

56,652 characters · 9 sections · 77 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Locally Robust Policy Learning: Inequality, Inequality of Opportunity and Intergenerational Mobility

abstractPolicy makers need to decide whether to treat or not to treat heterogeneous individuals. The optimal treatment choice depends on the welfare function that the policy maker has in mind and it is referred to as the policy learning problem. I study a general setting for policy learning with semiparametric Social Welfare Functions (SWFs) that can be estimated by locally robust/orthogonal moments based on U-statistics. This rich class of SWFs substantially expands the setting in athey2021policy and accommodates a wider range of distributional preferences. Three main applications of the general theory motivate the paper: (i) Inequality aware SWFs, (ii) Inequality of Opportunity aware SWFs and (iii) Intergenerational Mobility SWFs. I use the Panel Study of Income Dynamics (PSID) to assess the effect of attending preschool on adult earnings and estimate optimal policy rules based on parental years of education and parental income. \\ \\ JEL\ Classification: C13; C14; C21; D31; D63; I24 \\ Keywords: local robustness, U-statistics, Inequality, Intergenerational mobility, empirical welfare maximization. \\ R package (forthcoming): \url{https://joelters.github.io/home/code/}

Introduction

Whenever a treatment has heterogeneous effects it is important to decide carefully who should be treated. In the simplest case where we care about the average outcome, no budgetary limits exist and treatment effects are positive for everyone, the best policy is to treat everyone. However, we might have a limited budget, distributional concerns or negative treatment effects. Then, it is important to decide whether to treat or not to treat different individuals. This is the problem of policy learning.

In economics we might want to know whether to provide training to the unemployed or design rules to assign conditional cash transfers. In business, we might want to know whether to provide a discount to a customer. Judges have to decide whether to release someone on parole. Schools might want to know whether to provide extra-curricular lessons to some students. Certain medicines might be beneficial for some but detrimental for others.

The inherent distributional conscerns in these examples are quite different. Hence, we need a framework accommodating different SWFs. The framework has to be general, but also needs to allow for certain statistical guarantees. I provide a framework to estimate optimal rules for a rich class of semiparametric SWFs, estimable by U-statistics. Examples include the average outcome, Inequality aware SWFs, Inequality of Opportunity (IOp) aware SWFs and Intergenerational Mobility (IGM) SWFs.

To my knowledge, there is no prior work on IOp and IGM SWFs in the policy learning literature. IOp is the part of inequality explained by circumstances $X$ outside the control of the individual, e.g. sex, race or parental income. IOp SWFs do not penalize all inequality, only inequality explained by circumstances. Based on the seminal contributions in gaer1993, fleurbaey1995equal and roemer1998equality, the IOp literature has focused on measuring IOp. A popular measure is the Gini of the best predictions (in mean squared error sense) of the outcome $Y$ given the circumstances $X$, i.e. $G(\gamma(X))$ where $\gamma(X) = \mathbb{E}[Y|X]$ and $G(Z)$ denotes the Gini index of the random variable $Z$. To accommodate a possibly high-dimensional set of circumstances, IOp literature has started using machine learners to predict (e.g brunori2019inequality, brunori2019upward, brunori2021roots, brunori2021evolution, rodriguez2021inequality, carranza2022 or hufe2022fairness). The bias-variance trade-off in the prediction allows for bias which can creep into the IOp estimator. escanciano2023machine provide locally robust IOp estimators robust to such biases. I construct Neyman-orthogonal IOp aware SWFs.

Inequality SWFs have been studied in kasy2016partial, kitagawa2021equality or kock2023treatment. A popular SWF for an outcome $Y$ is $W = \mathbb{E}[Y](1-G(Y))$. This SWF values the average outcome but penalizes high inequality. leqi2021median propose to maximize average conditional quantiles. cui2024policy focus on a conditional quantile of the treatment effect using partial identification. wang2018quantile study quantile-optimal policies and adapt their theory to minimize Gini's mean difference. They use the U-statistics nature of the Gini mean difference and obtain asymptotic theory for a particular class of policy rules by using empirical U-process methods. I avoid U-processes theory by using a representation of U-statistics as sums-of-i.i.d. blocks introduced in hoeffding1963. This representation is key in proving the main result of the paper.

While inequality aware SWFs look at the distribution of $Y$, IOp aware SWFs focus on the distribution of $\gamma(X)$. An IOp aware SWF is $W = \mathbb{E}[Y](1-G(\gamma(X)))$, which penalizes IOp. This SWF adds an extra nuisance parameter, $\gamma(X)$, on top of the conditional expectations/propensity scores needed to identify treatment effects. Policy learning with semiparametric SWFs, which directly depend on additional unknown functions, has been little explored, with the exception of leqi2021median whose welfare depends on conditional quantiles.

IGM studies the relationship between child and parental outcomes. The Kendall-$\tau$ is a popular measure of mobility in the literature (see chetty2014land or kitagawa2018measurement). It looks at whether the parents of individual $i$ are richer than those of $j$ and $i$ is richer than $j$. An IGM aware SWF is $W = -|\tau - t|$ for some target $t \in [-1,1]$. For instance, we could decide the allocation of higher education scholarships to reduce dependence between parental and child's income.

The policy learning literature looks for optimal allocation rules $\pi$ mapping characteristics to binary treatment decisions. Optimal rules are searched within a class $\Pi$ of treatment rules to maximize welfare. Following manski2004statistical, I search for optimal policies in $\Pi$ so as to minimize regret, i.e. the expected difference between the best possible welfare and the welfare evaluated at the estimated policy. Other relevant work includes dehejia2005program, hirano2009asymptotics, stoye2009minimax,stoye2012minimax, chamberlain2011bayesian, bhattacharya2012inferring, tetenov2012statistical, kasy2016partial, kitagawa2018should,kitagawa2021equality, athey2021policy, or zhou2023offline.

Estimation of unknown functions in semiparametric SWFs challenges the statistical guarantees of estimated policy rules. This is due to slow convergence of non-parametric estimators, addressed in semiparametric methods through locally robust/orthogonal moments. These are moment conditions that identify the quantity of interest and allow for its estimation at $\sqrt{n}$ ($n$ is the sample size) rate. I expand previous work by considering any semiparametric SWF, possibly defined as a U-statistic, which can be estimated by locally robust/orthogonal scores. The main theoretical result provides an asymptotic upper bound to the regret of the estimated policy rule.

This paper is related to athey2021policy, leqi2021median and zhou2023offline in making use of the semiparametric literature on locally robust/orthogonal scores (e.g. chernozhukov2022locally) to obtain $\sqrt{n}$ rates of convergence even with nonparametric first steps. I build upon escanciano2023machine to expand policy learning results to SWFs defined by U-statistics. athey2021policy find rates of the regret that optimally depend on the complexity $\Pi$ and the sample size in observational settings where the propensity score is unknown. They do so for average-treatment-like SWFs. I generalize this setting by allowing general semiparametric SWFs, possibly defined as U-statistics.

Empirically, treatment allocation with inequality, IOp and IGM SWFs is hard. We need circumstances and parental income which are usually absent and to identify treatment effects. I look at the effect of preschool on adult earnings using the Panel Study of Income Dynamics (PSID) dataset. This empirical illustration has many advantages. Any variable that induces preschool attendance can be considered a circumstance under the (very reasonable) assumption that we cannot hold the kid responsible for these variables. Also, PSID has rich information on family background and it allows us to look at long-term outcomes. It also has limitations. Treatment is not randomly assigned, so I rely on the assumption of selection on observables. Preschool attendance is not a binary treatment since its quality varies. Furthermore, I have no information on the cost of treatment and allocating children to preschool based on their circumstances might not be enforceable or ethical. Observed preschool choices differ from estimated optimal rules, even when maximizing average outcomes, suggesting parents prioritize factors beyond future earnings.

The effect of preschool is heterogeneous. On average preschool has a positive effect on adult earnings but children with highly educated mothers and high parental income are negatively affected. This aligns with findings in psychology and economics (see fort2020cognitive). These heterogeneous effects have different implications for different SWFs. I estimate optimal treatment rules based on parental income and mother's education. Inequality aware SWFs treat individuals with negative treatment effects since the decrease in inequality compensated the decrease in average earnings. The same happens with the IGM welfare which has no average motive at all. The additive and IOp estimated optimal policy rules coincide. This coincidence is specific to the heterogeneous treatment effects in the data and not a general result. In this empirical illustration, maximizing the average already decreases IOp drastically.

I introduce the welfare objects in Section (ref). Section (ref) elaborates on the general theory for semiparametric SWFs which are linear on the distribution of the data (i.e. not U-statistics) and Section (ref) generalized to SWFs possibly defined as U-statistics. Section (ref) provides upper bounds on the regret and Section (ref) deals with the empirical illustration. All proofs are in the Appendix.

Welfare economics for inequality, IOp and rank correlations

The policy learning literature is at the intersection of welfare economics and econometrics. Before addressing the econometric problem, I introduce the key welfare objects of interest. For a continuous random outcome $Y_i \in \mathbb{R}^+$, the additive welfare is based on the average outcome: $W = \mathbb{E}[Y_i]$. Additive welfare does not care about distributional aspects other than the average. A first approach to include distributional concerns is to follow dalton1920measurement and atkinson1970measurement and consider increasing and concave transformations $u(\cdot)$ \footnote{With abuse of notation I call $W$ to all SWFs as they appear.}: $W = \mathbb{E}[u(Y_i)]$.

This SWF will rank two outcome distributions equally for all increasing and concave $u(\cdot)$ if the Lorenz curve of one of the distributions is everywhere above the Lorenz curve of the other distribution and has equal or higher mean; equivalently if one distribution second-order stochastically dominates the other. If we want to obtain a complete ordering we need to specify $u(\cdot)$ further. One popular choice is \[ u(y) =

cases\frac{y^{1-\theta}}{1-\theta} & if \theta \in (0,1) \\ \log(y) & if \theta = 1,

\] where $\theta$ captures the concavity of $u(\cdot)$ and can be interpreted as an inequality aversion parameter. I also focus on SWFs aware of Inequality of Opportunity (IOp). IOp is the part of total inequality which can be explained by circumstances, i.e. by variables that are outside the control of the individual such as parental education or parental income. Let $X_i \in \mathbb{R}^k$ be such a random vector of circumstances. Let also $\gamma(X_i) = \mathbb{E}[Y_i|X_i]$. By looking at the distribution of $\gamma(X_i)$ we get IOp averse SWFs: $W = \mathbb{E}[u(\gamma(X_i))]$.

If there is no IOp, circumstances are unable to predict the outcome and we have $\gamma(X_i) = \mathbb{E}[Y_i]$. In this case, $W = u(\mathbb{E}[Y_i])$ so we only care about the (transformed) average income. If we have maximum IOp, the outcome is a deterministic function of the circumstances and $\gamma(X_i) = Y_i$. Then, $W = \mathbb{E}[u(Y_i)]$. Since all inequality is IOp, we are back to the inequality averse SWF.

Alternatively, let $F_Y$ be the distribution of $Y$ and $F^{-1}_Y$ be the quantile function, for weights $w(\cdot)$, we might have \[ W = \int_0^1 F_Y^{-1}(\tau) w(\tau) d \tau. \] This welfare has been used in mehran1976linear, donaldson1980single, weymark1981generalized, donaldson1983ethically or aaberge2021ranking. Letting $w_k(\tau) = (k-1)(1-\tau)^{k-2}$ we get the extended Gini family of SWFs. I focus on $k = 3$, the standard Gini SWF which can be shown to be

align*[align* omitted — 92 chars of source]

where the second equality follows from writing the Gini of $Y_i$ as $G(Y_i) = \mathbb{E}[|Y_i-Y_j|]/\mathbb{E}[Y_i + Y_j]$ where $Y_j$ is an independent copy of $Y_i$. The welfare above is additive if there is no inequality and penalizes positive values of the Gini coefficient. If we only care about IOp, we can look at the distribution of $\gamma(X_i)$. In that case, we have

align*[align* omitted — 139 chars of source]

If $G(\gamma(X_i)) = 0$, we are back to the additive case. If there is full IOp, then $G(\gamma(X_i)) = G(Y_i)$ and we are back to the standard Gini SWF. I also consider the problem of intergenerational mobility. Let $X_{1i} \in \mathbb{R}$ be the parental outcome. A measure of association between $Y_i$ and $X_{1i}$ is the Kendall-$\tau$ \[ \tau = \mathbb{E}[sgn(Y_i-Y_j)sgn(X_{1i} - X_{1j})], \] where $sgn(a) = \mathds{1}(a > 0) - \mathds{1}(a< 0)$. This parameter is popular in the IGM literature (see chetty2014land or kitagawa2018measurement) where $X_{1i}$ is parental income and $Y_i$ is the child's income. It takes values between $1$ and $-1$. $\tau = 1$ means that whenever an individual has a higher income than another, she also has a higher parental income. $\tau = -1$ is the opposite. For some target Kendall-$\tau$ $t \in [-1,1]$ an IGM aware SWF is \[ W = -\biggl|\mathbb{E}\biggl[ sgn(Y_i-Y_j)sgn(X_{1i} - X_{1j}) \biggr] - t \biggr|. \] To my knowledge, this is a novel SWF. Note that it allows us to treat problems much more general than IGM. Setting $t = 0$, maximizing this SWF corresponds to allocating a treatment to minimize the dependence between two variables $Y_i$ and $X_{1i}$.

Policy learning with general orthogonal scores

Let $(Y_i(1),Y_i(0),D_i, X_i) \sim F_0$ where $(Y_i(1),Y_i(0)) \in \mathcal{Y} \times \mathcal{Y}$ are real-valued potential outcomes, i.e. $i$'s outcome under treatment and in the absence of treatment respectively. $D_i$ is a binary treatment and $X_i \in \mathcal{X}$ is a vector of pre-treatment covariates. Let $\gamma^{(j)}(X_i) = \mathbb{E}[Y_i(j)|X_i] \in \Gamma$ for $j = 0,1$ be potential predictions. We observe an i.i.d. sample $(Z_1,...,Z_n)$ with $Z_i = (Y_i,D_i,X_i) \in \mathcal{Z}$ and $Y_i = Y_i(1)D_i + Y_i(0)(1-D_i) \in \mathcal{Y}$. Let $\pi: \mathcal{X} \mapsto \{0,1\}$ be a binary treatment rule and $\Pi$ be a collection of such treatment rules. We are interested in choosing $\pi \in \Pi$ to maximize

align[align omitted — 130 chars of source]

For additive welfare, $g(Y_i(j),X_i,\gamma^{(j)}) = Y_i(j)$ for $j = 0,1$. Importantly, $g$ can depend on possibly infinite-dimensional nuisance parameters $\gamma$. $\gamma$ is a conditional expectation throughout the paper but this framework can be extended to allow for much more general first steps such as conditional quantiles (see ichimura2022influence).

example[IOp Atkinson] For an inequality averse SWF we can use $W(\pi) = \mathbb{E}[u(Y_i(1))\pi(X_i) + u(Y_i(0))(1-\pi(X_i))]$ with $u(\cdot)$ a concave function and $X_i$ a vector of circumstances. For an IOp averse SWF we look at the distribution of $\gamma(X_i)$: \[ W(\pi) = \mathbb{E}[u(\gamma^{(1)}(X_i))\pi(X_i) + u(\gamma^{(0)}(X_i))(1-\pi(X_i))]. \] This welfare has not been covered in the policy learning literature before. $\blacksquare$

((ref)) is not based on observables. To identify it, we need a sample from an experimental or observational experiment where the policy has already been implemented. Let $e(X_i) = \mathbb{P}(D_i = 1 | X_i)$ be the propensity score. I assume that the following holds.

assumptioni) $(Y_i(1),Y_i(0)) \perp D_i | X_i$, ii) $e(x) \in [\kappa, 1 - \kappa]$ for some $\kappa \in (0,1/2]$.

There are two ways of identifying welfare, the direct method (DM) based on conditional expectations or Inverse Propensity Score Weighting (IPW). I use the DM approach. All results in the paper for the IPW approach are in Appendix (ref). Let $\gamma(D_i,X_i) = \mathbb{E}[Y_i|D_i,X_i]$, $\gamma_j(X_i) = \gamma(j,X_i)$ for $j = 0, 1$ and $\varphi(D_i,X_i, \gamma) = \mathbb{E}[g(Y_i,X_i,\gamma)|D_i,X_i]$.

propositionUnder Assumption (ref), $W(\pi)$ is identified as \begin{align*} W(\pi) &= \mathbb{E}[\varphi(1,X_i,\gamma_1)\pi(X_i) + \varphi(0,X_i,\gamma_0)(1-\pi(X_i))]. \end{align*}

If $g$ does not depend on potential outcomes directly, i.e. $g(u, X_i,\gamma^{(j)}) = g(t, X_i,\gamma^{(j)}) \equiv g(X_i,\gamma^{(j)})$ for all $u,t \in \mathcal{Y}$ then $\varphi(D_i, X_i,\gamma) = g(X_i,\gamma)$. This happens in all IOp examples. Hence, we can have either $\gamma$ or $(\gamma,\varphi)$ as nuisance parameters. To enjoy local robustness to first steps, I first need the following assumption

assumptionThere exist $(\alpha_1,\alpha_0)$ such that for any $\tilde{\gamma}$ with $\mathbb{E}[\tilde{\gamma}(X_i)^2] < \infty$ and $j = 0,1$ \[ \frac{d}{d\tau} \mathbb{E}[\varphi(j,X_i,\bar{\gamma}_\tau)] \biggr|_{\tau = 0} = \frac{d}{d\tau} \mathbb{E}[\alpha_j(D_i,X_i)\bar{\gamma}_\tau(D_i,X_i)] \biggr|_{\tau = 0}, \] where $\bar{\gamma}_\tau = \gamma + \tau \tilde{\gamma}$ and $\mathbb{E}[\alpha_j(D_i,X_i)^2] < \infty$.

This is a common assumption in the semiparametric literature (e.g. (4.1) in newey1994asymptotic) and allows for $\varphi$ to depend non-linearly on $\gamma$, generalizing Assumption 1 in athey2021policy.

propositionThe orthogonal score is $\Gamma_i(\pi) = \Gamma_{1i}\pi(X_i) + \Gamma_{0i}(1-\pi(X_i))$, where \begin{align*} \Gamma_{1i} &= \varphi(1,X_i,\gamma) + \frac{D_i}{e(X_i)}(g(Y_i,X_i,\gamma_1) - \varphi(1,X_i,\gamma)) + \alpha_1(D_i,X_i)(Y_i-\gamma(D_i,X_i)), \\ \Gamma_{0i} &= \varphi(0,X_i,\gamma) + \frac{1-D_i}{1-e(X_i)}(g(Y_i,X_i,\gamma_0) - \varphi(0,X_i,\gamma)) + \alpha_0(D_i,X_i)(Y_i-\gamma(D_i,X_i)). \end{align*}

Orthogonal scores are formed by identifying scores and correction terms for nuisance parameters $\varphi$ and $\gamma$. Whenever $g$ does not depend on the potential outcomes directly we have that $g(Y_i, X_i,\gamma_j) - \varphi(j, X_i,\gamma) = 0$ for $j = 0,1$. To estimate the welfare for a given $\pi \in \Pi$ we employ cross-fitting as in chernozhukov2022locally. Let the data be split in $L$ groups $I_1,...,I_l$, then \[ \hat{W}_n(\pi) = \frac{1}{n}\sum_{l = 1}^L \sum_{i \in I_l} \hat{\Gamma}_{1i,l}\pi(X_i) + \hat{\Gamma}_{0i,l}(1-\pi(X_i)), \] where

align*[align* omitted — 418 chars of source]

and $(\hat{\varphi}_l,\hat{e}_l,\hat{\gamma}_l,\hat{\alpha}_{j,l})$, $j = 0,1$, are estimators of the nuisance functions which do not use observations in $I_l$.

examplecont{1}[IOp Atkinson (cont.)] For $\theta \in (0,1]$, let \[ U(\gamma(x)) = \begin{cases} \frac{\gamma(x)^{1-\theta}}{1-\theta} & \text{if } \theta \in (0,1) \\ \log(\gamma(x)) & \text{if } \theta = 1. \end{cases} \] In this case, $g = U$. The orthogonal score for $\theta \in (0,1]$ is \begin{align*} \Gamma_i(\pi) &= U(\gamma(1,X_i)) + \frac{\gamma(D_i,X_i)^{-\theta}D_i}{e(X_i)}(Y_i - \gamma(D_i,X_i))\pi(X_i) \\ &+ U(\gamma(0,X_i)) + \frac{\gamma(D_i,X_i)^{-\theta}(1-D_i)}{1-e(X_i)}(Y_i - \gamma(D_i,X_i))(1-\pi(X_i)), \end{align*} i.e. $\alpha_1(D_i,X_i) = e(X_i)^{-1}\gamma(D_i,X_i)^{-\theta}D_i$ and $\alpha_0(D_i,X_i) = (1-e(X_i))^{-1}\gamma(D_i,X_i)^{-\theta}(1-D_i)$. $\blacksquare$

The estimator of the optimal treatment rule among a class of rules $\Pi$ is $\hat{\pi} = \arg \max_{\pi \in \Pi} \hat{W}_n(\pi)$.

Policy learning with U-statistics

Let now $\pi_{ab}(X_i,X_j) = \mathds{1}(\pi(X_i) = a)\times\mathds{1}(\pi(X_j) = b)$ with $a,b \in \{0,1\}$. Now

align[align omitted — 167 chars of source]

Now, $W(\pi)$ depends on pairwise comparisons. Also, we are summing across $\{0,1\}^2$. This is because we have to take into account when both members of the pair are under treatment, or just one of them or none of them. Finally, we have $\pi_{ab}(X_i,X_j)$ instead of $\pi(X_i)$ since we need to account for when both members of the pair are allocated to treatment, just one of them or none of them.

example[Inequality] We can accommodate the standard Gini SWF with \[ g(Y_i(a),Y_j(b)) = (1/2)(Y_i(a) + Y_j(b) - |Y_i(a) - Y_j(b)|). \] $\blacksquare$
example[Inequality of Opportunity IOp] $\mathbb{E}[\gamma(X_i)](1-G(\gamma(X_i)))$ fits our setting by letting \[ g(X_i,X_j,\gamma^{(a)},\gamma^{(b)}) = (1/2)(\gamma^{(a)}(X_i) + \gamma^{(b)}(X_j) - |\gamma^{(a)}(X_i) - \gamma^{(b)}(X_j)|). \] $\blacksquare$
example[Kendal-$\tau$] To allocate a treatment targeting a specific Kendall-$\tau$, say $t \in \mathbb{R}$, we have to extend our setting to transformations of the right-hand side of (ref) \begin{align*} g(Y_i(a),X_{1i},Y_j(b),X_{1j}) &= sgn(Y_i(a)-Y_j(b))sgn(X_{1i} - X_{1j}), \\ W(\pi) &= -\biggl|\mathbb{E}\biggl[ \sum_{(a,b) \in \{0,1\}^2}g(Y_i(a),X_{1i},Y_j(b),X_{1j}) \pi_{ab}(X_i,X_j) \biggr] - t \biggr|. \end{align*} $\blacksquare$

For $a,b \in \{0,1\}$ let now $\varphi(a,X_i,b,X_j,\gamma_a,\gamma_b) = \mathbb{E}[g(Y_i,X_i,Y_j,X_j,\gamma_a,\gamma_b)|D_i = a,X_i,D_j = b,X_j]$ and $e_{ab}(X_i,X_j) = e_a(X_i)e_b(X_j)$ where for $c \in \{0,1\}$, $e_c(X_i) = \mathbb{P}(D_i = c |X_i)$.

propositionUnder Assumption (ref), $W(\pi)$ in ((ref)) is identified as \begin{align*} W(\pi) &= \mathbb{E}\biggl[\sum_{(a,b)\in \{0,1\}^2} \varphi(a,X_i,b,X_j,\gamma_a,\gamma_b)\pi_{ab}(X_i,X_j).\biggr], \end{align*}

Now I apply Proposition (ref) to identify the welfare in each of our three main examples.

examplecont{2}[Inequality (cont.)] In this example, welfare is identified by \[ W(\pi) = \mathbb{E}\biggl[\frac{1}{2} \sum_{(a,b) \in \{0,1\}^2} \mathbb{E}(Y_i + Y_j - |Y_i - Y_j| \mid D_i = a,X_i,D_j = b,X_j)\pi_{ab}(X_i,X_j)\biggr]. \] $\blacksquare$
examplecont{3}[IOp (cont.)] In this example, welfare is identified by \[ W(\pi) = \frac{1}{2}\mathbb{E}\biggl[\sum_{(a,b)\in \{0,1\}^2}\biggl(\gamma_a(X_i) + \gamma_b(X_j) - |\gamma_a(X_i) - \gamma_b(X_j)|\biggr)\pi_{ab}(X_i,X_j)\biggr]. \] $\blacksquare$
examplecont{4}[IGM (cont.)] In this example, welfare is identified by \[ W(\pi) = - \biggl| \mathbb{E}\biggl[\frac{1}{2} \sum_{(a,b) \in \{0,1\}^2} \mathbb{E}(sgn(X_{1i} - X_{1j}) sgn(Y_i - Y_j) \mid D_i = a,X_i,D_j = b,X_j)\pi_{ab}(X_i,X_j)\biggr] - t \biggr|. \] $\blacksquare$

To compute orthogonal scores we need to assume a linearization property as in Assumption (ref).

assumptionThere exist $\alpha_{ab,p}$, $P<\infty$, and $(c_{1p}, c_{2p})$ for $p = 1,...,P$, such that for all $(a,b) \in \{0,1\}^2$ the following linearization holds \[ \frac{d}{d\tau} \mathbb{E}[\varphi(a,X_i,b,X_j,\bar{\gamma}_\tau)] = \frac{d}{d\tau} \mathbb{E}\biggl[\sum_{p=1}^P\alpha_{ab,p}^\gamma(D_i,X_i,D_j,X_j)(c_{1p}\bar{\gamma}_\tau(D_i,X_i) + c_{2p}\bar{\gamma}_\tau(D_j,X_j))\biggr], \] where $\bar{\gamma}_\tau$ is defined as in Assumption (ref) and $\mathbb{E}[\alpha_j(D_i,X_i)^2] < \infty$.
propositionThe orthogonal score is $\Gamma_{ij}(\pi) = \sum_{(a,b) \in \{0,1\}^2} \Gamma_{ij}^{ab}\pi_{ab}(X_i,X_j),$ with \begin{align*} \Gamma_{ij}^{ab} = \varphi(a,X_i,b,X_j,\gamma_a,\gamma_b) &+ \phi_{ab}^\varphi(D_i,X_i,D_j,X_j,\varphi,\alpha^{\varphi}) + \phi_{ab}^\gamma(D_i,X_i,D_j,X_j,\gamma,\alpha^{\gamma}), \\ \phi_{ab}^\gamma(D_i,X_i,D_j,X_j,e,\alpha^{\gamma}) &= \sum_{p=1}^P \alpha_{ab,p}^\gamma(D_i,X_i,D_j,X_j)(c_{1p} Y_i + c_{2p} Y_j - c_{1p}\gamma(D_i,X_i) - c_{2p} \gamma(D_j,X_j)), \\ \phi_{ab}^\varphi(D_i,X_i,D_j,X_j,\varphi,\alpha^{m}) &= \alpha_{ab}^\varphi(D_i,X_i,D_j,X_j)(g(Y_i,X_i,Y_j,X_j,\gamma_a,\gamma_b) - \varphi(D_i,X_i,D_j,X_j,\gamma_a,\gamma_b)), \end{align*} and $\alpha_{ab}^\varphi(D_i,X_i,D_j,X_j) = D_{ij}^{ab}/e_{ab}(X_i,X_j)$ and $D_{ij}^{ab} = \mathds{1}(D_i = a)\mathds{1}(D_j = b)$.
examplecont{2}[Inequality (cont.)] In this example, we have that \begin{align*} \Gamma_{ij}^{ab} &= \frac{1}{2}\mathbb{E}(Y_i + Y_j - |Y_i - Y_j| \mid D_i = a,X_i,D_j = b,X_j) \\ &+ \frac{D_{ij}^{ab}}{2e_{ab}(X_i,X_j)}(Y_i + Y_j - |Y_i - Y_j| -\mathbb{E}(Y_i + Y_j - |Y_i - Y_j| \mid D_i = a,X_i,D_j = b,X_j)). \end{align*} $\blacksquare$
examplecont{3}[IOp (cont.)] I give the orthogonal score for IOp as a Proposition. \begin{proposition} Assume for all $(a,b) \in \{0,1\}^2$ that either (i) $\mathbb{P}(\gamma_a(X_i) - \gamma_b(X_j) = 0) = 0$ or that (ii) $x_i \neq x_j \implies \gamma_a(X_i) - \gamma_b(X_j) \neq 0$ and let $\delta_{ij}^{ab} = sgn(\gamma_a(X_i) - \gamma_b(X_j))$, then \begin{align*} \Gamma_{ij}^{ab} &= \frac{1}{2}\biggl(\gamma_a(X_i) + \gamma_b(X_j) - |\gamma_a(X_i) - \gamma_b(X_j)| \\ &+ \frac{\mathds{1}(D_i = a)}{e_a(X_i)}(1-\delta_{ij}^{ab})(Y_i - \gamma(D_i,X_i)) + \frac{\mathds{1}(D_j = b)}{e_b(X_j)}(1+\delta_{ij}^{ab})(Y_j - \gamma(D_j,X_j))\biggr). \end{align*} \end{proposition} These assumptions deal with the point of non-differentiability of the absolute value. For a thorough discussion see escanciano2023machine. $\blacksquare$

I use a cross-fitting algorithm used in escanciano2023machine. I split the pairs $\{(i,j) \in \{1,...,n\}^2: i < j\}$ in $L$ groups $I_1,...,I_l$, then

align[align omitted — 142 chars of source]

where $\hat{\Gamma}_{ij,l}$ is the same as $\Gamma_{ij}$ but with all nuisance parameters replaced by estimators which do not use observations in $I_l$. The estimator of the optimal treatment rule is \[ \hat{\pi} = \arg \max_{\pi \in \Pi} \hat{W}_n(\pi). \] For the IGM example, the estimation is slightly different.

examplecont{4}[IGM (cont.)] The orthogonal score is given by \begin{align*} \Gamma_{ij}^{ab} &= \mathbb{E}(sgn(X_{1i} - X_{1j}) sgn(Y_i - Y_j) \mid D_i = a,X_i,D_j = b,X_j) \\ &+ \frac{D_{ij}^{ab}}{e_{ab}(X_i,X_j)}(sgn(X_{1i} - X_{1j}) sgn(Y_i - Y_j) -\mathbb{E}(sgn(X_{1i} - X_{1j}) sgn(Y_i - Y_j) \mid D_i = a,X_i,D_j = b,X_j)). \end{align*} The estimator of the welfare for a given $\pi \in \Pi$ and target $t$ is \begin{equation} \hat{W}_n(\pi) = -\biggl|\binom{n}{2}^{-1}\sum_{l = 1}^L \sum_{(i,j) \in I_l} \sum_{(a,b) \in \{0,1\}^2} \hat{\Gamma}_{ij,l}^{ab}\pi_{ab}(X_i,X_j) - t\biggr|. \end{equation} $\blacksquare$

Asymptotic statistical guarantees

Now it is useful to make clear the dependence of the scores $\Gamma_{ij}^{ab}$ on the data and the nuisance parameters. Hence, I let now $\Gamma_{ij}^{ab} = \psi_{ab}(Z_i,Z_j,\gamma,\varphi,\alpha)$, where \[ \psi_{ab}(Z_i,Z_j,\gamma,\varphi,\alpha) = \varphi(a,X_i,b,X_j,\gamma_a,\gamma_b) + \phi_{ab}^{\gamma}(Z_i,Z_j,\gamma,\alpha^{\gamma}) + \phi_{ab}^{\varphi}(Z_i,Z_j,\varphi,\alpha^{\varphi}). \] $\psi_{ab}$ is the sum of an identifying function ($\varphi_{ab}$) and correction terms ($(\phi_{ab}^{\gamma},\phi_{ab}^\varphi$) needed to achieve orthogonality. For a given treatment rule $\pi$, orthogonal scores are given by \[ \Gamma_{ij}(\pi) = \sum_{(a,b) \in \{0,1\}^2}\psi_{ab}(Z_i,Z_j,\gamma,\varphi,\alpha)\pi_{ab}(X_i,X_j). \] This framework accommodates SWFs in Section (ref) if $\psi_{ab}$ does not depend on $Z_j$ and only depends on $a$ so that it can be written as $\psi_{a}(Z_i,\gamma,\varphi,\alpha)$ for $a \in \{0,1\}$. Hence, I only state conditions and results for SWFs in this Section. The IGM example does not fit in the general setting, however, results extend to this example by Corollary in Section C of the Online Appendix. Next, I give conditions on the convergence of the nuisance estimators and the complexity of $\Pi$ that allow me to prove asymptotic statistical guarantees for the estimator of treatment rules.

Conditions on the nuisance parameter estimators

I give high-level conditions for the estimators of nuisance parameters.

assumption$\mathbb{E}[|\psi(Z_{i},Z_{j},\gamma,\varphi,\alpha)|^{2}]<\infty$, $\omega \in \{\gamma,\varphi\}$ and for $(a,b) \in \{0,1\}^2$ \begin{enumerate} [(i)] • $n^{\lambda_\gamma}\sqrt{\mathbb{E}(|\varphi(a,X_{i},b,X_{j},\hat{\gamma}_{l})-\varphi(a,X_{i},b,X_{j},\gamma)|^{2})} = o(1) $ ; • $n^{\lambda_\varphi}\sqrt{\mathbb{E}(|\hat{\varphi}_l(a,X_{i},b,X_{j},\gamma)-\varphi(a,X_{i},b,X_{j},\gamma)|^{2})} = o(1) $ ; • $n^{\lambda_\gamma}\sqrt{\mathbb{E}(|\phi_{ab}^\gamma(Z_{i},Z_{j},\hat{\gamma}_l,\alpha^\gamma)- \phi_{ab}^{\gamma}(Z_{i},Z_{j},\gamma,\alpha^{\gamma})|^{2})}= o(1)$; • $n^{\lambda_\varphi}\sqrt{\mathbb{E}(|\phi_{ab}^\varphi(Z_{i},Z_{j},\hat{\varphi}_l,\alpha^\varphi)- \phi_{ab}^{\varphi}(Z_{i},Z_{j},\varphi,\alpha^{\varphi})|^{2})} = o(1)$; • $n^{\lambda_{\alpha}}\sqrt{\mathbb{E}(|\phi_{ab}^\omega(Z_{i},Z_{j},\omega,\hat{\alpha}_{l}^\omega)- \phi_{ab}^\omega(Z_{i},Z_{j},\omega,\alpha^{\omega})|^{2})} = o(1)$, \end{enumerate} where $1/4 < \lambda_\gamma, \lambda_\varphi, \lambda_\alpha$.

These are mean-square consistency conditions for $\hat{\gamma}_{l}$, $\hat{\varphi}_l$ and $\hat{\alpha}_{l}$. Assumption (ref) often follows from the L2 convergence rates of the nuisance estimators. The non-parametric literature gives rates for kernel regression and sieves/series (e.g. chen2007large). For $L_{1}$-penalty estimators such as Lasso see belloni2011 and belloni2013least. For low-level conditions for shrinkage and kernel estimators see Appendix B in sasaki2021estimation. Rates for $L_{2}$-boosting in low dimensions are in zhang2005boosting, and kueck2023estimation find rates for $L_{2}$-boosting with high dimensional data. For results on versions of random forests see wager2015adaptive and athey2019generalized. For single-layer, sigmoid-based neural networks see chen1999improved and for a modern setting of deep neural networks see farrell2021deep. Note that $\hat{\varphi}_l$ estimates conditional expectations where both the dependent variables and the conditioning ones are indexed by both $i$ and $j$. stute1991conditional calls such objects conditional U-statistics and studies the asymptotic properties of Nadaraya-Watson nonparametric estimators. I run the machine learning algorithms on the stacked pairs. Unfortunately, not much is known about rates for such machine learning regressions which are computationally demanding. Define now the following interaction terms for $\omega \in \{\gamma,\varphi\}$ and let $||\cdot||$ denote the L2 norm.

align*[align* omitted — 450 chars of source]
assumptionFor each $l =1,...,L$ \begin{enumerate}[(i)] • $\int \int \phi_{ab}^{\gamma}(z_i,z_j,\gamma,\hat{\alpha}_l^{\gamma}) F(dz_i)F(dz_j) = 0$ and $\int \int \phi_{ab}^{\varphi}(Z_i,Z_j,\varphi,\hat{\alpha}_l^{\varphi}) F(dz_i)F(dz_j) = 0$. • $\mathbb{E}(||\hat{\gamma}_l - \gamma||^2) = o(n^{-2\lambda_\gamma})$, $\mathbb{E}(||\hat{\varphi}_l - \varphi||^2) = o(n^{-2\lambda_\varphi})$ and \begin{align*} |\mathbb{E}[(\varphi(a,X_i,b,X_j,\tilde{\gamma}) + \phi_{ab}^{\gamma}(Z_i,Z_j,\tilde{\gamma},\alpha^{\gamma}))\pi_{ab}(X_i,X_j)]| &\leq C ||\tilde{\gamma} - \gamma||^2 \\ |\mathbb{E}[(\tilde{\varphi}(a,X_i,b,X_j,\gamma) + \phi_{ab}^{\varphi}(Z_i,Z_j,\tilde{\varphi},\alpha^{\varphi}))\pi_{ab}(X_i,X_j)]| &\leq C ||\tilde{\varphi} - \varphi||^2. \end{align*} \end{enumerate}

Assumption (ref) (i) is usually verified from visual inspection and (ii) requires L2 convergence rates and some smoothness. $C$ is a constant so the right-hand-sides above do not depend on $\pi \in \Pi$.

assumptionFor each $l = 1,...,L$: $\sqrt{n}\mathbb{E}(\hat{\xi}_{ij,ab,l}^\omega\pi_{ab}(X_i,X_j)) = o(1).$

These are rate conditions on the remainder terms $\hat{\xi}_{l}^\omega(w_{i},w_{j})$. Often, Assumption (ref) follows if $\sqrt{n} ||\hat{\alpha}_{l}^{\omega}-\alpha||||\hat{\omega}_{l}-\omega|| = o(1)$. In essence, it is enough for the product of the nonparametric estimators to go to zero at a $\sqrt{n}$ rate.

Conditions on the complexity of the policy class

The complexity of the policy class must also be restricted. If all sorts of subsets of $\mathcal{X}$ are allowed to decide who should be treated then we get overfitted policy rules. As in athey2021policy I measure the policy class complexity with its VC dimension (see for instance wainwright2019high) which is allowed to grow with the sample size.

assumptionThere are constants $0< \beta < 1/2$ and $n^*\geq 1$ such that for all $n \geq n^*$, $VC(\Pi_n) < n^\beta$.

Examples of finite VC-dimension classes are linear eligibility scores or generalized eligibility scores. Policy classes whose VC-dimension can increase with the sample size are for example decision trees which get deeper with sample size.

Upper bounds

Let now

align*[align* omitted — 480 chars of source]

$W(\pi)$ and $\widetilde{W}_n(\pi)$ are the welfare and the infeasible estimator of the welfare at policy rule $\pi$ when all nuisance parameters are known. $\hat{W}_n(\pi)$ is the feasible estimator which we already introduced in ((ref)). Let $W_{\Pi_n}^* = \sup_{\pi \in \Pi_n} W(\pi)$ be the best possible welfare. I give upper bounds to the regret: $\mathbb{E}[W_{\Pi_n}^* - W(\hat{\pi})]$. I start bounding the regret by

align[align omitted — 253 chars of source]

The second term above is a standard centered U-process indexed by $\pi \in \Pi_n$. I start as in athey2021policy by showing the rate of convergence of this second term. I work for some fixed $(a,b) \in \{0,1\}^2$ and let $\Pi_{ab,n} = \{\pi_{ab}: \pi \in \Pi_n\}$. The first step is to bound it by the Rademacher complexity which I define as \[ \mathcal{R}_n(\Pi_{ab,n}) = \mathbb{E}_\varepsilon \biggl(\sup_{\pi \in \Pi_n} \biggl| \lfloor n/2 \rfloor^{-1} \sum_{i=1}^{\lfloor n/2 \rfloor}\varepsilon_i \Gamma_{i,\lfloor n/2 \rfloor + i}^{ab} \pi_{ab}(X_{i},X_{\lfloor n/2 \rfloor + i})\biggr| \biggr), \] where $F_\varepsilon$ is the distribution of Rademacher random variables taking value $1$ and $-1$ with equal probability. $(\varepsilon_1,...,\varepsilon_{\lfloor n/2 \rfloor})$ are independent draws from $F_\varepsilon$. The next result gives a bound for the Rademacher complexity which relies on a characterization of U-statistics as dependent sums of independent sums introduced in hoeffding1963

lemma\[ \mathbb{E}\biggl[\sup_{\pi \in \Pi_n} |\widetilde{W}_n(\pi) - W(\pi)|\biggr] \leq \mathbb{E}[2 \mathcal{R}_n(\Pi_{ab,n})]. \]

We want an asymptotic upper bound for $\mathbb{E}[\mathcal{R}_n(\Pi_{ab})]$. While kitagawa2018should and others provide bounds in terms of the maximum of the (bounded) scores, athey2021policy provide bounds based on the variance. The next result provides a bound on the Rademacher complexity based on $S_{ab} = \mathbb{E}[ \Gamma_{i,j}^{2 \, ab}]$.

lemmaUnder Assumptions (ref) and (ref), if $\Gamma_{ij}^{ab}$ has bounded support, then \[ \mathbb{E}[\mathcal{R}_n(\Pi_{ab,n})] = \mathcal{O}\biggl(\sqrt{\frac{S_{ab} \cdot VC(\Pi_{ab,n})}{\lfloor n/2 \rfloor}} \biggr). \]

Lemma (ref) can be generalized to sub-Gaussianity. However, it comes at the cost of making the (already involved) proofs substantially less tractable. Now I provide asymptotic upper bounds for the first term in ((ref)). escanciano2023machine show that for given $\pi \in \Pi_n$, $\sqrt{n}(\hat{W}_n(\pi) - \widetilde{W}_n(\pi)) \to_p 0$. The next result makes this uniform in $\pi \in \Pi_n$.

lemma[Uniform coupling] Under Assumptions (ref) and (ref) \[ \sqrt{n}\mathbb{E}[\sup_{\pi \in \Pi_n}|\hat{W}_n(\pi) - \widetilde{W}_n(\pi)|] = \mathcal{O}\biggl(1 + \frac{VC(\Pi_{ab,n})}{\lfloor n/2 \rfloor^{\min(\lambda_\gamma,\lambda_\varphi,\lambda_\alpha)}}\biggr). \]

Finally, using Lemmas (ref) and (ref) the following holds.

theoremIf Assumptions (ref), (ref) and (ref) hold with $\beta < \min(\lambda_\gamma,\lambda_\varphi,\lambda_\alpha)$. Then \[ \mathbb{E}[W_{\Pi_n}^* - W(\hat{\pi})] = \mathcal{O}\biggl(\sqrt{\frac{S_{ab} \cdot (2VC(\Pi_{n})-1)}{\lfloor n/2 \rfloor}} \biggr). \]

Empirical illustration

I study the optimal allocation of children to preschool. I make use of the Panel Study of Income Dynamics (PSID) database. This survey contains a rich set of circumstances and long-term outcomes. In 1995, adults between 18-30 years old were asked about their participation in preschool. The outcome is the average earnings from 25 to 35 years old. I assume selection on sex, birthyear, average parental income in the 5 years before birth, mother's education, father's education, father's occupation and whether the individual is black. In Table (ref) we see the results of estimating the Average Treatment Effect (ATE), Gini, IOp and IGM measured by the Kendall-$\tau$.

table[table omitted — 358 chars of source]

To estimate the ATE, I use the doubly robust Augmented Inverse Propensity weighting from robins1994estimation using Conditional Inference Forests (CIF) to estimate the regression functions and propensity scores. I chose CIF by cross-validation among a pool of different machine learners. Under the assumption of no selection on observables, we observe a positive effect of attending preschool of 4,622\$ of added annual earnings. Dollars have been adjusted by the CPI to 2010 dollars. The Gini coefficient is $0.39$ and IOp is $0.17$, i.e. almost 44% of total inequality can be explained by the circumstances we observe. The Kendall-$\tau$ is around 0.17 which indicates a positive association between parental and child incomes.

I compute optimal treatment rules based on parental income and the mother's years of education. I set the target in the Kendall-$\tau$ welfare to zero, meaning that the aim is to completely erase intergenerational persistence. As the policy class, I use 2-depth decision trees. The U-statistic nature of the SWF prevents me from using the computational shortcuts in athey2021policy since the sub-trees are not independent optimization problems. Instead, I use the deciles of parental income as cutting points instead of all the observed values of parental income.

Optimal treatment allocation is the same for additive and IOp welfare. Although this seems surprising, it is possible if decreasing inequality of opportunity leads to sizeable reductions of the average. In fact, as reported in Table (ref), the estimated rule maximizing the average already drastically decreases IOp. I show the optimal rule under these two welfares in Figure (ref). At the terminal nodes, I report the number of observations, the conditional average treatment effect (CATE) in the node and the proportion of observations node that are treated in the data ($\hat{p})$. For additive/IOp welfare, the first cutting point is whether parental income is below or above the 40th percentile (51,515\$). If an observation is below this cut-off the tree splits according to the education of the mother. If parental income is below the 40th percentile and the mother's education is less than college (below 13 years) the tree allocates the observation to treatment. The CATE in this node is positive so, as we would expect, an additive policy maker treats these observations. If parental income is below the 40th percentile but the mother is highly educated the CATE is negative and hence the additive policy maker does not allocate the individual to treatment. For high parental income, we split next on the mother's education but at a higher level. If your parental income is higher than the 40th percentile and your mother attended college or less (16 years of education or less) you are allocated to treatment and in this node, we have large positive effects. However, we do not allocate kids with high parental income and high maternal education to preschool since the CATE in this group is negative.

figure[figure omitted — 1,223 chars of source]

If we take parental income to be a proxy for the quality of preschool, it is enough for the mother to have more than 13 years of education for the child to be better off without preschool. However, for children who attend better preschools (have higher parental income), the mother has to have more than 16 years of education for the child to be better off without preschool. This is in line with results in the psychology and economics literature (see fort2020cognitive). To decrease IOp further, we need to treat advantaged kids who do not benefit from preschool. The penalization of IOp is not severe enough to do this. The observed treatment allocation deviates significantly from the optimal rule, this is likely as parents consider more than future earnings when deciding on preschool.

In figure (ref), we see the optimal policy rule for the inequality SWF. The tree is the same except for the first cutting point on parental income. We first divide individuals into those with parental income lower and higher than the 20th percentile (37,699\$). Then, the division based on the mother's education is the same. Hence, compared to the previous tree, we shift 20% of the population to the right side subtree. For instance, a kid who has a parental income of 40,000\$ and whose mother has 16 years of education is not treated under the additive/IOp welfare but is treated under the inequality based optimal rule. Although masked by other observations in the node, this 20% of the population who is switched to treatment has an estimated negative CATE. When we penalize all sorts of inequalities, it starts becoming optimal to decrease the average outcome to decrease inequality.

figure[figure omitted — 1,216 chars of source]

In Figure (ref) we see the results for an IGM aware SWF. Notice that in the IGM welfare there is no efficiency motive and we target a zero Kendall-$\tau$. Hence, the optimal policy is even more controversial since individuals with positive treatment effects are not treated and individuals with negative treatment effects are treated.

figure[figure omitted — 1,208 chars of source]

Finally, in Table (ref) we see a summary of the results and compare the estimated optimal treatments with situations in which no one or everyone is treated. For additive, IOp and inequality welfares, treating no one gives the worst welfare. In the IGM case, treating no one and treating everyone give basically the same welfare (note that maximal welfare in the IGM case is 0). The additive and IOp welfares have an estimated optimal policy rule that attains the highest average outcome and the lowest IOp. The decrease in IOp under this rule is drastic. While in the sample we can explain 44% of total inequality with circumstances (IOp/Gini), at the optimal additive/IOp rule we explain 35%. This explains why both rules coincide. In other settings, maximizing the average might increase IOp. The estimated optimal treatment rule for IGM gives a similar Gini and IOp as the the estimated optimal policy rule for IOp or inequality welfare, but at a much larger cost in average outcome. The estimated optimal policy rule under the inequality welfare gives the lowest Gini compared to additive and IOp welfares. Interestingly, it gives the highest IOp across all welfares. IGM rule gives the lowest Kendall-$\tau$.

table[table omitted — 1,931 chars of source]

Finally, comparing the results with what we observe with the treatment allocation in the sample, we achieve a higher mean with all other welfares except with the IGM welfare. The Gini in the sample is the same as the one under the estimated optimal additive and IOp rule. The observed IOp in the sample is higher than the one achieved under the estimated rules of all other welfares. IGM observed in the sample is the lowest (highest Kendall-$\tau$) compared to all welfares. Finally, the share of treated in the sample is also lower than the one achieved under the estimated optimal rule of all other welfares. However, this could be explained by not taking into account costs of treatment.