Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
63,979 characters · 17 sections · 0 citation commands
Inequality Sensitive Optimal Treatment Assignment
Starting with the work of Manski (2000), Manski (2004) and following Wald (1939), Wald (1945), Wald (1971), a large and growing literature in economics has considered the problem of optimal policy learning and treatment assignment, rooted in statistical decision theory. For example, according to Bayesian statistical decision theory, a researcher represents uncertainty about treatment effects using a prior, identifies a function that maps incomes to welfare, and selects the policy or treatment that leads to the highest expected welfare, where the expectation is taken with respect to that prior. An alternative decision-theoretic approach avoids the specification of priors and instead studies rules for learning which policies would be uniformly satisfactory in terms of welfare (Manski (2004), Manski (2005), Manski (2007b), Manski (2007a); Schlag (2006); Hirano and Porter (2009); Stoye (2009), Stoye (2012); Tetenov (2012), Manski and Tetenov (2016), and Kitagawa and Tetenov (2018)).
The optimal policy learning or treatment assignment problem has also been considered in parallel literatures developed in both statistics (Luedtke and Laan (2016), Qian and Murphy (2011), Zhang et al. (2012), Zhao et al. (2012)), machine learning (Beygelzimer and Langford (2009), Dudík, Langford, and Li (2011), Li et al. (2012), Jiang and Li (2016), Thomas and Brunskill (2016), Kallus and Zhou (2018), Strehl et al. (2010), Swaminathan and Joachims (2015)), and recently at the intersection of all these fields (Athey and Wager (2021)).
Despite all these developments, the primary approach for policy evaluation and treatment assignment used both in academia, industry and policy circles remains estimating average treatment effects of randomized A/B experiments assessed using the theory of hypothesis testing. Unfortunately, it is well-known that treatment selection based on this approach is problematic: First, the approach offers no rationale for the conventionally used Type I error probabilities (5% or less) and Type II error probabilities (10% to 20%); second, the approach pays no attention to the magnitude and distribution of losses to welfare when errors occur (Manski and Tetenov (2016)). Policymakers selecting from among different policy options need to account for their distributional impacts, and existing analyses of randomized policy experiments often either neglect distributional issues or lack an economic framework for evaluating them.
The literature on welfare economics is ideally suited for providing the requisite economic framework, but it has seldom informed the theory and practice of program evaluation. A notable exception is Kitagawa and Tetenov (2021), who write: “{[}The{]} rich and insighful works in welfare economics have not yet well linked to econometrics and empirical analysis for policy design.” In the proposed project, and in line with the aim of Kitagawa and Tetenov (2021), I establish connections between the theoretical welfare economics literature and statistical decision theory.\footnote{An early reference is Dehejia (2005). See also the literature reviewed in Section 1.1 of Kitagawa and Tetenov (2021).} While Kitanawa and Tetenov focus on the class of rank-dependent Social Preferences (SP) (Blackorby and Donaldson (1978)) and assume that their welfare criterion is point-identified by the sampling process, here my focus is specifically on the special case of additively separable SP (Atkinson (1970)) but consider sampling processes that allows for both point and partial identification, as in Manski (2024). One should then view these lines of research as complementary.
The present work is part of a series of papers that seek to develop an integrated statistical framework for analyzing distributional impacts in interventions, building on modern welfare economic theory. First, Fleurbaey and Zambrano (2024) develops a methodology that can help an evaluator decide among different kinds of SP under certainty, guided by the tradeoffs between the well-being of individuals the evaluator would find acceptable. Once the nature of the tradeoffs is precisely identified, this determines the specific SP one should use. Second, Zambrano (2024) extends work by Fleurbaey (2010) and identifies the normative assumptions that an inequality averse social planner, drawing inferences based on data, would want to maintain in their evaluation of the treatment effects. The main result in Zambrano (2024) indicates that these assumptions determine the specific cardinal representation of the SP under certainty that should be brought into the decision problem under uncertainty: the egalitarian equivalent (ee) representation, a concept similar to the certainty equivalent in expected utility theory. Third, Flores, Kairy, and Zambrano (2024) focuses on developing the methodology for properly estimating inequality sensitive treatment effects, and their bounds, together with corresponding uncertainty estimates for these.
In this paper, I extend the optimal treatment choice apparatus described in Manski (2024) to the case where the welfare evaluation is made using the egalitarian equivalent measures described above, and explore the optimal treatment rules that arise (Section (ref), Theorem (ref) and Theorem (ref)). I also created a companion website to the paper, available at \url{https://osf.io/wv5jt/}, which contains interactive visuals, apps and narratives that explain and motivate our results in simple and intuitive terms.
Consider a fixed and finite population of \(n\) individuals, where each individual \(i\) has a known income \(y_{i} \ge 0\). For ease of exposition, this Section considers income as an index of individual advantage, but the results extend immediately to any cardinal, interpersonally comparable variable. In particular, it is possible to adjust income for nonmarket aspects of quality of life that individuals enjoy or endure, and use this adjusted income (usually called “equivalent income”) as the relevant index instead of ordinary income.
An evaluator has SP that can be represented by a social welfare function \(W(y_1,...,y_n):=\sum_{i=1}^{n}f\left(y_{i}\right)\) (Atkinson (1970)), where \(f\) is an increasing function with values taking an interval in \(\mathbb{R}\cup\left\{ -\infty,+\infty\right\}\). Within this class of SP, selecting \(f\) pins down the specific SP of the evaluator. This class of SP is sometimes called generalized-utilitarian, where \(f\) being and affine function corresponds to the utilitarian SP and \(f\) being strictly concave corresponds to the prioritarian SP. See, e.g., Adler (2022) for discussion.
It is well-known that a given SP has many different cardinal representations. For instance, for any monotone transformation \(g\), \(W(y_1,...,y_n)\) and \(g\left(W(y_1,...,y_n)\right)\) generate the same ranking over income distributions. One such representation is the egalitarian equivalent representation \(\mathcal{EE}(y_1,...,y_n):=f^{-1}\left(W(y_1,...,y_n)\right)\), which denotes the level of income \(ee = \mathcal{EE}(y_1,...,y_n)\) such that, the evaluator would be indifferent between the distribution \((y_1,...,y_n)\) and \((\underbrace{ee,..., ee}_{n-times}).\) It is less known, however, that the choice of representation may matter in practice, that is, in the context of solving an empirical welfare maximization problem.
To investigate this matter, Zambrano (2024) considers an evaluator that faces uncertainty about what income distribution arises with a given treatment. Assume that the evaluator represents the uncertainty through a finite set of states of the world, \(S = \{s_{1},...,s_{m}\}\). Each treatment is associated with an income distribution \(y^{s}\) under each state of the world and a prospect, \(y\), collects the income distributions induced by a given treatment across the \(m\) states of the world.
The evaluator's problem is then to rank the prospects \(y = (y_{i}^{s})_{i \in N, s \in S }\), where \(y_{i}^{s}\) describes the income attained by individual \(i\) in state \(s\), \(y_{i} = (y_{i}^{s})_{s \in S}\), and \(y^{s} = (y_{i}^{s})_{i \in N}\). Let \(\Upsilon \subseteq \mathbb{R}^{nm}\) denote the relevant set of such prospects over which the evaluation must be made. In this setting, the SP (a complete, transitive, binary relation) over the set \(\Upsilon\) is denoted \(R\), with strict preference \(P\) and indifference \(I\).
In what follows it will be useful to refer to the preferences over prospects the evaluator has for the special case when \(n=1\), and preferences over the restricted set \(\Upsilon_{1} \subseteq \mathbb{R}^{m}\) are denoted by \(\succeq\). One can think of these preferences as those held by an evaluator that is unconcerned by distributional considerations, and that treats society as though it consists of a single or representative individual. Therefore, when comparing two treatments, this evaluator is solely concerned with the effects this treatment has on this one individual. We take these preferences as known, perhaps arising from considerations regarding how the single-person decision problem under uncertainty or ambiguity should be approached (see, e.g., Stoye (2011) for discussion).
Theorem 3.1 in Zambrano (2024), following Fleurbaey (2010), shows that, under standard continuity, dominance, and Pareto conditions, the following is true: if, in the absence of inequality, the evaluator acts as though there is only one (representative) individual then, in the presence of both inequality and uncertainty, the evaluator chooses among treatments by comparing the profiles \((\mathcal{EE}(y^1),...,\mathcal{EE}(y^m))\) among treatments according to \(\succeq\).
The interpretation is that the social evaluation can be done as the single-person evaluation, but one applies the single-person decision methodology under uncertainty or ambiguity to the \(m-\)dimensional vector of egalitarian equivalents \((\mathcal{EE}(y^1),...,\mathcal{EE}(y^m))\). Therefore, given some SP about known distributions of income in society, one incorporates those into an optimal treatment assignment framework under uncertainty and ambiguity, given a set of states of the world, by first aggregating across individuals for each state, using the egalitarian equivalent function, and then aggregating across states in whichever way the evaluator normally does so in single-person problems.
This is significant because it turns out that the choice of representation of the evaluator's preferences can be consequential in cases with statistical uncertainty: one would obtain different rankings over treatments depending on what representation one was using. Understanding this is crucial before we embark on the development of a theory of inequality averse optimal treatment assignment, so let's examine an example in some detail that makes precisely this point.
Consider a situation, reproduced from Zambrano (2024), with three individuals, two states of the world, and two prospects. Prospect \(a\) is given by \(\left[
\right]\) whereas Prospect \(b\) is given by \(\left[
\right]\), where rows correspond to individuals and columns corresponds to states.
Consider, in addition, an inequality averse evaluator who summarizes the distribution of income in any given state by the geometric mean (GM). This means that, given a prospect \(y\), the evaluator computes, for each state \(s\), the magnitude \((y_1^s y_2^s y_3^s)^{\frac{1}{3}}\). Assume further that the evaluator is a risk neutral Bayesian decision maker, with priors of \((\frac{2}{3},\frac{1}{3})\) on states 1 and 2 respectively. Then this evaluator would prefer Prospect \(b\) over Prospect \(a\) (the expected certainty equivalent for prospect \(a\) is \(3.95\) whereas for prospect \(b\) is \(4.48\)). Table (ref) summarizes these calculations.
However, an equivalent representation of those geometric mean social preferences in any given state is the arithmetic mean of the logs (AML) of the incomes. If we were to use this representation, this means that, given a prospect \(y\), the evaluator computes, for each state \(s\), the magnitude \({\frac{1}{3}}(\ln y_1^s + \ln y_2^s + \ln y_3^s)\), and the expected welfare is then \(1.29\) for Prospect \(a\) and \(1.22\) for Prospect \(b\). The change in representation would thus make us change how we order the prospects. Table (ref) summarizes these calculations.
In a situation like the one described in Table (ref), the choices made by a minimax regret evaluator also vary depending on whether one uses the GM or the AML as the representation for the social preferences, but the choices made by a maximin evaluator will not depend on the choice of representation. I introduce and discuss these preferences methods in the context of the welfare evaluation problem under uncertainty and ambiguity in Section (ref).
When contemplating which representation to use, it bears noticing that working with the AML representation of the evaluator's social preferences under certainty amounts to imputing a degree of risk aversion to the evaluation that the evaluator does not really have. This can be seen most easily in the case of a risk neutral Bayesian evaluator in a single-person evaluation. In this case, using the AML representation amounts to applying a concave transformation of the data coming from a single individual before taking expectations, and this would make the decision maker act as though they are risk averse.\footnote{With Bernoulli utility function given by the log function.} The GM representation, on the other hand, makes no such imputation.
Statistical decision theory adds to the above structure by assuming that the evaluator observes data generated by some sampling distribution and uses these data to guide their decision. The main difference between this earlier work and my current analysis is that social preferences in the present context are represented in accordance with Theorem 3.1 in Zambrano (2024). This turns out to be of crucial importance under partial identification and a decision must be made under uncertainty or ambiguity. I consider here the choices made by an evaluator with preferences over prospects as in Section (ref) who is considering two treatments, \(d \in \{a, b\}\) (e.g. control and treatment), which respectively lead to the prospects \(y(a)\) and \(y(b)\) in \(\Upsilon\).
Let \(f\) be a continuous, monotone, strictly concave function. Then the egalitarian equivalent of the income distribution associated with \(y^{s}(d)\) is \(ee^{s}(d):= \mathcal{EE}(y^{s}(d)) = f^{-1}\left(\frac{1}{n} \sum_{i=1}^{n} f(y^{s}_{i}(d))\right)\) for \(d \in \{a, b\}\).
Assume that the evaluator observes data \(Z^{t}=(Z_{1},\dots,Z_{t})\) that are independent and identically distributed with \(Z_j \sim P_s\) on some space \(\mathcal{Z}\). Let \(\mathcal{P}=\{P_s: s\in S\}.\)
As in Section 3 in Manski (2024), I first consider decision making with the knowledge that econometricians have assumed in identification analysis. Then the evaluator knows at state \(s\) the probability distribution \(P_s \in \mathcal{P}\) but not necessarily the underlying state. In this context, a treatment rule is a mapping \(\delta: \mathcal{P} \rightarrow [0,1]\), which gives the probability of a (future) individual being assigned to treatment \(b\), given knowledge of \(P_s\). Then the egalitarian equivalent of treatment rule \(\delta\) in state \(s\), given \(P_s\), is
\[ ee^{s}(\delta(P_s)) = f^{-1}\left[\delta(P_s)\left(\frac{1}{n}\sum_{i=1}^{n} f\left(y_{is}(b)\right)\right)+(1-\delta(P_s))\left(\frac{1}{n}\sum_{i=1}^{n} f\left(y_{is}(a)\right)\right)\right] \]
which simplifies to \(ee^{s}(\delta(P_s)) = f^{-1}\left[\delta(P_s)f(ee^{s}(b))+(1-\delta(P_s))f(ee^{s}(a))\right].\) For expositional clarity, I label the point-mass assignments \(ee^{s}(0)\) and \(ee^{s}(1)\) by \(ee^{s}(a)\) and \(ee^{s}(b)\), respectively.
As a consequence of 3.1 in Zambrano (2024), the treatment assignment problem boils down to selecting \(\delta(P_s)\) in order to obtain the most favorable profile \[ (ee^{1}(\delta(P_s)),...,ee^{m}(\delta(P_s))) \]
according to the preferences \(\succeq\) over the restricted set \(\Upsilon_{1}\) discussed in Section (ref), and given \(P_s\). Let \(\pi\) be a prior probability distribution on \(S\) and, for each \(P_s \in \mathcal{P}\), let \(S(P_s) \subset S\) denote the truncated state space obtained with knowledge of \(P_s\). Below I consider three versions of \(\succeq\), and investigate the characteristics of the optimal treatment rules according to these criteria.
Equation (ref) corresponds to the Bayesian criterion, Equation (ref) corresponds to the maximin criterion and Equation (ref) corresponds to the minimax regret criterion, and where the regret of a treatment rule at a state is defined as the diffference between the most favorable welfare obtainable given knowledge of the state and the welfare associated with the treatment rule at that state. Each of these treatment rules encode a different way of incorporating uncertainty and ambiguity into the analysis of welfare. A maximin evaluator selects the treatment rule with the largest worst-case welfare across states; a minimax regret evaluator anticipates the worst regret of a treatment rule across states, and focuses on selecting the treatment rule with the smallest worst regret; and a Bayesian evaluator allows what happens to welfare in the entire state space to inform the evaluation of a treatment rule, weighted by the relative likelihood of the different states, as captured by the evaluator's prior on the state space. For any of these, I assume, for simplicity, that ties are broken in favor of treatment \(a\).
Two cases of interest arise. In the first case, \(S(P_s) = \{s\}.\) In this case, we say that the true state is point identified. The true state is partially identified if \(S(P_s)\) is a proper non-singleton subset of \(S\). Define
the egalitarian equivalent treatment effect (EETE) at \(s\).
If the true state is point identified, Equation (ref) and Equation (ref) both reduce to the optimization problem \(\max_{\delta(P_s) \in [0,1]} ee^{s}(\delta(P_s))\). Equation (ref) reduces to \(\min_{\delta(P_s) \in [0,1]} [ \max \{ee^{s}(a), ee^{s}(b)\}\) \(- ee^{s}(\delta(P_s))]\) which is equivalent to \(\max_{\delta(P_s) \in [0,1]} ee^{s}(\delta(P_s))\). In any of these three cases, the optimal solution is therefore the same: \(\delta(P_s) = 1 \left(\tau_{ee}(s) > 0 \right)\), that is, to assign individuals to treatment \(b\) if \(\tau_{ee}(s) > 0\) and to treatment \(a\) otherwise. In this case, true egalitarian equivalent welfare at every state \(s\) is maximized.
If the true state is partially identified, all these criteria generally yield different answers, and those answers may be sensitive to what representation of the evaluator's social preferences under certainty one chooses to adopt. The first point was already made by Manski (2024). The second point is novel.
Suppose there exist states \(s_{w}, s_{a}\) and \(s_{b}\) such that \(ee^{s_{w}}(a) = ee^{s_{b}}(a) = \min_{s \in S(P_s)}ee^{s}(a)\), \(ee^{s_{w}}(b) = ee^{s_{a}}(b) = \min_{s \in S(P_s)}ee^{s}(b)\), and \(ee^{s_{d}}(d) = \max_{s \in S(P_s)}ee^{s}(d)\) for \(d \in \{a, b\}\).
The interpretation is that state \(s_{w}\) is a worst state for both treatments, state \(s_a\) is a best state for treatment \(a\) and a worst state for treatment \(b\), and state \(s_b\) is a best state for treatment \(b\) and a worst state for treatment \(a\). Assuming the existence of these states is not needed for the result below, but the assumption facilitates exposition.
Theorem (ref) below extends the results in Sections 5-1-5.3 of Manski (2024) to the present setting.
All proofs are in the Appendix.
In this Section, I aim to illustrate how the optimal treatment rules used by inequality averse evaluators compare to those used by their inequality neutral counterparts in a concrete setting under partial identification. My starting point below is a variation on the binary outcome example in Manski (2004), p. 1226, and Manski (2019), p. 301, also studied in Stoye (2009), p. 72.
An evaluator with preferences over prospects in \(\Upsilon\) as in Section (ref) is considering two treatments, \(d \in \{a, b\}\), with \(y_{i}^{s}(d) \in \{\underline{p},\bar{p}\}\) for \(d = a, b\) and \(\underline{p},\bar{p} \in (0,1]\). The evaluator knows the outcome distribution of the status quo treatment, \(y(a)\), but does not know the outcome distribution of the innovation, \(y(b)\). Let \(P(y^{s}) = \frac{1}{n} \sum_{i=1}^{n} 1_{\{y_{i}^{s}=\bar{p}\}}\), \(p(a) = P(y(a))\) and \(p^{s}(b) = P(y^{s}(b))\) for \(s = 1,...,m\). Further, assume that \(s_{a}, s_{b} \in S\), where \(0 = P(y^{s_{a}}(b)) < p(a) < P(y^{s_{b}}(b)) = 1\).
Let \(\mathcal{EE}(x_{1},...,x_{n}) = f^{-1}(\frac{1}{n} \sum_{i=1}^{n} f(x))\), with \(f\) continuous, strictly increasing, strictly concave, with \(f(\underline{p}) = 0\) and \(f(\bar{p}) = 1\). Then the egalitarian equivalent of the outcome distribution associated with \(y^{s}\) is \(\mathcal{EE}(y^{s}) = f^{-1}(P(y^{s})f(\bar{p}) + (1 - P(y^{s}))f(\underline{p})) = f^{-1}(P(y^{s})).\)
In this context, the interpretation is that \(\mathcal{\mathcal{EE}}(y^{s})\) is the notional level of the outcome variable such that, if everyone in society had that outcome level, then the evaluator would be indifferent between that notional society and the actual society, where \(P(y^{s})\) of the individuals have an outcome of \(\bar{p}\) and \(1 - P(y^{s})\) of the individuals have an outcome of \(\underline{p}\).
Let \(\delta\) be the fraction of the population to be assigned to the treatment. Then the egalitarian equivalent of treatment rule \(\delta\) in state \(s\) is \(ee^{s}(\delta) = f^{-1}(p(a) + (p^{s}(b) - p(a))\delta)\), where I simply write \(ee(a)\) for \(ee^{s}(a).\)
As in Section (ref) above, the evaluation boils down to the statistical comparison of the profiles \((ee^{1}(\delta),...,ee^{m}(\delta))\) as one varies \(\delta\) from zero to one. I now show how the three statistical decision theories discussed above would approach this evaluation.
An inequality averse (\(IA\)) Bayesian evaluator would choose \(\delta\) to maximize \(\sum_{s=1}^{m} ee^{s}(\delta)\pi(s)\) for some prior \(\pi\) on \(S\). For reference, in this context an inequality neutral (\(IN\)) Bayesian evaluator chooses \(\delta\) to maximize \(\sum_{s=1}^{m} (p(a) + (p^{s}(b) - p(a))\delta) \pi(s)\). The solution to the \(IN\) Bayesian evaluator's problem is: \(\delta^*_{in} = 1\) if \(E_{\pi}[p(b)] := \sum_{s=1}^{m} p^{s}(b) \pi(s) > p(a)\) and \(\delta^*_{in} = 0\) if \(E_{\pi}[p(b)] \le p(a)\) (Manski (2004), p. 1228). The following is true:
Figure (ref) provides the intuition for the \(m = 2\) case, given a prior \(\pi\) over the two states. The axes measure the proportions \(P(y)\) in both states of the world for any prospect \(y\). The gray line denotes all prospects \(y\) with \(E_{\pi}[P(y)] = p(a).\) Therefore, the \(IN\) Bayesian evaluator accepts any prospect in the shaded gray area. The blue line denotes all prospects \(y\) with \(E_{\pi}[\mathcal{EE}(y)] = ee(a)\). Therefore, the \(IA\) Bayesian evaluator accepts any prospect in the shaded blue area. The segment connecting \(p(a)\) (corresponding to \(\delta = 0\)) to \(p(b)\) (corresponding to \(\delta = 1\)) contains all the treatment rules the evaluator chooses from. Thus, in the top panel all the feasible treatment rules are in the yellow segment, and both Bayesian evaluators set \(\delta^*_{in} = 1\). In the middle panel, all the feasible treatment rules are in the green segment, the \(IN\) Bayesian evaluator rejects the treatment, setting \(\delta^*_{in} = 0\), but the \(IA\) Bayesian evaluator accepts the treatment, setting \(\delta^*_{ia} = 1\). In the bottom panel, all the feasible tretament rules are in the orange segment, and both evaluators reject the treatment, setting \(\delta^*_{in} = \delta^*_{ia} = 0\).
A minimax regret evaluator would choose \(\delta\) to minimize \(R(\delta):= max_{s \in S} R(\delta,s)\), where the \(IA\) evaluator defines \(R(\delta,s)\) by \(R_{ia}(\delta,s) := \max \{ee(a), ee^{s}(b)\} - ee^{s}(\delta)\) whereas \(R(\delta,s)\) for the \(IN\) evaluator simplifies to \(R_{in}(\delta,s) := (1 - \delta)[p^{s}(b) - p(a)] 1_{\{p^{s}(b) \ge p(a)\}} + \delta[p(a) - p^{s}(b)] 1_{\{p^{s}(b) \le p(a)\}}\). The solution to the \(IN\) minimax regret problem is: \(\delta^{*}_{in} = 1 - p(a)\). (Stoye (2009), p. 73).
Figure (ref) provides the intuition for the result. In the left panel, the dashed yellow line corresponds to \(R_{in}(\delta,s_{a})\), the dashed blue line corresponds to \(R_{in}(\delta,s_{b})\), the dashed gray line corresponds to \(R_{in}(\delta)\). This function is minimized at \(\delta^{*}_{in} = 1 - p(a)\). In the right panel, the solid yellow line corresponds to \(R_{ia}(\delta,s_{a})\), the solid blue line corresponds to \(R_{ia}(\delta,s_{b})\), the solid gray line corresponds to \(R_{ia}(\delta)\). This function is minimized at \(\delta^{*}_{ia} > 1 - ee(a)\). Since \(ee(a) < p(a)\), the result follows.
An \(IA\) maximin evaluator chooses \(\delta\) to maximize \(min_{s \in S} \{ee^{s}(\delta)\}\), whereas the \(IN\) maximin evaluator maximizes instead \(min_{s \in S} \{(p(a) + (p^{s}(b) - p(a))\delta)\}\). The solution for the \(IN\) maximin evaluator is \(\delta^{*} = 0\) (Manski (2004), p. 1228) but, because \(f^{-1}\) is monotone, both problems share the same solution.
JobCorps (Schochet, Burghardt, and McConnell (2008)) is a widely studied education and training program for disadvantaged youth. Despite it being a randomized intervention, estimating the effect of the program on applicant's wages is difficult due to the fact that the evaluator only observes wages for those individuals who are employed. The implication of this, in the context of the present paper, is that, to the extent that the outcome variable of interest for the evaluator is wages, one may not be able to point identify at state \(s\) the objects \(ee^{s}(a)\) and \(ee^{s}(b)\) and therefore \(\tau_{ee}(s)\) is also not point identified at \(s\). However, partial identification may be achievable, under relatively mild assumptions.
Flores, Kairy, and Zambrano (2024) adapts the bounds analysis of Horowitz and Manski (2000), Lee (2009), and Chen and Flores (2015) in order to arrive at a relevant set of lower and upper bounds for \(ee^{s}(a)\) and \(ee^{s}(b)\). While the goal in Flores, Kairy, and Zambrano (2024) is to use those magnitudes to obtain bounds on the \(EETE\), below I use them to determine the optimal treatment assignment according to the minimax regret criterion when the evaluator is inequality averse. I abstract from estimation problems in this sub-section and treat these bounds as recoverable in a large sample and therefore known.\footnote{The proper estimation metodology of these effects, together with their bounds and uncertainty estimates is the focus of Flores, Kairy, and Zambrano (2024).}
In order to arrive at inequality sensitive estimations of the effect of the treatment, one needs to specify the function \(f\) that captures the attitudes towards inequality of the evaluator. In this application, I consider an evaluator with \(f(y) = \frac{y^{1-\gamma}}{1-\gamma}\), where \(\gamma \ge 0\) is an inequality aversion parameter. I allow for two kinds of evaluators: \(\gamma = 0\) and \(\gamma = 2\). To interpret these choices, consider that, when \(\gamma = 2\), the evaluator would wish, in a two-person evaluation, to protect at least \(50\%\) of an individual's wage regardless of what happens to the wages of the other individual. For reference, the egalitarian equivalent measure that corresponds to \(\gamma = 2\) is the harmonic mean, and \(\gamma \rightarrow 1\) and \(\gamma =0\) correspond, respectively, to the egalitarian equivalent measures given by the geometric and arithmetic means. Either of those last two choices corresponds to a level of protected wages equal to zero. Loosely, values of \(\gamma\) greater than one essentially protect those with the lowest wages from losing everything should the better off in terms of wages gain disproportionately from an intervention. See Fleurbaey and Zambrano (2024) for details.
Table (ref) reports Horowitz and Manski, Lee, and Chen and Flores bounds for \(ee^{s}(a)\) and \(ee^{s}(b)\), for \(\gamma = 0\) and \(\gamma = 2\), when the outcome variable is hourly wages at week 208 after random assignment. Those bounds can then be used to obtain the corresponding optimal treatment assignments, using Theorem (ref). Figure (ref) describes how these assignments are obtained. In each graph, the blue line denotes \(R(\delta, s_b):=ee^{s_b}(b) - ee^{s_b}(\delta)\), with \(R(0, s_b) = ee^U(b) - ee^L(a)\) (computed using the estimates in Table (ref)), and \(R(1, s_b)=0\). In turn, the yellow line denotes \(R(\delta, s_a):=ee^{s_a}(a) - ee^{s_a}(\delta)\) with \(R(0, s_a)=0\) and \(R(1, s_a) = ee^U(a) - ee^L(b)\) (also computed using the estimates in Table (ref)). The intersecion of these lines in the four leftmost graphs in Figure (ref) corresponds to the treatment rule that minimizes worst regret. On the two graphs on the right in Figure (ref), regret is always worst along \(R(\delta, s_b)\), as the blue line is uniformly above the yellow line, and in those two cases worst regret is minimized at \(\delta = 1\). Table (ref) reports the resulting optimal treatment assignments.
To better understand the results from Table (ref) note that, as we move, for each row, from the leftmost to the rightmost column in the table, the fraction of individuals assigned to treatment \(b\) grows, until the fraction reaches 1. The reason why this happens is that the move from the leftmost to the rightmost column in Table (ref) corresponds to a progressive reduction in the size of the identified set, until it no longer contains zero (Chen and Flores (2015)).
From Table (ref) we also learn that, unlike in the application discussed in Section (ref), it is not always the case that the inequality averse evaluator assigns a comparative larger fraction of the population to treatment \(b\), relative to the inequality neutral evaluator. The differences between these applications stem from the following: In the application from Section (ref), the introduction of inequality aversion increases the worst regret of treatment \(a\) but decreases the worst regret of treatment \(b\), as shown on the right panel in Figure (ref). This has the unambiguous effect, as discussed in Section (ref), of increasing the fraction assigned to treatment \(b\) as inequality aversion grows. In the present application, however, inequality aversion may increase or decrease the worst regret of either treatment. This is so because worst regret is computed as the difference of two egalitarian equivalent measures, each of which shrink with inequality aversion. Which of the two measures shrinks faster then determines whether worst regret grows or shrinks for a given treatment option as inequality aversion increases.
Let's take a look at how this plays out in practice in the context of the JobCorps application. First, consider the case of the Horowitz and Manski bounds. Given those bounds, the worst regret from treatment \(a\) shrinks from \(\$12-\$5.4=\$6.6\) per hour to \(\$8.6-\$4.1=\$4.5\) per hour as we move from \(\gamma=0\) to \(\gamma=2\) (Table (ref)). This is driven by the fact that \(ee^U(b)\) drops by more (\(\$12-\$8.6=\$3.4\)) than what \(ee^L(a)\) drops by (\(\$5.4-\$4.1=\$1.3\)) as we move from \(\gamma=0\) to \(\gamma=2\). On the other hand, the worst regret from treatment \(b\) grows from \(\$10.7-\$6.1=\$4.6\) per hour to \(\$9-\$3.8=\$5.2\) per hour as we move from \(\gamma=0\) to \(\gamma=2\). This is driven by the fact that \(ee^U(a)\) drops by less (\(\$10.7-\$9=\$1.7\)) than what \(ee^L(b)\) drops by (\(\$6.1-\$3.8=\$2.3\)) as we move from \(\gamma=0\) to \(\gamma=2\). The drop in worst regret for treatment \(a\), coupled with the rise in worst regret for treatment \(b\), as we move from \(\gamma=0\) to \(\gamma=2\), therefore causes the optimal treatment assignment to move from \(\delta=0.59\) to \(\delta = 0.42\) (Table (ref)).
The case with Lee bounds illustrates the opposite situation, where the worst regret from treatment \(a\) grows from \(\$8.7-\$7.9=\$0.8\) per hour to \(\$7.7-\$6.6=\$1.1\) per hour as we move from \(\gamma=0\) to \(\gamma=2\) (Table (ref)). On the other hand, the worst regret from treatment \(b\) shrinks from \(\$7.9-\$7.5=\$0.4\) per hour to \(\$6.6-\$6.5=\$0.1\) per hour as we move from \(\gamma=0\) to \(\gamma=2\). The worst regret for treatment \(b\) is much lower than the worst regret for treatment \(a\) for \(\gamma=0\), which is why the optimal treatment assignment in this case is above \(0.5\), at \(\delta=0.66\) (Table (ref)). The rise in worst regret for treatment \(a\) coupled with the drop in worst regret for treatment \(b\), as we move from \(\gamma=0\) to \(\gamma=2\), causes the optimal treatment assignment to move further up, from \(\delta=0.66\) to \(\delta = 0.95\).
Finally, the case with the Chen and Flores bounds nicely illustrates that both the inequality neutral and inequality averse evaluators may agree in many cases of interest. In particular, whenever both the upper and lower bounds of the egalitarian equivalent treatment effect share the same sign, both evaluators will assign the same treatment: treatment \(a\) if they share a negative sign, and treatment \(b\) if they share a positive sign. Given the treatment effect bounds implied by Table (ref), both evaluators will assign everyone to treatment \(b\) in this case (Table (ref)).
In this Section, I consider the situation where the evaluator does not know the sampling distribution but observes finite data, \(Z^t \in \mathcal{Z^t}\), that are informative about \(s\). A statistical treatment rule \(\delta: \mathcal{Z^t} \rightarrow [0,1]\), gives the probability of a (future) individual being assigned to treatment \(b\), given \(Z^t\). To the extent that knowledge of finite data does not shrink the state space, the only decision problem worth revisiting is that of the Bayesian evaluator, which becomes
and has as solution \(\delta^{B}(Z^t) = 1 \left(E_{\pi}\left[\tau_{ee}(s)|Z^t\right] > 0\right)\), a solution analogous to that obtained in Theorem (ref) for the Bayesian evaluator.
To illustrate how the solution of the problem described by Equation (ref) can be used in practice, I now turn to an application where the evaluator is Bayesian and has access to the outcome of several related randomized experiments. In particular, I examine Meager (2022), who estimates posterior distributions of the effect of microcredit interventions on profit, consumption and other variables using data from randomized trials that expand access to microcredit in seven countries. I take the distribution of consumption before and after treatment to be the primary object of analysis and for this reason I focus below on the five countries in Meager's meta study for which consumption data is available.
Meager reports considerable treatment effect heterogeneity, with large segments of the distribution of consumption nearly unaffected by the policy (from the 5-th to the 75-th percentiles), together with large yet uncertain differences on the upper tails of the distribution of consumption of the treatment and control groups, especially within the group of households with previous business experience. Meager states that, given that the treatment will probably increase inequality, “the social welfare effects of microcredit are likely to be complex.” (Meager (2022), p. 1821). A description of Meager's consumption model for non-zero consumption levels follows.
Let \(d_{ik} \in \{a,b\}\) denote the treatment assignment to individual \(i\) in site \(k\), where \(k = 1,...,5.\) Let MvN represent the multivariate normal distribution, and let \(I\) denotes the identity matrix. We then have:
The interpretation is that, for every site \(k\), consumption is lognormally distributed, with log-mean \(\mu_{k}\) and log-standard deviation \(\sigma_{k}\) for \(y_{ik}(a)\), and log-mean \(\mu_{k} + \zeta_{k}\) and log-standard deviation \(\sigma_{k} \lambda_{k}\) for \(y_{ik}(b)\).\footnote{This description is contained in the file tailored-hierarchical-pdf-log-normal-1-\ tail.stan, which can be found in the repository for Meager (2022), available at \url{https://bitbucket.org/rmeager/aggregating-distributional-treatment-effects/src/master/}.} Let \(y_k(d)\) denote the vector of consumption in site \(k\) given treatment \(d \in \{a,b\}\). Then, with this structure in place, I am able to obtain a closed form solution for the egalitarian equivalent of \(y_k(d)\):
I consider here, as in Section (ref), an evaluator with inequality aversion parameter of \(\gamma = 2\). I then compute mean treatment effects and egalitarian equivalent treatment effects using Meager's Markov Chain Monte Carlo (MCMC) output, denoted \(\hat{\pi}\), which contains three chains with four thousand draws per chain,\footnote{These draws are contained in the file microcredit_consumption_lognormal_tailored\ _hierarchical_pdf_output_5000_iters.RData, which can be found in the paper's repository reported above.} and where I am able to use Equation (ref) to calculate the egalitarian equivalent measures in every draw.\footnote{The mean income measures can also be calculated in every MCMC draw using Equation (ref) with \(\gamma = 0\).}
Figure (ref) reports the posterior distributions of the treatment effects \(\tau(y_k):=E[y_k(b)] - E[y_k(a)]\) and \(\tau_{ee}(y_k):=\mathcal{EE}(y_k(b)) - \mathcal{EE}(y_k(a))\), and Table (ref) reports estimates for \(\hat{\tau}(k):=E_{\hat{\pi}}[\tau(y_k)]\) and \(\hat{\tau}_{ee}(k) := E_{\hat{\pi}}[\tau_{ee}(y_k)]\) for the five countries (\(k=1,...,5\)), where the expectations are computed using the MCMC draws \(\hat{\pi}\). Table (ref) also reports the posterior probabilities, \(P_{\hat{\pi}}[\tau(y_k)>0]\) and \(P_{\hat{\pi}}[\tau_{ee}(y_k)>0]\), that the average and egalitarian equivalent treatment effects are positive according to those draws.
Using the results from Table (ref) one reaches the conclusion that incorporating inequality considerations into the analysis can plausibly reverse the policy recommendation one would make if one focuses solely on what happens on average across the distribution of outcomes, and sharpen the recommendations one would make if one focuses on the quantile treatment effects without further aggregation. To see this, Table 5 presents Meager (2022)'s Bayesian quantile treatment effect estimates on consumption.
We see that the treatment effects on mean consumption reported in Table (ref) are driven by large changes that take place at the top quantiles of the distribution of consumption in Table 5, whereas the egalitarian equivalent treatment effect estimates reported in Table (ref) are driven by small changes that take place at the bottom quantiles of the distribution of consumption in Table 5. The egalitarian equivalent treatment effect approach executes a principled aggregation of the treatment effect heterogeneity across quantiles, according to the attitudes towards inequality of the evaluator, and the aggregation tools developed through Theorem 3.1 in Zambrano (2024) and Theorem (ref) in the present paper.
Given a statistical treatment rule, and before the realization of the sample data, the profile \((ee^{1}(\delta(Z^t)),...,ee^{m}(\delta(Z^t)))\) is a random vector. In order to study this profile, ex ante, we therefore need to extend the preferences of the evaluator over \(\Upsilon_{1}\) to account for the additional sources of uncertanity (sampling uncertanity) that the evaluator faces. As in Manski (2004) (and following Wald (1971)), one can measure the performance of \(\delta(\cdot)\) in state \(s\) by its expected welfare across samples, which in this case means \(E_{s} \left[ee^{s}(\delta(Z^t))\right]\) \(:=\int{ee^{s}(\delta(Z^{t}))dP_{s}^{t}(Z^{t})}\). Not knowing the true state, the evaluator then assesses \(\delta(\cdot)\) by the state-dependent expected egalitarian equivalent vector \(\left(E_{s}\left[ee^{s}(\delta(Z^{t}))\right]\right)_{s \in S}\). Let \(\Gamma\) denote the set of statistical treatment rules and and assume that \(S\) is an open subset of \(\mathbb{R}^{J}\), \(J>0\). Let \(\pi\) here denote a prior measure on \(S\).
The ex-ante versions of Equation (ref), Equation (ref) and Equation (ref) in this setting are
I focus below on obtaining asymptotically optimal solutions to Equation (ref) and Equation (ref). To do so, I consider a sequence of experiments \(\{P^{t}_{s}, s \in S\}\) as the sample size \(t\) grows.
Let \(s_0\) be such that \(\tau_{ee}(s_0) = 0\), noting that the presence of sampling uncertainty may make it difficult to distinguish between treatments in terms of their egalitarian equivalents if the true state is close to \(s_0\). Following Hirano and Porter (2009), I reparametrize the state space and consider parametric sequences of the form \(s_0 + \frac{h}{\sqrt{t}}\) for \(h \in \mathbb{R}^{J}\). The intuition behind the choice is the following: for \(\ddot{s}\) such that \(ee^{\ddot{s}}(a) \ne ee^{\ddot{s}}(b)\), the treatment that is better at \(\ddot{s}\) will be better for all local alternatives \(\ddot{s} + \frac{h}{\sqrt{t}}\) asymptotically. Therefore, a local reparametrization centered around \(s_0\) ensure that we are looking a the cases where it is hardest to determine which is the better treatment even as the sample size grows large. In what follows, I will employ the following assumptions, which are standard in the limit of experiments framework (Le Cam (2012), Van der Vaart (1998), Hirano and Porter (2020)):
The result below is a consequence of Proposition 3.1 and Theorems 3.2, 3.4 and 3.5 in Hirano and Porter (2009):
The interpretation is that, under the structure provided by the assumptions in Theorem (ref), one can act as if all relevant uncertanity has been resolved in large samples, and make a decision that would be optimal if the point estimate \(\hat{s}_t\) of \(s\) were accurate.
Below I revisit the application in Section (ref) under the assumptions behind Theorem (ref), and taking the sample sizes of the five RCTs to be large. From Meager's MCMC output we can obtain Bayesian point estimates of the profile \(\left(\mu_k, \zeta_k, \sigma_k, \lambda_k\right)\) from the Bayesian hierarchical model from Section (ref), and for the five countries \((k = 1, ..., 5)\). Plugging these estimates into Equation (ref), we can then obtain estimates of \(\mathcal{EE}(y_k(d))\), \(d \in \{a,b\}\), and then of \(\tau_{ee}(s)\). Table (ref) below reports the results from this analysis.
These results are essentially identical to those reported in Table (ref) and the conclusions one reaches about the optimal treatment assignment for each of these countries are therefore the same.
My aim with this paper is to contribute towards the development of an integrated theory of how account for inequality in the distribution of treatment effects in experimental and observational settings. To adopt the tools developed above in an applied setting, the recommended workflow would be as follows: First, use the results in Fleurbaey and Zambrano (2024) to help determine which SP under certainty to bring into the analysis. Second, use Theorem 3.1 in Zambrano (2024) to determine how to extend that SP to a world where risk, uncertainty or ambiguity play a prominent role. Third, use the tools in Flores, Kairy, and Zambrano (2024) to properly estimate egalitarian equivalent treatment effects, and bounds, together with corresponding uncertainty estimates for these. Fourth, and last, use the results in Section (ref), Theorem (ref) and Theorem (ref), to identify the optimal treatment assignment rule needed for the statistical decision problem that an inequality sensitive evaluator may want to solve.
An important missing ingredient in the present analysis is that I make no use of covariate information. Introducing covariates is known to help the evaluator obtain better bounds on treatment effects (as in Lee (2009) and Semenova (2023)), but it also opens up the question as to whether the evaluator wishes to account solely for the inequality generated by a treatment conditional on covariates, or also incorporate the inequality generated by a treatment across the multiple values that the covariates may take. Taking a look at both of these important issues is left for future work.