Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
99,804 characters · 18 sections · 106 citation commands
Welfare at Risk: Distributional impact of policy interventions
{JEL classification: C35, C61, D90.} \\ {\bf Keywords:} Welfare, latent utility models, quantiles, superquantiles, compensated variation, treatment effect, generalized Roy models.
\thispagestyle{empty}
\textwidth 5.95in \textheight 600pt {1.1}
Samuelson-Bergson welfare analysis of policy interventions typically focuses on average gains or losses, often overlooking their distributional consequences (e.g., bhattacharya2024nonparametric). Although mean welfare changes provide informative aggregate measures, they can mask substantial heterogeneity across individuals and groups. For instance, a policymaker may want to learn which subpopulations are most affected by a new intervention. Similarly, a policymaker may wish to know whether specific populations of interest are treated fairly in terms of utility gains. Addressing this limitation requires methods that capture distributional variation, such as quantile-based welfare analysis, heterogeneous-agent frameworks, or inequality measures such as the Atkinson index and the Gini coefficient.
A key requirement for assessing distributional heterogeneity in welfare analysis is knowledge of the distribution of individual welfare. Unfortunately, in many settings, this distribution is unknown to the analyst and cannot be inferred from available data. There are at least two reasons for this lack of identification. First, in many environments, the analyst observes only aggregate market data, so only average welfare changes can be identified (Berry1994,BerryLevinsohnPakes1995,Berry_Haile_ECMA2014). Second, even when individual (micro) data are available, outcomes are determined in a potential-outcomes framework in which each agent is observed in only one of two states—treated or untreated. This feature generates a fundamental identification problem: because the same individual cannot be observed in both states simultaneously, the individual welfare change, and hence its distribution, is unidentifiable (HECKMAN2007_Handbook_partI).
In this paper, we develop a distributional welfare framework that accounts for the heterogeneity in welfare generated by policy interventions. In particular, we derive quantile-based bounds that allow the analyst to learn about the distribution of individual welfare changes, even when this distribution is unobserved or unidentified.
Our approach builds on the perturbed utility model (PUM) of Allen_Rehbeck and McFadden_Fosgerau_2012, a latent choice model that accommodates general additive unobserved heterogeneity and nests several well-known frameworks, including the additive random utility model (ARUM) developed in mcfadden1972conditional, McFadden1978, mcf1.\footnote{Beyond ARUM, the PUM class also encompasses bundling models (Gentzkow_AER_2007; iaria2020identification; fox2017note), consumer choice models with latent constraints (agarwal2022demand), and demand systems with additive unobserved heterogeneity (brown2003strong; Matzkin_2007).} The ARUM framework is widely-used due to its tractability and its foundations in discrete choice theory, which facilitate counterfactual analysis (Berry1994; BerryLevinsohnPakes1995; Berry_Haile_ECMA2014). Furthermore, ARUMs play a central role in the treatment-effects literature, where self-selection is modeled as a discrete choice problem. This structure underpins the development of the marginal treatment effect (MTE), which makes explicit the roles of observables and unobservables in individuals’ decisions to participate in a treatment or social program (Bjorklund_Moffitt1987; Heckman_Vytlacil_ECMA_2005_MTE, HECKMAN2007_Handbook_partI; Heckman_Urzua_Vytlacil).
We consider environments in which the analyst observes only aggregate market data or estimates of conditional average welfare changes (such as average treatment effects). In such settings, the distribution of individual welfare changes is not identifiable, motivating the need for a framework that delivers informative bounds. As shown in Allen_Rehbeck, the generalizable PUM framework allows for the identification of conditional average treatment effects, a key requirement for our results.
Our framework relies on the concept of superquantiles, as developed in Rockafellar2000OptimizationOC, ROCKAFELLAR20021443 and Rockafellar_Royset. Superquantiles generalize quantiles by taking expectations over the distribution’s tail beyond a given quantile—for instance, the 20% superquantile reports the average welfare among the bottom 20%. They thus provide quantile-specific information while also summarizing average welfare across the distribution, yielding a richer view of distributional impacts.
We make several contributions. First, Theorem (ref) shows that the superquantile of individual welfare changes—unobservable to the analyst—is bounded above by the superquantile of conditional average welfare changes, which is identified. This result highlights that information on conditional average welfare changes (or average treatment effects) provides valuable insights into the unobserved distribution of individual welfare changes. Economically, this allows the analyst to characterize, for any $\beta$-quantile, the average welfare gain or loss among the $(100\times \beta)\%$ worst affected by a policy change. More importantly, it enables the analyst to identify which subpopulations are most adversely affected and to assess the efficacy and fairness of alternative policy interventions. In addition, Theorem (ref) shows that, under additional assumptions regarding the distribution of individual welfare changes relative to conditional average welfare changes, the superquantile of individual welfare changes is also bounded below. Computationally, we can compute both bounds by solving linear programming problems.
Our second contribution applies the framework to three economic environments in which distributional welfare is central to policy analysis. The first application examines compensated variation (CV) arising from changes in goods prices. Within the PUM class, we derive bounds on the individual CV distributions across subpopulations and show that these bounds reveal heterogeneity in the unobserved CV distribution. More importantly, we demonstrate that conditional average CV — often far easier to obtain than individual-level data — can be used to extract meaningful information about the underlying individual CV distribution. From a data perspective, the analysis highlights that even aggregate market data can be informative about the distributional consequences of price changes, as captured by the CV.
In our second application, we study treatment allocation and welfare maximization when participation is endogenous due to self-selection. Building on Sasaki_Ura_2024, who focus on average welfare maximization, we show how our distributional framework can be combined with the MTE to evaluate the distributional welfare consequences of self-selection and unobserved heterogeneity across relevant populations and subgroups. In particular, we show that a planner with distributional objectives can identify which subpopulations are most adversely affected by a given treatment choice. A key insight of our analysis is that it disentangles the distinct roles of observable characteristics and unobserved heterogeneity in shaping welfare outcomes.
Moreover, our framework enables meaningful comparisons across policies based on their distributional welfare implications and the outcomes of worst-affected subpopulations. This perspective links our approach to regret-based criteria studied by Manski2004, Kiatagawa_Tetenov_2018, and Athey_Wagner_2021, while going beyond average-welfare comparisons by explicitly accounting for potential harm to vulnerable groups and the role of endogenous participation.
In our final application, we analyze social program evaluation in generalized Roy models with subjective participation costs (Heckman_Vytlacil_ECMA_2005_MTE). We extend the cost–benefit framework of Eisenhauer_Heckman_Vytlacil_2015 by incorporating explicit distributional considerations. Although the distributions of benefits, costs, and individual welfare are not directly observable, we show that our bounds are informative about each of these objects. In particular, the identification results in Eisenhauer_Heckman_Vytlacil_2015 can be leveraged to construct bounds that recover informative features of the unobserved benefit, cost, and welfare distributions.
Moreover, we derive bounds characterizing the distributions of costs, benefits, and welfare for agents who participate in the program, thereby providing distributional information for the treated (treatment on the treated). To the best of our knowledge, these results are novel within the context of generalized Roy models with participation costs.
The rest of the paper unfolds as follows. Section (ref) discusses the related literature. Section (ref) briefly presents the relevant aspects of the PUM framework and formalizes the notion of superquantiles, discussing its key properties. Section (ref) develops upper and lower bounds that allow us to learn about the distribution of individual-level welfare changes, which are often unidentified. Section (ref) examines the applications to compensated variation, welfare maximization, treatment allocation, and program evaluation with subjective participation costs. Section (ref) concludes.
Welfare analysis in latent utility models has received significant attention. mcf1 introduced the general ARUM framework, demonstrating its aggregation properties and establishing the existence of a representative agent, making it well-suited for welfare analysis. Small_Rosen_1981 extended this by adapting traditional welfare economics methods to discrete choice models. Hanneman_1996 explored welfare changes in ARUM, emphasizing income effects, while Mcfadden_1996 discussed computational techniques to estimate mean CV changes. McFadden_Fosgerau_2012 further generalized demand and welfare analysis using the PUM framework. However, these studies primarily focus on average welfare and CV analysis, overlooking distributional aspects.
Much work has been done on non-parametric identification of latent utility models. In the context of the PUM model used in this paper, Allen_Rehbeck provides non-parametric identification methods for average welfare changes. Moreover, their results apply to aggregate data, a less demanding data requirement that our results also apply to. However, as with traditional welfare analysis, their results focus on average welfare measures. Therefore, our results are complementary to theirs, as we provide another tool for analyzing the distribution of welfare effects at the individual level that goes beyond their average analysis while leveraging their non-parametric identification results.
Bhattacharya_2015 studies welfare in non-additive random utility models (NARUMs) and identifies the marginal distributions of CV and EV using conditional choice probabilities. Although both their analysis and ours provide non-parametric tools for distributional welfare evaluation, the approaches differ in three main respects. First, their results are specific to NARUMs, whereas our framework applies to PUMs, allowing for richer demand patterns — including substitutability and complementarity. Second, Bhattacharya_2015 requires individual-level choice data, whereas our approach accommodates the standard setting in which the researcher observes only aggregate data (Berry_Haile_ECMA2014). We show that, by leveraging superquantiles, meaningful distributional welfare analysis remains feasible in this environment, making our results complementary to theirs. Third, our framework extends beyond demand to include treatment allocation, welfare maximization, and cost–benefit analysis in generalized Roy models. In the latter case, we show how our bounds deliver informative distributional insights for the subgroup that endogenously selects into treatment.
echenique2024utilitarian studies preference aggregation through Harsanyi’s utilitarian approach, establishing a foundation for distributional welfare measures based on quantiles of individual welfare effects. We differ in four key ways. First, rather than an axiomatic approach, we propose bounds that can be used when the researcher has aggregate data or access to a collection of estimates of conditional average welfare changes. Thus, we exploit available information to learn about the unknown distribution of individual welfare changes. Second, we leverage superquantiles to provide both quantile-specific welfare insights and the corresponding average welfare, thereby enhancing understanding of the welfare distribution. Third, whereas echenique2024utilitarian focuses on the ARUM class, our framework applies to PUMs and therefore to a broader class of models. Finally, they do not discuss either the problem of welfare maximization and treatment allocation, nor the cost-benefit analysis when agents have subjective participation costs.
Finally, our paper relates to the literature on distributional treatment effects. Qi_et_al_2023 and fan2025policylearningalphaexpectedwelfare also employ superquantiles for treatment allocation, but their analysis differs from ours in several ways. They study a binary-treatment environment without unobserved heterogeneity, so MTEs play no role - they restrict attention to linear regression models - and they do not address the problem of bounding the distribution of individual welfare changes or learning about welfare in settings such as compensated variation, welfare maximization with self-selection, or generalized Roy models with subjective participation costs.
The closest paper to ours is Kallus2023. Like us, he uses superquantiles to recover features of the individual welfare distribution from conditional average treatment effects. Nonetheless, the approaches diverge in essential respects. We work with the broader class of PUMs, covering a wide range of demand and behavioral models. We show that superquantiles can recover distributional welfare information even when the analyst observes only aggregate market data—an issue not treated in Kallus2023. We also provide bounds for welfare maximization and treatment choice under self-selection, and we develop a distributional analysis for generalized Roy models with subjective treatment costs, including the distribution of welfare among actual participants. These settings lie outside the scope of Kallus2023. Thus, while related, the two papers address distinct questions and deliver different types of distributional welfare insights.
This section serves two purposes. First, we present the perturbed utility model (PUM), following Allen_Rehbeck and McFadden_Fosgerau_2012. The PUM is a broad class of latent utility models with additively separable unobserved heterogeneity, encompassing widely used formulations such as additive random utility models (ARUMs), bundling models, and matching models. Second, we develop a distributional framework for analyzing welfare in PUMs, extending the analysis beyond average effects. To do so, we briefly review standard quantiles and introduce the less familiar but central concept of superquantiles, which will form the basis of our distributional welfare results.
Consider a decision-maker (DM) making a utility-maximizing choice among an unordered set of alternatives $j\in\mathcal{J}=\left\{1,\ldots, J\right\}$. The DM's optimal choice vector $Y$ satisfies:
where $B\subseteq\mathbb{R}^J$ denotes the DM's budget constraint and $X_j=(X_{j,1},\dots,X_{j,d_j})'$ denotes the observed characteristics of good $j$, which has $d_j$ observed characteristics. These good-specific regressors are collected into $X=(X_1',\dots,X_J')'$, and $X\subset \mathcal{X}$. We note that there can be common regressors across $X_j$ and $X_k$, $j\neq k$ and $j,k\in\mathcal{J}$. Therefore, $X$ can encode observed characteristics of both the goods and the DM. The vector $\vec{u}=(u_1,\dots,u_J)$ encodes how the desirability of a good varies with its regressors. Moreover, $\mathcal{R}$ is called the disturbance function and is a function of the DM's choice, $y$, and unobserved heterogeneity of the agent, $\varepsilon\in E$. In addition, the term $\mathcal{R}$ represents a regularizer term that smooths out choices.
To better see how PUM serves as an umbrella for a large class of models, consider a straightforward ARUM. In an ARUM environment, a DM must choose one of the choices $j\in\mathcal{J}$, whose utility is given by $$v_j=u_j(X_j)+\varepsilon_j$$ Setting $B=\Delta_J$, where $\Delta_J$ is the $J$-dimensional simplex defined as: $$\Delta_J=\Big\{y\in\mathbb{R}^J|\sum\limits_{j=1}^Jy_j=1,\quad y_j\geqslant0\forall j\in\mathcal{J}\Big\}$$ and disturbance function: $$\mathcal{R}(y,\varepsilon)=\sum\limits_{j=1}^Jy_j\varepsilon_j,$$ it follows that the ARUM framework corresponds to a particular instance of a PUM. Furthermore, in the ARUM case it is well-known that under the assumption that $\varepsilon$ is absolutely continuous, the DM's will choose one of the available options $j\in \mathcal{J}$, which corresponds to choosing one of the vertices of $\Delta_J$.
In general, we will focus on the case where $B=\Delta_J$ as above. However, most of our results apply to the general case of $B$ being a closed convex subset of $\mathbb{R}^J$. Following Allen_Rehbeck, we make use of the following assumption throughout the paper.
Throughout the paper, we also make use of the following notation: $$U(y;x,\varepsilon)\triangleq\sum_{j=1}^Jy_ju_j(x_j)+\mathcal{R}(y,\varepsilon)$$ for all $y\in B$ and $X=x$. When $X$ is treated as a random variable, we use $U(y;X,\varepsilon)$.
Allen_Rehbeck show that the PUM aggregates. Based on their aggregation result, we can define the conditional average welfare $W(x)$ as: $$W(x)\triangleq \mathbb{E}\left[\max_{y\in B}U(y;x,\varepsilon)\mid X=x\right].$$
The previous expression can be interpreted as a generalization of the social surplus function introduced by mcf1 in the context of ARUMs. More importantly for our purposes, Allen_Rehbeck show that welfare changes, measured by $W(x^1)-W(x^0)$, are nonparametrically identified using only aggregate market data. We exploit this result as a key building block of our analysis.
Instead of focusing on average welfare (or average welfare changes) across the entire distribution of $\varepsilon$, our focus will be on average welfare among a specific subset of the population, namely the $(100 \times \beta) \%$-worst affected. To formalize this notion, we use the concept of quantiles and the lesser-known concept of superquantiles introduced in Rockafellar2000OptimizationOC, ROCKAFELLAR20021443, and Rockafellar_Royset.
Let $Z$ be a random variable with a finite mean. The $\beta$-quantile corresponds to $F_Z^{-1}(\beta)=\inf \left\{\lambda: F_Z(\lambda) \geq \beta\right\},$ where $F_Z(z)=\mathbb{P}(Z \leq z)$. The $\beta$-superquantile Rockafellar2000OptimizationOC corresponds to the average among the $(100 \times \beta) \%$-lowest outcomes, which formally is defined as:
where $[t]_-\triangleq\min\{t,0\}$. The supremum is attained when $\lambda$ equals to the $\beta-$quantile, which corresponds to $F_Z^{-1}(\beta)=\inf \left\{\lambda: F_Z(\lambda) \geq \beta\right\}, \quad$ where $F_Z(z)=\mathbb{P}(Z \leq z)$.
Provided $F_Z\left(F_Z^{-1}(\beta)\right)=\beta$, i.e. assuming that $Z$ is continuous, then $\mathbb{S}_\beta(Z)=\mathbb{E}\left[Z \mid Z \leq F_Z^{-1}(\beta)\right]$. In the general case, where $Z$ may be discontinuous, we have the general inclusion:
Furthermore, $\mathbb{S}_\beta(Z)$ is continuous in $\beta$, concave, translation invariant, homogeneous, and monotone (Shapiro_et_al_2013). These properties make $\mathbb{S}_\beta(Z)$ the correct way to generalize the “average of the $(100 \times \beta) \%$-lowest values” when ambiguity occurs due to discontinuities. More importantly, as we shall see, the notion of superquantiles allows us to account for the distributional heterogeneity in welfare and in welfare changes.
To garner some intuition, Figure (ref) illustrates the superquantile of a normally distributed random variable, $Z$. Note that $f_Z(\cdot)$ denotes the probability density function of $Z$ while $F_Z(\ cdot) $ denotes the cumulative density function of $Z$. For a given quantile $\beta$, the lower superquantile $\mathbb{S}_\beta(Z)$ would (roughly) be the blue shaded area, normalized by multiplying by ${1\over\beta}.$ While this section has focused on lower superquantiles due to their relevance to the rest of the paper, it is straightforward to adapt (ref) to capture the right tail of the distribution associated with $Z$. The adaptation comes from the fact that the right superquantile can be defined as $\bar{\mathbb{S}}_{\alpha}(Z)=-\mathbb{S}_{1-\beta}(-Z),$ where $\alpha=1-\beta$. In particular, $\bar{\mathbb{S}}_{\alpha}(Z)$ can be written as: $$\overline{\mathbb{S}}_\alpha(Z)=\min_{\lambda\in \mathbb{R}}\left \{\lambda+{1\over 1-\alpha}\mathbb{E}(Z-\lambda)_+\right\}$$ where $(t)_{+}=\max\{t,0\}$.
Using the notion of superquantile, we define the distributional welfare $W_\beta(x)$ as:
The term $W_\beta(x)$ measures the welfare for the $(100\times\beta)\%$ populations worst affected.\footnote{To make this intuition transparent, assume that the distribution of the random variable $M(y;x,\varepsilon)\triangleq \max_{y\in B}U(y;x,\varepsilon)$ is continuous. Then, we can use this assumption to get $W_\beta(x)=\mathbb{E}(M(y;x,\varepsilon)\mid X=x,M(y;x,\varepsilon)\leq F_{M(y;x,\varepsilon)}).$ Accordingly, $W_\beta(x)$ can be expressed as the average welfare for the $(100\times\beta)\%$ at the bottom of the welfare distribution. In the online Appendix, we discuss some basic properties of $W_\beta(x).$}
In this section, we show how superquantiles allow us to move beyond purely average-based analysis. Our objective is to learn about the distribution of individual welfare changes without imposing strong structural assumptions or requiring detailed micro-level data. We focus on the empirically familiar setting in which only aggregate market data is available and demonstrate that the same information used to identify average welfare changes can, through superquantiles, yield informative insights into the distributional consequences of policy interventions.
We begin by defining the individual welfare change as the random variable
where $X \in \mathcal{X}$ denotes observable characteristics, and $t = (t_0, t_1)$ represents the values of a predetermined variable before and after an intervention. For example, $t_0$ and $t_1$ may correspond to prices observed before and after a policy intervention, such as the introduction of an ad valorem tax. Similarly, $t$ encapsulates information about whether an individual has been assigned to a social program or not, in which case $t_1=1$ and $t_0=0$, respectively.
It is worth noting that for a type $(X,\varepsilon)$, the realization of the random variable $\tau(X,t,\varepsilon)$ is unobservable to the analyst. More importantly, in many relevant economic settings, the distribution of $\tau(X,t,\varepsilon)$ is unknown, posing two key identification challenges when analyzing the distribution of individual welfare. The first arises in the typical situation in which the analyst observes only aggregate market data. Because individual decision-making is not available, the conventional approach in such settings is to work with the conditional average welfare change, defined as
where the expectations are taken with respect to $\varepsilon$.\footnote{We note that the second equality uses the assumption that $F_{\varepsilon\mid X,t_1}=F_{\varepsilon\mid X,t_0}=F_{\varepsilon\mid X}$ for all $X\in \mathcal{X}$.}
As shown in Allen_Rehbeck, $\tau(X,t)$ is non-parametrically identified for general distributions of $\varepsilon$ and flexible specifications of the deterministic utility component $u(X)$. However, such identification results are typically uninformative about the distributional implications of policy interventions.\footnote{See bhattacharya2024nonparametric for an excellent survey of approaches to achieve identification of welfare changes in demand models with aggregate and micro-data.}
The second issue arises even when the analyst has access to rich micro-level data. Even though the analyst may be able to observe individuals' decisions, identifying $\tau(X,t,\varepsilon)$ remains impossible due to the evaluation problem. To illustrate, suppose $t_1$ indicates that a worker participates in a training program aimed at increasing productivity, while $t_0$ represents the status quo of non-participation. In this environment, each worker is in only one of these states at a time, but never both. The treatment effect literature has traditionally addressed this challenge by also focusing on average treatment effects, which again does not address the distributional implications of interventions (HECKMAN2007_Handbook_partI).
Therefore, while the direct identification of $\tau(X,t,\varepsilon)$ remains infeasible in the two scenarios mentioned above, our goal is to gain some insight into the distribution of this random variable using the information contained in the conditional average welfare changes, $\tau(X,t)$. For this purpose, we define the $\beta$-superquantile associated with $\tau(X,t,\varepsilon)$ as follows:
Expression ((ref)) is concerned with the left tail of $\tau(X,t,\varepsilon)$, which depends on the joint distribution of $X$ and $\varepsilon$. In words, $\mathbb{S}_\beta(\tau(X,t,\varepsilon))$ quantifies the average change in utility among the bottom ($100\times\beta)\%$ of the population affected by an intervention that modifies $t_0$ to $t_1$. From an economic standpoint, expression ((ref)) can be interpreted as a welfare measure of how the bottom ($100\times\beta)\%$ of all individuals fare from this intervention.
While $\mathbb{S}_{\beta}(\tau(X,t,\varepsilon))$ is economically meaningful, fully identifying it first requires identifying the full distribution of $\tau(X,t,\varepsilon)$, which remains infeasible for aforementioned reasons. In fact, without parametric assumptions, the distribution of $\tau(X,t,\varepsilon)$, and in turn $\mathbb{S}_{\beta}(\tau(X,t,\varepsilon))$, is identifiable only with detailed, individual-level data in the absence of the missing outcome problem.\footnote{It is worth noting that the superquantile functional $\mathbb{S}_\beta(\tau(X,t,\varepsilon))$ is not additive. As a result $\mathbb{S}_\beta(\tau(X,t,\varepsilon))\neq \mathbb{S}_\beta\Big(\max_{y\in B}U(y;X,t_1,\varepsilon)\Big)-\mathbb{S}_\beta\Big(\max_{y\in B}U(y;X,t_0,\varepsilon)\Big)$. Consequently, differences in superquantiles of indirect utility cannot be used to infer the distribution of individual welfare changes.} However, the following theorem provides an upper bound for Eq. ((ref)) that does not rely on parametric assumptions nor detailed individual-level data.
{\em Proof.} All proofs are collected in the Appendix (ref).
Theorem (ref) establishes that the superquantile of individual utility changes can be bounded above by the superquantile of conditional average welfare changes. The significance of bound ((ref)) lies in its empirical feasibility. To reiterate, the left-hand side cannot be estimated without detailed individual-level data and/or strong parametric assumptions. In contrast, the right-hand side can be identified non-parametrically from aggregate data alone. For example, identification can be achieved within the PUM framework using the results of Allen_Rehbeck, or under the class of RUMs following the nonparametric approach of Berry_Haile_ECMA2014. Similarly, bound (ref) is also relevant where the analyst faces the problem of missing outcomes, as discussed above, which persists even in the case of rich micro-level data.\footnote{In sections (ref) and (ref) we discuss this problem in further detail.}
From an economic perspective, Theorem (ref) is valuable because it offers a framework for evaluating the distributional welfare effects of policy interventions. To illustrate, define the total average welfare change as $\tau(t) = \mathbb{E}_{F_{X}}(\tau(X,t))$. Suppose a policy yields $\tau(t) > 0$, but $\mathbb{S}_\beta(\tau(X,t)) < 0$ for a substantial fraction of the population (e.g., $\beta = 0.2$). Comparing these two quantities provides a first-order assessment of whether a policy generates overall welfare gains while imposing losses on specific subgroups. This information might indicate that the policy change yields some large “winners" from the policy, masking harm that other subpopulations may be experiencing. Furthermore, $\mathbb{S}_\beta(\tau(X,t))$ enables us to identify which subpopulations are adversely affected. Assuming continuity, $\mathbb{S}_\beta(\tau(X,t))$ corresponds to the average welfare change among agents with $\tau(X,t) \leq F^{-1}_{\tau(X,t)}(\beta)$. Since $\tau(X,t)$ is identifiable, this approach allows us to identify which observable subgroups experience welfare losses under the policy. It also provides insights into how to improve a policy and minimize these harms. In summary, the key fact behind the bound
Theorem (ref) can be expressed in an equivalent form in terms of the average of quantiles. The following result formalizes this observation.
It is worth remarking that $F^{-1}_{\tau(X,t,\varepsilon)}(\theta)$ corresponds to the $\theta$-quantile of the individual utility changes. In general, as we said earlier, this quantity is not identified. On the other hand, $F^{-1}_{\tau(X,t)}(\theta)$ gives us the $\theta$-quantile of the conditional average welfare changes, as measured by $\tau(X,t)$.
There is one final point worth mentioning regarding Theorem (ref). Notice, we did not exploit any structure of the PUM framework in the derivation of this bound. Therefore, the bound applies beyond the additive structure of the PUM class and is a general bound. However, the key requirement for this bound to be useful is the nonparametric identifiability of $\tau(X,t)$. The main advantage of using the PUM class is that it allows the identification of $\tau(X,t)$ with only aggregate market data. Therefore, while the PUM is not the only class of models for which Theorem (ref) applies, it is the most general class that we know of for which the bound can be non-parametrically estimated using only aggregate data. With rich micro-level data, nonparametric identification of $\tau(X,t)$ is again possible if utilities are additive (see, e.g., Bhattacharya_2015,bhattacharya2024nonparametric) or under certain conditions if utilities are non-additive (see e.g, imbens2009identification).
We now discuss how to lower bound $\mathbb{S}_\beta(\tau(X,t,\varepsilon))$. The following theorem provides a simple method to do so, assuming an additional condition.
The previous result shows that when the difference between $\tau(X,t,\varepsilon)$ and $\tau(X,t)$ is uniformly bounded by a constant $\gamma$, then $\mathbb{S}_\beta(\tau(X,t,\varepsilon)$ is lower bounded by $\mathbb{S}_\beta(\tau(X,t))-\gamma$, a quantity that is identified as long as $\gamma$ is known to the analyst (Manski1990).
It is worth emphasizing that Theorems (ref) and (ref) enable the analyst to learn about the distribution of individual welfare changes, even when only aggregate market data or average welfare measures are available. Theorem 1 does so without any knowledge of the distribution of unobservables nor any parametric assumptions; Theorem 2 requires an assumption on the range of treatment effects possible, conditional on observables. To fix ideas, consider the case in which the PUM reduces to a standard ARUM. In this setting, bhattacharya2024nonparametric argues that average welfare and observed choice probabilities are uninformative about the welfare change distribution. Our results overturn this conclusion: average welfare changes do contain information about the underlying individual welfare distribution. The source of this difference is methodological. We extract information from average welfare changes themselves, whereas bhattacharya2024nonparametric relies solely on choice probabilities, thereby discarding welfare-relevant variation.
We conclude this section by discussing a simple approach to compute $\mathbb{S}_\beta(\tau(X,t))$. Suppose the analyst has access to a collection of K estimates $\hat{\tau}(X_1), \ldots, \hat{\tau}(X_K)$, where, for ease of exposition, we omit the dependence on $t$. Using these estimates, a direct approach to computing $\mathbb{S}_\beta(\tau(X,t))$ is to employ the variational representation (ref) together with a plug-in method, yielding the following program:
Problem (ref) is concave but, unfortunately, non-smooth in $\lambda$. However, we can avoid the analytical complications introduced by non-smoothness by observing that program (ref) can be equivalently expressed as a linear programming problem:
Program (ref) offers a straightforward way to compute $\hat{\lambda}$ and $\hat{\mathbb{S}}_\beta(\hat{\tau}(X,t))$. The problem involves a linear objective function with only $2K$ linear constraints, making it computationally simple and easy to implement with standard optimization software. The value of $\hat{\lambda}$ can be interpreted as an estimator of the $\beta$-quantile of the random variable $\tau(X,t)$, while the value of the objective function at the optimum, $V(\hat{\lambda})$, provides an estimate of $\mathbb{S}_\beta(\tau(X,t))$.
Because this approach is based on a set of pre-estimated values $\hat{\tau}(X_1), \ldots, \hat{\tau}(X_K)$, it effectively turns the estimation of $\mathbb{S}_\beta(\tau(X,t))$ into a linear program with estimated inputs. This feature introduces sampling variability that must be accounted for in inference. To handle this, one can apply the inferential methods developed in Shum_2022.
While the results in the preceding sections are general, this section illustrates how they apply to standard economic settings. We consider three applications. First, we examine the distributional properties of CV and discuss how information in conditional-average CV helps infer the distributional impacts of exogenous price changes. Second, we apply our framework to the analysis of distributional welfare in treatment allocation problems, demonstrating how our results account for unobserved heterogeneity and self-selection. Finally, we analyze treatment effects in a setting with endogenous participation and subjective treatment participation costs. In this context, we focus on the benefits, costs, and welfare implications of social programs.
We begin by discussing how to specialize the results from the previous section to the problem of compensated variation ($\operatorname{CV}$). In the ARUM literature, the problem of $\operatorname{CV}$ is well studied. Most studies of CV have two particular features. First, they focus on the average case. Second, they focus on the traditional ARUM. We aim to relax these two assumptions while fixing the constant marginal utility of income assumption.
We consider a setting with $J \geq 2$ goods. Let $X_j$ denote the observed non-price characteristics of good $j$, and let $\mathcal{X}_j$ represent the space of covariates associated with that good. The vector of covariates contains the qualities and prices of the different goods. The overall space of non-price covariates is $\mathcal{X} \triangleq \prod_{j=1}^J \mathcal{X}_j$. Let $p = (p_1, \ldots, p_J)\in \mathbb{R}_{+}^J$ denote the price vector, where $p_j$ is the price of good $j$ with $j = 1, \ldots, J$. For a given pair $(X,p)\in \mathcal{X}\times\mathbb{R}_{+}^J$, we assume that observable utility corresponds to the term $h_j(X_j)+\gamma(I-p_j)$ where $I>0$ is consumer's income and $\gamma$ denotes the marginal utility of income.
Our goal is to understand the distributional implications of an intervention that changes the prices of goods. In particular, let $p_0 = (p_{10}, \ldots, p_{J0})$ and $p_1 = (p_{11}, \ldots, p_{J1})$ denote the vectors of prices before and after the intervention, respectively. Note that in terms of our original notation $p_1=t_1$ and $p_0=t_0$ respectively. Accordingly, we use the notation $t=(p_0,p_1).$
Throughout the analysis, we assume that the non-price characteristics $X$ remain fixed and do not respond to the intervention.
Accordingly, at the individual level, we have the following condition that monetarily quantifies the impact of changing prices:
where $\operatorname{CV}$ is the CV necessary to keep the consumer with characteristics $X$ and unobserved tastes $\varepsilon$ with the same indirect utility level. Exploiting the quasi-linearity of the utility function, we obtain a simple analytical expression for the individual $\operatorname{CV}$, which we denote as $\operatorname{CV}(X,t, \varepsilon)$ to emphasize the dependence on $X$, $t$, and $\varepsilon$:
where $U(y;X,p_l,\varepsilon)\triangleq \sum_{j=1}^Jy_ju_j(X_j,p_{jl})+D(y,\varepsilon)$ and $u_j(X_j,p_{jl})=h_j(X_j)-\gamma p_{jl}$ for $l=0,1.$\\
Note that due to the quasilinear structure, income effects do not impact $\operatorname{CV}(X,t,\varepsilon)$. However, $\operatorname{CV}(X,t,\varepsilon)$ can capture rich patterns of complementarity and substitutability across goods, as well as general additive forms of unobserved heterogeneity. Furthermore, $\operatorname{CV}(X,t,\varepsilon)$ is not restricted to the case of discrete choice problems. In fact, when $\mathcal{R}(y,\varepsilon)=\sum_{j=1}^Jy_j\varepsilon_j$ and $B=\Delta_J$, expression ((ref)) yields the $\operatorname{CV}$ in ARUMs. In other words, (ref) applies when vector $y$ refers to continuous quantities, without imposing a particular structure on $B$, being the discrete choice model a particular case.
Taking conditional expectation with respect to $\varepsilon$ in ((ref)) we find:
The average $\operatorname{CV}(X)$ is commonly used when the analyst has access only to aggregate market data (see, for example, Berry_Haile_ECMA2014). However, this measure is uninformative about the potential distributional consequences of price changes. The following result shows how Theorems (ref) and (ref) in Section (ref) can be applied in the $\operatorname{CV}$ context to learn about the distributional consequences of price changes.
Some remarks are in order. First, the left-hand side in (ref) corresponds to the $\beta$-superquantile of individual compensating variation across all realizations of $X$ and $\varepsilon$; the right-hand side is the $\beta$-superquantile of the average conditional compensating variation across all realizations of $X$ only. Therefore, Proposition (ref)$(i)$ informs us how the individual $\operatorname{CV}$ across an entire population can be bounded using the information contained in the $\operatorname{CV}(X,t)$s. In doing so, our analysis is refined to identify harmed (or worst-affected) subpopulations defined by specific covariates $X$, which is relevant for fairness and equity considerations in market interventions. This follows from interpreting $\mathbb{S}_\beta(\operatorname{CV}(X,t))$ as a summary of heterogeneity along realizations of the relevant regressors in $X$. Moreover, identifying the worst-affected subpopulations enables more targeted interventions.
A second important observation is that the bound ((ref)) can be identified using aggregate data. For instance, Berry_Haile_ECMA2014 and Allen_Rehbeck show that welfare averages can be identified in demand models with unobserved heterogeneity. Thus, bound ((ref)) is informative about the superquantile of the distribution of individual $\operatorname{CV}$ under minimal data and distributional assumptions. Proposition (ref)(ii) complements this finding by providing a lower bound ((ref)) to the unobserved $\operatorname{CV}$ distribution, which depends on observables where the parameter $\mu$ can represent an upper bound in the $\operatorname{CV}$ that the social planner can allocate (Manski1990).
At a higher level, Proposition (ref) is informative for learning about the distribution of individual $\operatorname{CV}$. As discussed earlier, our analysis offers a direct counterpoint to bhattacharya2024nonparametric, who argues that average welfare measures and conditional choice probabilities are uninformative about the distribution of individual welfare changes—even under additively separable unobserved heterogeneity, as in ARUMs. Proposition (ref) demonstrates that, within the PUM class, aggregate market data are sufficient to conduct meaningful distributional $\operatorname{CV}$ analysis.
In this section, we apply our framework to the analysis of welfare in the allocation of social programs when participation decisions are endogenous due to self-selection. Following the approach in Sasaki_Ura_2024, our objective is to demonstrate how the MTE can be used to assess the distributional welfare consequences of self-selection and unobserved heterogeneity across different populations and groups of interest. Below, we present the model as described in Sasaki_Ura_2024.
We consider the following causal model:
where $V$ denotes an observed outcome variable, $D$ denotes an observed binary treatment variable, $Z$ denotes a vector of observed exogenous variables, $V_0$ and $V_1$ denote unobserved potential outcomes under no treatment and under treatment, respectively, and $\tilde{\varepsilon}$ denotes an unobserved factor (unobserved heterogeneity) of the treatment selection. Let $\mathcal{Z}$ denote the set of all observables covariates $Z.$
Equation ((ref)) models the outcome production using the potential-outcome framework, while equation ((ref)) models the treatment selection via a binary RUM (threshold-crossing) model. Note that ((ref)) is a particular instance of the PUM.
The function $\tilde{u}$ in the assignment model ((ref)) is nonparametric and is unknown to the econometrician. The model allows for endogeneity (unobserved confoundedness) in the sense that $(V_0, V_1)$ and $\tilde{\varepsilon}$ may be statistically dependent even when conditioned on $Z$. To achieve identification, we assume the vector $Z$ to contain excluded exogenous variables (i.e., excluded instruments) as well as included exogenous variables. Similar to Sasaki_Ura_2024, we use the following assumption:
Part (i) concerns the treatment assignment model ((ref)) solely, and this is the only independence assumption to be imposed on the model, implying that we can allow for an arbitrary statistical dependence between the potential outcomes $(V_0, V_1)$ and $\tilde{\varepsilon}$, even conditional on $Z$. Part (ii) states the exclusion restriction of the random subvector $Z_0$ of $Z$, and bounded second moments of the potential outcomes $\left(V_0, V_1\right)$. Part (iii) rules out point masses and holes in the conditional distribution of $\tilde{\varepsilon}$ given $X$. For ease of exposition, we define $\mathcal{X}$ as the set of all observable covariates $X$.
Following the literature on the marginal treatment effect (MTE), it is standard to apply normalizing transformations ($\varepsilon \equiv F_{\tilde{\varepsilon} \mid X}(\tilde{\varepsilon})$ and $u(Z) \equiv F_{\tilde{\varepsilon} \mid X}(\tilde{u}(Z))$) in the threshold crossing model ((ref)). An important implication of Assumption (ref) is that $D=1\{u(Z)-\varepsilon \geq 0\}$ and $\varepsilon$ is distributed uniformly over $[0,1]$ conditional on $Z$. Accordingly, and without loss of generality, the threshold-crossing treatment selection model ((ref)) can be equivalently expressed as
As a consequence we will use ((ref)) in place of the original model ((ref)).
Now we study the case where the planner chooses a non-randomized policy/rule that maps $Z$ to treatment status. Formally we consider maps $\pi:\mathcal{Z}\mapsto \{0,1\}$;.The set of all Borel measurable functions from $\mathcal{Z}$ to $\{0,1\}$ is denoted by $\Pi$. For $\pi\in \Pi$, $V(\pi(Z))$ denotes the utility that an individual with covariates $Z$ derives from the policy $\pi$. In particular, we can express $V(\pi(Z))$ as
Expression ((ref)) represents the individual outcome of an individual with observables $Z$ when planner chooses $\pi\in \Pi.$ The role of the treatment assignment rule is to assign an individual with covariates $Z$ to the treatment, i.e., $D=1$, whenever $\pi(Z)=1$.
It is well known that the distribution of $V(\pi)$ is unidentified, a direct consequence of the missing-outcome problem. As a result, the policy-learning and welfare-maximization literature has largely focused on average welfare as the criterion for selecting an optimal rule $\pi \in \Pi$. Formally, we define the average welfare function $\mathcal{W}:\Pi\mapsto\mathbb{R}$
Let $\mathcal{W}(\pi, z)\triangleq\mathbb{E}\left[\mathcal{W}(\pi)\mid Z=z\right]=\mathbb{E}(V_0\mid Z=z)+\pi(Z)\mathbb{E}(V_1-V_0\mid Z=z)$ denote the conditional welfare associated with $\pi$ for a fixed realization $Z = z$.\footnote{ This definition makes explicit the fact that $\mathcal{W}(\cdot,\cdot)$ depends on the choice of $\pi$ and in the conditional value $z$. However, in our formal statements and proofs we work with the random variable $\mathcal{W}(\pi,Z)\triangleq\mathbb{E}\left[\mathcal{W}(\pi)\mid Z\right]$.}
Finally, we define the MTE as follows: $$\operatorname{MTE}( x, \bar{\varepsilon})=\mathbb{E}\left[V_1-V_0 \mid X=x, \varepsilon=\bar{\varepsilon}\right].$$
Under Assumption (ref), Sasaki_Ura_2024 shows that
Representation ((ref)) characterizes the average welfare associated with policy $\pi$. However, it is silent about the distributional welfare implications of implementing such a rule.
The following result shows that conditional average welfare contains meaningful information about the distributional consequences of policy $\pi$. Specifically, it establishes that the conditional average welfare is informative for bounding the distribution of individual welfare effects and, therefore, for assessing the potential harm faced by different subpopulations under the policy.
The previous result provides a tractable way to bound the distributional consequences of choosing a policy $\pi$. In particular, Theorem (ref) allows the analyst to characterize the $(100\times \beta)\%$ worst-off populations under policy $\pi$. A key feature of bound ((ref)) is that it highlights how unobserved heterogeneity $\varepsilon$ and the $\operatorname{MTE}$ shape the distributional impact of the chosen policy, even though the full distribution of $V(\pi)$ is not identified.
Another important implication of our framework is that it facilitates comparisons across policies regarding the potential harm that particular groups may experience. To formalize this idea, consider two policies $\pi$ and $\pi’ \in \Pi$. We are interested in the change in individual utility resulting from switching from $\pi$ to $\pi’$, given by
Because the distribution of $V(\pi’) - V(\pi)$ is not identified—due to the missing outcome problem—direct learning about its distribution is infeasible. However, the following result shows that the information contained in the MTE, together with the policies $\pi$ and $\pi’$, allows the analyst to bound the distributional consequences of moving from $\pi$ to $\pi’$.
The previous result provides a simple condition that allows the planner to compare two policies. In particular, Proposition (ref) provides a bound on the distribution of individual welfare changes associated with switching from policy $\pi$ to $\pi’$. These bounds enable the policymaker to assess the distributional consequences of alternative policies, with a particular focus on identifying the groups most affected. From an econometric standpoint, bound ((ref)) can be implemented using the results in byambadalai2022welfaregains.
It is worth noting that bound ((ref)) is closely related to regret-based criteria studied by Manski2004, Kiatagawa_Tetenov_2018, and Athey_Wagner_2021. To see this, let $\pi^\ast$ be the policy that maximizes welfare in the lower tail of the distribution. By applying the bound ((ref)) with $\pi^\prime = \pi^\ast$, Proposition (ref) allows us to bound the welfare loss from choosing an alternative policy $ \pi$ relative to $\pi^\ast$. In this sense, our framework provides distributional analogues of regret measures, informing policymakers about regret in terms of worst-case welfare outcomes for certain subpopulations of interest rather than average regret.
Finally, we note that Theorem (ref) and Proposition (ref) extend straightforwardly to settings with multi-valued treatments.
In this section, we consider the welfare framework introduced by Eisenhauer_Heckman_Vytlacil_2015, which, in the context of treatment effects, examines the marginal benefits and marginal costs of policies. Their framework extends the modern treatment effect literature by providing a method to identify both the marginal benefits and the marginal costs of policy interventions. In particular, they incorporate agents’ subjective costs associated with participation in social programs, which allows them to analyze the benefits, costs, and surplus (benefits minus costs) of specific social programs in average terms. Our goal is to integrate their results with our framework to explore how policymakers can assess not only the average benefits, costs, and surpluses of different policies, but also their distributional implications.
We focus in the traditional binary case. As in Section (ref), we assume that are two potential outcomes $\left(V_0, V_1\right)$ and a choice indicator $D$, with $D=1$ if the agent selects into treatment so that $V_1$ is observed and $D=0$ if the agent does not select into treatment so that $V_0$ is observed. Similar to section (ref), we use the potential outcome equation to denote the value of $V$ as $V=DV_1+(1-D)V_0$. Following Eisenhauer_Heckman_Vytlacil_2015, we assume a separable structure in the outcomes where $\mathbb{E}\left(V_j \mid X\right)=\mu_j(X)$ and
As in previous sections, $X$ is a (random) vector of covariates observed by the analyst, while ($\nu_0, \nu_1$) are unobserved (to the analyst) heterogeneity terms. Combining $V=DV_1+(1-D)V_0$ with ((ref)) we get
Let $B=V_1-V_0$ denote the individual gross benefit of treatment, defined as the causal effect on $V$ of moving an otherwise identical individual from state 0 to state 1. Thus, $B$ measures the ceteris paribus change in the outcome induced by treatment.
Let $C$ denote the agent’s subjective treatment cost, defined as:
where $Z$ represents an observed random vector of cost shifters and $\nu_C$ is a random variable unobserved by the analyst. Under this specification, the conditional expectation in Eq. ((ref)) satisfies $\mathbb{E}[C \mid Z] = \mu_C(Z)$.
Individuals choose to participate in the treatment if the perceived benefit from participation is greater than the subjective cost:
where $\mathcal{W}$ is the individual welfare (surplus), that is, the net benefit, from treatment:
$$
$$ with $\mu_\mathcal{W}(X, Z)=\left[\mu_1(X)-\mu_0(X)\right]-\mu_C(Z)$ and $\varepsilon_{\mathcal{W}}=\nu_C-\left(\nu_1-\nu_0\right)$.
Our distributional analysis of treatment costs and welfare does not impose functional-form restrictions on $\mu_0$, $\mu_1$, or $\mu_C$, nor does it require parametric assumptions on the distributions of $\nu_0$, $\nu_1$, or $\nu_C$. Instead, our results build on the identification framework of Eisenhauer_Heckman_Vytlacil_2015, which enables the analyst to recover the relevant conditional average objects needed for our bounds.
Similar to section (ref), it is easy to see that the choice model ((ref)) corresponds to a particular instance of the PUM. To formalize this, let $ P(X, Z)\triangleq \operatorname{Pr}(D=1 \mid X$, $Z)$ denote the probability of selecting into treatment given $(X,Z)$. Given the structure of this cross-threshold model, we note that $P(X, Z)=F_\varepsilon\left(\mu_\mathcal{W}(X, Z)\right)$, where $F_{\varepsilon_{\mathcal{W}}}(\cdot)$ denotes the distribution of $\varepsilon_{\mathcal{W}}$. For ease of exposition, we denote $P(X, Z)$ by $P$, suppressing the $(X, Z)$ argument. In addition, we use the fact that $U_\mathcal{W}=$ $F_{\varepsilon_\mathcal{W}}(\varepsilon_{\mathcal{W}})$ is a uniform random variable. In particular, different values of $\overline{\varepsilon}_\mathcal{W}$ denote different quantiles of $\varepsilon$. Given our previous assumptions, $F_{\varepsilon_\mathcal{W}}$ is strictly increasing, and $P(X, Z)$ is a continuous random variable conditional on $X$. Throughout this section, we assume the following.
Our goal is to learn about the distributions of benefits, costs, and welfare—objects that are unobserved by the analyst. To do so, we rely on several key parameters introduced by Eisenhauer_Heckman_Vytlacil_2015, which provide identification of the relevant conditional average objects that our analysis builds upon.
The first parameter is the conditional average treatment effect (ATE) benefit given by: $$ B^{\operatorname{ATE}}(x) \triangleq \mathbb{E}\left(Y_1-Y_0 \mid X=x\right)=\mu_1(x)-\mu_0(x) . $$
$B^{\operatorname{ATE}}(x)$ denotes the ATE for individuals with characteristics $X = x$: that is, the causal effect of assigning treatment randomly to all individuals of type $x$, under full compliance and abstracting from general equilibrium or spillover effects.
The second parameter of interest is the average treatment benefit for individuals who actually receive the treatment, commonly referred to as the benefit of treatment on the treated. $$
$$
Heckman_Vytlacil_1999,Heckman_Vytlacil_ECMA_2005_MTE show that a uniform approach to $B^{\mathrm{ATE}}(x)$ and $B^{\mathrm{TT}}(x)$ is possible by using the MTE parameter, which is defined as:
$$
$$ The function $B^{\mathrm{MTE}}\left(x, u_\mathcal{W}\right)$ is the treatment effect parameter that conditions the unobserved desire to select into treatment.
Eisenhauer_Heckman_Vytlacil_2015 note that conventional treatment-effects analysis does not define, identify, or estimate any component of treatment costs. To fill this gap, they introduce three cost parameters: the average cost of treatment, the average cost of treatment for those who select into treatment, and the marginal cost of treatment.
$$
$$
Finally, we define a set of welfare parameters. Recalling that $\mathcal{W}=B-C=\mu_\mathcal{W}(X, Z)-\varepsilon_{\mathcal{W}}$ we get: $$
$$ and $$
$$
Eisenhauer_Heckman_Vytlacil_2015 show how to identify the previous parameters. Proposition (ref) below establishes that their average identification results help us to learn about the distributional aspects of the welfare distribution.
The result in Proposition (ref) shows that $\mathcal{W}^{\operatorname{ATE}}(X,Z)$ and $\mathcal{W}^{\operatorname{MTE}}(X,Z,U_{\mathcal{W}})$ constitute the best available approximations to the unobserved welfare effect $\mathcal{W}$. Consequently, they can be used to study the distributional behavior of $\mathcal{W}$ and to bound potential welfare losses for specific subpopulations. Intuitively, $\mathbb{S}_{\beta}(\mathcal{W})$ captures the welfare effect among the worst-off $(100\times\beta)\%$ of individuals, while $\mathbb{S}_{\beta}(\mathcal{W}^{\operatorname{ATE}}(X,Z))$ in bound ((ref)) captures the worst outcomes only across groups defined by $(X,Z)$. Similarly, in bound ((ref)), $\mathbb{S}_{\beta}(\mathcal{W}^{\operatorname{MTE}}(X,Z,U_{\mathcal{W}}))$ captures the worst outcomes only across groups defined by $\left(X, Z, U_{\mathcal{W}}\right)$. This distinction highlights the role of unobserved heterogeneity in assessing distributional impacts across subgroups. To the best of our knowledge, bounds in Proposition (ref) are new in the context of generalized Roy models with participation costs.
This distinction highlights the role of unobserved heterogeneity in assessing distributional impacts across subgroups. To the best of our knowledge, bounds in Proposition (ref) are new in the context of generalized Roy models with participation costs.
It is worth remarking that the bounds ((ref)) and ((ref)) explicitly incorporate the cost of treatment participation. Consequently, the result provides information about the potential welfare losses faced by different subpopulations as a function of observables $(X,Z)$, the cost $C$, and unobserved heterogeneity $U_{\mathcal{W}}$. Furthermore, our framework extends Eisenhauer_Heckman_Vytlacil_2015's results by providing lower bounds that identify which groups incur the highest costs. To do so, we focus on the right superquantile $\overline{\mathbb{S}}_\alpha(\cdot)$, which—as discussed earlier—captures the average outcome in the upper tail of the distribution and is therefore the appropriate object for assessing individuals who face the highest participation costs.
The previous result provides information on the ATE and MTE costs of the $(100\times (1-\alpha ))\%$ worst affected. Expression ((ref)) provides a bound in terms of observables $Z$. Intuitively, this lower bound allows the analyst learn which particular subgroups are bearing a higher cost. Accordingly, this bound can guide the reduction of treatment costs in target subgroups, which can be welfare-improving. Similarly, bound ((ref)) complements the previous analysis by incorporating the role of unobservables.
Finally, we point out that Eisenhauer_Heckman_Vytlacil_2015's analysis is conducted in terms of average costs. Proposition (ref) shows that their results are informative about the distributional implications of treatment costs, considering both observables and unobservables.
Our final result concerns the distributional properties of $\mathcal{W}^{\operatorname{TT}}$ and $C^{\operatorname{TT}}$.\footnote{For ease of exposition, we omit the analysis of $B^{\operatorname{TT}}$. However, all arguments extend directly to that case.} Because of the missing-outcome problem, neither distribution is observable to the analyst (HECKMAN2007_Handbook_partI). Our objective is to exploit the identification results in Eisenhauer_Heckman_Vytlacil_2015 for the parameters $\mathcal{W}^{\operatorname{TT}}(X,Z)$ and $C^{\operatorname{TT}}(Z)$ in order to recover informative bounds on the unknown distributions of $\mathcal{W}^{\operatorname{TT}}$ and $C^{\operatorname{TT}}$.
In doing so, we make use of the following notation:
In a similar way, for the cost $C$ we define :
Intuitively, $\mathbb{S}_\beta\!\left(\mathcal{W}^{\mathrm{TT}}\right)$ is the superquantile of the welfare distribution for the treated. It equals the average welfare among the $(100\times \beta)\% $ of treated individuals who experience the lowest welfare gains, including welfare losses. Similarly, $\overline{\mathbb{S}}_\alpha(C^{\mathrm{TT}})$ is the superquantile of the treatment-cost distribution for the treated. It corresponds to the average treatment cost for the $(100\times (1-\alpha))\%$ of treated individuals who incur the highest program costs.
The following is the main result of this section.
Some remarks are in order. First, part (i) bounds the welfare of the $(100\times \beta)\%$ worst-affected individuals in the population. To see the relevance of bound ((ref)), assume without loss of generality that $\mathcal{W}$ has a continuous distribution. In this case, $\mathbb{S}_\beta(\mathcal{W}^{\mathrm{TT}}) = \mathbb{E}\!\left[\mathcal{W}\mid \mathcal{W}\le F_{\mathcal{W}}^{-1}(\beta),\, D=1\right]$. However, because of the fundamental problem of missing outcomes, neither the distribution of $\mathcal{W}$ nor its conditional distribution (conditional on D=1) is identified. Thus $\mathbb{E}\!\left[\mathcal{W}\mid \mathcal{W}\le F_{\mathcal{W}}^{-1}(\beta), D=1\right]$ is not observed by the analyst. By contrast, using the identified function $\mathcal{W}^{\mathrm{TT}}(X,Z)$, we can construct the upper bound $\mathbb{E}\!\left[\mathcal{W}(X,Z)\mid \mathcal{W}(X,Z)\le F_{\mathcal{W}(X,Z)}^{-1}(\beta),\, D=1\right]$, which provides an informative measure of the potential gains or losses experienced by the bottom $(100\times \beta)\%$ of the welfare distribution among treated individuals. Thus, based on observables, the policymaker can identify which treated subgroups are likely to experience declines in welfare.
Part (ii) establishes a lower bound on the treatment cost among individuals who received the treatment. As with $\mathcal{W}^{\mathrm{TT}}$, the conditional distribution of $C^{\mathrm{TT}}$ is not identified. However, using the results in Eisenhauer_Heckman_Vytlacil_2015, the lower bound ((ref)) is identified and can be computed using the linear programming procedure described in Section (ref). From an economic perspective, the lower bound ((ref)) is informative along at least two dimensions. First, it enables the analyst to determine which subpopulations bear higher treatment costs among the treated. Second, it sheds light on the fairness of treatment costs across groups—for example, whether a particular intervention is regressive or progressive.
We close this section by noting that, to the best of our knowledge, Theorem (ref) is new in the cost-benefit analysis of treatment effects.
This paper develops a distributional framework for analyzing welfare heterogeneity across individuals. By combining the concept of superquantiles with observable-group average welfare effects, we show how to construct informative bounds on individual welfare changes without relying on strong structural assumptions or detailed micro-level data.
Although the framework is broadly applicable, much of the analysis is conducted within the class of perturbed utility models (PUMs), which offer a empirically tractable setting for nonparametrically estimating average welfare effects using standard data. We illustrate the usefulness of the framework in three economic environments: compensating variation from price changes, treatment decisions with endogenous participation, and social programs modeled through generalized Roy models with subjective participation costs. Across these applications, the core result established in Section (ref) extends naturally, even though the relevant welfare concepts and identification strategies differ.